A hyperspectral anomaly detection method based on fractional field perception and contrast statistical regularization network

CN122821166APending Publication Date: 2026-09-25UNIV OF SCI & TECH BEIJING +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611039582.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

尽管取得了这些进展,大多数基于深度学习的高光谱异常检测方法仍工作在像元级别或依赖相对较浅的空谱融合,缺乏对频域的显式建模,因此无法利用对有效异常背景判别至关重要的频域结构差异

Benefits of technology

[0046]首先,设计了LFrFT模块,通过可学习分数阶参数自适应调制光谱频率分量,有效捕获弱异常的细微频域特性。其次,引入MSSE模块,通过深度可分离卷积提取多尺度空间上下文信息并融合光谱特征,增强异常目标的空谱相关性。第三,采用结合基于对比学习的投影头的SACL动态缓解特征分布偏移并提高跨场景背景与异常特征的语义可分性。在六个真实高光谱数据集上的广泛实验证明了所提FDCRN的有效性和鲁棒性。该方法在所有评估数据集上持续取得最高的AUC值,超过0.97,在弱目标检测精度和复杂场景泛化性能方面优于七种最先进的对比方法,同时保持竞争力的计算效率。未来工作将探索FDCRN与高光谱超分辨率技术的集成以实现亚像元异常检测,以及将所提框架扩展到多时相高光谱数据分析。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821166A_ABST
    Figure CN122821166A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of hyperspectral image anomaly detection, in particular to a hyperspectral anomaly detection method based on a fractional domain perception and contrast statistical regularization network, which is based on a deep background suppression framework and optimizes anomaly detection performance from three dimensions of frequency domain enhancement, space-spectrum fusion and distribution correction. First, a lightweight fractional Fourier transform module is designed, the spectral frequency components are adaptively modulated through a learnable fractional order parameter, and frequency domain features are fused to enhance the capture ability of subtle differences of weak small targets; second, a multi-scale space-spectrum enhancement module is proposed, deep separable convolution is used to extract multi-scale spatial information and fuse spectral features, and the space-spectrum correlation of weak small targets is enhanced; third, a statistical adaptive and contrast learning module is introduced, feature distribution deviation is dynamically corrected, and the semantic distance between the background and the anomaly is increased. The application is superior to the contrast method in terms of weak target detection precision and complex scene generalization performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hyperspectral image anomaly detection technology, and in particular to a hyperspectral anomaly detection method based on fractional domain perception and contrastive statistical regularization network, which is especially suitable for the detection and identification of weak and small anomalous targets in complex background scenes. Background Technology

[0002] Hyperspectral remote sensing technology plays a crucial role in fields such as ecological environment monitoring, geological disaster early warning, and military target identification. Among these, hyperspectral anomaly detection is one of the core technologies supporting remote sensing Earth observation and resource exploration. Hyperspectral images capture ground feature information across dozens to hundreds of continuous spectral bands, with each pixel corresponding to a unique spectral curve, accurately representing the material composition and physical properties of ground features. This capability makes it possible to detect subtle differences in ground features that are difficult to distinguish using traditional imaging methods.

[0003] However, hyperspectral images face challenges such as high dimensionality and abundant redundant information, complex land cover, mixed pixels containing multiple land features, and limited spatial resolution. These characteristics make it difficult to directly apply traditional detection methods to hyperspectral data. Therefore, developing anomaly detection techniques specifically targeting hyperspectral characteristics is of great significance for improving the efficiency of remote sensing data utilization and expanding practical applications.

[0004] In its early stages, hyperspectral anomaly detection primarily focused on statistical modeling and representation learning. The classic benchmark in this field is the Reed-Xiaoli (RX) algorithm. However, the RX algorithm's dependence on global statistics makes it perform poorly in scenes with non-uniform backgrounds, as its assumed Gaussian distribution often deviates from the actual ground cover distribution. To address these limitations, researchers have proposed several improved methods, such as the Local Window Adaptive RX (LAIRX) algorithm, which focuses on local background statistics to mitigate distribution heterogeneity, and the Kernelized RX (KRX) method, which handles nonlinear relationships through kernel mapping.

[0005] Although RX and its variants are based on statistical modeling of the background distribution, they struggle to adapt to the nonlinear and heterogeneous distributions common in real hyperspectral images.

[0006] In traditional representation learning-based methods, Low-Rank Sparse Matrix Factorization (LSMAD) relies solely on the low-rank and sparsity properties of the spatial domain for modeling, failing to leverage frequency domain features to enhance the spectral discriminative power of weak targets. In contrast, Kernel Isolated Forest Detection (KIFD) isolates anomalous targets but lacks adaptive modeling of the frequency domain structure. Furthermore, the iterative optimization of spatial information through local windows fails to establish a multi-scale spatial-spectral joint enhancement mechanism, resulting in insufficient capture of the spatial correlation of weak targets. These methods do not fully utilize the frequency domain structural information embedded with subtle target-background differences, leading to reduced sensitivity to high-frequency perturbations and limited representation capabilities for weak targets.

[0007] In recent years, deep learning has become a research hotspot due to its powerful feature extraction and nonlinear modeling capabilities. Data-driven methods based on architectures such as autoencoders (AEs) and generative adversarial networks (GANs) have significantly improved background suppression and anomaly detection in complex scenes by learning deep latent representations or utilizing reconstruction errors. Furthermore, low-rank regularization and total variational anomaly detection (LARTVAD) methods preserve spatial-spectral structure through tensor modeling and introduce total variational regularization to enhance structural consistency. Despite these advances, most deep learning-based hyperspectral anomaly detection methods still operate at the pixel level or rely on relatively shallow spatial-spectral fusion, lacking explicit modeling in the frequency domain and thus failing to leverage the frequency domain structural differences crucial for effective anomaly background detection.

[0008] Meanwhile, some early methods, especially those based on pure spectral autoencoders, did not fully incorporate spatial context information and relied mainly on pixel-by-pixel spectral features. This resulted in the spatial correlation of weak targets being easily overwhelmed by background noise, leading to insufficient detection response strength and poor stability.

[0009] In recent years, methods such as High-Frequency Feature-Guided Diffusion Model (HFGDM), Transformer-based Autoencoder Framework (TAEF), Feature Enhancement Backdistillation Network (FERD), and Multi-Scale Dual-Domain Reconstruction Mask Network (MrsNet) have achieved significant performance improvements. However, they have failed to fundamentally resolve the core contradiction between insufficient adaptiveness of background modeling and weak discriminative power of anomalous features. Background Suppression Diffusion Model (BSDM) introduces diffusion modeling into hyperspectral image processing, improving robustness in cluttered scenes by treating complex backgrounds as random noise. However, for scenes with weak targets, it fails to effectively address issues such as insufficient frequency domain feature mining, inadequate utilization of spatial-spectral context correlation, and imbalanced background distribution. Therefore, achieving robust and reliable detection in complex and realistic hyperspectral remote sensing scenes remains a challenging task. Summary of the Invention

[0010] To address the aforementioned problems, this invention proposes a hyperspectral anomaly detection method based on Fractional Domain Sensing and Contrastive Statistical Regularized Network (FDCRN).

[0011] This method, based on a deep background suppression framework, optimizes anomaly detection performance from three dimensions: frequency domain enhancement, spatial-spectral fusion, and distribution correction. First, a Lightweight Fractional Fourier Transform (LFrFT) module is designed. This module adaptively modulates spectral frequency components using learnable fractional-order parameters and fuses frequency domain features to enhance the capture of subtle differences in small targets. Second, a Multi-scale Spatial-Spectral Enhancement (MSSE) module is proposed. This module utilizes deep separable convolutions to extract multi-scale spatial information and fuses spectral features to enhance the spatial-spectral correlation of small targets. Third, a Statistical Adaptive and Contrastive Learning (SACL) module is introduced to dynamically correct feature distribution shifts and increase the semantic distance between the background and anomalies.

[0012] This invention provides a hyperspectral anomaly detection method based on fractional domain sensing and contrastive statistical regularization networks, the specific steps of which include:

[0013] S1. A deep background suppression framework is constructed as the basic background suppression mechanism. Pseudo-background Gaussian noise is gradually introduced through the forward diffusion process to model the complex background distribution, and a denoising network is trained to predict the noise injected at different diffusion stages.

[0014] S2. Design a lightweight fractional Fourier transform module. For each pixel spectral vector of the input hyperspectral image, a continuous rotation from the spectral domain to the fractional domain is achieved through a learnable fractional parameter θ. Frequency domain feature enhancement is performed using fast Fourier transform, fractional phase modulation, and amplitude modulation, and the enhanced features are fused with the original spectrum through a residual path.

[0015] S3 is designed as a multi-scale spatial-spectral enhancement module. 3×3 and 5×5 depthwise separable convolutions are used to capture local fine-grained structures and mesoscale contextual relationships, respectively. Channel fusion is performed through 1×1 pointwise convolutions, a spectral mixer is introduced to mix cross-dimensional spatial-spectral features, and the enhanced features are embedded into the backbone network through residual connections.

[0016] S4, Design a statistical adaptive and contrastive learning module. Calculate the spatial pixel mean and standard deviation of each spectral band of the input features to construct a statistical descriptor. Generate an adaptive bias embedding through a multilayer perceptron. Dynamically correct the feature distribution through element-wise addition. Increase the semantic distance between background and anomalous features through a contrastive learning projector.

[0017] S5, Training and Inference. During the training phase, network parameters are updated through joint optimization of denoising loss and contrastive learning loss; during the inference phase, clear images are estimated from noisy observations through a backdiffusion process to obtain anomaly detection results.

[0018] The specific methods for each step are as follows:

[0019] S1, Construct a deep background suppression framework.

[0020] For input hyperspectral images H, W, and B represent the height, width, and number of spectral bands of the hyperspectral image, respectively. First, a correlation matrix of all pixels in the hyperspectral image consistent with background statistics is generated. (False background noise) The noise matrix Each pixel in the sample is sampled independently as follows: .

[0021] in The mean is variance is Gaussian distribution, and Let be the mean and variance of all spatial pixels in the k-th band, respectively. (Complete pseudo-background noise tensor) By along the spectral dimension Stacking form.

[0022] Due to the nonlinearity and highly complex distribution of the hyperspectral background, single-step injection of Gaussian noise cannot faithfully approximate its intrinsic statistical properties. This abrupt perturbation severely distorts the background structure and may obscure weak anomalous features. Furthermore, this contradicts the Markov chain representation of the diffusion model, which relies on gradual noise accumulation and a smoothing distribution. Therefore, a forward diffusion process is constructed to progressively introduce pseudo-background noise to model the background distribution. Let the total number of diffusion steps be T, and the linear noise scheduling is defined as: ,in Indicates the first The noise injection rate of the step. and Let represent the initial and final noise injection rates, respectively, and T be the total number of diffusion steps. The cumulative product is defined as: ,in Let represent the cumulative noise product of the first t steps. The noise tensor at step t is: .

[0023] The goal of the diffusion model is to progressively mix the original hyperspectral image with Gaussian noise and train a denoising network. Predict the noise injected at different diffusion stages, among which This represents the learnable parameters of the denoising network. It is a time-step embedding. Since it is equivalent to Gaussian maximum likelihood estimation, The loss function enables the network to accurately model the Gaussian noise distribution by minimizing the squared Euclidean distance between the predicted noise and the actual noise. The training objective is stated as: .

[0024] During the inference process, the trained network predicts noise. A clear picture of estimating time step 0 using backdiffusion: .in This represents the residual image after background suppression. With background suppression applied, this operation subtracts the predicted noise component from the noisy observations and rescales the result back to the original signal amplitude. The structure naturally retains the distribution that deviates from the learning context.

[0025] S2, Design a lightweight fractional Fourier transform module.

[0026] For the spectral vector of any pixel in a hyperspectral image The spectral and frequency characteristics of hyperspectral anomalies differ significantly across different scenarios. To enable the model to adaptively match the optimal analysis dimensions for different scenarios, learnable parameters are introduced. [0, 1], enabling continuous rotation of the spectrum from the spectral domain to the fractional domain. The physical meaning of this operation is to dynamically switch the analytical dimensions of the spectrum. When When = 1, it degenerates into a standard Fourier transform, i.e., complete frequency domain analysis; when When = 0, it degenerates to the original spectrum, i.e., full spectral domain analysis; when Taking the median value allows for the simultaneous capture of both the spectral shape in the spectral domain and the characteristics in the frequency domain. Its mathematical definition is: .in Represents the spectral vector The fractional order of the transformation is The LFrFT operator, The imaginary unit, and These are the trigonometric cotangent and cosecant functions, respectively. Let represent the spectral domain variable corresponding to the original spectral dimension, and u represent the fractional domain variable in the transformed LFrFT domain. The integral realizes the mapping from the spectral domain to the fractional domain.

[0027] In practical implementation, to adapt to the parallel computing of GPUs, the LFrFT numerical approximation algorithm is adopted. First, the spectral vector... Perform a standard Fast Fourier Transform (FFT) to obtain the frequency domain representation. : .in This represents a one-dimensional FFT operator for transforming the spectral vector. Then, spectral-frequency rotation is achieved through learnable fractional-order phase modulation. .in This represents the spectral-frequency mixing characteristics after the fractional Fourier transform. It is by The controlled phase slope acts as a differential phase shift applied to different frequency components, thereby highlighting the frequency components associated with the anomaly.

[0028] To accurately filter and enhance frequency components correlated with anomalies, a lightweight multilayer perceptron (MLP) consisting of two fully connected layers is introduced to generate amplitude modulation weights. This lightweight design adaptively learns the importance of different frequency components while avoiding excessive computational load. Specifically, the frequency domain features after fractional Fourier transform... That is, the spectral-frequency mixing characteristics after fractional Fourier transform. An initial weight vector is obtained through a nonlinear mapping. A tanh activation function is then used to restrict the weight range to [-1, 1], finally generating amplitude modulation weights. : The core function of tanh activation is to prevent excessively large weight values ​​from over-modulating spectral features, while allowing weights to be positive or negative. This is achieved by generating amplitude modulation weights. Frequency domain characteristics Element-wise multiplication enhances anomalous correlation frequencies and suppresses background redundancy frequencies. .in This represents the frequency domain characteristics after amplitude modulation. This represents an element-wise multiplication operation. To preserve the basic form of the original spectrum and avoid losing key spectral contours through frequency domain transformation, the original spectrum is fused with the enhanced features after inverse frequency domain transformation via a residual path: .in This represents the final enhanced spectral vector after residual fusion. This represents the inverse fractional Fourier transform. MLP and The parameters can be adaptively adjusted during model training, automatically focusing on the most discriminative spectral-frequency dimensions of hyperspectral data in different scenarios, ultimately enhancing the spectral quality. It can more accurately depict the subtle spectral differences of small targets.

[0029] S3, design a multi-scale spatial spectrum enhancement module.

[0030] Spatial neighborhood context means that the representation of a given cell is related to the representations of its surrounding cells, and anomalous targets often exhibit spatial inconsistencies with the local background. In hyperspectral images, such anomalies can appear at different spatial scales, ranging from isolated point perturbations to mesoscale regional structures.

[0031] To accommodate these multi-scale spatial features and avoid scale bias introduced by a single convolutional kernel, the MSSE module employs multi-scale depthwise separable convolutions to capture both fine-grained local structure and broader contextual information. Specifically, This represents a cube of a hyperspectral image processed by LFrFT. First, a 3×3 depthwise separable convolution is used to extract local fine structures corresponding to point-like or small-sized anomalies: .in This represents a 3×3 depthwise separable convolution operation, which decomposes standard convolution into depthwise convolution and pointwise convolution. The depthwise convolution step independently convolves each spectral band with a 3×3 kernel, while the pointwise convolution step fuses spectral information through 1×1 convolution.

[0032] Then, a mesoscale context association for small-region anomalies is adapted by accumulating an accumulation set using 5×5 depth separable volumes: .in This represents a 5×5 depthwise separable convolution operation used to capture the spatial context associations at a medium scale.

[0033] Depthwise separable convolution decouples spatial convolution from channel fusion. While maintaining spatial feature extraction capabilities, the number of parameters is only a fraction of that of ordinary convolution. This effectively avoids computational redundancy caused by the high number of channels in hyperspectral data, ensuring the efficiency and effectiveness of multi-scale feature extraction. Deep fusion of spatial-spectral features is achieved through lightweight convolution operations. Firstly, multi-scale spatial features... and Channel concatenation is performed. Then, 1×1 point convolution is used to fuse spatial features through channel fusion, compressing redundant dimensions while fusing spatial information at different scales. in The cube representing the fused spatial spectral features. This represents a 1×1 pointwise convolution operation used for channel fusion and dimensionality compression. Channel dimension splicing operation representing multi-scale spatial features.

[0034] Building upon this, a 1×1 point convolution along the spectral direction is introduced. While maintaining the spatial dimension, channel transformation is applied to the spectral vector at each spatial location, achieving cross-dimensional mixing of spatial structure information and spectral difference information to obtain enhanced spatial-spectral features. : .

[0035] Introducing batch normalization ( ) and Gaussian error linear unit ( Activation can alleviate the training instability caused by fluctuations in the distribution of hyperspectral data, making the features after spatial-spectral fusion more discriminative.

[0036] To avoid the loss of basic spectral information caused by directly replacing the original input with fused spatial spectral features, and to prevent the model from being insensitive to anomalies with significant spectral differences but weak spatial features, residual connections are used to embed enhanced features into the backbone network, enabling parallel flow of original and enhanced features: .

[0037] Final output It possesses both spectral discriminative power and spatial consistency.

[0038] S4, Design a statistical adaptive and contrastive learning module.

[0039] First, calculate the input hyperspectral features. Mean of all spatial pixels in each spectral band and standard deviation A statistical vector is constructed by concatenating the band-level mean and standard deviation. Then, an adaptive bias embedding matching the spectral feature dimension is generated through nonlinear mapping. : The learnable parameters of an MLP can adaptively adjust the mapping relationship during training, allowing the bias embedding to better match the background distribution characteristics of different scenes. Subsequently, through element-wise addition, the dynamically corrected feature distribution is embedded to obtain the distribution-corrected features. : .in Is with Spatial dimension matching A matrix of all ones. This indicates the outer product operation, which will... 3D bias vector broadcasting formation and Dimension matching Empty spectrum feature cube. The data is sent to the depth encoder, which consists of residual layers. By combining the time-step embedding information of the diffusion model, deep encoding of features is completed. Finally, a hyperspectral image with the same dimensions as the original hyperspectral image is generated through the output layer.

[0040] To alleviate feature overlap and insufficient semantic discriminativeness between weak anomaly features and background features in high-dimensional space, the modified representation is projected onto a normalized embedding space. An unsupervised contrastive learning strategy that does not require prior partitioning of background and anomaly samples is employed. Two different viewpoints with different noise perturbations are generated for the same hyperspectral pixel. Let the generated feature pair be represented as... and The contrast loss is defined as: .in This indicates the number of pixels in a single training batch. Represents cosine similarity. This is the temperature parameter (set to 0.1 in the experiment). Indicates corresponding to Positive samples are obtained by applying different noise perturbations to the same pixel. This represents negative samples, i.e., feature embeddings of different pixels in a batch. In actual training, this method employs a dual-loss collaborative optimization strategy. The contrastive learning loss mentioned above... With denoising loss The weighted combination represents the total loss, jointly driving network parameter updates. The formula for the total loss is: .in It is the core of the denoising network loss, The weights for the learning loss are used for comparison. During training, the gradient of the total loss is backpropagated to all learnable parameters of the network.

[0041] This module employs a collaborative design integrating statistical adaptive distribution correction with contrastive learning-based semantic discrimination. By mitigating the heterogeneity of feature distribution in complex backgrounds, it enhances the discriminative power of weak anomalies and maintains stable detection performance under high background noise.

[0042] S5, Training and Reasoning.

[0043] A dual-loss collaborative optimization strategy is employed during the training phase. Hyperspectral image data pairs are sampled in batches, and noise is progressively introduced through a forward diffusion process, while simultaneously generating enhanced views with different noise perturbations for the same pixel. This is achieved through denoising loss. The constrained network accurately predicts the injected noise by comparing the learning loss. Increase the semantic distance between the background and anomalous features. Total loss Drive network parameter updates.

[0044] The inference phase employs a reverse diffusion process. For the input test hyperspectral image, the trained network progressively denoises the image, estimating a clear image at time step 0. . As a residual image after background suppression, it naturally preserves the anomalous structures that deviate from the learned background distribution, thus obtaining the final anomaly detection result.

[0045] The technical effects of this invention are as follows:

[0046] First, an LFrFT module was designed to adaptively modulate spectral frequency components using learnable fractional-order parameters, effectively capturing subtle frequency domain characteristics of weak anomalies. Second, an MSSE module was introduced to extract multi-scale spatial context information and fuse spectral features through depthwise separable convolution, enhancing the spatial-spectral correlation of anomalous targets. Third, a SACL combining a contrastive learning-based projection head was employed to dynamically mitigate feature distribution shifts and improve the semantic separability of background and anomalous features across scenes. Extensive experiments on six real hyperspectral datasets demonstrate the effectiveness and robustness of the proposed FDCRN. This method consistently achieves the highest AUC values ​​(over 0.97) on all evaluation datasets, outperforming seven state-of-the-art contrastive methods in terms of weak target detection accuracy and generalization performance in complex scenes, while maintaining competitive computational efficiency. Future work will explore the integration of FDCRN with hyperspectral super-resolution techniques for sub-pixel anomaly detection and extend the proposed framework to multi-temporal hyperspectral data analysis. Attached Figure Description

[0047] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0048] Figure 1 This is a schematic diagram of the overall framework of the FDCRN of the present invention;

[0049] Figure 2 This is a diagram of the LFrFT module architecture of the present invention;

[0050] Figure 3 This is a diagram of the MSSE module architecture of the present invention;

[0051] Figure 4 This is a diagram of the SACL module architecture of the present invention;

[0052] Figure 5 The pseudo-color images and ground truth images of six hyperspectral datasets used in the example;

[0053] Figure 6 A comparison of the detection results of FDCRN and the comparison method in the example;

[0054] Figure 7 For example, a comparison of ROC curves on Gulfport;

[0055] Figure 8 For example, a comparison of ROC curves on the Luoyang Vehicle;

[0056] Figure 9 For example, a comparison of ROC curves on the Countryside;

[0057] Figure 10 For example, a comparison of ROC curves on a SAU Aircraft;

[0058] Figure 11 For example, comparing ROC curves on BIT Rock;

[0059] Figure 12 For comparison of ROC curves on Wasteland in the example;

[0060] Figure 13 Box plot of anomaly-background separation performance on Gulfport for example;

[0061] Figure 14 Box plot of anomaly-background separation performance on the Luoyang Vehicle in the example;

[0062] Figure 15 Box plot of anomaly-background separation performance on Countryside in the example;

[0063] Figure 16 Box plot of anomaly-background separation performance on an example SAU Aircraft;

[0064] Figure 17 Box plot of anomaly-background separation performance on BIT Rock in the example;

[0065] Figure 18 The box plot shows the anomaly-background separation performance on Wasteland in this example. Detailed Implementation

[0066] This embodiment employs a hyperspectral anomaly detection algorithm framework based on fractional-domain sensing and contrastive statistical regularization networks, as follows: Figure 1 As shown. Given an input hyperspectral image X∈R^(H×W×B), where H, W, and B represent the height, width, and number of spectral bands of the hyperspectral image, respectively. A Deep Background Suppression Framework (DBSF) is used as the basic background suppression mechanism to model complex backgrounds. The overall processing flow consists of three main stages. First, a background suppression mechanism is introduced, such as... Figure 2 The LFrFT module shown enhances the spectral representation in the fractional frequency domain, thereby facilitating the discrimination of weak targets. Secondly, based on the enhanced spectral features, methods such as... Figure 3 The MSSE module shown extracts joint spatial-spectral information and enhances the spatial consistency of potential anomalies. Third, it introduces... Figure 4 The SACL mechanism shown adaptively adjusts the feature distribution and generates the final anomaly map, improving robustness in complex scenarios.

[0067] This embodiment is as follows: Figure 5Experimental validation was performed on six real-world hyperspectral datasets. The first dataset, Gulfport, was acquired using a Reflective Optical Systems Imaging Spectroradiometer (ROSIS) sensor over an airport tarmac area, with a spatial resolution of 100×100 pixels and 205 spectral bands. The second dataset, Luoyang Vehicle, was acquired using a GaiaSky-mini2 dual-channel hyperspectral sensor in Luoyang, Henan Province. The scene is dominated by dense vegetation, with vehicles designated as anomalous targets, and the spatial size is 100×100 pixels with 176 spectral bands. The third dataset, Countryside, was acquired in a rural scene, with anomalous targets corresponding to small-scale military camouflage targets, and the spatial size is 200×200 pixels with 48 spectral bands. The fourth dataset, SAU Aircraft, was acquired using a GaiaSky-mini2 dual-channel hyperspectral sensor at Shenyang Aerospace University, covering an airport tarmac area, with parked aircraft designated as anomalous targets, and the spatial size is 200×200 pixels with 126 spectral bands. The fifth dataset is BIT Rock, collected by the GaiaSky-mini2 dual-channel hyperspectral sensor on the campus of Beijing Institute of Technology. It covers bare land terrain, with scattered small rock patches as anomalies. The spatial size is 256×256 pixels, with 176 spectral bands. The sixth dataset is Wasteland, collected in a wasteland scene. The spatial size is 512×512 pixels, with 198 spectral bands. The anomalous targets are tiny camouflaged military vehicles scattered throughout the scene.

[0068] The experiments were conducted on a workstation equipped with an NVIDIA RTX 5060 GPU, an Intel Core i7-14650HX CPU, 16GB of RAM, and a 1TB PC SN500S WD solid-state drive. The operating system was Windows 11 24H2. The deep learning framework used was PyTorch 2.8.0, combined with CUDA 12.8 for GPU-accelerated computation. Data preprocessing and visualization tools included NumPy 1.26.4, OpenCV 4.12.0, and Matplotlib 3.10.7.

[0069] To verify the effectiveness of our proposed method, seven comparative methods were selected, covering traditional statistical modeling, kernel-based isolation, and advanced deep learning methods. The classic Reed-Xiaoli (RX) algorithm is a statistically representative method that detects anomalies by measuring the spectral deviation of each pixel relative to the global background distribution. KIFD enhances anomaly-background separability through kernel mapping and isolated forests, and optimizes spatial information through iterative local windowing. TAEF captures long-range spatial-spectral dependencies through Transformer attention and detects anomalies through reconstruction errors. HFGDM constructs a frequency-aware diffusion framework to decouple high-frequency anomaly features from low-frequency background components to achieve progressive background suppression. FERD utilizes backdistillation feature enhancement to mine subtle anomaly differences and improve the ability to discriminate weak targets. MrsNet employs a multi-scale dual-domain mask reconstruction paradigm to constrain consistent background representation and suppress false alarms. BSDM utilizes pseudo-background noise to learn the background distribution, improving the background suppression performance of diffusion inference. For evaluation metrics, the Receiver Operating Characteristic (ROC) curve represents the trade-off in detection performance under different decision thresholds, with the false positive rate plotted along the horizontal axis and the true positive rate plotted along the vertical axis. The area under the ROC curve (AUC) is a core quantitative indicator for measuring overall detection accuracy, with a value range of [0.5, 1.0]. A value closer to 1 indicates that the algorithm can achieve a higher Pd at a lower Pf. Box plots assess separability by comparing the detection response distributions of anomaly and background pixels. Runtime measures the total time (in seconds) required for the complete detection process, reflecting the actual efficiency of the method.

[0070] like Figure 6 The quantitative comparison results of AUC shown indicate that FDCRN achieved the highest AUC values ​​on all six datasets, reaching 0.9911, 0.9983, 0.9982, 0.9897, 0.9997 and 0.9909 on Gulfport, Luoyang Vehicle, Countryside, SAU Aircraft, BIT Rock and Wasteland respectively, all exceeding 0.9890, with five datasets exceeding 0.9900, demonstrating stable and reliable detection capabilities across different scenarios.

[0071] like Figures 7-12 As shown, the horizontal axis of the ROC curve represents the False Positive Rate (FPR), and the vertical axis represents the True Positive Rate (TPR). The closer the curve is to the upper left corner, the better the detection performance. The ROC curve analysis results show that FDCRN achieves excellent overall performance on all six datasets, and its ROC curve is always closest to the upper left corner.

[0072] like Figure 7As shown, in the Gulfport Airport apron scenario, FDCRN achieves a rapid increase and reaches a high detection rate in the low false positive rate region.

[0073] like Figure 8 As shown, in the dense vegetation scene of Luoyang Vehicle, FDCRN significantly outperforms all the comparison methods, with the curve rising sharply.

[0074] like Figure 9 As shown, in the Countryside camouflage small target scene, the FDCRN curve is almost close to the top left corner, and the detection rate is close to 0.99.

[0075] like Figure 10 As shown, FDCRN maintains a stable lead in the SAU Aircraft airport scenario.

[0076] like Figure 11 As shown, in the BIT Rock bare soil scene, FDCRN achieved the highest detection rate at all thresholds.

[0077] like Figure 12 As shown, in large-scale Wasteland scenarios containing tiny camouflaged military vehicle targets, FDCRN achieves a steep rise in the low false positive rate region and quickly approaches perfect detection, while most contrasting methods have a slow curve rise and ultimately a limited detection rate.

[0078] like Figures 13-18 As shown, the horizontal axis of the box plot represents different methods and background / anomaly categories, while the vertical axis represents the RX detector response value. The better the separation between the boxes, the better the separability between the anomaly and the background. The box plot results show that FDCRN achieves excellent anomaly-background separation on all six datasets.

[0079] like Figure 13 As shown, in the Gulfport scenario, the background response distribution of FDCRN is closely concentrated at zero and has a narrow range, while the median and upper quartile of the anomalous response are significantly higher than all the comparison methods.

[0080] like Figure 14 As shown, in the Luoyang Vehicle scenario, FDCRN achieves the clearest anomaly-background statistical interval.

[0081] like Figure 15 As shown, in the Countryside camouflaged target scene, FDCRN's abnormal responses are concentrated and far from the background distribution, while the comparison methods show significant overlap.

[0082] like Figure 16 As shown, in the SAU Aircraft scene, FDCRN has the most thorough background suppression.

[0083] like Figure 17 As shown, FDCRN exhibits the highest anomaly-background separation in the BIT Rock scene.

[0084] like Figure 18 As shown, in the Wasteland large-scale wasteland scene, FDCRN has only a very small distribution overlap even under challenging conditions, while other methods all have obvious overlapping areas.

[0085] The runtime comparison demonstrates that the proposed FDCRN achieves a favorable balance between detection performance and computational efficiency. Traditional statistical RX and kernel-based KIFD are the fastest on all datasets, but their detection accuracy is insufficient in complex scenes. Deep learning-based methods such as TAEF, HFGDM, FERD, MrsNet, and BSDM typically introduce higher computational overhead. In contrast, FDCRN maintains a moderate and competitive inference speed across all test scenarios, requiring only 276.93 seconds to handle the largest Wasteland scenario, significantly outperforming TAEF (480.62 seconds), HFGDM (480.62 seconds), FERD (876.01 seconds), MrsNet (701.48 seconds), and BSDM (396.03 seconds), validating its feasibility and deployment potential.

[0086] Ablation experiments validated the necessity and synergistic effect of the core modules. Quantitative results show that the LFrFT module plays a crucial role in capturing subtle spectral differences in weak targets. Removal of the LFrFT module significantly reduced the AUC across all datasets, particularly on the Countryside dataset, where it dropped from 0.9982 to 0.9747, demonstrating the effectiveness of LFrFT in uncovering the frequency-domain spectral differences of camouflaged weak targets. The MSSE module provided the most significant performance gain, but its removal resulted in the most severe performance loss on the Wasteland dataset, decreasing from 0.9909 to 0.9134, indicating that multi-scale spatial-spectral aggregation is crucial for complex backgrounds and large-scale scenes. Removal of the SACL module generally caused a performance degradation, dropping from 0.9982 to 0.9618 on the Countryside dataset. This is attributed to SACL's ability to calibrate imbalanced background distributions through contrastive projection and enhance the discriminative power of anomaly-background features. The complete FDCRN achieves the best AUC on all six datasets, verifying the synergistic complementarity of LFrFT, MSSE, and SACL: LFrFT optimizes spectral frequency characteristics, MSSE enriches multi-scale feature representation, and SACL enhances feature separability, forming a synergistic optimization pipeline with a significant complementary and promoting effect.

Claims

1. A hyperspectral anomaly detection method based on fractional domain sensing and contrastive statistical regularization network, characterized in that, Based on a deep background suppression framework, anomaly detection performance is optimized from three dimensions: frequency domain enhancement, spatial-spectral fusion, and distribution correction. First, a lightweight fractional Fourier transform module is designed to adaptively modulate the spectral frequency components through learnable fractional parameters and fuse frequency domain features to enhance the ability to capture subtle differences in weak targets. Secondly, a multi-scale spatial-spectral enhancement module is proposed, which uses depthwise separable convolution to extract multi-scale spatial information and fuse spectral features to enhance the spatial-spectral correlation of weak targets. Third, a statistical adaptive and contrastive learning module is introduced to dynamically correct feature distribution shifts and increase the semantic distance between the background and the anomaly.

2. The hyperspectral anomaly detection method based on fractional domain sensing and contrastive statistical regularization network according to claim 1, characterized in that, The specific steps include: S1. Construct a deep background suppression framework as the basic background suppression mechanism; gradually introduce pseudo-background Gaussian noise through the forward diffusion process to model the complex background distribution, and train a denoising network to predict the noise injected at different diffusion stages. S2, Design a lightweight fractional Fourier transform module; For each pixel spectral vector of the input hyperspectral image, a continuous rotation from the spectral domain to the fractional domain is achieved through a learnable fractional parameter θ. Frequency domain feature enhancement is performed using fast Fourier transform, fractional phase modulation and amplitude modulation, and the enhanced features are fused with the original spectrum through the residual path. S3, a multi-scale spatial spectrum enhancement module is designed; 3×3 and 5×5 depthwise separable convolutions are used to capture local fine-grained structures and mesoscale contextual relationships respectively, channel fusion is performed by 1×1 pointwise convolution, a spectral mixer is introduced to mix cross-dimensional spatial spectrum features, and the enhanced features are embedded into the backbone network through residual connections. S4, Design a statistical adaptive and contrastive learning module; Calculate the spatial pixel mean and standard deviation of each spectral band of the input feature to construct a statistical descriptor, generate an adaptive bias embedding through a multilayer perceptron, dynamically correct the feature distribution through element-wise addition, and increase the semantic distance between the background and abnormal features through a contrastive learning projection head; S5, Training and Inference: During the training phase, network parameters are updated through joint optimization of denoising loss and contrastive learning loss; during the inference phase, clear images are estimated from noisy observations through a backdiffusion process to obtain anomaly detection results.

3. The hyperspectral anomaly detection method based on fractional domain sensing and contrastive statistical regularization network according to claim 2, characterized in that, The specific method for S1 is as follows: For input hyperspectral images H, W, and B represent the height, width, and number of spectral bands of the hyperspectral image, respectively; first, a correlation matrix of all pixels in the hyperspectral image consistent with the background statistics is generated; pseudo-background noise. The noise matrix Each pixel in the sample is sampled independently as follows: , in The mean is variance is Gaussian distribution, and Let be the mean and variance of all spatial pixels in the k-th band, respectively; and let be the complete pseudo-background noise tensor. By along the spectral dimension Stacking form; The forward diffusion process is constructed by gradually introducing pseudo-background noise to model the background distribution; let the total number of diffusion steps be T, and the linear noise scheduling is defined as follows: ; in Indicates the first The noise injection rate of the step. and These represent the initial and final noise injection rates, respectively, with T being the total number of diffusion steps; the cumulative product is defined as: ; in Let represent the cumulative noise product of the first t steps; the noise tensor of the t-th step is: ; The goal of the diffusion model is to progressively mix the original hyperspectral image with Gaussian noise and train a denoising network. Predict the noise injected at different diffusion stages, among which This represents the learnable parameters of the denoising network. It is a time step embedding; The training objective is stated as follows: ; During the inference process, the trained network predicts noise. A clear picture of estimating time step 0 using backdiffusion: ; in This represents the residual image after background suppression.

4. The hyperspectral anomaly detection method based on fractional domain sensing and contrastive statistical regularization network according to claim 3, characterized in that, The specific method for S2 is as follows: For the spectral vector of any pixel in a hyperspectral image Introducing learnable parameters [0, 1], realizing continuous rotation of the spectrum from the spectral domain to the fractional domain; when When = 1, it degenerates into a standard Fourier transform, i.e., complete frequency domain analysis; when When = 0, it degenerates to the original spectrum, i.e., full spectral domain analysis; when When taking the median value, both the spectral shape in the spectral domain and the characteristics in the frequency domain are captured simultaneously; its mathematical definition is: ; in Represents the spectral vector The fractional order of the transformation is The LFrFT operator, The imaginary unit, and These are the trigonometric cotangent and cosecant functions, respectively. Let represent the spectral domain variable corresponding to the original spectral dimension, and u represent the fractional domain variable in the transformed LFrFT domain. The integral realizes the mapping from the spectral domain to the fractional domain. The numerical approximation algorithm of LFrFT is adopted; firstly, the spectral vector is... Perform a standard Fast Fourier Transform (FFT) to obtain the frequency domain representation. : ; in This represents a one-dimensional FFT operator for transforming spectral vectors; then, spectral-frequency rotation is achieved through learnable fractional-order phase modulation. ; in This represents the spectral-frequency mixing characteristics after the fractional Fourier transform. It is by Controlled phase slope; To accurately filter and enhance frequency components related to anomalies, a lightweight multilayer perceptron (MLP) consisting of two fully connected layers is introduced to generate amplitude modulation weights; specifically, the frequency domain features after fractional Fourier transform... That is, the spectral-frequency mixing characteristics after fractional Fourier transform. An initial weight vector is obtained through nonlinear mapping; a tanh activation function is then used to restrict the weight range to [-1, 1], and finally, amplitude modulation weights are generated. : ; The core function of tanh activation is to prevent excessively large weight values ​​from over-modulating spectral features, while allowing weights to be positive or negative; this is achieved by generating amplitude-modulated weights. Frequency domain characteristics Element-wise multiplication enhances anomalous correlation frequencies and suppresses background redundancy frequencies. ; in This represents the frequency domain characteristics after amplitude modulation. This represents an element-wise multiplication operation; the original spectrum is fused with the enhanced features after inverse frequency domain transformation via the residual path. ; in This represents the final enhanced spectral vector after residual fusion. Indicates the inverse fractional Fourier transform; MLP and The parameters are adaptively adjusted during model training, automatically focusing on the most discriminative spectral-frequency dimensions of hyperspectral data in different scenarios, ultimately enhancing the spectral quality. It can depict subtle spectral differences in small targets.

5. The hyperspectral anomaly detection method based on fractional domain sensing and contrastive statistical regularization network according to claim 4, characterized in that, The specific method for S3 is as follows: To address the introduced scale bias, the MSSE module employs multi-scale depthwise separable convolutions to capture both fine-grained local structure and broader contextual information; specifically, This represents a cube of a hyperspectral image processed by LFrFT; firstly, a 3×3 depthwise separable convolution is used to extract local fine structures corresponding to point-like or small-sized anomalies: ; in This represents a 3×3 depthwise separable convolution operation, which decomposes the standard convolution into depthwise convolution and pointwise convolution. The depthwise convolution step uses a 3×3 kernel to independently convolve each spectral band, while the pointwise convolution step fuses spectral information through 1×1 convolution. Then, a mesoscale context association for small-region anomalies is adapted by accumulating an accumulation set using 5×5 depth separable volumes: ; in This represents a 5×5 depthwise separable convolution operation used to capture mid-scale spatial context associations; Depthwise separable convolution decouples spatial convolution from channel fusion; it achieves deep fusion of spatial-spectral features through lightweight convolution operations; firstly, it addresses multi-scale spatial features. and Channel concatenation is performed; then, 1×1 point convolution is used to fuse spatial features, compressing redundant dimensions while fusing spatial information at different scales. ; in The cube representing the fused spatial spectral features. This represents a 1×1 pointwise convolution operation used for channel fusion and dimensionality compression. Channel dimension splicing operation representing multi-scale spatial features; Building upon this, a 1×1 point convolution along the spectral direction is introduced; while maintaining the spatial dimension, the spectral vector at each spatial location is transformed through channel transformation to achieve cross-dimensional mixing of spatial structure information and spectral difference information, resulting in enhanced spatial-spectral features Y: ; Introducing batch normalization Gaussian error linear unit Activation mitigates the training instability caused by fluctuations in the distribution of hyperspectral data, making the features after spatial-spectral fusion more discriminative; Residual connections are used to embed enhanced features into the backbone network, enabling parallel flow of original and enhanced features: ; Final output It possesses both spectral discriminative power and spatial consistency.

6. The hyperspectral anomaly detection method based on fractional domain sensing and contrastive statistical regularization network according to claim 5, characterized in that, The specific method for S4 is as follows: First, calculate the input hyperspectral features. Mean of all spatial pixels in each spectral band and standard deviation A statistical vector is constructed by concatenating the band-level mean and standard deviation. Then, an adaptive bias embedding matching the spectral feature dimension is generated through nonlinear mapping. : ; The learnable parameters of the MLP adaptively adjust the mapping relationship during training, enabling the bias embedding to better match the background distribution characteristics of different scenes; subsequently, through element-wise addition, the dynamically corrected feature distribution is embedded to obtain the distribution-corrected features. : ; in Is with Spatial dimension matching A matrix of all ones. This indicates the outer product operation, which will... 3D bias vector broadcasting formation and Dimension matching Spatial spectral feature cube; The data is sent to the depth encoder, which consists of residual layers. By combining the time-step embedding information of the diffusion model, deep encoding of features is completed; finally, a hyperspectral image with the same dimensions as the original hyperspectral image is generated through the output layer. An unsupervised contrastive learning strategy is employed that does not require prior knowledge to distinguish between background and anomalous samples; two different viewpoints with different noise perturbations are generated for the same hyperspectral pixel; [The remaining text appears to be incomplete and requires further context.] The generated feature pairs are represented as and The contrast loss is defined as: ; in This indicates the number of pixels in a single training batch. Represents cosine similarity. This is the temperature parameter (set to 0.1 in the experiment). Indicates corresponding to Positive samples are obtained by applying different noise perturbations to the same pixel. Negative samples represent feature embeddings of different pixels in a batch; contrastive learning loss. With denoising loss The weighted combination represents the total loss and jointly drives network parameter updates; the total loss formula is: ; in It is the core of the denoising network loss, These are the weights of the contrastive loss; during training, the gradient of the total loss is backpropagated to all learnable parameters of the network.

7. The hyperspectral anomaly detection method based on fractional domain sensing and contrastive statistical regularization network according to claim 6, characterized in that, The specific method for S5 is as follows: During the training phase, a dual-loss collaborative optimization strategy is employed. Hyperspectral image data pairs are sampled in batches, and noise is gradually introduced through a forward diffusion process. Simultaneously, enhanced views with different noise perturbations are generated for the same pixel. This is achieved through denoising loss. The constrained network accurately predicts the injected noise by comparing the learning loss. Increase the semantic distance between the background and anomalous features; total loss Drive network parameter updates; The inference phase employs a reverse diffusion process; for the input test hyperspectral image, noise is progressively denoised using a trained network, and a clear image is estimated at time step 0. ; As a residual image after background suppression, it naturally preserves the anomalous structures that deviate from the learned background distribution, thus obtaining the final anomaly detection result.