Heterogeneous perception hyperspectral anomaly detection method based on homogeneous characteristic aggregation memory
By constructing a heterogeneous sensing autoencoder and a homogeneous feature aggregation memory module, the problem of insufficient detection accuracy in hyperspectral anomaly detection is solved, achieving efficient identification of anomaly targets in complex backgrounds and improving the accuracy and stability of detection.
Patent Information
- Application Number
- CN202511057130.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
AI Technical Summary
Existing hyperspectral anomaly detection methods suffer from insufficient detection accuracy and poor robustness under complex backgrounds and strong spectral variations, making it difficult to effectively distinguish between anomalous targets and the background. In particular, they are prone to false detections or missed detections in scenarios where the similarity between the target and the background is high.
A heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory is adopted. By constructing a heterogeneous sensing autoencoder and a homogeneous feature aggregation memory module, combined with a deformable convolutional encoder, a surrogate attention decoder and a deep clustering framework, the feature representation learning is optimized, and the model's ability to model spatial structures and anomalous targets is enhanced.
It significantly improves the model's generalization ability and stability in anomaly detection tasks, enabling it to more accurately identify anomalous targets in complex backgrounds, improving detection accuracy and robustness, and making it suitable for various remote sensing scenarios.
Smart Images

Figure CN120953802A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of remote sensing image processing and intelligent sensing technology, and particularly relates to a heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory. Background Technology
[0002] Hyperspectral Anomaly Detection (HAD) provides an effective means of identifying potential anomalies, helping to reveal the underlying causes of anomalous phenomena and providing a basis for further intervention and response. However, due to the high dimensionality, complex spectral characteristics, and strong correlations between bands of hyperspectral images (HSIs), this task faces many technical challenges in modeling and inference. To address these difficulties, researchers have proposed various improvement strategies and methods, continuously driving the improvement of anomaly detection performance. Existing HAD technologies can be broadly classified into two categories: those based on traditional methods and those based on deep learning methods.
[0003] As one of the earliest and most widely used methods in the field of hyperspectral anomaly detection, the RX algorithm, proposed by Reed et al. in 1990, constructs an anomaly discrimination model based on statistical distribution. This method is based on the assumption that background pixels follow a multivariate Gaussian distribution in high-dimensional space. It uses the covariance matrix to model background features and quantifies the spectral difference between the tested pixels and the background model using Mahalanobis distance, achieving the identification of anomalous pixels without prior target information. However, in real-world HSI, the background distribution is often complex and variable, making it difficult to satisfy the above statistical assumptions, resulting in insufficient accuracy in modeling methods based on the overall image mean and covariance matrix. Subsequently, to improve the adaptability of the RX framework, researchers have developed various improvement strategies. Although these improvement strategies have advanced hyperspectral anomaly detection in different dimensions, with the increasing complexity of application scenarios, traditional modeling paradigms still face performance bottlenecks in dealing with complex backgrounds, weak targets, and unstructured anomalies.
[0004] In recent years, the development of deep learning has injected new vitality into hyperspectral anomaly detection. Lin et al. combined low-rank and sparse representation (LRSR) with deep neural networks to propose an interpretable hyperspectral anomaly detection network, LRSR-I2Net. This method expands the ADMM solution process of the LRSR model into an end-to-end trainable structure, achieving adaptive learning of regularization parameters, and introduces total variation regularization to enhance spatial information modeling capabilities. Mu et al. proposed an unsupervised detection method based on a multivariate probability distribution autoencoder. It reconstructs the entire image through a multilayer autoencoder with energy-weighted skip connections and combines a backpropagation mechanism to suppress the reconstruction of anomalous regions. At the same time, it uses a multivariate skewed t-distribution to model the residuals to achieve anomaly localization. Fan et al. designed a robust graph autoencoder detector, introducing superpixel segmentation and graph regularization to improve robustness to noise and anomalies in unsupervised training. The above two types of methods have different focuses: the former focuses on full image reconstruction and residual recognition, but has limited ability to perceive local details; the latter emphasizes local structure modeling, but is insufficient in modeling the global background. To address this, Liu et al. proposed a self-supervised learning-based spatial-spectral dual-domain adaptive network that integrates a spatial state module and a spectral domain self-attention module to construct a multi-scale feature fusion and frequency suppression mechanism, achieving more refined background reconstruction and anomaly detection. It should be noted that while self-supervised learning methods achieve hyperspectral anomaly detection without labels, they still have certain limitations: these methods mainly rely on low-level feature differences for judgment and lack high-level semantic understanding capabilities. When faced with complex background distributions and target feature similarities, the detection accuracy and stability still need improvement.
[0005] In summary, existing HAD methods suffer from insufficient detection accuracy, poor robustness, and limited feature representation capabilities under complex backgrounds and strong spectral variations. Their performance is particularly prone to significant degradation when dealing with real-world scenarios involving weak targets coupled with complex backgrounds or drastic background changes. The specific reasons for this are as follows:
[0006] (1) Some traditional methods are still based on the assumption of Gaussian distribution of the background, which is difficult to adapt to the complex and varied background structure in actual HSI.
[0007] (2) Most deep learning methods only focus on local structure or pixel-level spectral reconstruction, and fail to achieve effective joint modeling of global and local spatial-spectral features;
[0008] (3) Existing methods rely on low-level features for anomaly identification and lack high semantic level guidance. When spectral features are similar or the background is complex, false detection or missed detection is likely to occur.
[0009] This invention aims to address the insufficient accuracy of anomaly detection in hyperspectral images due to high spectral correlation and complex background structures, especially in scenarios with strong spectral variability and high similarity between targets and background. Existing methods generally suffer from reconstruction errors that fail to effectively distinguish between anomalous targets and the background. This invention proposes a heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory. By introducing a heterogeneous sensing autoencoder and a homogeneous feature aggregation memory module, it effectively enhances the model's ability to model spatial structures and suppress anomalous targets, thereby improving the accuracy and stability of anomaly detection. Summary of the Invention
[0010] The purpose of this invention is to provide a heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory, so as to solve the problems of insufficient robustness and limited reconstruction ability of existing hyperspectral anomaly detection methods proposed in the background art in handling practical scenarios with strong spectral variability and complex background.
[0011] To achieve the above objectives, the present invention employs the following technical solution:
[0012] In its first aspect, this invention proposes a heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory, comprising the following steps:
[0013] A heterogeneous perceptual autoencoder is constructed, which includes a deformable convolutional encoder and a surrogate attention decoder. The deformable convolutional encoder enables the convolutional kernel to dynamically adjust the sampling position by introducing an additional offset at each position of the convolutional kernel. The surrogate attention decoder performs global feature aggregation and context broadcasting by introducing a surrogate token.
[0014] A homogeneous feature aggregation memory module is constructed and combined with a heterogeneous sensing autoencoder to obtain a heterogeneous sensing hyperspectral anomaly detection model. The homogeneous feature aggregation memory module guides the heterogeneous sensing autoencoder to learn background feature representation under unsupervised conditions by combining deep clustering strategy with memory mechanism.
[0015] A loss function is designed to optimize the heterogeneous sensing hyperspectral anomaly detection model as a whole during training; the prototype representation of memory update in the iterative homogeneous feature aggregation memory module is optimized by minimizing the cross-entropy loss function, and the reconstruction process is optimized by the reconstruction loss function.
[0016] An optimized heterogeneous sensing hyperspectral anomaly detection model was used for anomaly detection.
[0017] Preferably, the deformable convolutional encoder performs the following specific actions during the encoding phase:
[0018] Three encoder layers with consistent structure are used, each employing a deformable convolutional network architecture to process the input image. hsi m hsi n hsi c These represent width, height, and number of spectral bands, respectively.
[0019] Each layer of the deformable convolutional network architecture first performs two-dimensional convolution operations, followed by the application of the LeakyReLU activation function; then, the BatchNorm technique is introduced to optimize the network's learning process; further, the dropout strategy is used to randomly discard some network connections; and deformable convolutions combined with the SoftMax activation function are integrated.
[0020] Deformable convolution introduces additional learning parameters, allowing the sampling points of the convolution kernel to be dynamically adjusted according to the features of the input image;
[0021]
[0022] Where h is the output feature, x is the original data, p0 is the position of the center pixel of the convolution kernel, and p n It is relative to the position of the center pixel within the convolution kernel, which is nine pixels in total. ω(p n ) is located in p in the grid according to the rules. n The corresponding weights of the pixels; Δp n It is the offset parameter of all sampling positions in the convolution kernel;
[0023] The implicit feature representation of the input image is obtained through the processing of a deformable convolutional encoder.
[0024] h1 = DeformConv(Conv(X))
[0025] h2 = DeformConv(Conv(h1))
[0026] h3 = DeformConv(Conv(h2))
[0027] Here, h1, h2, and h3 are the hidden features of the first, second, and third layers of the encoder, respectively.
[0028] Preferably, the homogeneous feature aggregation memory module is combined with a heterogeneous perceptual autoencoder, as follows:
[0029] The deep clustering framework is combined with a heterogeneous perceptual autoencoder. The deep clustering framework includes a clustering layer that is introduced after the deformable convolutional encoder. The deep clustering framework and the deformable convolutional encoder share the network feature extraction part to output the category to which the hyperspectral image pixels belong.
[0030] The clustering predictions output by the clustering layer are applied to the memory mechanism to learn the homogeneity of features; the core of the memory mechanism is the memory storage unit, which has the functions of memory retrieval and memory update.
[0031] The category prediction information produced by the clustering layer is used in conjunction with the feature representation generated by the heterogeneous perceptual autoencoder to iteratively update the prototype representation in the homogeneous feature aggregation memory module by minimizing the cross-entropy loss function.
[0032] Then through the weighted combination of memory prototypes Perform feature reconstruction.
[0033] Preferably, the proxy attention decoder, during the decoding phase, specifically performs the following:
[0034] Output of the front-end network The process begins by obtaining the query Q, key K, and value V through a linear mapping, and then constructing the proxy token A from Q using pooling or convolution operations.
[0035]
[0036] A = Pooling(Q)
[0037] Among them, W Q W K W V For a learnable mapping matrix, LN denotes layer normalization;
[0038] The proxy attention decoder introduces a proxy attention mechanism, setting proxy token A as the information intermediary. The proxy attention mechanism consists of two Softmax attention processes: proxy aggregation and proxy broadcasting.
[0039] V A =Softmax(AK) T V (Agent Aggregation)
[0040]
[0041] Global context modeling is performed using proxy tokens.
[0042] Preferably, the loss function is as follows:
[0043] The loss function includes the cross-entropy loss function L. c and reconstruction loss function L m Specifically:
[0044] L = L c +L m
[0045] Encoder f and clustering layer h c Collaborative optimization improves performance by minimizing the cross-entropy loss function; let the pixel... Its corresponding category label is y i The clustering prediction is expressed as follows:
[0046]
[0047] Category label y i Defined as the posterior probability q(y|x) i The process is represented as follows:
[0048]
[0049] Where N = hsi m ×hsi n , Indicates a cascade operation, y i ∈{y i} C Represents C clusters, h c This represents the clustering layer.
[0050] Preferably, the memory mechanism is as follows:
[0051] Given a cluster prediction p i At that time, an attention-based strategy is used to perform the memory retrieval operation r(·), specifically:
[0052]
[0053] in, For memory storage content, d is the dimension of the memory prototype;
[0054] M = w(h, P)
[0055] Where w(·) is the memory update operation, used to update the memory prototype; This is the feature encoding set for all hyperspectral image samples.
[0056] Preferably, the weighted combination of the read memory prototypes The feature encoding of the current hyperspectral image is used for reconstruction, and the reconstruction loss is calculated as follows:
[0057]
[0058] Where, x i Represents a pixel; L m This represents the reconstruction loss function.
[0059] Preferably, the anomaly detection is performed as follows:
[0060] Mahalanobis distance is used to evaluate the anomaly score of the tested pixels, thereby detecting anomalous pixels in hyperspectral images; the specific calculation process is as follows:
[0061] S(x i )=(x i -μ) T Γ -1 (x i -μ)
[0062] in, These represent the mean vector and covariance matrix, respectively; x i Represents a pixel.
[0063] In a second aspect, this invention proposes a heterogeneous sensing hyperspectral anomaly detection system based on homogeneous feature aggregation memory, comprising:
[0064] Heterogeneous perceptual autoencoders include deformable convolutional encoders and surrogate attention decoders.
[0065] A deep clustering framework, including clustering layers;
[0066] The homogeneous feature aggregation memory module includes a memory storage unit, which has memory retrieval and memory update functions.
[0067] Compared with the prior art, the beneficial effects of the present invention are:
[0068] (1) This invention proposes a heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory. By introducing a deformable convolutional encoder and a surrogate attention decoder, it enhances the modeling ability of spatial structure and long-range dependency information in hyperspectral images. Furthermore, the designed homogeneous feature aggregation memory module achieves effective aggregation of background features and suppression of anomalous responses through deep clustering and memory mechanisms, thereby enhancing the model's ability to perceive complex anomaly patterns. This invention significantly improves the model's generalization ability and stability in anomaly detection tasks by optimizing feature representation learning through an implicit guidance mechanism without requiring supervised information. It is applicable to various remote sensing scenarios with hyperspectral characteristics and has strong engineering application potential.
[0069] (2) The method in this invention introduces a heterogeneous perceptual autoencoder structure, employing a combination design of a deformable convolutional encoder and a surrogate attention decoder. Compared to the traditional fixed sampling structure, this design can adaptively adjust the sampling position of the convolutional kernel, enhancing the sensitivity to object deformation and spatial features, thereby improving the ability to identify abnormal targets. This structure compensates for the shortcomings of existing methods in extracting complex spatial structures and global semantic features.
[0070] (3) The method of this invention innovatively designs a homogeneous feature aggregation memory module, which guides the model to learn compact background feature representation under unsupervised conditions by combining deep clustering with a memory mechanism. Unlike traditional methods that rely on residual detection, this module can effectively suppress the reconstruction of abnormal regions, improve the separability of abnormalities and background, and thus enhance the robustness and stability of detection.
[0071] (4) The method in this invention introduces the surrogate attention mechanism into the field of hyperspectral anomaly detection for the first time, and achieves efficient global context modeling through surrogate tokens. Unlike directly using the self-attention mechanism, the surrogate mechanism significantly reduces computational complexity while maintaining the ability to model long-distance dependencies, thus improving the practicality of the model in large-scale image processing. Attached Figure Description
[0072] Figure 1 This is a flowchart of the heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory in this invention;
[0073] Figure 2 This is a schematic diagram of the heterogeneous sensing autoencoder structure in this invention;
[0074] Figure 3 This is a structural block diagram of the homogeneous feature aggregation memory module in this invention;
[0075] Figure 4 This is a structural block diagram of the proxy attention decoder in this invention;
[0076] Figure 5 The images show the detection results of the heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory in this invention on different hyperspectral image datasets. Detailed Implementation
[0077] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0078] Example 1:
[0079] like Figure 1 As shown, the heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory consists of three parts: a heterogeneous sensing autoencoder, a homogeneous feature aggregation memory module, a training loss, and anomaly detection. The main steps include:
[0080] Part 1: Constructing a heterogeneous perceptual autoencoder.
[0081] like Figure 2 As shown, the heterogeneous perceptual autoencoder structure includes a deformable convolutional encoder and a surrogate attention decoder. The deformable convolutional encoder introduces an additional offset at each position of the convolutional kernel, enabling the kernel to dynamically adjust its sampling position and more flexibly adapt to changes in target shape and size, thus extracting key information more effectively. The surrogate attention decoder introduces surrogate tokens to achieve global feature aggregation and context broadcasting. This effectively reduces computational complexity while maintaining high sensitivity to long-range dependencies, allowing the decoder to efficiently integrate encoder output features and achieve full information fusion and processing.
[0082] During the encoding phase, three structurally consistent encoder layers are employed, each using a deformable convolutional network architecture to process the input image. hsi m hsi n hsi c These represent width, height, and the number of spectral bands, respectively. The processing flow of each layer begins with a 2D convolution operation using a 3x3 convolution kernel; then, the LeakyReLU activation function is applied to provide non-saturating activation, ensuring the network can effectively learn and transmit complex features; next, BatchNorm is introduced to optimize the network's learning process; to reduce the risk of overfitting, a dropout strategy is further employed, which enhances the model's generalization ability by randomly discarding some network connections. Furthermore, to improve the model's ability to capture dynamic spatial features, deformable convolutions combined with the SoftMax activation function are integrated, improving the shape of conventional convolution kernels and extracting richer spatial information.
[0083] The core characteristic of deformable convolution lies in the dynamic adaptability of the convolution kernel's sampling points, a characteristic achieved by introducing learnable offset parameters. Traditional 3x3 standard convolution kernels, when performing convolution operations, follow the mathematical form defined in Equation 1.
[0084]
[0085] Where h is the output feature, x is the original data, p0 is the position of the center pixel of the convolution kernel, and p n It is relative to the position of the center pixel within nine pixels of the convolution kernel, following a regular grid. ω(p n ) is located at p n The corresponding weights of the pixels.
[0086] Unlike traditional convolution operations, deformable convolution introduces additional learning parameters, allowing the sampling points of the convolution kernel to be dynamically adjusted based on the features of the input data, thereby more effectively capturing the spatial information of the image. As shown in Equation 2:
[0087]
[0088] Where, Δp n It is the offset parameter of all sampling positions in the convolution kernel. Through processing by a deformable convolutional encoder, the implicit feature representation of the input data is obtained.
[0089]
[0090] Here, h1, h2, and h3 are the hidden features of the first, second, and third layers of the encoder, respectively.
[0091] In the decoding stage, in order to achieve global aggregation and efficient reconstruction of features, this invention introduces a proxy attention mechanism. By setting a proxy token with a number far less than the input token as an information intermediary, the model can effectively model long-distance dependencies while reducing computational complexity, thereby improving the ability to represent anomalous targets in HSI.
[0092] Specifically, such as Figure 2 , Figure 4 As shown, the output of the front-end network The process begins by obtaining the query Q, key K, and value V through a linear mapping, and then constructing the proxy token A from Q using pooling or convolution operations.
[0093]
[0094] Among them, W Q W K W V LN represents the layer normalization, which is a learnable mapping matrix.
[0095] Agent attention consists of two Softmax attention processes: agent aggregation and agent broadcasting.
[0096]
[0097] Part Two: Constructing a Homogeneous Feature Aggregation Memory Module.
[0098] During the training of autoencoder networks, if there is a lack of explicit constraints on anomaly categories, the network may learn feature codes containing anomaly information, leading to over-reconstruction of anomalous situations. To avoid this problem, this invention introduces a homogeneous feature aggregation memory module to control the reconstruction capability of the autoencoder network, ensuring that it can only efficiently reconstruct background data. The core of this module is the memory prototype, which represents the semantic center of each category, thus effectively guaranteeing the network's learning of background features. To accurately obtain the semantic center of each category, it is necessary to identify the category information of each sample in the image. For this purpose, a deep clustering framework is organically combined with the autoencoder network to achieve the learning of discriminative features of samples and accurate classification under unsupervised conditions.
[0099] Specifically, such as Figure 2 , Figure 3 As shown, a clustering layer is introduced immediately after the encoder to output the category to which a sample belongs. Within this deep clustering framework, the K-means clustering algorithm is used to generate pseudo-labels for the samples, which guide the training process of the clustering layer. Simultaneously, the category prediction information produced by the clustering layer is jointly used with the feature representation generated by the autoencoder to iteratively update the prototype representation in the memory module. This deep clustering strategy not only extracts highly discriminative features but also forms corresponding cluster predictions, laying the foundation for homogeneous learning of feature representations in the subsequent memory model.
[0100] Let pixels Its corresponding category label is y i This label reflects the cluster affiliation of a given sample. In the model proposed in this invention, the clustering model and the encoder share the network feature extraction part. The clustering prediction is represented as shown in Equation 6:
[0101]
[0102] Where N = hsi m ×hsi n , Indicates a cascade operation, y i ∈{y i} C Represents C clusters, h c This represents the clustering layer.
[0103] Encoder f and clustering layer h c Collaborative optimization improves performance by minimizing the cross-entropy loss function. In anomaly detection tasks, due to the class label y... i It is usually unknown, therefore y i Defined as the posterior probability q(y|x) i The process is represented as shown in Formula 7:
[0104]
[0105] Clustering predictions obtained through deep clustering methods are applied to the memory mechanism for learning feature homogeneity. The core of the memory mechanism is a memory storage unit with two main functions: memory retrieval and memory update. Using this mechanism, not only can memory prototypes be retrieved from memory storage, but existing memory prototypes can also be updated based on new input data.
[0106] Given a cluster prediction p i At this time, an attention mechanism strategy is used to perform memory retrieval operations r(·), thereby effectively retrieving relevant information from memory.
[0107]
[0108] in, d represents the memory storage content, and d represents the dimension of the memory prototype.
[0109] Unlike conventional methods that directly utilize the encoded features of the samples themselves for reconstruction, this algorithm uses a weighted combination of memorized prototypes. The feature encoding representing the current sample is used for reconstruction. In this process, the reconstruction loss is calculated as shown in Equation 9:
[0110]
[0111] The memory update operation w(·) is responsible for updating the memory prototype, and its process can be represented by Equation 10.
[0112] M = w(h,P) (10)
[0113] in, This is the feature encoding set for all samples. Specifically, for each category, its memory prototype is constructed using a weighted combination method. This method comprehensively considers the feature encodings of all samples belonging to that category and their respective predicted probabilities, thereby ensuring that the memory prototype can more accurately reflect the characteristics of that category.
[0114] Part 3: Training Loss and Anomaly Detection.
[0115] The overall optimization process of the model during training can be divided into two parts. First, the category prediction information generated by the clustering layer in the homogeneous feature aggregation memory module is combined with the feature representation produced by the autoencoder. The prototype representation within the memory module is iteratively updated by minimizing the cross-entropy loss function to improve the accuracy of category prediction. Second, for each training sample, the model reconstructs the model using the homogeneous encoding output from the homogeneous feature aggregation memory module. This process is optimized by the reconstruction loss function to ensure that the reconstruction result is as close as possible to the original sample. The overall loss function comprehensively considers the losses from both parts, thereby comprehensively optimizing the model's performance.
[0116] L = L c +L m (11)
[0117] When performing model inference and prediction, this invention uses Mahalanobis distance to evaluate the anomaly score of the tested pixel, thereby detecting anomalous pixels in the difference image. The specific calculation process is shown in Formula 12:
[0118] S(x i )=(x i -μ) T Γ -1 (x i -μ) (12)
[0119] in, These represent the mean vector and the covariance matrix, respectively.
[0120] This invention employs a heterogeneous perceptual autoencoder for validation on various hyperspectral image datasets, including the Gulfport dataset, Pavia dataset, Texas Coast dataset, and San Diego-2 dataset. The detection results are as follows: Figure 5 As shown. According to Figure 5 The detection images obtained from different datasets show that the heterogeneous perceptual autoencoder in this invention accurately distinguishes abnormal targets from the background during the hyperspectral image reconstruction process.
[0121] The key features of this invention include two core innovative modules: a heterogeneous perceptual autoencoder structure and a homogeneous feature aggregation memory module. Together, they construct a novel unsupervised framework suitable for HAD detection tasks, which differs significantly from existing technologies in both structural design and feature modeling mechanism.
[0122] In existing technologies, most methods focus on anomaly detection based on local or global reconstruction errors, but often struggle to simultaneously model spatial structure and spectral features. In contrast, the heterogeneous perceptual autoencoder proposed in this invention employs a deformable convolutional encoder combined with a surrogate attention decoder, which can flexibly adapt to changes in object shape and scale, improve key feature extraction capabilities, effectively aggregate long-distance dependent features, and reduce computational complexity.
[0123] Furthermore, unlike traditional autoencoders that directly utilize input sample features for reconstruction, this invention introduces a homogeneous feature aggregation memory module that combines deep clustering to dynamically update homogeneous background prototypes in the memory space. It then uses an attention mechanism to retrieve background semantic representations from memory for reconstruction. This mechanism significantly enhances the consistency of background reconstruction while suppressing the reconstruction of anomalous regions, thereby improving anomaly detection accuracy.
[0124] In summary, this invention significantly improves the accuracy, stability, and computational efficiency of anomaly detection by introducing a heterogeneous perceptual coding structure, a memory-enhanced clustering mechanism, and a lightweight proxy attention strategy. It effectively overcomes the shortcomings of existing methods in terms of robustness under complex background conditions and insensitivity of reconstruction errors to anomalies, and has stronger feature discrimination capabilities and adaptability. It is an improved solution with practical value and promising prospects in this technical field.
[0125] The above description is only for the purpose of helping to understand the method and core essence of the present invention, but the scope of protection of the present invention is not limited thereto. For those skilled in the art, any equivalent substitutions or modifications made to the technical solution and inventive concept disclosed in the present invention within the scope of the technology disclosed in the present invention should be covered within the scope of protection of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory, characterized in that, Includes the following steps: A heterogeneous perceptual autoencoder is constructed, which includes a deformable convolutional encoder and a surrogate attention decoder. The deformable convolutional encoder enables the convolutional kernel to dynamically adjust the sampling position by introducing an additional offset at each position of the convolutional kernel. The surrogate attention decoder performs global feature aggregation and context broadcasting by introducing a surrogate token. A homogeneous feature aggregation memory module is constructed and combined with a heterogeneous sensing autoencoder to obtain a heterogeneous sensing hyperspectral anomaly detection model. The homogeneous feature aggregation memory module guides the heterogeneous sensing autoencoder to learn background feature representation under unsupervised conditions by combining deep clustering strategy with memory mechanism. A loss function is designed to optimize the heterogeneous sensing hyperspectral anomaly detection model as a whole during training; the prototype representation of memory update in the iterative homogeneous feature aggregation memory module is optimized by minimizing the cross-entropy loss function, and the reconstruction process is optimized by the reconstruction loss function. An optimized heterogeneous sensing hyperspectral anomaly detection model was used for anomaly detection.
2. The heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory according to claim 1, characterized in that, The deformable convolutional encoder, in the encoding phase, specifically as follows: Three encoder layers with consistent structure are used, each employing a deformable convolutional network architecture to process the input image. hsi m hsi n hsi c These represent width, height, and number of spectral bands, respectively. Each layer of the deformable convolutional network architecture first performs two-dimensional convolution operations, followed by the application of the LeakyReLU activation function; then, the BatchNorm technique is introduced to optimize the network's learning process; further, the dropout strategy is used to randomly discard some network connections; and deformable convolutions combined with the SoftMax activation function are integrated. Deformable convolution introduces additional learning parameters, enabling the sampling points of the convolution kernel to be dynamically adjusted according to the features of the input image; through the processing of the deformable convolution encoder, the latent feature representation of the input image is obtained.
3. The heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory according to claim 2, characterized in that, The homogeneous feature aggregation memory module is combined with the heterogeneous perceptual autoencoder, as follows: The deep clustering framework is combined with a heterogeneous perceptual autoencoder. The deep clustering framework includes a clustering layer that is introduced after the deformable convolutional encoder. The deep clustering framework and the deformable convolutional encoder share the network feature extraction part to output the category to which the hyperspectral image pixels belong. The clustering predictions output by the clustering layer are applied to the memory mechanism to learn the homogeneity of features; The core of the memory mechanism is the memory storage unit, which has the functions of memory retrieval and memory update; The category prediction information produced by the clustering layer is used in conjunction with the feature representation generated by the heterogeneous perceptual autoencoder to iteratively update the prototype representation in the homogeneous feature aggregation memory module by minimizing the cross-entropy loss function. Then through the weighted combination of memory prototypes Perform feature reconstruction.
4. The heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory according to claim 3, characterized in that, The proxy attention decoder, in the decoding phase, specifically as follows: Output of the front-end network The process begins by obtaining the query Q, key K, and value V through a linear mapping, and then constructing the proxy token A from Q using pooling or convolution operations. The proxy attention decoder introduces a proxy attention mechanism, setting proxy tokenA as an information intermediary. The proxy attention mechanism consists of two Softmax attention processes: proxy aggregation and proxy broadcasting, which are used for global context modeling.
5. The heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory according to claim 4, characterized in that, The loss function includes the cross-entropy loss function L. c and reconstruction loss function L m Specifically: L=L c +L m Encoder f and clustering layer h c Collaborative optimization improves performance by minimizing the cross-entropy loss function; let the pixel... Its corresponding category label is y i The clustering prediction is expressed as follows: Category label y i Defined as the posterior probability q(y|x) i The process is represented as follows: Where N = hsi m ×hsi n , Indicates a cascading operation, y i ∈{y i } C Represents C clusters, h c This represents the clustering layer.
6. The heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory according to claim 5, characterized in that, The memory mechanism is as follows: Given a cluster prediction p i At that time, an attention-based strategy is used to perform the memory retrieval operation r(·), specifically: in, For memory storage content, d is the dimension of the memory prototype; M = w(h, P) Where w(·) is the memory update operation, used to update the memory prototype; This is the feature encoding set for all hyperspectral image samples.
7. The heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory according to claim 6, characterized in that, Weighted combination of the read memory prototypes The feature encoding of the current hyperspectral image is used for reconstruction, and the reconstruction loss is calculated as follows: Where, x i Represents a pixel.
8. The heterogeneous sensing hyperspectral anomaly detection method based on homogeneous feature aggregation memory according to claim 1, characterized in that, The anomaly detection process is as follows: Mahalanobis distance is used to evaluate the anomaly score of the tested pixels, thereby detecting anomalous pixels in hyperspectral images; the specific calculation process is as follows: S(x i )=(x i -m) T C -1 (x i -m) in, These represent the mean vector and covariance matrix, respectively; x i Represents a pixel.
9. A heterogeneous sensing hyperspectral anomaly detection system based on homogeneous feature aggregation memory applied to the method of claim 1, characterized in that, include: Heterogeneous perceptual autoencoders include deformable convolutional encoders and surrogate attention decoders. A deep clustering framework, including clustering layers; The homogeneous feature aggregation memory module includes a memory storage unit, which has memory retrieval and memory update functions.