Hyperspectral image open set spectral spatial feature extraction and classification method

Through the hyperspectral image open set classification method of multi-stage progressive network and multimodal feature fusion, the problem of difficulty in identifying new categories in open set scenarios is solved, and the accurate distinction between known categories and unknown categories is achieved, which significantly improves classification performance.

CN120163995APending Publication Date: 2025-06-17HARBIN UNIV OF SCI & TECH
View PDF 0 Cites 13 Cited by

Patent Information

Application Number
CN202510328817.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing hyperspectral classification methods perform well in closed set scenarios, but in open set scenarios, the misclassification rate has increased sharply due to the emergence of new categories, making it difficult to effectively identify unknown categories.

Method used

A method of extracting and classification of open-set spectral spatial features of hyperspectral images is proposed. Through the fusion of multi-stage progressive network and multimodal feature, the joint characteristics of the empty spectrum are extracted, and the intra-class feature density and inter-class distinction are optimized through attention mechanism and comparison learning, forming a joint optimization framework that takes into account both open-set and closed-set scenarios.

Benefits of technology

It significantly improves the ability to identify known categories, enhances the adaptability and robustness to unknown categories, significantly improves the comprehensive performance of classification, and reduces the misclassification rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure QLYQS_3
    Figure QLYQS_3
Patent Text Reader

Abstract

The invention provides a hyperspectral image open set spectral spatial feature extraction and classification method, and aims to improve the classification precision of a known category and an unknown category in a hyperspectral image and the robustness of a classification model. According to the method, input data are preprocessed through fractional Fourier transform, and rotation analysis of a signal time-frequency plane is achieved. Multi-scale features of multi-branch cavity convolution are fused through an enhanced spectrum space residual module, deep separable convolution enhancement nonlinear representation of a lightweight convolution enhancement block is combined, and self-adaptive separation and weighted fusion of high / low frequency features are realized by innovating a dual-frequency enhancement module. Proposing a class perception comparison loss function, integrating anchoring loss, triple loss and regularization terms, and collaboratively optimizing intra-class compactness, inter-class separability and an open set decision boundary. The model framework provided by the invention is combined with the improved loss function, so that the classification precision of the known class and the unknown class of the hyperspectral image is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The method for extracting spectral spatial features and classifying hyperspectral images of the present invention belongs to the technical field of image processing. Specifically, it belongs to the technical field of hyperspectral open-set image classification. Background Art

[0002] This paper proposes an end-to-end classification model for complex scenes, which improves the classification performance by jointly optimizing time-frequency analysis and multi-modal feature fusion. In the preprocessing stage, an adaptive time-frequency transformation method is used to optimize the signal energy distribution and suppress noise interference. In the feature extraction stage, a multi-stage progressive network is designed, which combines three-dimensional convolution and spectral feature enhancement mechanism to extract spatio-spectral joint features, uses a multi-scale context-aware structure to expand the adaptability of the model to complex ground object distributions, and balances the parameter efficiency and deep spatial representation ability through lightweight design. Aiming at the characteristics of frequency-domain features, a high-low frequency decoupling and enhancement strategy is proposed to dynamically adjust the frequency band weights to highlight effective detail information. The attention mechanism is used to realize the dynamic fusion of multi-modal features, and combined with contrast learning and boundary constraint strategies, the intra-class feature compactness, inter-class distinguishability and decision boundary clarity are synchronously optimized to form a joint optimization framework that takes into account both open-set and closed-set scenarios. Experiments show that this method has significant advantages in feature expression ability and anti-interference performance.

[0003] Hyperspectral imaging technology, with its nanometer-level spectral resolution (usually covering the electromagnetic spectrum range of 400 - 2500 nm), can continuously and densely sample the reflection characteristics of ground objects to form spectral curves with diagnostic features. This technology breaks through the limitation of the band discreteness of traditional multispectral imaging. The continuous spectral information it obtains directly reflects the microscopic physical and chemical processes such as molecular vibration and electronic transition of substances, making different substances show a unique "spectral fingerprint" effect in the range of 400 - 2500 nm. For example, the hydroxyl group in minerals has a characteristic absorption peak near 2200 nm, and the change in chlorophyll content in vegetation will significantly affect the spectral morphology in the red edge region. This data characteristic of "combining image and spectrum" enables hyperspectral technology not only to identify the macroscopic morphology of targets, but also to reveal the material composition of hidden targets through sub-pixel-level spectral mixture decomposition, showing irreplaceable advantages in fields such as camouflage material identification in military reconnaissance, pollutant source tracing in environmental monitoring, and crop stress detection in precision agriculture, promoting a revolutionary transformation of remote sensing analysis from traditional spatial form identification to quantitative inversion of material composition.

[0004] Although deep learning-based hyperspectral classification methods have made significant progress in closed-set scenarios (with classification accuracy exceeding 90% in typical experimental environments), their performance improvement is based on the strong assumption that the class spaces of the training set and the test set completely overlap. In practical applications, due to factors such as the dynamic changes in surface cover and the continuous emergence of new artificial materials, test data often contains new classes that were not labeled during the training phase. Experimental studies have shown that when the proportion of unknown classes in the test set reaches 15%, the misclassification rate of traditional closed-set classification models will rise sharply to over 25%, and the misclassified samples are mostly concentrated in the class boundary regions with similar spectral patterns. Open-set classification technology reconstructs the decision-making mechanism of the classifier to effectively reject unknown classes while maintaining the classification accuracy of known classes. Its core breakthrough points are reflected in three aspects: First, by contrastive learning and feature compactness constraints, a feature distribution of known classes with high cohesion is constructed to compress the intra-class feature variance; Second, a dynamic rejection region is established around the feature space, and the extreme value theory or generative adversarial network is used to simulate the distribution of unknown classes to form an expandable decision boundary; Finally, domain adaptation and feature disentanglement techniques are introduced to enhance the robustness of the model to cross-domain interference factors such as lighting conditions and atmospheric transmission, ensuring the stability of feature expression in open environments.

[0005] This patent adopts the Hyperspectral Open-Set Spectral-Spatial Feature Extraction and Classification (HOS-SSFC) method, which divides hyperspectral data into n - 1 known classes and 1 unknown class. Among them, the known classes are divided into a training set and a test set according to a certain proportion, while the unknown class is only used in the test phase. This method constructs a hyperspectral image open-set classification framework, and its technical process is as follows: First, the data is divided into n - 1 known classes (closed set) and 1 unknown class (open set). The known classes are divided into a training set and a test set according to a proportion, and the unknown class is only used in the test phase to verify the open-domain recognition ability of the model. In the preprocessing stage, the fractional Fourier transform (FrFT) is used for time-frequency joint analysis. By optimizing the fractional order parameter, the energy concentration of the key frequency band is enhanced and the noise interference is suppressed. The feature extraction network adopts a multi-stage progressive optimization architecture, and extracts spectral-spatial joint features through 3D convolution and the Squeeze-and-Excitation Module (SE). Among them, the SE module recalibrates the features of the spectral dimension with a channel compression ratio of 16:1 to strengthen the discriminative band response; on this basis, the Enhanced Spectral-Spatial Residual Module (ESSRM) captures multi-scale context space features through a multi-branch dilated convolution structure, expands the model receptive field to adapt to the complex ground object distribution; subsequently, the Lightweight Convolutional Enhancement Block (LCEB) uses cascaded dilated convolution and gating mechanism to enhance the deep spatial non-linear representation ability while reducing the number of parameters; further, the Dual-Frequency Enhancement Module (DFE) is used to decouple the high / low frequency features in the frequency domain, and an adaptive weighting strategy is adopted to suppress the low-frequency redundant noise and enhance the high-frequency detail components. After the above multi-modal features are dynamically fused through the channel-spatial dual attention mechanism, the classification results are output through the fully connected layer. During model training, a composite objective function that fuses class-aware contrast loss and triplet loss, combined with boundary constraints and regularization terms, collaboratively optimizes the intra-class feature aggregation degree, inter-class separation degree, and the clarity of the open-set decision boundary, forming an end-to-end joint optimization framework. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the present invention proposes a hyperspectral image open set classification method, which realizes the accurate distinction between known and unknown categories through the combination of feature extraction and optimization. A multi-scale feature extraction framework is constructed for the spectral and spatial characteristics of hyperspectral images to jointly model complex spectral information and spatial structures. The extraction of spectral features focuses on the effective fusion of multi-band information, enhancing the discriminative power and expressive ability of the features, while spatial features capture key edge and structural information through specific enhancement strategies to optimize the model's perception ability. Through the fusion of multi-level features and the classification optimization mechanism, the model effectively improves the recognition ability of known categories and enhances the adaptability and robustness to unknown categories, thus significantly improving the comprehensive performance of classification.

[0007] The hyperspectral image open set classification algorithm based on multi-scale convolution and feature fusion described in the present invention includes:

[0008] First, a complex analytic signal is constructed for the input real signal, and an orthogonal imaginary part component is generated through Hilbert transform to construct a complex domain representation. Subsequently, zero-mean preprocessing is performed to eliminate the interference of the signal DC component on the frequency domain analysis. In the frequency domain processing stage, after mapping the signal to the frequency domain through fast Fourier transform, a frequency shift correction operation is used to align the spectral energy distribution, and then frequency domain weighting processing is performed based on the discretized fractional-order rotation kernel function constructed by Hermite eigen-decomposition to achieve controllable angular rotation of the signal in the time-frequency plane (the rotation angle is controlled by the fractional-order parameter α). Finally, the time-domain signal is reconstructed through inverse Fourier transform to obtain an enhanced feature representation with the characteristics of rotational time-frequency energy aggregation. This process dynamically adjusts the time-frequency resolution by adjusting the value of α, effectively improving the ability to jointly capture transient and steady-state features in non-stationary signals.

[0009] This method constructs a dataset using the standard open set classification protocol: for the original dataset containing n categories, it is first re-divided into n - 1 known categories (closed set) and 1 unknown category (open set). The specific implementation process includes: 1) The data of known categories are divided into a training set and a test set according to a preset ratio for model training and closed set performance verification; 2) The data of unknown categories are excluded from the training stage throughout, and only combined with the test set of known categories to form the final test set to construct a complete open set test environment. This division strategy strictly simulates the open set recognition requirements of new categories in the real scenario by controlling the appearance of unknown category data only in the test stage.

[0010] Input the features after preprocessing into the spectral shallow feature extraction module. First, the spectral shallow feature extraction module performs an empty-spectrum joint feature modeling on the input training samples. This module uses a 7×7×7 three-dimensional convolutional kernel for initial feature extraction. The design of its large-scale convolutional kernel can synchronously capture the local empty-spectrum features within the spatial neighborhood and the global correlation across the spectral dimensions. Subsequently, a channel attention mechanism is introduced, and the spectral dimension feature recalibration is realized through the SE module. Specifically, the global average pooling layer is used to compress the spatial information, and a channel dependence relationship is constructed by combining two fully connected layers (the channel compression ratio is set to 16:1), effectively enhancing the response intensity of the discriminative spectral channels while suppressing the interference of irrelevant bands. Through the synergistic effect of multi-scale convolution and dynamic channel weighting, the module can adaptively fuse the empty-spectrum features of different scales and finally output a multi-scale empty-spectrum feature vector with strong discriminative power, laying a robust feature representation foundation for subsequent open-set classification tasks.

[0011] First, perform initial empty-spectrum feature extraction on the training samples through a 7×7×7 three-dimensional convolutional kernel. The design of this large-scale convolutional kernel synchronously captures the local geometric patterns (7×7 spatial window) within the spatial neighborhood and the long-range dependence relationships (7-layer spectral segment associations) across the spectral dimensions. In the feature optimization stage, the SE channel attention mechanism compresses the spatial dimensions of the empty-spectrum features output by the three-dimensional convolution through global average pooling to generate a channel description vector to characterize the importance distribution of each spectral band. Subsequently, this vector undergoes cross-channel non-linear interaction through two fully connected layers: the first layer reduces the computational complexity of redundant channels through a compression ratio of 16:1 (C→C / 16), and the second layer restores the original channel dimension to retain spectral discriminability. Finally, a channel weight vector is generated through the Sigmoid activation function to directionally enhance the response intensity of key spectral segments (such as the channels corresponding to characteristic absorption peaks) and suppress noise interference. This process is combined with the cross-dimensional feature fusion of the three-dimensional convolution, and through the dynamic recalibration of the empty-spectrum joint features, a multi-scale empty-spectrum feature tensor with dimensions of H×W×C×D is output, which deeply integrates the spatial structure of ground objects, spectral fingerprint characteristics, and cross-scale context relevance, constructing a highly discriminative feature expression space for open-set classification tasks.

[0012] Subsequently, input the features into the enhanced spectral-spatial residual module (ESSRM). The module uses a 7×1×1 three-dimensional convolutional kernel for feature encoding in the spectral dimension. After stabilizing the feature distribution through the batch normalization layer, the LCEB module realizes feature enhancement through depthwise separable convolution and the channel attention mechanism, and the MBB module fuses the spectral-spatial joint features through a parallel multi-scale convolution path. Finally, the original input and the enhanced features are fused through a residual connection, and a non-linear mapping is realized using the LeakyReLU activation function.

[0013] The features (input feature tensor) after feature extraction by combining a 3D convolutional kernel with an SE module are input into the ESSRM module. First, they pass through a 3D convolution with a kernel size of (7×1×1), and the weight matrix is W ∈ R C×C×7×1×1 , without a bias term:

[0014]

[0015] The processed features are input into the LCEB module. First, they pass through a depthwise separable convolution:

[0016]

[0017] (1) Pointwise convolution expansion:

[0018]

[0019] (2) GeLU activation:

[0020]

[0021] (3) Pointwise convolution compression:

[0022]

[0023] The output features after LCEB processing are input into the MBB module:

[0024] (1) The input features are processed through Branch 1 (3×3×3 convolution): The weight W b1 ∈R C×C×3×3×3 .

[0025]

[0026] (2) The input features are processed through Branch 2 (1×1×1 convolution): The weight W b2 ∈R C×C×1×1×1 .

[0027] Y b2 =W b2 [c,c′,0,0,0]·Y l [b,c′,d,h,w]

[0028] (3) The features obtained from Branch 1 and Branch 2 are fused:

[0029]

[0030] The tensor Y output by the MBB module MBB is subjected to a residual connection with the original input Y in Y res =Y MBB +Xin , and finally, the output feature Y is obtained through the LeakyReLU activation.

[0031] The processed features are input into the spatial feature extraction module, which adopts a multi-stage optimization architecture to process the input spectral features: First, the initial spatio-spectral joint features are extracted through a three-dimensional convolutional kernel (3×3×3) to synchronously model the local spatial structure and spectral context association; Subsequently, an efficient channel attention module is introduced to establish cross-channel dependencies through lightweight one-dimensional convolutions, dynamically enhancing the responses of key feature channels; At the same time, a high-frequency enhancement unit is embedded to separate high-frequency detail components based on frequency-domain filtering and apply non-linear reinforcement to enhance the saliency of edge and texture features. On this basis, two-dimensional residual convolutional blocks are stacked to achieve deep spatial feature abstraction, and the gradient decay problem is alleviated through skip connections; Then, the SE channel attention module is cascaded to globally recalibrate the feature channels, and the ECA module is applied twice to optimize cross-channel interactions. The cascaded collaborative processing of multiple modules effectively realizes the multi-scale progressive enhancement of spatial features, and finally outputs a hierarchical spatial feature representation with strong discriminative power.

[0032] The processed features are input into the dual-frequency enhancement module. The dual-frequency enhancement module optimizes the spatial feature representation based on the frequency-domain decomposition and feature reconstruction strategy: First, 3×3 average pooling is performed on the input features to obtain the low-frequency components, and bilinear interpolation is used to restore the original resolution to maintain spatial topological consistency; Then, the high-frequency detail components are separated through the residual calculation between the original features and the low-frequency components. The two branches respectively use 3×3 convolutions for feature transformation to strengthen the global structural stability of the low-frequency branch and the local detail saliency of the high-frequency branch. Finally, the dual-frequency features are weighted and fused to significantly improve the edge contour resolution and texture feature discriminative power while maintaining the integrity of the main structure, achieving the multi-scale frequency-domain collaborative enhancement of spatial features.

[0033] Finally, the extracted features are further processed by the spatial residual extraction module after spatial residual processing of the frequency-enhanced features, and then the spatial features and spectral features are concatenated and input into the classification module to realize the classification function of the dataset. Description of the Drawings

[0034] Figure 1 is the flowchart of the HOS-SSFC model in the method of the present invention;

[0035] Figure 2 is the model structure diagram of the HOS-SSFC model in the method of the present invention;

[0036] Figure 3 is the ground pseudo-color map and true value map of the PaviaU dataset used in the method of the present invention;

[0037] Figure 4They are the ground pseudo-color map and the true value map of the Houston2013 dataset used in the method of the present invention;

[0038] Figure 5 They are the pseudo-color maps of the classification results of the PaviaU dataset of the HOS-SSFC model and other models of the present invention.

[0039] Figure 6 They are the pseudo-color maps of the classification results of the Houston2013 dataset of the HOS-SSFC model and other models of the present invention. Detailed implementation manners

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention.

[0041] The hyperspectral image open set spectral space feature extraction and classification method under this detailed implementation manner has a flowchart as Figure 1 shown, and includes the following steps:

[0042] Step 1: Input the original hyperspectral dataset to be classified. Divide the n categories in the data into n-1 known categories and 1 unknown category. Among them, the data of the known categories is divided into a training set and a test set according to a certain proportion, and the data of the unknown category does not participate in the training and is directly merged with the test set of the known categories as the final test set.

[0043] Step 2: Input the training samples into the spectral feature extraction module for spectral feature extraction. First, synchronously extract the spatial-spectral joint features through a three-dimensional convolution kernel to construct an initial feature basis; then introduce the SE channel attention mechanism to dynamically adjust the channel weights based on global average pooling and a fully connected layer to strengthen the discriminative spectral band response. Subsequently, cascade the three-dimensional residual convolution module to ensure the stable propagation of gradients through skip connections, and at the same time use a multi-branch parallel structure to capture the multi-scale spatial-spectral correlation features under different receptive fields. To further improve the non-linear representation ability, embed a convolution enhancement module, and use a gating mechanism and dilated convolution combination to enhance local feature interaction. Finally, achieve feature aggregation across the spatial dimension through a global average pooling layer, and output a strongly discriminative spectral feature vector to provide a robust spatial-spectral joint representation basis for open set classification.

[0044] Step 2-1: Input the training samples into the 3D convolutional cascaded SE module. First, perform cross-dimensional feature fusion through a 7×7×7 3D convolutional kernel (a 7×7 neighborhood in the spatial dimension and a 7-band window in the spectral dimension) to synchronously capture local spatial texture features and cross-spectral long-range correlations. Then, introduce the SE channel attention mechanism, use global average pooling to generate a channel statistical description vector, construct a bottleneck structure with a compression ratio of 16:1 through a fully connected layer, dynamically recalibrate the spectral channel weight coefficients, and enhance the response intensity of discriminative spectral bands such as absorption peaks. The wide-field perception ability of the 3D convolution and the dynamic channel screening of the SE module complement each other, and finally generate a feature tensor that fuses multi-scale spatial context and optimizes spectral discriminability, providing a robust spatio-spectral joint representation basis for subsequent classification tasks.

[0045] Step 2-2: Input the shallow features extracted in Step 2-2 into the residual spectral feature extraction module composed of two 3×3×3 convolutional kernels and residual connections. This module processes the features through two layers of 3×3×3 3D convolution and residual connections to alleviate the problem of gradient disappearance.

[0046] Step 2-3: Input the features extracted in Step 2-2 into the LCEB module composed of a 7×7 depthwise separable convolution and a 1×1 pointwise convolution. The depthwise convolution independently performs convolution operations for each input channel to capture the local high-frequency detail features of the hyperspectral image; the pointwise convolution fuses information in the channel dimension through a 1×1 convolution to enhance the feature expression ability between channels.

[0047] Step 2-4: Input the features obtained in Step 2-3 into the multi-branch block composed of a 1×1 convolution and a 3×3 convolution. Build a dual-branch processing flow by parallelly deploying 1×1 and 3×3 2D convolutional kernels. The 1×1 convolution focuses on modeling the spectral correlations between local bands, and the 3×3 convolution captures the cross-band spatial context dependencies. In view of the high spectral correlation characteristics of the hyperspectral image, this module extracts fine-grained spectral difference features and spatially structured-preserving features through two paths respectively, and dynamically fuses the dual-path features using a learnable channel attention mechanism - automatically assigning spectral and spatial weight coefficients based on feature saliency. This heterogeneous convolution strategy effectively balances spectral discriminability and spatial representation integrity, and finally outputs enhanced features that fuse multi-scale spectral correlations and spatial context information, significantly improving the model's ability to distinguish complex ground object boundaries.

[0048] Step 3: Input the spectral features obtained in Step 2 into the spatial feature extraction module for spatial feature extraction. The spatial feature extraction module adopts a multi-stage collaborative optimization strategy to process the input spectral features: First, it realizes joint spatio-spectral feature modeling through a three-dimensional convolutional kernel, synchronously capturing local spatial neighborhood information and cross-spectral dimension context correlation; then it embeds an efficient channel attention module to establish dynamic weight relationships between channels based on lightweight one-dimensional convolution, enhancing the response intensity of key feature channels. At the same time, a frequency domain enhancement unit is introduced to enhance the significant expression of edge and texture features through high-frequency component separation and non-linear enhancement operations. On this basis, multiple two-dimensional residual convolutional blocks are cascaded to achieve deep spatial feature abstraction, ensuring stable gradient propagation through skip connections, and alternately applying the SE channel recalibration module and the ECA module to achieve cross-scale optimization of feature channels. The synergistic effect of multiple mechanisms effectively realizes the hierarchical progressive enhancement of spatial features, and finally outputs a feature tensor that fuses multi-granularity spatial context and enhanced detail representation.

[0049] Step 3-1: Input the spectral features obtained in Step 2 into the spatial feature extraction module composed of a 49×3×3 three-dimensional convolution and ECA.

[0050] Step 3-2: Input the spatial features extracted in Step 3-1 into the high-frequency enhancement module. This module first performs 3×3 average pooling to extract low-frequency features, and then restores the original size through 3×3 upsampling. The low-frequency and high-frequency features are processed separately through 3×3 convolutions, and finally the two are fused to enhance the detail information and improve the recognition of spatial features.

[0051] Step 3-3: Input the features extracted in Step 3-2 into the two-dimensional residual extraction module. This module first performs spatial feature extraction through 3×3 two-dimensional convolution combined with residual connections to enhance the ability to express spatial information, and then combines the SE module and the ECA module to further optimize the feature expression.

[0052] Step 4: After splicing the spatial features extracted in Step 3 and the spectral features extracted in Step 2, send them into the fully connected layer for final classification. The classification module adopts a joint optimization strategy to achieve open-set classification: splice the spatial features and spectral features in the channel dimension, and then map them to the deep feature embedding space through the fully connected layer. Based on the Euclidean distance metric in the feature embedding space, construct a class-aware contrast loss function to constrain the compactness of the feature distributions of the same-class samples, and at the same time expand the inter-class distance of the different-class samples. For the open-set recognition requirement, set a dynamic distance threshold mechanism - when the minimum distance between the test sample and all known class anchors exceeds the preset threshold, it is determined as an unknown class. The classification layer uses the Softmax activation function to output the probability distribution, synchronously fuse the cross-entropy loss and the contrast loss for end-to-end optimization, prevent overfitting by regularizing the network parameters, and introduce a boundary loss function to strengthen the inter-class interval of the decision boundary. This multi-loss collaborative mechanism improves the classification accuracy of known classes while significantly enhancing the model's rejection ability for unknown samples in the open domain.

[0053] To verify the effectiveness of the model proposed in this patent for hyperspectral data classification, this method selects the internationally recognized hyperspectral remote sensing benchmark datasets Pavia University and Houston2013 for algorithm verification. The Pavia University dataset was collected by the ROSIS (Reflective Optics System Imaging Spectrometer) pushbroom imaging spectrometer from the urban area of Pavia, Italy. After denoising preprocessing, 115 effective spectral bands are retained, with the characteristic of a high spatial resolution of 1.3 meters per pixel; the other benchmark dataset, Houston2013, was obtained by the ITRES CASI-1500 airborne hyperspectral imager, covering the campus of the University of Houston and its surrounding areas in the United States, with a spatial resolution of 2.5 meters per pixel and 144 spectral bands. As the benchmark data for the 2013 IEEE GRSS Data Fusion Contest, its complex urban ground object distribution and multi-class vegetation coverage provide a more challenging test scenario for open-set classification. The two datasets, through complementary spatial resolutions (1.3m vs 2.5m), band coverage ranges (115 vs 144), and scene complexities (urban area vs mixed land use), provide a multi-dimensional standard test environment for algorithm evaluation. Their rich spectral features and fine spatial detail verification capabilities have obtained extensive academic consensus in the field of hyperspectral intelligent interpretation.

[0054] The experiment compared the classification performance of three methods, namely the model proposed in this invention (HOS-SSFC), open-set classification based on supervised contrastive learning (OSC-SCL), and multitask deep learning method that simultaneously conducts classification and reconstruction in the open world (MDL4OW), on the PaviaU dataset and the Houston2013 dataset, as shown in Table 1 and Table 2. MDL4OW is a model for hyperspectral image classification, especially suitable for open-set scenarios with unknown classes. It adopts a multitask learning method, combines the classification task and the reconstruction task, and identifies unknown classes by comparing the reconstructed data with the original data, while improving the classification performance of known classes. OSC-SCL enhances the ability to distinguish between known and unknown classes by combining spectral contrast learning with an anchor optimization strategy. Its feature is to use the contrast constraint of spectral features to enhance the discriminability in the spatial-spectral joint features, which is suitable for the open-set classification problem of hyperspectral images.

[0055] Table 1 Classification Results on PaviaU Dataset

[0056]

[0057] In the open-set classification task of the PaviaU dataset, HOS-SSFC demonstrated an overall leading performance advantage. From the overall indicators, the overall classification accuracy (OA) of HOS-SSFC reached 97.92% ± 0.19%, which was 3.83 percentage points higher than the sub-optimal OSC-SCL method (94.09% ± 0.21%). At the same time, the average accuracy (AA) and F1 score reached 97.03% ± 0.18% and 98.38% ± 0.19% respectively, significantly better than other comparison methods. This performance improvement was mainly due to the deep joint optimization of the spatial-spectral features by the model: through the synergistic effect of the multi-scale convolution structure and the dual-frequency enhancement module, the model effectively captured the subtle spectral differences and spatial texture features in the hyperspectral image, especially showing stronger feature discrimination ability in the complex ground object boundary areas (such as the transition zone between buildings and vegetation).

[0058] In terms of classification stability, the standard deviations of all indicators of HOS-SSFC are lower than those of other methods (for example, the standard deviation of the KAPPA coefficient is only 0.25%, while that of OSC-SCL is 2.02%), indicating that it has more robust generalization performance under different training set division scenarios. Specifically at the class level, HOS-SSFC achieves an accuracy close to or exceeding 98% in most land cover classes (such as the F1 value of class 2 is 98.56% ± 0.61%, and that of class 4 is 98.49% ± 0.56%). Especially in the classification of "bare soil" and "asphalt pavement" (class 3), which are easily confused by traditional methods, the F1 value of 94.1% ± 1.83% far exceeds the 91.26% ± 1.16% of the second-place OSC-SCL. This is attributed to the class-aware contrast loss function adopted by the model, which significantly reduces the misjudgment probability between different-class samples by strengthening the intra-class feature compactness.

[0059] Table 2 Classification Results of Houston2013 Dataset

[0060]

[0061]

[0062] In the open-set classification task of the Houston2013 dataset, HOS-SSFC demonstrates significant comprehensive advantages and scenario adaptability. From the perspective of global indicators, HOS-SSFC surpasses all comparison models with an overall classification accuracy (OA) of 89.67% and an average accuracy (AA) of 91.03%. It improves by 0.97 and 0.63 percentage points respectively compared to the sub-optimal OSC-SCL method. Especially in terms of the KAPPA coefficient (88.69%), it leads the second-place MDL40W (87.45%) by 1.24 percentage points, indicating that the consistency between its classification results and the true labels reaches a higher level. Although the F1 score (91.11%) is slightly lower than that of OSC-SCL (91.56%), this small difference is due to the trade-off of the known-class accuracy caused by the model's enhanced ability to detect unknown classes in the open-set scenario, which is more in line with the core requirements of open-set classification in practical applications.

[0063] In the recognition of key land object categories, HOS-SSFC showed strong feature capture ability and robustness. For example, for the classification tasks of Category 1 (healthy grassland) and Category 3 (artificial grassland), HOS-SSFC achieved F1 values ​​of 99.04±0.59% and 99.90±0.09% respectively, which is not only significantly better than other models (such as Category 3 OSC-SCL is 99.93±0.11%), but its standard deviation is compressed to 0.09%, verifying the model's sensitivity to subtle spectral differences and anti-interference ability. This advantage stems from the synergy of multi-scale convolution and dual-frequency enhancement modules: through the joint feature extraction of spatial spectrum and the adaptive weighting of high and low frequency components in the frequency domain, the model effectively enhances the spectral response characteristics of vegetation-covered areas, while suppressing low-frequency noise interference such as water reflection and building shadows that are common in urban environments. In addition, in complex mixed terrain scenes (such as the fifth type of parking lot), HOS-SSFC maintains its leading position with an F1 value of 97.39%, which is 4.5 percentage points higher than traditional methods (such as CROSR's 92.89%), highlighting the optimization effect of the class-aware contrast loss function on the compactness of intra-class features.

[0064] In order to subjectively evaluate the classification effect, Figure 5 , Figure 6 The true value map of the PaviaU dataset, Houston2013 dataset and the pseudo-color map of the classification results of each method are shown respectively. It can be seen that the method in this paper is closer to the real distribution of objects and the area of ​​misclassification is greatly reduced. The experimental results of the model proposed in this paper on the PaviaU dataset prove its significant advantages in the open set classification task of hyperspectral images. Its multi-scale convolution and feature fusion strategy greatly improves the recognition ability of new categories while extracting the features of known categories, providing a new and effective solution for the field of hyperspectral image classification.

Claims

1. A method for extracting and classifying spectral spatial features of an open set of hyperspectral images, characterized by: The data is divided into n-1 known categories and 1 unknown category. The known categories are divided into training and test sets in proportion, and the unknown categories are only used for testing. In the preprocessing stage, the fractional Fourier transform (FrFT) is applied to enhance the spectral representation to capture key features while minimizing noise. The preprocessed features are first extracted by a 3D convolution kernel combined with a squeeze-and-excitation module (SE). The initially extracted features are then passed through an enhanced spectral-spatial residual module (ESSRM), where multi-branch blocks further enhance the learning of spatial features. In addition, a lightweight convolutional enhancement block (LCEB) uses 3D convolution to extract deep spatial features and improve their nonlinear representation. Subsequently, the spectral features are refined through three-dimensional convolution, efficient channel attention module (ECA) and dual-frequency enhancement module (DFE). DFE is introduced to optimize high-frequency information. DFE separates and enhances high-frequency components while suppressing low-frequency noise, further refining the capture of spectral features. Finally, the spectral features and spatial features are concatenated and then classified through a fully connected layer. The loss function combines class-aware contrast loss and traditional contrast loss to balance intra-class compactness and inter-class separability. Anchor loss, triplet loss, boundary loss and a regularization term are integrated to ensure that the model effectively distinguishes known and unknown categories in the latent space.

2. A method for extracting and classifying spectral spatial features of an open set of hyperspectral images, characterized by: Import the original hyperspectral dataset to be classified. For the n categories in the dataset, re-divide it into n-1 known categories and 1 unknown category. The specific operation is to divide the data of known categories into a training set and a test set according to a specific ratio; the data of unknown categories are not used in the training process, but will be directly integrated with the test set of known categories to finally form a complete test set for classification evaluation.

3. A method for extracting and classifying spectral spatial features of an open set of hyperspectral images, characterized by: The input features are first subjected to discrete fractional Fourier transform, and the angle rotation analysis of the signal time-frequency plane is realized through frequency domain phase rotation. The algorithm first converts the input real signal into complex form and performs centralization, performs fast Fourier transform on the signal and applies frequency shift alignment operation, performs multiplication operation of discretized fractional rotation kernel function in the frequency domain, and finally reconstructs the time domain signal through inverse Fourier transform. The input features are first subjected to discrete fractional Fourier transform (FrFT), and the angle rotation analysis of the signal time-frequency plane is realized by frequency domain phase rotation. Suppose the input real signal is x(n), n = 0, 1, ..., N-1, where N is the length of the signal. First, convert it to complex form, which can be obtained by adding the imaginary part to zero. c (n), namely: x c (n)=x(n)+j·0,n=0,1,...N-1 Then the signal is centralized to move the mean of the signal to the origin, which helps in subsequent processing. z (n) can be expressed as: For the centralized signal x z (n) Perform fast Fourier transform to obtain its frequency domain representation X(k): X(k)=FFT{x z (n)},k=0,1,…,N-1 In order to achieve phase rotation in the frequency domain, a frequency shift alignment operation is required. This step usually involves multiplying the frequency domain signal by a frequency shift factor. The frequency domain signal X after frequency shift alignment s (k) is: Then multiply the frequency domain discrete fractional-order rotation kernel function, the discrete fractional-order rotation kernel function F α (k,l) is defined as: Where α is the order of the fractional Fourier transform, and δ() is the unit impulse function. Multiply the frequency domain signal after frequency shift alignment with the discrete fractional rotation kernel function to obtain the processed frequency domain signal: Y(k)=X s (k)·F α (k,l),k=0,1,…,N-1 Finally, the processed frequency domain signal is inverse fast Fourier transformed to reconstruct the time domain signal y(n): y(n)=IFFT{Y(k)},n=0,1,…,N-1.

4. A method for extracting and classifying spectral spatial features of an open set of hyperspectral images, characterized by: The preprocessed features are input into the spectral shallow feature extraction module. First, the input training samples are modeled with spatial-spectral joint features through the spectral shallow feature extraction module. This module uses a 7×7×7 three-dimensional convolution kernel for initial feature extraction. Its large-scale convolution kernel design can simultaneously capture local spatial-spectral features in the spatial neighborhood and global correlations across spectral dimensions. Subsequently, the channel attention mechanism is introduced, and the spectral dimension feature recalibration is realized through the SE module. Specifically, the global average pooling layer is used to compress spatial information, and the two fully connected layers are combined to construct the channel dependency (the channel compression ratio is set to 16:1), which effectively enhances the response strength of the discriminative spectral channel and suppresses the interference of irrelevant bands. Through the synergy of multi-scale convolution and dynamic channel weighting, the module can adaptively fuse spatial-spectral features of different scales, and finally output a multi-scale spatial-spectral feature vector with strong discriminative power, which builds a robust feature representation basis for subsequent open set classification tasks.

5. A method for extracting and classifying spectral spatial features of an open set of hyperspectral images, characterized by: The features extracted in step 4 are then passed through an enhanced spectral spatial residual module, which realizes hyperspectral feature extraction through a three-dimensional convolutional architecture. Its core consists of a dual enhancement mechanism consisting of a lightweight convolution enhancement block and a multi-branch feature fusion module. The module uses a 7×1×1 three-dimensional convolution kernel to encode features in the spectral dimension. After stabilizing the feature distribution through a batch normalization layer, the lightweight convolution enhancement block module realizes feature enhancement through a deep separable convolution and a channel attention mechanism, while the multi-branch feature fusion module fuses spectral-spatial joint features through a parallel multi-scale convolution path. Finally, the original input is fused with the enhanced features through a residual connection, and the LeakyReLU activation function is used to realize nonlinear mapping. This design effectively solves the problem of gradient attenuation in deep networks, and the axial specificity of the three-dimensional convolution kernel is particularly suitable for spatial-spectral joint feature parsing of hyperspectral data cubes. The features extracted by the three-dimensional convolution kernel combined with the SE module (input feature tensor X in ) is input into the ESSRM module and first undergoes a three-dimensional convolution with a kernel size of (7×1×1) and a weight matrix W∈R C×C×7×1×1 , without bias: The processed features are input into the LCEB module and first undergo a depthwise separable convolution: (1) Point-by-point convolution expansion: (2) GeLU activation: (3) Point-by-point convolution compression: The output features after LCEB processing are input into the MBB module: (1) Input feature is input to branch 1 (3×3×3 convolution) for processing: weight W b1 ∈R C×C×3×3×3 . (2) Input feature is input to branch 2 (1×1×1 convolution) for processing: weight W b2 ∈R C×C×1×1×1 . Y b2 =W b2 [c,c′,0,0,0]·Y l [b,c′,d,h,w] (3) Fusion of features obtained from branch 1 and branch 2: The tensor Y output by the MBB module MBB With the original input Y in Perform residual connection Y res =Y MBB +X in , and finally the output feature Y is obtained through LeakyReLU activation. The purpose of the multi-branch block is to capture the multi-scale spatial information of the input features through convolution operations at different scales. In the multi-branch structure, each branch uses a different convolution kernel, which enables the model to learn features at different scales and integrate these features by summing them. The MBB module enhances the sensitivity to high-frequency information (such as edges and textures) and improves the model's ability to capture detailed features through a combination of deep convolution and point convolution.

6. A method for extracting and classifying spectral spatial features of an open set of hyperspectral images, characterized by: The spectral features extracted in 5 are input into the spatial feature extraction module, which extracts preliminary features through three-dimensional convolution and further enhances the spatial features by combining ECA and dual-frequency enhancement modules. Subsequently, multi-level feature extraction is achieved through the joint operation of the two-dimensional residual extraction module, SE module and ECA module.

7. A method for extracting and classifying spectral spatial features of an open set of hyperspectral images, characterized by: The extracted spatial features are input into the dual-frequency enhancement module. The dual-frequency enhancement module optimizes the spatial feature representation through multi-scale frequency domain decomposition and feature reconstruction mechanism. The specific process is: first, a 3×3 average pooling operation is performed on the input feature map to extract the low-frequency spatial components containing the main structure; then, the original resolution is restored through bilinear interpolation upsampling to retain the spatial consistency of the low-frequency information. In order to enhance the feature representation capability, the residual of the upsampled low-frequency component and the original feature is calculated to separate the high-frequency component containing the detail texture, and a 3×3 convolution kernel is used for nonlinear transformation respectively - the low-frequency branch focuses on enhancing the global structural robustness, and the high-frequency branch focuses on local detail sharpening. Finally, the dual-channel features are integrated through an adaptive weight fusion strategy, which significantly improves the edge clarity and texture resolution while retaining the main contour, and realizes the multi-frequency domain collaborative enhancement of spatial features. First, extract the low-frequency information in the input features: Where i∈[0,H / 2-1],j∈[0,W / 2-1], output low-frequency component X L ∈R B×C×H / 2×W / 2 . Next, the low-frequency component is upsampled to the original size by bilinear interpolation, and the original feature is compared with the obtained low-frequency feature X L The high-frequency features are obtained by subtraction, and the obtained high-frequency features and low-frequency features are enhanced by the MBB module. Finally, the enhanced high-frequency and low-frequency components are added element by element to obtain the output features.

8. A method for extracting and classifying spectral spatial features of an open set of hyperspectral images, characterized by: The features extracted in 7 are input into the spatial residual extraction module to further process the features after high-frequency enhancement, and finally the spatial features and the spectral features extracted in 5 are spliced ​​together.

9. A method for extracting and classifying spectral spatial features of an open set of hyperspectral images, characterized by: The features extracted in 8 are input into the classification module. The classification method adopted by the present invention is based on metric learning and contrastive learning methods, and the classification performance of the model is optimized by calculating the Euclidean distance between the sample and the class anchor point and using contrastive loss. The proposed loss function is specifically designed to enhance feature representation and improve classification performance, especially for hyperspectral open set classification tasks. It combines class-aware contrast loss and contrastive loss to achieve a balance between intra-class compactness and inter-class separability. These components synergistically optimize the latent space, enabling the model to effectively handle both known and unknown classes. The anchor loss L that forms the basis of the loss function proposed in the present invention a Defined as: where f i represents the feature vector of the i-th sample, N represents the batch size, is the anchor point of the real class. To further enhance the separability between classes, an anchor loss is also introduced. This ensures that the features are not only aligned with their true class anchors, but also keep a sufficient distance from other class anchors, thereby reducing misclassification. t Defined as: d ij For sample x i to its true category anchor The Euclidean distance. The metric loss introduces an additional constraint to enforce a minimum threshold between the distance of a sample to its true class anchor and the nearest false class anchor. This constraint ensures a sharp boundary in the latent space, which is especially important for dealing with ambiguous samples near the decision boundary. m Defined as: Regularization term L r It is mainly used to balance the model's sensitivity to category features and global distribution stability. The core idea is: for each sample, first calculate the sum of its exponential similarities with all category anchors, and then subtract its similarity with its own true category anchor. This operation forces the model to keep samples close to the anchor of the category they belong to (enhancing intra-class compactness) while avoiding too high overall similarity with other category anchors (reducing the risk of inter-class confusion) during training. By constraining the distribution of similarity in the feature space, this regularization term effectively alleviates the model's over-reliance on a few easily distinguishable features, thereby improving the model's generalization ability under complex data distributions and preventing overfitting. Regularization term L r Defined as: The total loss function L is: L=αL a +βL t +γL m +λL r The present invention also uses contrastive loss to refine the spectral and spatial feature representations. This ensures that features from the same class are close in the joint spectral-spatial feature space, while features from different classes are pushed further away. By integrating spectral and spatial dimensions, contrastive loss enhances the discriminative power of the network. The proposed loss function enables the model to achieve a refined latent space characterized by intra-class compactness and inter-class separability. This comprehensive loss design plays a vital role in addressing the complexity of the open set hyperspectral classification task.

Citation Information

Cited By

  • Partial discharge identification and positioning method based on multi-modal calibration and blind source separation

    CN121049664A

  • Hyperspectral image classification method based on multi-scale space and spectrum enhancement fusion

    CN121121319A

  • Remote sensing image semantic segmentation method fusing frequency domain modeling and lightweight linear attention

    CN121170289A

  • Mixed noise-oriented sperm head classification method and system

    CN121259822A

  • A method and system for classifying sperm heads in the presence of mixed noise

    CN121259822B