Hyperspectral image classification method based on spectral feature reconstruction

By constructing a spectral feature reconstruction network, the problems of uneven distribution of categories, mixed cell interference and feature transfer loss in hyperspectral image classification are solved, and high-precision image classification is achieved, especially in the recognition of small sample data and complex land object.

CN120356015AActive Publication Date: 2025-07-22UNIV OF SHANGHAI FOR SCI & TECH

Patent Information

Application Number
CN202510839214.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

The existing hyperspectral image classification technology is difficult to achieve high-precision classification when facing problems such as uneven category distribution, mixed cell interference and feature transfer loss, especially in poor results in small sample data and complex land object recognition.

Method used

The spectral feature reconstruction method is adopted to construct a reversible fusion network for spectral spatial feature reconstruction, including feature information reconstruction module, spectral spatial reversible fusion module, semantic marking and position embedding module and Transformer encoder module. The dynamic interaction between spectral and spatial features is achieved through reversible structure, and the global dependence is captured in combination with the self-attention mechanism to enhance feature expression and transmission.

Benefits of technology

Effectively reduce the loss of mixed cell information, improve the distinction between difficult samples and the classification balance of few samples, and improve classification accuracy and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356015A_ABST
    Figure CN120356015A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral image classification method based on spectral feature reconstruction, and belongs to the technical field of hyperspectral image processing, and the method comprises the steps: S1, collecting an original hyperspectral image block # imgabs0 #; s2, constructing a spectrum space feature reconstruction reversible fusion network, wherein the spectrum space feature reconstruction reversible fusion network comprises FIR, SSIF, SP and TE; s3, inputting the # imgabs1 # into FIR (Finite Impulse Response) to generate reconstruction features; s4, inputting the reconstructed features into the SSIF to generate enhanced features; s5, inputting the enhanced features into the SP to generate a semantic mark sequence; s6, inputting the semantic marking sequence into TE, and generating # imgabs2; according to the hyperspectral image classification method based on spectral feature reconstruction, lossless transmission of difficult sample features is realized, distinguishing feature expression of mixed pixels is enhanced, and classification balance of few sample categories is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hyperspectral image processing, and in particular to a hyperspectral image classification method based on spectral feature reconstruction. Background Art

[0002] Hyperspectral images (HSIs) play a crucial role in fields such as geological exploration, agricultural monitoring, and environmental remote sensing due to their rich spectral information and spatial details. Among them, accurate image classification is the core link to realize their application value. However, current HSI classification technologies still face many core challenges that need to be urgently solved in practical applications. Firstly, the problem of unbalanced class distribution significantly restricts classification performance. In HSI data, the number of samples of some key ground objects (such as rare minerals) is small, forming a huge gap with the number of samples of the majority classes. This unbalanced class distribution will cause the classifier to overfit the features of the majority class samples during training, resulting in a significant decline in the recognition ability of the minority class samples, making it difficult to accurately detect and classify key ground objects, and seriously affecting the application effect of HSI in fields such as resource exploration and ecological protection. Secondly, mixed pixel interference is another major obstacle to HSI classification. In the edge area of HSI, the spectral features of different ground objects are mixed with each other and the boundaries are blurred, forming mixed pixels. Traditional classification methods are difficult to effectively distinguish the features of different ground objects in mixed pixels. Furthermore, the problem of feature transfer loss during the training process of deep networks is prominent. When dealing with complex HSI data, although deep neural networks have powerful feature extraction capabilities, during the training process, for difficult samples with complex and indistinguishable spectral features, their effective features are easily interfered by noise or lost during the transfer in multiple layers of the network, and cannot provide a reliable basis for classification decisions, greatly limiting the performance improvement of the classification model. Existing HSI classification technologies have obvious deficiencies in dealing with the above challenges. Traditional methods such as SVM and PCA can only utilize shallow spectral features and completely ignore the value of image spatial information, making it difficult to meet the high-precision classification requirements of HSI; although 3D-CNN attempts to capture spatial and spectral information, its computational complexity is extremely high, and it is difficult to achieve efficient spectral dependence modeling when dealing with HSIs containing long-sequence spectral data. In recent years, emerging deep learning methods also have limitations. For example, MorphFormer (literature: S.K. Roy et al., IEEE TGRS, 2023) enhances spatial feature extraction through morphological operations, but has insufficient ability to retain key features when dealing with few-sample data; SSFTT (literature: L. Sun et al., IEEE TGRS, 2022) combines the CNN and Transformer architectures for feature fusion, however, its irreversible fusion method results in a large amount of key information loss in mixed pixels, seriously affecting the classification effect. In summary, the existing HSI classification techniques are difficult to effectively solve problems such as unbalanced class distribution, mixed pixel interference, and feature transfer loss. There is an urgent need for a new classification method and system to improve the accuracy and reliability of HSI classification. Summary of the Invention

[0003] The purpose of the present invention is to provide a hyperspectral image classification method based on spectral feature reconstruction to solve the problems existing in the above background technology.

[0004] To achieve the above object, the present invention provides a hyperspectral image classification method based on spectral feature reconstruction, including the following steps: S1. Collect the original hyperspectral image block ; S2. Construct a spectral-spatial feature reconstruction reversible fusion network, including a feature information reconstruction module, a spectral-spatial reversible fusion module, a semantic marking and position embedding module, and a Transformer encoder module; S3. Input the into the feature information reconstruction module to generate a reconstructed feature ; S4. Input the reconstructed feature into the spectral-spatial reversible fusion module, and realize the dynamic interaction of spectral and spatial features through a reversible structure to generate an enhanced feature ; S5. Input the into the semantic marking and position embedding module to map the shallow features to deep semantic markings and introduce position information, thereby generating a semantic marking sequence ; S6. Input the into the Transformer encoder module, capture global dependencies through the self-attention mechanism, and generate .

[0005] Preferably, the original hyperspectral image block in step S1 is: ; wherein, is the window size, and is the number of spectral channels.

[0006] Preferably, step S3 specifically includes: S31. Spectral channel interaction: Use 1×1 point convolution to expand the spectral channels to 256 dimensions for enhancing information interaction between different bands, which is: ; S32. Local spatial feature extraction: Use 3×3 convolution to extract local spatial features: ; S33. Feature splicing: Splice the original input and the convolution result along the channel dimension to generate 768-dimensional reconstructed features : .

[0007] Preferably, the output dimensions in step S31, step S32, and step S33 are respectively: , , .

[0008] Preferably, step S4 specifically includes: S41. Three-branch splitting: Divide into three sub-parts along the channel dimension: ; S42. Set a bottleneck feature extraction unit to perform feature extraction on , , , and the output is , specifically: S421. Perform 3×3 depthwise separable convolution on each sub-part, which is: ; Among them, the number of convolution kernels of DWConv is 64, the stride is 1, and the output dimension is the same as the input; S422. Enhance important features through the attention weight matrix : ; Among them, is the number of channels; Add the enhanced features to the original input: ; S43. Splice the outputs of the three sub-parts along the channel dimension and reduce the dimension to the original spectral channel number B through 1×1 point convolution, and add the dimension-reduced result to the original input to retain the original information: ; .

[0009] Preferably, step S5 specifically includes: S51. Feature flattening: ; S52. Define two weight matrices that follow a Gaussian distribution, where is the number of labels; Generate semantic tags : ; ; ; S53. Introduce learnable classification tags , and splice them with semantic tags: ; Add trainable position embeddings : ; And apply Dropout to prevent overfitting: .

[0010] Preferably, step S6 specifically includes: S61. Multi-head self-attention: Initialize the query , key , value matrices, where the number of heads , and the dimension of each head = 16: ; Single-head attention calculation: ; Multi-head result splicing: ; Among them, is the output projection matrix; S62. Cross-layer adaptive fusion: Fuse the output of the current layer with the output of the th layer : ; Among them, is a learnable parameter, and the initial value is 0.5; S63. The multi-layer perceptron includes two fully connected layers and the GELU activation function: ; Among them, , respectively represent the weight matrices; , represent the biases.

[0011] Therefore, the present invention adopts the above-mentioned hyperspectral image classification method based on spectral feature reconstruction, and has the following beneficial effects: (1) Implement lossless interaction of spectral-spatial features through the SSIF module to reduce the loss of mixed pixel information; (2) Feature weighting mechanism based on the Hadamard product to enhance the discriminability of difficult samples; (3) The CAF strategy strengthens the inter-layer feature transfer to alleviate the problem of gradient disappearance; (4) Computational efficiency: Depthwise separable convolution (DWConv) reduces the number of parameters.

[0012] The technical solution of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings

[0013] Figure 1 It is a schematic structural diagram of the spectral-spatial feature reconstruction reversible fusion network according to the embodiment of the present invention; Figure 2 It is a schematic diagram of the three-branch structure and the bottleneck feature extraction unit details of the spectral-spatial reversible fusion module according to the embodiment of the present invention; Figure 3 It is a schematic diagram of the multi-head self-attention and CAF fusion of the Transformer encoder module according to the embodiment of the present invention. Specific Embodiments

[0014] The following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0015] A hyperspectral image classification method based on spectral feature reconstruction includes: Collecting the original hyperspectral image patches , where is the window size, is the number of spectral channels.

[0016] Constructing a spectral-spatial feature reconstruction reversible fusion network (SSRIFT), as shown in Figure 1 shown, including: I. Feature Information Reconstruction Module (FIR): Enhance the expression ability of spectral-spatial features and solve the problem of sparse features with few samples.

[0017] Input: The original hyperspectral image patches ; Objective: Reconstruct highly discriminative features through spectral channel interaction and spatial feature extraction.

[0018] Specifically including: 1. Spectral Channel Interaction: Use a 1×1 pointwise convolution (PWConv1) to expand the spectral channels to 256 dimensions for enhancing information interaction between different bands: ; Among them, the number of convolutional kernels of PWConv1 is 256, and the stride is 1.

[0019] Output Dimension: .

[0020] 2. Local Spatial Feature Extraction: Use a 3×3 convolution (Conv) to extract local spatial features, with the number of convolutional kernels being 256 and the stride being 1: ; Output Dimension: .

[0021] Activation Function: ReLU (to suppress gradient explosion), followed by batch normalization (BatchNorm).

[0022] 3. Feature Concatenation: Concatenate the original input and the convolution result along the channel dimension to generate a 768-dimensional reconstructed feature : .

[0023] Output Dimension: . Through multi-scale feature fusion, enhance the feature density of few-shot classes.

[0024] II. Spectral-Spatial Invertible Fusion Module (SSIF): Achieve lossless feature transfer through an invertible network structure, retaining mixed pixel information, as Figure 2 shown.

[0025] Input: ; Objective: Achieve dynamic interaction between spectral and spatial features through an invertible structure to avoid feature loss.

[0026] Specifically include: 1. Three-Branch Segmentation: Divide along the channel dimension into three sub-parts

[0027] ; 2. Bottleneck Feature Extraction Unit (BFE): (1) Depthwise Separable Convolution (DWConv): Perform 3×3 depthwise separable convolution on each sub-part to extract local features while reducing the number of parameters: ; Among them, the number of convolutional kernels of DWConv is 64, the stride is 1, and the output dimension is the same as the input; (2)Feature enhancement and suppression: Hadamard product ( ): Enhance important features through the attention weight matrix : ; Among them, is the number of channels; Element-wise addition (+): Add the enhanced features to the original input to suppress noise: ; (3)Feature fusion and dimensionality reduction: Concatenate the outputs of the three sub-parts along the channel dimension and reduce the dimension to the original spectral channel number B through 1×1 pointwise convolution (PWConv2): ; ; Residual connection: Add the dimensionality reduction result to the original input to retain the original information: .

[0028] The lossless transfer of features is achieved through a reversible structure, and the classification accuracy of mixed pixel samples is improved by 12.6%.

[0029] III. Semantic tagging and position embedding module (SP): Map shallow features to deep semantic tags and fuse them with position information.

[0030] Input: output by the SSIF module; Objective: Map shallow features to deep semantic tags and introduce position information.

[0031] Specifically include: 1. Feature flattening: ; 2. Gaussian weight mapping: Define two weight matrices that follow a Gaussian distribution , where is the number of tags; Generate semantic tags : ; ; ; 3. CLS Token and position embedding: Introduce a learnable classification token and concatenate it with the semantic tags: ; Add trainable position embeddings : ; And apply Dropout (rate 0.1) to prevent overfitting: .

[0032] IV. Transformer Encoder Module (TE): Captures global sequence dependencies and combines Cross-Layer Adaptive Fusion (CAF) to optimize feature flow, as Figure 3 shown.

[0033] Input: ; Objective: Capture global dependencies through self-attention mechanism and optimize classification features.

[0034] Specifically includes: 1. Multi-Head Self-Attention (MSA): Initialize query , key , value matrices, where the number of heads , and the dimension of each head = 16: ; Single-head attention calculation: ; Multi-head result concatenation: ; Among them, is the output projection matrix; 2. Cross-Layer Adaptive Fusion (CAF): Fuse the output of the current layer with the output of the -th layer : ; Among them, is a learnable parameter with an initial value of 0.5; S63. Multi-Layer Perceptron (MLP) contains two fully connected layers (dimension 64 → 256 → 64) and GELU activation function: ; Among them, , respectively represent weight matrices; , represent biases.

[0035] Example 1 The above method is used to classify the Indian Pines dataset, including: Data preprocessing: intercept 7×7 pixel blocks.

[0036] Network construction: FIR module: 1×1 convolution (256 cores) → 3×3 convolution (256 cores) → splicing.

[0037] SSIF module: Split ratio =256, =512, number of DWConv cores 64.

[0038] TE module: 4-head self-attention, hidden layer dimension 64.

[0039] Training parameters: Adam optimizer (lr=5e-4), Batch Size=64, Epoch=400.

[0040] For example, there are only two training samples for category 7 and category 9, but the classification accuracy is 100% on the test samples. Good classification accuracy can also be achieved on categories with uneven sample distribution.

[0041] Embodiment 2 Applying the Botswana small sample scenario Improvements: The number of tokens in the SP module is n=64n=64, and the CAF fusion weight is adjusted dynamically.

[0042] Results: OA reaches 99.16%, which is 0.26% higher than MorphFormer. In this dataset, the sample distribution is relatively uniform, but the number of training samples is small, with only 4 to 15 samples per class. High classification accuracy can be achieved on such small sample datasets.

[0043] Therefore, the present invention adopts the above-mentioned hyperspectral image classification method based on spectral feature reconstruction to achieve lossless transmission of difficult sample features, enhance the discriminative feature expression of mixed pixels, and improve the classification balance of few sample categories.

[0044] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A hyperspectral image classification method based on spectral feature reconstruction, characterized in that, It includes the following steps: S1. Collect the original hyperspectral image blocks ; S2. Construct a spectral-spatial feature reconstruction reversible fusion network, including a feature information reconstruction module, a spectral-spatial reversible fusion module, a semantic marking and position embedding module, and a Transformer encoder module; S3. Reconstruct the input feature information reconstruction module to generate reconstructed features ; S4. Input the reconstructed features into the spectral-spatial reversible fusion module, and through the reversible structure, realize the dynamic interaction between spectral and spatial features to generate enhanced features ; S5. Feed the input semantic markers into the position embedding module to map the shallow features into deep semantic markers and introduce position information, thereby generating a sequence of semantic markers ; S6. Input into the Transformer encoder module, capture global dependencies through the self-attention mechanism, and generate . through the self-attention mechanism, and generate .

2. The hyperspectral image classification method based on spectral feature reconstruction according to claim 1, wherein The original hyperspectral image block in step S1 is as follows: ; Among them, is the window size, is the number of spectral channels.

3. A hyperspectral image classification method based on spectral feature reconstruction according to claim 1, characterized in that, Step S3 specifically includes: S31. Spectral channel interaction: Use 1×1 point convolution to expand the spectral channels to 256 dimensions for enhancing the information interaction between different bands, as: ; S32. Local spatial feature extraction: Use 3×3 convolution to extract local spatial features: ; S33. Feature splicing: Splice the original input and the convolution result along the channel dimension to generate 768-dimensional reconstructed features : 。 4. A hyperspectral image classification method based on spectral feature reconstruction according to claim 3, characterized in that The output dimensions in step S31, step S32, and step S33 are respectively: , , .

5. A hyperspectral image classification method based on spectral feature reconstruction according to claim 1, characterized in that, Step S4 specifically includes: S41. Three-branch splitting: Divide into three sub-parts along the channel dimension: ; S42. Set up a bottleneck feature extraction unit to perform feature extraction on , , specifically as follows: S421. Perform 3×3 depthwise separable convolution on each sub-part, as: ; Among them, the number of convolution kernels of DWConv is 64, the stride is 1, and the output dimension is the same as the input; S422. Enhance important features through the attention weight matrix Enhance important features: ; Among them, is the number of channels; Add the enhanced feature to the original input: ; S43. Concatenate the outputs of the three sub-parts along the channel dimension and reduce the dimension to the original number of spectral channels B through 1×1 point convolution, as: ; ; Add the dimensionality reduction result to the original input and retain the original information: 。 6. The hyperspectral image classification method based on spectral feature reconstruction according to claim 1, wherein, Step S5 specifically includes: S51. Feature flattening: ; S52. Define two weight matrices that follow Gaussian distribution , where is the number of labels Generate semantic tags : ; ; ; S53. Introduce learnable classification tags , and splice them with semantic tags: ; Add trainable positional embeddings : ; And apply Dropout to prevent overfitting: 。 7. A hyperspectral image classification method based on spectral feature reconstruction according to claim 1, characterized in that, Step S6 specifically includes: S61. Multi-head self-attention: Initialize query , key , value matrices, where the number of heads , and the dimension of each head = 16: ; Single-head attention calculation: ; Multi-head result concatenation: ; Among them, is the output projection matrix; S62. Cross-layer adaptive fusion: Fuse the output of the current layer with the output of the layer for fusion: ; Among them, is a learnable parameter with an initial value of 0.5; S63. The multi-layer perceptron includes two fully connected layers and a GELU activation function: ; Among them, and respectively represent weight matrices; and represent biases.

Citation Information

Patent Citations

  • Hyperspectral image classification method based on multi-scale cavity convolution attention network

    CN113963182A

  • Hyperspectral image classification method based on spectrum enhanced cyclic consistency Transform

    CN118781384A

  • Low-light image enhancement method based on reversible neural network

    CN119540088A

  • Hyperspectral and multispectral image fusion method based on adaptive multi-scale features

    CN120047326A

  • Robust HSI classification composite deep network construction method

    CN120068977A

Cited By

  • Few-sample hyperspectral image classification method based on semantic anchor enhanced prototype learning

    CN121837783A