Hyperspectral image classification method and system based on three-dimensional mixed frequency domain attention network

By adopting a three-dimensional mixed frequency domain attention network in hyperspectral image classification, combined with the frequency domain attention module and attention block, the shortcomings of traditional 3D-CNN in feature frequency distribution and information redundant noise processing are solved, and classification accuracy and robustness are improved.

CN120147734APending Publication Date: 2025-06-13YANCHENG INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510232223.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional 3D-CNNs tend to ignore the frequency distribution of features when processing hyperspectral images, resulting in inaccurate information extraction and ineffective in alleviating the interference of information redundancy and noise on the model.

Method used

A hyperspectral image classification method based on three-dimensional mixed frequency domain attention network is adopted, and a dual-branch 3D convolutional neural network structure is constructed, combined with the frequency domain attention module and attention block, spectral and spatial features are extracted, and feature fusion and classification decisions are made.

Benefits of technology

It improves classification accuracy and robustness, enhances the model's perceived ability of multi-scale information, effectively balances the importance of spectral and spatial characteristics, improves gradient flow and avoids gradient vanishing problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147734A_ABST
    Figure CN120147734A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of hyperspectral image classification, in particular to a hyperspectral image classification method and system based on a three-dimensional mixed frequency domain attention network, and the method comprises the following steps: carrying out the preprocessing of input hyperspectral image data; the method comprises the following steps: constructing a double-branch 3D convolutional neural network structure which comprises a spectral branch and a spatial branch, and respectively extracting spectral features and spatial features; wherein the spectral branch extracts spectral features through a plurality of 3D convolution layers, and the spatial branch extracts spatial features through the 3D convolution layers and the frequency domain attention module; fusing the features extracted from the spectral branches and the spatial branches to obtain a fused feature map; and inputting the fused feature map into a classifier, carrying out classification decision, and outputting a classification result. According to the method, by introducing a frequency domain attention mechanism, key features can be focused more accurately, and the perception ability of the model to multi-scale information is enhanced, so that the classification accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hyperspectral image classification, and particularly to a hyperspectral image classification method and system based on a three-dimensional hybrid frequency domain attention network. Background Technique

[0002] Hyperspectral image classification has important applications in fields such as environmental monitoring, geological exploration, military reconnaissance, and deep space exploration. Hyperspectral remote sensing technology can obtain high-resolution image data from visible light to mid-infrared, thermal infrared, etc. These data contain information in multiple bands, providing rich spatial, radiation, and spectral information for ground object classification, target recognition, and environmental analysis. However, hyperspectral data has high dimensions and large computational amounts, and traditional classification methods often have difficulty effectively extracting and utilizing this information, resulting in limited classification accuracy and robustness.

[0003] In recent years, with the development of deep learning technology, especially the application of convolutional neural networks (CNNs) in hyperspectral image classification, the classification performance has been significantly improved. Among them, three-dimensional convolutional neural networks (3D-CNNs) can better extract features from the spectral dimension through cube convolutional kernels, which are superior to two-dimensional convolutional neural networks (2D-CNNs). However, when traditional 3D-CNNs process spatial and spectral information, they are prone to ignoring the frequency distribution of features, resulting in inaccurate information extraction and being unable to effectively alleviate the interference of information redundancy and noise on the model.

[0004] In order to further improve the accuracy and robustness of hyperspectral image classification, the present invention proposes a hyperspectral image classification method and system based on a three-dimensional hybrid frequency domain attention network. Summary of the Invention

[0005] The purpose of the present invention is to solve the shortcomings existing in the prior art, and to propose a hyperspectral image classification method and system based on a three-dimensional hybrid frequency domain attention network.

[0006] To achieve the above purpose, the present invention adopts the following technical solution: A hyperspectral image classification method based on a three-dimensional hybrid frequency domain attention network, comprising the following steps:

[0007] a) Data preprocessing: Perform preprocessing operations on the input hyperspectral image data;

[0008] b) Feature extraction: By constructing a dual-branch 3D convolutional neural network structure, including a spectral branch and a spatial branch, respectively extract spectral features and spatial features; wherein, the spectral branch extracts spectral features through multiple 3D convolutional layers, and the spatial branch extracts spatial features through 3D convolutional layers and a frequency domain attention module;

[0009] c) Feature fusion: Fuse the features extracted by the spectral branch and the spatial branch to obtain a fused feature map;

[0010] d) Classification decision: Input the fused feature map into a classifier for classification decision and output the classification result.

[0011] Preferably, in the above solution, the frequency-domain attention module includes a Fourier transform unit, a frequency-domain feature adjustment unit, and an inverse Fourier transform unit; among them, the Fourier transform unit converts the input features into the frequency domain, the frequency-domain feature adjustment unit analyzes and adjusts different frequency components, and the inverse Fourier transform unit maps the adjusted frequency-domain features back to the spatial domain.

[0012] Preferably, in the above solution, the dual-branch 3D convolutional neural network structure further includes a spectral attention block and a spatial attention block; among them, the spectral attention block is used to dynamically assign weights to different spectral bands, and the spatial attention block is used to dynamically assign weights to different spatial positions.

[0013] Preferably, in the above solution, the classifier is any one of a support vector machine, a random forest, a logistic regression, or a deep neural network.

[0014] The present invention also provides a hyperspectral image classification system based on a three-dimensional hybrid frequency-domain attention network, including:

[0015] Data preprocessing module: Used to preprocess the input hyperspectral image data;

[0016] Feature extraction module: Used to construct a dual-branch 3D convolutional neural network structure and extract spectral features and spatial features;

[0017] Feature fusion module: Used to fuse spectral features and spatial features;

[0018] Classification decision module: Used to make a classification decision on the fused features and output the classification result.

[0019] Preferably, in the above solution, the frequency-domain attention module in the feature extraction module includes a Fourier transform unit, a frequency-domain feature adjustment unit, and an inverse Fourier transform unit.

[0020] Preferably, in the above solution, the feature extraction module further includes a spectral attention block and a spatial attention block.

[0021] Preferably, in the above solution, the classification decision module uses any one of a support vector machine, a random forest, a logistic regression, or a deep neural network as the classifier.

[0022] Preferably, in the above solution, the result output module further includes a visualization module for presenting the detection results to the user in a visual manner.

[0023] The present invention has the following beneficial effects:

[0024] 1. In the present invention, by introducing the frequency-domain attention mechanism, the present invention can more precisely focus on key features, enhance the model's perception ability of multi-scale information, and thus improve the classification accuracy.

[0025] 2. In the present invention, the set frequency-domain attention mechanism helps to balance the importance of spectral and spatial features, and improve the classification performance after feature fusion. At the same time, the use of the Mish activation function also helps to improve the gradient flow and avoid the gradient vanishing problem, enhancing the robustness of the model.

[0026] 3. In the present invention, the method and system of the present invention are applicable to various hyperspectral image classification tasks, including environmental monitoring, geological exploration, military reconnaissance, and deep space exploration. Description of the Drawings

[0027] Figure 1 It is the network structure diagram of the hyperspectral image classification system based on the three-dimensional hybrid frequency-domain attention network in the present invention;

[0028] Figure 2 It is the classification diagram of each method of the IP dataset in the present invention;

[0029] Figure 3 It is the classification diagram of each method of the UP dataset in the present invention;

[0030] Figure 4 It is the classification diagram of each method of the SV dataset in the present invention. Detailed Embodiments

[0031] A hyperspectral image classification method based on a three-dimensional hybrid frequency-domain attention network includes the following steps:

[0032] a) Data preprocessing: Perform preprocessing operations on the input hyperspectral image data;

[0033] b) Feature extraction: By constructing a dual-branch 3D convolutional neural network structure, including a spectral branch and a spatial branch, extract spectral features and spatial features respectively; among them, the spectral branch extracts spectral features through multiple 3D convolutional layers, and the spatial branch extracts spatial features through 3D convolutional layers and a frequency-domain attention module;

[0034] c) Feature fusion: Fuse the features extracted by the spectral branch and the spatial branch to obtain a fused feature map;

[0035] d) Classification decision: Input the fused feature map into a classifier for classification decision and output the classification result.

[0036] As Figure 3 shown, the hyperspectral image classification algorithm with frequency-domain attention combines a 3D convolutional neural network and an attention mechanism, which is divided into a spectral branch and a spatial branch to extract spectral and spatial feature maps respectively. These feature maps are fused and finally used for classification. The model uses the Mish activation function, which has the characteristics of no upper bound and a lower bound, can effectively avoid gradient vanishing and activation function saturation, and at the same time enhance the gradient flow of negative value inputs, improving accuracy and generalization ability.

[0037] In the spatial branch, the input is processed through a 3D convolutional layer and the batch normalization layer of the Mish activation function to generate a feature map of size (9×9×97, 24). Subsequently, the spatial features are extracted through a frequency-domain attention module, and an output of 1×60 is obtained through a global average pooling layer. Similarly, in the spectral branch, spectral features are extracted through multiple convolutional layers (with sizes of 3×3 and channel numbers of 150, 100, and 60 in sequence), and finally a feature map of 9×9×60 is output, which is processed through a frequency-domain attention module and a global average pooling layer.

[0038] The feature maps extracted by the spatial and spectral branches are concatenated, and the final classification result is obtained through a fully connected layer and a Softmax layer, thus improving the model performance.

[0039] The frequency-domain attention module includes a Fourier transform unit, a frequency-domain feature adjustment unit, and an inverse Fourier transform unit. Among them, the Fourier transform unit converts the input features to the frequency domain, the frequency-domain feature adjustment unit analyzes and adjusts different frequency components, and the inverse Fourier transform unit maps the adjusted frequency-domain features back to the spatial domain.

[0040] The dual-branch 3D convolutional neural network structure also includes a spectral attention block and a spatial attention block. Among them, the spectral attention block is used to dynamically assign weights to different spectral bands, and the spatial attention block is used to dynamically assign weights to different spatial positions.

[0041] The classifier can be any one of a support vector machine, a random forest, a logistic regression, or a deep neural network.

[0042] A hyperspectral image classification system based on a three-dimensional hybrid frequency-domain attention network, comprising:

[0043] Data preprocessing module: used to preprocess the input hyperspectral image data; the data preprocessing module is the starting link of the system, and its main task is to perform necessary preprocessing operations on the input hyperspectral image data. These operations include but are not limited to denoising, calibration, normalization, etc., aiming to improve the data quality and provide a solid foundation for subsequent feature extraction and classification decision-making;

[0044] Specifically, the denoising operation can help eliminate random noise in the image and improve the image clarity; the calibration operation can perform geometric and radiometric calibration on the image to ensure the accuracy of the image data; the normalization operation can adjust the pixel values of the image data to an appropriate range for subsequent processing.

[0045] Feature extraction module: used to construct a dual-branch 3D convolutional neural network structure and extract spectral features and spatial features; construct a dual-branch 3D convolutional neural network structure and extract spectral features and spatial features.

[0046] The dual-branch 3D convolutional neural network structure includes a spectral branch and a spatial branch.

[0047] Spectral branch: focuses on extracting the spectral features of the hyperspectral image. This branch processes the input hyperspectral image data through a series of components such as 3D convolutional layers, pooling layers, and activation functions to extract spectral features reflecting the material properties.

[0048] Spatial branch: focuses on extracting the spatial features of the hyperspectral image. This branch also adopts a 3D convolutional neural network structure but pays more attention to capturing information such as the spatial distribution law and adjacent relationships in the image.

[0049] Feature fusion module: used to fuse spectral features and spatial features to form a more comprehensive and rich feature representation, which can be achieved through simple feature concatenation, weighted summation, etc. The fused features will be used as the input of the classification decision module.

[0050] Classification decision module: used to make classification decisions on the fused features and output the classification results.

[0051] The frequency-domain attention module in the feature extraction module includes a Fourier transform unit, a frequency-domain feature adjustment unit, and an inverse Fourier transform unit.

[0052] The frequency domain attention module is a key component in the feature extraction module. It combines Fourier transform and inverse Fourier transform techniques, as well as frequency domain feature adjustment strategies, aiming to enhance the feature extraction ability through frequency domain information. The Fourier transform unit converts the input features from the spatial domain to the frequency domain for frequency domain analysis. The frequency domain feature adjustment unit analyzes and adjusts different frequency components in the frequency domain to highlight key features and suppress noise or unimportant components. The inverse Fourier transform unit maps the adjusted frequency domain features back to the spatial domain for subsequent processing.

[0053] The feature extraction module also includes a spectral attention block and a spatial attention block.

[0054] Spectral attention block: By dynamically allocating weights, the model can focus more on important spectral bands while suppressing redundant or useless band information.

[0055] Spatial attention block: Calculates the weights for each position, highlighting the features of the target region and helping the model focus on important spatial regions in the image.

[0056] The classification decision module uses any one of support vector machine, random forest, logistic regression, or deep neural network as the classifier.

[0057] Support vector machine (SVM): A binary classifier based on the maximum margin principle, which can also be extended to multi-class problems. It classifies by finding a hyperplane to maximize the margin between two classes of samples.

[0058] Random forest (RF): An ensemble learning method that improves classification accuracy and robustness by constructing multiple decision trees and integrating their output results.

[0059] Logistic regression (LR): A generalized linear model for binary classification problems. It maps the output of linear regression to probability values between 0 and 1 by applying a logistic function (usually the sigmoid function).

[0060] Deep neural network (DNN): A neural network structure with multiple hidden layers that can learn complex non-linear mapping relationships. It can optimize network parameters through the backpropagation algorithm to improve classification performance.

[0061] Experimental example:

[0062] Three publicly available hyperspectral datasets were selected for the experiments: Indian Pines (IP), University of Pavia (UP), and Salinas Valley (SV). The IP dataset contains 224 bands (200 effective bands), with a resolution of 145×145 pixels and a total of 16 ground object classes. The UP dataset has 103 bands, an image size of 610×340 pixels, and covers 9 ground object classes. The SV dataset has 204 effective bands, an image size of 512×217 pixels, and contains 16 ground object classes. These three datasets provide diverse experimental data for hyperspectral image classification.

[0063] 4.2 Experimental Setup

[0064] In the experiment, the batch size of the deep learning-based comparison method was set to 16, the Adam optimizer was selected, and the learning rate was 0.0005. The model was selected based on the highest accuracy on the validation set. If the accuracies were the same, the model with the smallest loss was chosen. The best model was saved at each iteration, and if the performance in the next iteration was better, the original model was replaced.

[0065] 4.3 Experimental Comparison Results

[0066] In the Indian Pines dataset, the total number of samples is 10,249, with 307 samples each in the training set and the validation set, and 9,635 samples in the test set. Table 2 shows that the algorithm proposed in this paper is superior to other algorithms in terms of OA, with a 5.19% improvement compared to other HybridSN algorithms.

[0067] The results show that introducing the frequency domain attention module effectively improves the classification accuracy.

[0068] Table 1 Classification Results of the IP Dataset

[0069]

[0070]

[0071] In the Pavia University dataset, the total number of samples is 42,776, with 210 samples each in the training set and the validation set, and 42,356 samples in the test set. Table 2 shows that the algorithm in this paper is superior to other algorithms in terms of OA, with a 3.66% improvement compared to other HybridSN algorithms. The results show that the frequency domain attention module effectively improves the classification accuracy.

[0072] Table 2 Classification Results of the UP Dataset

[0073]

[0074] In the Salinas Valley dataset, the total number of samples is 54,129, with 263 samples each for the training set and the validation set, and 53,603 samples for the test set. Table 3 shows that the algorithm in this paper is superior to other algorithms in OA, with a 3.34% improvement compared to other HybridSN algorithms. The results indicate that the frequency-domain attention module effectively improves the classification accuracy.

[0075] Table 3 Classification results of the SV dataset

[0076]

[0077]

[0078] This paper proposes a hyperspectral image classification algorithm with frequency-domain attention (3D-MixFANet). This network adopts a spectral branch and a spatial branch. The spectral branch combines with the spectral attention block through the frequency-domain module, which can effectively capture the dependencies between spectral bands in hyperspectral images, thereby improving the extraction and classification accuracy of spectral information. The experimental results show that the algorithm proposed in this paper achieves better classification performance on three publicly available hyperspectral datasets (IP, UP, SV) compared to other classification methods. As mentioned above, the above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.

Claims

1. A hyperspectral image classification method based on a three-dimensional hybrid frequency domain attention network, characterized in that: The following steps are involved: a) Data preprocessing: preprocessing the input hyperspectral image data; b) Feature extraction: By constructing a dual-branch 3D convolutional neural network structure, including a spectral branch and a spatial branch, spectral features and spatial features are extracted respectively; the spectral branch extracts spectral features through multiple 3D convolutional layers, and the spatial branch extracts spatial features through 3D convolutional layers and frequency domain attention modules; c) Feature fusion: The features extracted by the spectral branch and the spatial branch are fused to obtain a fused feature map; d) Classification decision: Input the fused feature map into the classifier, make a classification decision, and output the classification result.

2. The hyperspectral image classification method based on three-dimensional hybrid frequency domain attention network according to claim 1 is characterized in that: The frequency domain attention module includes a Fourier transform unit, a frequency domain feature adjustment unit and an inverse Fourier transform unit; wherein the Fourier transform unit converts the input features into the frequency domain, the frequency domain feature adjustment unit analyzes and adjusts different frequency components, and the inverse Fourier transform unit maps the adjusted frequency domain features back to the spatial domain.

3. The hyperspectral image classification method based on three-dimensional hybrid frequency domain attention network according to claim 1 is characterized in that: The dual-branch 3D convolutional neural network structure also includes a spectral attention block and a spatial attention block; wherein the spectral attention block is used to dynamically assign weights to different spectral bands, and the spatial attention block is used to dynamically assign weights to different spatial positions.

4. The hyperspectral image classification method based on three-dimensional hybrid frequency domain attention network according to claim 1 is characterized in that: The classifier is any one of a support vector machine, a random forest, a logistic regression or a deep neural network.

5. A hyperspectral image classification system based on a three-dimensional hybrid frequency domain attention network, characterized in that: include: Data preprocessing module: used to preprocess the input hyperspectral image data; Feature extraction module: used to construct a dual-branch 3D convolutional neural network structure and extract spectral features and spatial features; Feature fusion module: used to fuse spectral features and spatial features; Classification decision module: used to make classification decisions on the fused features and output the classification results.

6. The hyperspectral image classification system based on three-dimensional hybrid frequency domain attention network according to claim 5 is characterized in that: The frequency domain attention module in the feature extraction module includes a Fourier transform unit, a frequency domain feature adjustment unit and an inverse Fourier transform unit.

7. The hyperspectral image classification system based on three-dimensional hybrid frequency domain attention network according to claim 5 is characterized in that: The feature extraction module also includes a spectral attention block and a spatial attention block.

8. The hyperspectral image classification system based on three-dimensional hybrid frequency domain attention network according to claim 5 is characterized in that: The classification decision module uses any one of support vector machine, random forest, logistic regression or deep neural network as a classifier.

Citation Information

Cited By

  • Medical hyperspectral image enhancement method and system based on multi-domain fusion

    CN121582077A