Hyperspectral image classification method based on spatial-spectral adaptive learning and pixel-level filtering
By designing a pixel-level filtering method with spatial-spectral adaptive learning and utilizing pixel-level feature extraction and adaptive filter kernels, the problem of insufficient precision in hyperspectral image classification is solved, and higher classification accuracy is achieved.
Patent Information
- Application Number
- CN202310718384.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-15
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2043-06-15
AI Technical Summary
Existing hyperspectral image classification methods fail to fully utilize pixel-level features, resulting in insufficient classification accuracy, especially in edge and difficult-to-distinguish areas.
A pixel-level filtering method based on spatial-spectral adaptive learning is designed. By constructing a spatial-spectral feature extraction network and a filtering dictionary, the filter kernel is adaptively selected for pixel-level filtering to fully utilize the discriminative information of the pixel points.
The classification accuracy of hyperspectral images is improved, especially in the classification accuracy of detail features and edge areas, showing an overall classification effect that is better than existing methods.
Smart Images

Figure CN116824238B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing image processing, and in particular relates to a hyperspectral image classification method based on spatial-spectral adaptive learning and pixel-level filtering. Background Art
[0002] Hyperspectral images have attracted significant attention in the field of remote sensing due to their rich spatial information and the presence of hundreds of continuous narrowband spectral information. Therefore, the information contained in hyperspectral images can be used to represent the distribution of objects and distinguish between categories. In recent years, hyperspectral technology has been widely used in various fields such as Earth observation, resource management, chemical imaging, and environmental monitoring. The goal of hyperspectral image classification is to predict the category of each pixel.
[0003] In the past, a large number of methods used hand-crafted features to achieve HSI classification. However, traditional feature extraction methods are based on the design of data features themselves and therefore cannot match different types of data. Deep learning models are trained in an end-to-end manner, capable of learning the features of different types of input data and outputting prediction results to complete the hyperspectral image classification task. In the early stages, people's attention was focused on the spectral information of hyperspectral images, and a large number of pixel point vector-based methods were proposed for hyperspectral image classification. However, studies have shown that considering spatial information can further improve the accuracy of classification. Some methods focus on the spectral or spatial dimensions and achieve hyperspectral image classification through attention-weighted features.
[0004] However, these methods primarily investigate features within a spatial neighborhood or the entire spatial range, rather than at the pixel level. In reality, every pixel in a hyperspectral image contains rich, discriminative features, known as pixel-level features. Pixel-level features can be viewed as describing global information about all pixels in space. Therefore, fully leveraging the information contained in all spatial pixels can improve the classification accuracy of edges and difficult-to-distinguish areas in hyperspectral images. Summary of the Invention
[0005] The present invention proposes a hyperspectral image classification method based on spatial-spectral adaptive learning and pixel-level filtering. In order to predict the category to which a pixel point belongs, a pixel-level filtering structure based on spatial-spectral adaptive learning is proposed. By selecting an adaptive filter kernel at the pixel level, pixel-level classification of hyperspectral images is achieved.
[0006] The technical solution for achieving the purpose of the present invention is as follows: In a first aspect, the present invention provides a hyperspectral image classification method based on pixel-level filtering of spatial-spectral adaptive learning, comprising the following steps:
[0007] The first step is to design a spatial-spectral pixel-level feature extraction network, which takes hyperspectral image patches as input and obtains discriminative feature maps through a multi-layer convolutional structure.
[0008] The second step is to construct an overcomplete dictionary consisting of multiple filter bases to form the filter kernel corresponding to each pixel;
[0009] In the third step, the features from the first step are used as a guide to adaptively select the filter kernel for each pixel.
[0010] The fourth step is to perform a filtering operation on the filtered hyperspectral image using the pixel-level filter kernel selected in the third step;
[0011] In the fifth step, the filtered hyperspectral feature map is sent to the fully connected layer, and the classifier obtains the final pixel category prediction.
[0012] In a second aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in the first aspect when executing the program.
[0013] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.
[0014] Compared with the existing technology, the present invention has the following significant features: (1) a novel spatial-spectral feature extraction structure is designed, which fully considers the spatial and spectral information of the original hyperspectral image. In addition, the extracted pixel-level features are used as a guide to adaptively select the filter kernel from the generated filter dictionary; (2) the adaptively selected filter kernel is used to achieve classification of the hyperspectral image, making full use of the discriminative information contained in different pixels.
[0015] This paper uses hyperspectral images to learn network parameters in a supervised manner. The hyperspectral images are used as input to a spatial-spectral pixel-level feature extraction process to learn the discriminative features corresponding to each pixel. The output of this process serves as a guide for adaptively selecting filter kernels from a predefined dictionary.
[0016] The hyperspectral image feature map obtained through a simple convolution operation uses an adaptive filter kernel to achieve pixel-level adaptive filtering. Ablation experiments demonstrate the effectiveness of our proposed module, which can capture cross-channel correlation, local neighborhood similarity, and more diverse features after the filtering operation. Furthermore, the proposed pixel-level filtering method is also expected to be useful in the fields of image denoising and deblocking.
[0017] The present invention is further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flow chart of the method of the present invention.
[0019] Figure 2 This is a content coding network structure diagram proposed by the present invention.
[0020] Figure 3 Classification results of different methods on the dataset, including (a) pseudo-color image (b) ground truth image (c) SVM (d) residual network (ResNet) (e) CBAM (f) ECA-Net (g) A2S2K-ResNet (h) regularized spectral–spatial global learning (RSSGL) (i) the method of the present invention. DETAILED DESCRIPTION
[0021] The present invention proposes a pixel-level filtering method of spatial-spectral adaptive learning to fully consider the discriminant information of pixels at different spatial positions, thereby solving the problem of hyperspectral image classification. It mainly consists of a spatial-spectral adaptive learning module and a pixel-level filtering module. Specifically, the purpose of the former is to obtain the joint spatial-spectral discriminant features of each pixel point and use it as a guide for adaptive selection of filter kernels. The latter uses an adaptive filter kernel to implement pixel-level filtering of hyperspectral images to learn the discriminant features contained in different pixel points and use them for classification tasks. Figure 1 、 Figure 2 As shown, the steps of this method are as follows:
[0022] The first step is to design a spatial-spectral pixel-level feature extraction network, which takes a hyperspectral image block as input and obtains a discriminative feature map through a multi-layer convolutional structure. The specific process is as follows:
[0023] (1) Design a spatial-spectral pixel-level feature extraction network, which consists of four stages, each of which contains 3, 4, 6, and 3 bottleneck blocks respectively. For each bottleneck block, it consists of three convolutional layers and a spatial-spectral pixel-level feature extraction process. We define the input data of the spatial-spectral pixel-level feature extraction process as y, and then define the characteristic function of this process as follows:
[0024] X bottleneck,i =SSP(X bottle,i θ s )
[0025] Among them, X bottle,i Represents the input feature map, SSP(X bottle,i θ s ), that is, X bottleneck,irepresents the output feature map after the spatial-spectral pixel-level feature extraction process, θ s is a parameter, i represents the number of bottleneck blocks contained in each stage.
[0026] (2) In order to learn spatial-spectral discriminative features, we introduce the attention mechanism, the steps are as follows:
[0027] ① Input feature map X group After the GAP operation, the integrity of the channel information is maintained by compressing the spatial dimension of each channel:
[0028]
[0029] Where S is X group Length / width of space, X GAP is the output feature after GAP operation;
[0030] ② In order to learn the attention-weighted interaction information across channels, the channel attention mechanism is used:
[0031]
[0032] in, represents a 3-D convolution operation with a kernel size of 1×1×k, X spectral is the weighted channel attention feature, Indicates the product operation according to the channel dimension of the feature map, σ represents the sigmoid activation function;
[0033] ③ In order to learn the attention-weighted features of the spatial dimension, the spatial attention mechanism is used:
[0034]
[0035] Among them, f 7×7 represents a 2-D convolution operation with a kernel size of 7×7, GN represents a group norm group regularization operation, GMP represents a global maximum pooling global maximum pooling operation, cat represents a channel-by-channel splicing operation, X spatial is the weighted spatial attention feature;
[0036] ④ Use BN normalization to fuse the spatial and channel weighted features to obtain the final output features:
[0037] regroup=BN(X spectral +X spatial )
[0038] Among them, BN represents batch norm batch normalization operation, regroup represents the fused spatial-spectral features, which serves as the output of the entire spatial-spectral pixel-level feature extraction process.
[0039] The second step is to construct an overcomplete filter dictionary. Considering the properties of the filters, as well as their complexity and effectiveness, we chose a filter dictionary consisting of 51 Gaussian filter kernels and 21 Difference of Gaussians (DoG) filter kernels. Gaussian filter kernels are generated by scaling, rotating, and stretching a Gaussian, while Difference of Gaussians (DoG) are formed by the differences of Gaussians. These two kernels, respectively, have strong image representation capabilities and the ability to preserve detailed information such as object edges and textures.
[0040] The third step is to adaptively select the filter kernel based on the features. The specific process is as follows:
[0041] (1) Using the regularized pixel feature Λ learned in the first step as a guide, adaptively select the pixel-level filter kernel F from the filter dictionary D by linear combination. i :
[0042]
[0043] Among them, i represents the i-th pixel point in the space, L represents the number of filter bases in the filter dictionary, Λ i,l Represents the lth channel of the i-th pixel in the feature Λ, which is used as the coefficient of the selection filter dictionary, D l Represents the lth filter basis in the dictionary D;
[0044] (2) Reshape the filter kernel of the pixel point into a form that is convenient for subsequent operations:
[0045]
[0046] Among them, k is the spatial size of the filter kernel, and F is the reshaped pixel filter kernel.
[0047] The fourth step is to use the selected pixel-level filter kernel to perform the filtering operation on the filtered hyperspectral image. The specific process is as follows:
[0048]
[0049] Among them, p′ i′,j′ represents the pixel point at position (i′, j′) of the filtered hyperspectral image, F i′,j′ is the adaptive filter kernel at the corresponding position, W and H are the width and height of the space, y filter is the filtering result.
[0050] In the fifth step, the filtered hyperspectral feature map is sent to the fully connected layer, and the classifier obtains the final pixel category prediction.
[0051] The effect of the present invention can be further illustrated by the following simulation experiments:
[0052] Simulation conditions
[0053] Simulation experiment Pavia University (UP) dataset: consists of 610×340 pixels and 115 spectral bands with a wavelength range of 0.43-0.86μm. Twelve of the bands were removed due to noise, so the remaining 103 spectral bands are generally used. It includes trees, asphalt roads (Asphalt), bricks (Bricks), pastures (Meadows), etc. The ground truth includes 9 categories, and they are not all mutually exclusive. We selected 1% of the labeled samples as training samples and the other samples as test samples. Here, we use the UP dataset to verify the feasibility of the proposed method in hyperspectral image classification, and compare the performance of the six classification methods in classification accuracy. The simulation experiments are all configured under the Windows 11 operating system with an AMD Ryzen 5600X (3.7GHz) and RTX 3060GPU environment, and the program is written in Python and PyCharm 2021.
[0054] The evaluation indicators used in the present invention are overall accuracy (OA), average accuracy (AA) and statistical kappa coefficient (κ).
[0055] Simulation content
[0056] The proposed method uses the UP dataset to test its performance. To test the performance of the proposed method, we compared the proposed hyperspectral image classification method based on spatial-spectral adaptive learning and pixel-level filtering with currently popular hyperspectral classification methods. These methods include SVM, ResNet, CBAM, ECA-Net, A2S2K-ResNet, and RSSGL.
[0057] Analysis of simulation experiment results
[0058] Table 1 shows the comparison results of different evaluation indicators of the UP dataset under different classification methods. It can be seen from Table 1 that in the UP dataset, the hyperspectral image classification method based on spatial-spectral adaptive learning and pixel-level filtering proposed in the present invention has a good effect on detail features and overall classification. Compared with SVM, ResNet, CBAM, ECA-Net, A2S2K-ResNet, and RSSGL, the accuracy of the classification results is improved. The effect diagram of the method of the present invention on the UP dataset is shown in the figure. Figure 3 As shown in FIG, the simulation experimental results of the above real data sets demonstrate the effectiveness of the method of the present invention.
[0059] Table 1. Quantitative evaluation of different algorithms on the UP dataset (OA, AA, kappa)
[0060]
Claims
1. A hyperspectral image classification method based on pixel-level filtering with spatial-spectral adaptive learning, characterized in that: The following steps are involved: The first step is to design a spatial-spectral pixel-level feature extraction network, which takes hyperspectral image patches as input and obtains discriminative feature maps through a multi-layer convolutional structure. The attention mechanism is introduced to learn the features of all pixels through the spatial-spectral pixel-level feature extraction process. The specific process is as follows: (1) Design a spatial-spectral pixel-level feature extraction network, which consists of 4 stages, each stage contains 3, 4, 6, and 3 bottleneck blocks respectively; for each bottleneck block, it consists of three convolutional layers and a spatial-spectral pixel-level feature extraction process; define the input data of the spatial-spectral pixel-level feature extraction process as X, then define the characteristic function of the process as follows: X bottleneck,i =SSP(X bottle,i ;θ s ) Among them, X bottle,i Represents the input feature map, SSP(X bottle,i θ s ), that is, X bottleneck,i represents the output feature map after the spatial-spectral pixel-level feature extraction process, θ s is a parameter, i represents the number of bottleneck blocks contained in each stage; (2) Introducing the attention mechanism to learn spatial-spectral discriminative features, the steps are as follows: ① Input feature map X group After the global average pooling operation, the integrity of the channel information is maintained by compressing the spatial dimension of each channel: Where S is X group Length / width of space, X GAP is the output feature after GAP operation; ②Use the channel attention mechanism to learn the attention-weighted interaction information across channels: in, represents a 3-D convolution operation with a kernel size of 1×1×k, X spectral is the weighted channel attention feature, Indicates the product operation according to the channel dimension of the feature map, σ represents the sigmoid activation function; ③Use the spatial attention mechanism to learn the attention-weighted features of the spatial dimension: Among them, f 7×7 represents a 2-D convolution operation with a kernel size of 7×7, GN represents a group norm group regularization operation, GMP represents a global maximum pooling global maximum pooling operation, cat represents a channel-by-channel splicing operation, X spatial is the weighted spatial attention feature; ④Use batch norm to normalize and fuse the spatial and channel weighted features to obtain the final output features: regroup=BN(X spectral +X spatial ) Among them, BN represents batch norm batch normalization operation, regroup represents the fused spatial-spectral features, which serves as the output of the entire spatial-spectral pixel-level feature extraction process; The second step is to construct an overcomplete dictionary consisting of multiple filter bases to form the filter kernel corresponding to each pixel; In the third step, the features from the first step are used as a guide to adaptively select the filter kernel for each pixel. The fourth step is to perform a filtering operation on the filtered hyperspectral image using the pixel-level filter kernel selected in the third step; In the fifth step, the filtered hyperspectral feature map is sent to the fully connected layer, and the classifier obtains the final pixel category prediction.
2. The hyperspectral image classification method based on pixel-level filtering based on spatial-spectral adaptive learning according to claim 1 is characterized in that: The second step is to construct an over-complete filter dictionary, which consists of 51 Gaussian filter kernels and 21 DoG filter kernels.
3. The hyperspectral image classification method based on pixel-level filtering based on spatial-spectral adaptive learning according to claim 1 is characterized in that: The third step is to adaptively select the filter kernel at the pixel point. The specific process is as follows: (1) Using the regularized pixel feature Λ learned in the first step as a guide, adaptively select the pixel-level filter kernel F from the filter dictionary D by linear combination. i : Among them, i represents the i-th pixel point in the space, L represents the number of filter bases in the filter dictionary, Λ i,l Represents the lth channel of the i-th pixel in the feature Λ, which is used as the coefficient of the selection filter dictionary, D l Represents the lth filter basis in the dictionary D; (2) Reshape the filter kernel of the pixel point into a form that is convenient for subsequent operations: Among them, k is the spatial size of the filter kernel, and F is the reshaped pixel filter kernel.
4. The hyperspectral image classification method based on pixel-level filtering based on spatial-spectral adaptive learning according to claim 3 is characterized in that: The fourth step is to use the selected pixel-level filter kernel to perform the filtering operation on the filtered hyperspectral image. The specific process is as follows: Among them, p' i',j' Indicates the pixel point at position (i', j') of the filtered hyperspectral image, F i',j' is the adaptive filter kernel at the corresponding position, W and H are the width and height of the space, y filter is the filtering result.
5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 4 are implemented.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Hyperspectral image classification method based on adaptive spatial spectrum attention kernel generation network
CN114998725A
Unsupervised Latent Low-Rank Projection Learning Method for Feature Extraction of Hyperspectral Images
US20230114877A1