A lightweight tensor convolutional long short-term memory network for hyperspectral classification

By designing a lightweight tensor convolutional length and short-time memory network, the overfitting problem and high storage complexity of existing deep learning algorithms in small samples are solved, and efficient hyperspectral image classification performance is achieved, which is suitable for the field of intelligent processing of airborne remote sensing.

CN116958709BActive Publication Date: 2025-05-13BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311073728.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-24
Publication Date
2025-05-13
Estimated Expiration
2043-08-24

AI Technical Summary

Technical Problem

The existing hyperspectral image classification algorithm based on deep learning has problems of insufficient training and overfitting, especially in small samples, and its complex structure and high-dimensional parameters lead to high storage complexity, which limits its integration and practical engineering applications on airborne platforms.

Method used

A lightweight tensor convolutional length and short-time memory network is designed. By constructing a lightweight deep tensor null spectrum convolutional length and short-time memory network, combined with the tensor full connection layer, the compression of network parameters and storage complexity is achieved, and the problem of depth overfitting is alleviated.

Benefits of technology

It effectively reduces the parameter quantity and storage complexity of deep networks, improves the classification performance of small sample hyperspectral images, and provides an efficient classification method suitable for the field of airborne remote sensing intelligent processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958709B_ABST
    Figure CN116958709B_ABST
Patent Text Reader

Abstract

The present invention provides a hyperspectral classification method of a lightweight tensor convolutional long short-term memory network. In view of the problems of large number of weight parameters and high storage complexity of spatial-spectral convolutional long short-term memory units, two lightweight tensor spatial-spectral convolutional long short-term memory units are designed based on tensor chain decomposition: (1) each convolution operation in each gate structure is respectively extended to the tensor domain; (2) the four convolution kernels corresponding to the input data at the current moment and the output data at the previous moment are respectively integrated, stacked into two large-size convolution kernels and extended to the tensor domain. With the two lightweight units as the basic structure, two lightweight deep tensor spatial-spectral convolutional long short-term memory networks are respectively designed for hyperspectral image classification. The present invention can maintain good spatial-spectral feature extraction capability, and effectively alleviate the overfitting problem of the entire model at a lower network parameter amount and storage complexity, thereby improving the small sample hyperspectral image classification performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of airborne remote sensing intelligent processing, and in particular relates to a hyperspectral classification method of a lightweight tensor convolution long short-term memory network. Background Art

[0002] Hyperspectral remote sensing technology is an earth observation technology that combines imaging technology and spectral technology. The images obtained have nanometer-level spectral resolution and a data structure characteristic of "graph-spectrum integration", which has unique advantages in material identification. To this end, the EU's "100 Major Innovation Breakthroughs for the Future" report, "Medium- and Long-Term Development Plan for Civil Space Infrastructure", "Medium- and Long-Term Development Plan for Ecological and Environmental Satellites", etc., all clearly stated that hyperspectral remote sensing technology should be vigorously developed, and the acquisition capacity of hyperspectral images has been greatly improved. Among them, hyperspectral image intelligent analysis and processing technology, especially feature extraction and classification technology, is the most critical information acquisition technology.

[0003] With the rapid development of computer vision and artificial intelligence technology, deep learning has been widely used in hyperspectral remote sensing image classification tasks. However, due to the special imaging mechanism, hyperspectral images usually have problems such as limited labeled samples and noise / abnormal data interference, which limits the classification performance of deep learning algorithms. Existing hyperspectral image classification algorithms based on deep learning have the following main shortcomings: (1) Although deep learning algorithms have achieved classification performance far superior to traditional algorithms, the characteristics of complex structures and high-dimensional parameters often lead to insufficient training and overfitting problems, which limit the further improvement of their classification performance, especially in the case of small samples. The network overfitting problem is particularly serious. (2) While achieving excellent classification performance, deep learning algorithms are accompanied by deeper, wider, and more complex network structures, resulting in a large number of network training parameters and high storage complexity, which limits the integration and practical engineering application of such algorithms on embedded platforms such as airborne platforms. Therefore, how to design a deep neural network compression method suitable for hyperspectral image ground object classification, or even a lossless compression method, is a key problem that technicians in the field of airborne remote sensing intelligent processing urgently need to solve. Summary of the invention

[0004] The purpose of the present invention is to provide a hyperspectral image classification method based on a lightweight deep tensor spatial-spectral convolutional long short-term memory network, with the goal of fully reducing the parameters and storage complexity of deep networks and improving the accuracy of ground object classification under small samples. It is applied to the field of airborne remote sensing intelligent processing, alleviates the problem of deep overfitting, and improves the classification performance of small sample hyperspectral images, so as to solve the key problems mentioned in the above background technology.

[0005] To achieve the above object, the present invention provides the following technical solution: a hyperspectral classification method based on a lightweight tensor convolutional long short-term memory network, comprising the following steps:

[0006] S1. Take out the local neighborhood window centered on each pixel of the hyperspectral image to obtain a three-dimensional spatial spectrum data cube, and build a training set and a test set in combination with the corresponding category information;

[0007] S2. Taking the space-spectral convolutional long short-term memory unit as the basic structure, two lightweight tensor space-spectral convolutional long short-term memory units are constructed from two perspectives;

[0008] S3. Using two types of tensor space-spectral convolutional long short-term memory units as the basic structure and combining them with tensor fully connected layers, two lightweight deep tensor space-spectral convolutional long short-term memory networks are designed;

[0009] S4. Using the constructed training set, the lightweight deep tensor spatial-spectral convolutional long short-term memory network is trained to obtain a hyperspectral image spatial-spectral joint classification model;

[0010] S5. Use the trained hyperspectral image spatial-spectral joint classification model to classify the test set and predict the corresponding classification results.

[0011] Furthermore, in the step S2, tensor chain decomposition is utilized to expand the spatial-spectral convolution long short-term memory unit to the tensor domain from two perspectives: 1) each convolution operation of each gated structure in the unit is expanded to the tensor domain respectively; 2) the four convolution kernels corresponding to the input data at the current moment and the output data at the previous moment in the unit are respectively taken as a whole, stacked into two large-size convolution kernels and expanded to the tensor domain respectively, so as to realize the compression of the weight parameter amount in the spatial-spectral convolution long short-term memory unit.

[0012] In step S3, a backbone feature extraction network is built with the two lightweight tensor space-spectral convolutional long short-term memory units designed in step S2 as the basic structure, a classification network is built in combination with the tensor fully connected layer, and two lightweight deep tensor space-spectral convolutional long short-term memory networks are designed.

[0013] In step S4, two lightweight deep tensor space-spectral convolution long short-term memory networks are end-to-end trained based on the training set obtained by executing step S1, and the network parameters are updated until convergence to obtain a trained hyperspectral image classification model.

[0014] In step S5, the test set obtained by executing step S1 is classified on the hyperspectral image spatial-spectral joint classification model obtained by executing step S4, the corresponding classification results are predicted, and the feasibility of the proposed lightweight classification method is verified.

[0015] The lightweight hyperspectral image spatial-spectral joint classification algorithm proposed in the present invention can maintain good spatial-spectral feature extraction capability, and realize effective compression of network parameters and storage complexity, thereby improving its small sample hyperspectral image classification performance, and thus providing a solution for remote sensing image intelligent processing tasks in scenarios with limited computing resources and scarce samples on airborne platforms.

[0016] The present invention provides a hyperspectral classification method based on a lightweight tensor convolution long short-term memory network, with the goal of reducing the number of deep network parameters and storage complexity and improving the classification accuracy of ground objects under small samples. Taking into account the data characteristics of hyperspectral images in the field of airborne remote sensing, a new classification model based on a lightweight deep tensor space-spectral convolution long short-term memory network is constructed. Compared with the existing classification model, the classification algorithm proposed in the present invention can better realize the classification of ground objects in hyperspectral images under the premise of greatly reducing the number of network parameters and storage complexity, and improves the classification performance of the entire algorithm under small sample conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly explain the purpose, design ideas and innovation of the lightweight hyperspectral remote sensing image classification method proposed in the present invention, the present invention will be described in detail with reference to the accompanying drawings and attached tables.

[0018] Figure 1 This is a flow chart of the lightweight hyperspectral remote sensing image classification method proposed in this invention.

[0019] Figure 2 This is the internal structure diagram of the two lightweight tensor space-spectral convolution long short-term memory units proposed in the present invention.

[0020] Figure 3 Structural diagram of the two lightweight deep tensor spatial spectral convolution long short-term memory networks proposed in this invention. DETAILED DESCRIPTION

[0021] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] Example 1

[0023] This embodiment is a hyperspectral classification method of a lightweight tensor convolutional long short-term memory network provided by the present invention, such as Figure 1 As shown, the following steps are included:

[0024] S1. Take out the local neighborhood window centered on each pixel of the hyperspectral image to obtain a three-dimensional spatial spectrum data cube, and combine it with the corresponding category information to construct a training set and a test set.

[0025] Assume that the original hyperspectral image has a dimension of W×H×D, where W, H, and D are the width, height, and number of spectral bands, respectively. First, for the i-th pixel, select the s×s local neighborhood centered on it as the spatial information, and combine the spectral information to obtain a three-dimensional data cube and refactor it into Among them, τ is the time step dimension, and in the experiment, the value is 1. Among them, K is the number of spectral bands after the image spectrum is reduced in dimension by the principal component analysis algorithm.

[0026] Then, construct the training data separately:

[0027]

[0028] And the test data:

[0029]

[0030] The corresponding training labels are for

[0031] and test tags

[0032] Get the training set and test set and is the number of training samples and the number of test samples.

[0033] S2. Taking the space-spectral convolution long short-term memory unit as the basic structure, two lightweight tensor space-spectral convolution long short-term memory units are constructed from two perspectives, such as Figure 2 shown.

[0034] As shown in formula (1), this is the internal calculation formula of the original spatial spectral convolutional long short-term memory unit (3-D Convolutional Long Short-term Memory, ConvLSTM3D):

[0035]

[0036] in, Input data to the ConvLSTM3D unit, and is the output data and state of the ConvLSTM3D unit at the previous moment, w, h and a are the width, height and spectral dimensions. C and S are the number of input channels and output channels. i t 、f t and t are input, forget and output gates, corresponding to weights and The dimensions are k×k×k×C×S or k×k×k×S×S. In theory, the value of k in width, height, and spectral dimensions can be different. · is x, h, and c. ※ is a three-dimensional convolution operation. is the Haddon code product operation.

[0037] According to formula (1), the ConvLSTM3D unit includes the input data And the output data of the previous moment Corresponding to two different convolution kernels: and A total of 8 convolution kernels (not considering ). To reduce the number of unit parameters and storage complexity, construct Figure 2 Lightweight tensor space-spectral convolution long short-term memory unit based on tensor decomposition theory.

[0038] First, the two input data and They are reconstructed into 3+d-order high-order tensors along the channel dimension, respectively. and

[0039] for Figure 2 (a), based on 3D Tensor-Train (3DTT), we first transform and Corresponding to 8 convolution kernels and Decompose into d+1 tensor chain decomposition kernels along each channel dimension:

[0040]

[0041] in, and represents the 3DTT core, j w , j h , j d =1,2,...,k,c d =1,2,...,C d ,s d =1,2,...,S d . is the 3DTT rank, r0=r d+1 =1. To facilitate subsequent analysis, the remaining ranks [r1, ..., r d ] are all the same, represented by the variable r. Then, according to the 3DTT calculation principle, complete and and and The convolution-tensor modular product operation of and And reconstruct them into 4-order tensors respectively and As the output result of the 8 convolution operations in the unit. To facilitate the distinction from the subsequent content, the above calculation process is expressed as and The first tensor space-spectral convolution long short-term memory unit (named TTConvLSTM3D-1 unit) is obtained, and formula (1) is updated as follows:

[0042]

[0043] for Figure 2 (b) Different from the TTConvLSTM3D-1 unit, in order to further compress the number of parameters of the entire unit, while reducing the computational complexity of formula (2), we explore the correlation between the weights of different gate structures. First, and The corresponding 8 convolution kernels and Cascading along the channel dimension, we get:

[0044]

[0045] Among them, [·,·,·,·] represents a cascade operation, and and Then, similar to formula (2), and Decompose into d+1 tensor chain decomposition kernels along the channel dimension:

[0046]

[0047] in, and For the 3DTT core, Q=4S,q d =1,2,...,Q d Then, we can get the same formula (3) and and and The result of the convolution-tensor modular product operation: and And reconstructed into three-order tensors respectively and Finally, and Split equally into 4 4th-order tensors along the channel dimension and 4 rank-4 tensors As the output result of the 8 convolution operations in the unit. To distinguish it from the above content, the above calculation process is expressed as and Therefore, the second tensor space-spectral convolution long short-term memory unit (named TTConvLSTM3D-2 unit) is obtained, and formula (1) is updated as follows:

[0048]

[0049] Among them, split(·) represents the average split operation along the channel dimension.

[0050] S3. Using two types of tensor space-spectral convolutional long short-term memory units as the basic structure and combining them with tensor fully connected layers, two lightweight deep tensor space-spectral convolutional long short-term memory networks are designed, such as Figure 3 shown.

[0051] Firstly, the two types of tensor space-spectral convolutional long short-term memory units in formulas (3) and (6) are used as the basic structures to obtain two types of tensor space-spectral convolutional long short-term memory network layers (TTConvLSTM3D-1 network layer and TTConvLSTM3D-2 network layer).

[0052] Then, the three-dimensional data cube constructed in S1 As input data, we extract its deep spatial-spectral features by alternately stacking l layers of TTConvLSTM3D-1 network layers and maximum pooling layers. By alternately stacking l layers of TTConvLSTM3D-2 network layers and maximum pooling layers, its deep spatial spectral features are extracted. Then, the two deep spatial spectral features are mapped to and And input to a tensor chain decomposition fully connected layer (TT Fully Connected Layer, TTFC Layer), mapped into feature vectors and The calculation formula is as follows:

[0053]

[0054]

[0055] in, is the TTFC core, U and V are the number of input and output channels of TTFC Layer, U = w l h l a l S l .u d=1,2,...,U d , v d =1, 2, ..., V d .

[0056] Finally, the feature vector and Input into a traditional fully connected layer (FC Layer) and mapped into feature vectors and Input into the Softmax function and predict the corresponding classification category y i1 and i2 . Where N represents the number of ground object categories.

[0057] Based on the above design, we can get Figure 3 Two lightweight deep tensor space-spectral convolutional long short-term memory networks (named SSTTCL3DNN-1 and SSTTCL3DNN-2, respectively) are shown.

[0058] S4. Use the constructed training set to train the lightweight deep tensor spatial-spectral convolutional long short-term memory network to obtain the hyperspectral image spatial-spectral joint classification model.

[0059] First, the network loss functions of SSTTCL3DNN-1 and SSTTCL3DNN-2 are defined as follows:

[0060]

[0061]

[0062] Among them, y i For input data The corresponding true category label, y i1 and i2 Predicted labels output by the SSTTCL3DNN-1 and SSTTCL3DNN-2 networks.

[0063] Then, based on the training set constructed in step S1 The loss functions in formulas (9) and (10) are respectively optimized by the adaptive momentum optimization algorithm. and End-to-end training and optimization are performed to enable the SSTTCL3DNN-1 and SSTTCL3DNN-2 networks to effectively achieve spatial-spectral joint classification of hyperspectral images with low network parameters and low storage complexity.

[0064] S5. Use the trained hyperspectral image spatial-spectral joint classification model to classify the test set and predict the corresponding classification results.

[0065] First, the test set Input them into the above-trained SSTTCL3DNN-1 and SSTTCL3DNN-2 classification models respectively to obtain the corresponding predicted labels and Then, the predicted label and and the test set label Y test A comparison is made to evaluate the hyperspectral image classification performance of the proposed lightweight deep tensor spatial-spectral convolutional long short-term memory network.

[0066] Example 2

[0067] Based on Example 1, this embodiment selects a public airborne hyperspectral remote sensing image dataset (University of Pavia dataset) to conduct simulation experiments on the lightweight hyperspectral remote sensing image space-spectrum joint classification method proposed in the present invention to verify its feasibility and effectiveness. The dataset is a hyperspectral image of the University of Pavia campus in northern Italy taken by an airborne reflective optical system imaging spectrometer with a spatial resolution of 1.3 meters and a wavelength range of 0.43-0.86 microns, with a spatial resolution of 610×340 pixels. After removing some invalid bands and background pixels, the entire dataset contains 103 spectral bands, 42,779 pixels and 9 types of ground objects for experimental research and comparative analysis. Among them, 10 label samples are randomly selected from each type of ground object target for model training, and the remaining samples are used for testing and verification. In addition, this embodiment selects 3 novel hyperspectral image classification algorithms in the past 5 years as comparison methods. Including: Spatial-Spectral 2-Dand 3-D and Convolutional Long Short-Term Memory Neural Network (SSCL2DNN and SSCL3DNN.IEEE Trans.Geosci.Remote Sens.2020), Spatial-Spectral Tensor-train 2-D Convolutional Long Short-Term Memory Neural Network (SSTTCL2DNN.IEEEJ.Sel.Topics Signal Process.2021).

[0068] Table 1 Network parameter count (pcs) and model storage size (MB) of different classification algorithms under the University of Pavia dataset

[0069]

[0070] Table 2 Classification results of different classification algorithms on the University of Pavia dataset (%)

[0071]

[0072] Table 1 gives a comparative analysis of the network parameter quantity (number) and model storage size (MB) of the algorithm of the present invention and the above-mentioned comparative method. Table 2 gives the classification results (%) of the algorithm of the present invention and the above-mentioned comparative method under the data set, including the classification accuracy of each category, the overall classification accuracy (Overall Accuracy, OA), the average classification accuracy (Average Accuracy, AA) and the Kappa coefficient, and the results shown are the average values ​​of 10 random experiments.

[0073] According to Table 1, in terms of network parameter amount and model storage size, the two lightweight hyperspectral image spatial-spectral joint classification models (SSTTCL3DNN-1 and SSTTCL3DNN-2) proposed in the present invention compress the model size of the original SSCL3DNN from 17.30MB to 0.45MB and 0.34MB, respectively, which is compressed by about 38.44 times and 50.88 times. Further, according to Table 2, in terms of hyperspectral image ground object classification performance, compared with the original SSCL3DNN model, the SSTTCL3DNN-1 and SSTTCL3DNN-2 models proposed in the present invention improve the OA index by 2.31% and 4.37%, respectively. The experimental results prove the effectiveness of the lightweight deep tensor spatial-spectral convolution long short-term memory network proposed in the present invention in compressing the network parameter amount and storage complexity, while improving the classification performance of small sample hyperspectral images.

[0074] Aiming at the field of airborne remote sensing intelligent processing, the present invention proposes a hyperspectral classification method based on a lightweight tensor convolutional long short-term memory network. Through network structure derivation, experimental results and comparative analysis, the feasibility and effectiveness of the algorithm proposed in the present invention in reducing the number of deep network parameters and model storage complexity and improving the classification performance of small sample hyperspectral images are proved, thereby providing a solution reference for remote sensing image intelligent processing tasks in scenarios with limited computing resources and scarce samples on airborne platforms.

Claims

1. A hyperspectral classification method based on a lightweight tensor convolutional long short-term memory network, characterized in that: The following steps are involved: S1. Take out the local neighborhood window centered on each pixel of the hyperspectral image to obtain a three-dimensional spatial spectrum data cube, and build a training set and a test set in combination with the corresponding category information; S2. Taking the space-spectral convolutional long short-term memory unit as the basic structure, two lightweight tensor space-spectral convolutional long short-term memory units are constructed from two perspectives; By using tensor chain decomposition, the spatial spectral convolution long short-term memory unit is expanded to the tensor domain from two perspectives: 1) Each convolution operation of each gated structure in the unit is expanded to the tensor domain respectively; 2) The four convolution kernels corresponding to the current moment input data and the previous moment output data in the unit are stacked as a whole into two large-size convolution kernels and expanded to the tensor domain respectively, so as to realize the compression of the weight parameter quantity in the spatial spectral convolution long short-term memory unit; S3. Using two types of tensor space-spectral convolutional long short-term memory units as the basic structure and combining them with tensor fully connected layers, two lightweight deep tensor space-spectral convolutional long short-term memory networks are designed; S4. Using the constructed training set, a lightweight deep tensor spatial-spectral convolutional long short-term memory network is trained to obtain a hyperspectral image spatial-spectral joint classification model; S5. Use the trained hyperspectral image spatial-spectral joint classification model to classify the test set and predict the corresponding classification results.

2. The hyperspectral classification method of a lightweight tensor convolutional long short-term memory network according to claim 1 is characterized in that: In step S3, a backbone feature extraction network is built with the two lightweight tensor space-spectral convolutional long short-term memory units designed in step S2 as the basic structure, a classification network is built in combination with the tensor fully connected layer, and two lightweight deep tensor space-spectral convolutional long short-term memory networks are designed.

3. The hyperspectral classification method of a lightweight tensor convolutional long short-term memory network according to claim 1 is characterized in that: In step S4, two lightweight deep tensor space-spectral convolution long short-term memory networks are end-to-end trained based on the training set obtained by executing step S1, and the network parameters are updated until convergence to obtain a trained hyperspectral image classification model.

4. The hyperspectral classification method of a lightweight tensor convolutional long short-term memory network according to claim 1 is characterized in that: In step S5, the test set obtained by executing step S1 is classified on the hyperspectral image spatial-spectral joint classification model obtained by executing step S4, the corresponding classification results are predicted, and the feasibility of the proposed lightweight classification method is verified.

Citation Information

Patent Citations

  • Method for improving forecasting precision of short temporary rainfall

    CN114462578A

  • Hyperspectral image classification method of lightweight hybrid tensor neural network

    CN116051896A