Hyperspectral image classification method based on tensor width learning network
Through the tensor width learning network-based method, the problems of random feature mapping instability and three-dimensional data structure damage in hyperspectral imaging are solved, and high-precision and efficient hyperspectral image classification are achieved.
Patent Information
- Application Number
- CN202510434324.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-17
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-18
AI Technical Summary
Existing width learning methods have problems in hyperspectral imaging that random feature mapping leads to unstable classification performance and destroys three-dimensional data structures when processing higher-order tensor data.
Using a method based on tensor width learning network, CP decomposition is performed through null spectrum feature extraction and alternating least squares method to generate a hyperspectral null spectrum double dictionary representation, and a standard orthogonal activation function is used to randomly map and enhance feature nodes, and finally the classification results are obtained through decision-level fusion.
It realizes the stable processing of high-order tensor data in different scenarios, captures spatial and spectral structure information, improves classification accuracy and stability, and reduces training time.
Smart Images

Figure CN120339839A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image classification, and particularly relates to a hyperspectral image classification method based on a tensor wide learning network. Background Art
[0002] Hyperspectral imaging (HSI), as one of the key technologies in the field of remote sensing, has become an important way for humans to achieve earth observation. Compared with traditional remote sensing images, hyperspectral remote sensing images not only have higher spectral resolution, but also have further improved spatial resolution. However, while the high spatial and spectral resolution remote sensing images bring us richer spectral information and stronger spatial visual impact, they also bring a series of new problems in information extraction and pattern recognition. Due to the three-dimensional cube data structure of hyperspectral remote sensing images, traditional vector forms cannot effectively retain the inherent geometric structure of hyperspectral remote sensing data, usually resulting in the loss of local structure information. Secondly, the large number of bands in hyperspectral remote sensing images usually leads to too high algorithm computational complexity.
[0003] Due to its advanced feature learning ability, deep learning has developed rapidly and is applied to hyperspectral remote sensing image processing. Compared with traditional classification methods, deep learning has strong non-linear representation ability and can distinguish and extract deeper and different features.
[0004] Using a convolutional neural network (CNN) for image classification can effectively utilize the global data and multi-scale information of hyperspectral imaging. However, these improvements mainly benefit from a large number of parameter optimizations. In order to avoid overfitting, solving this high-dimensional optimization problem usually requires a large number of well-labeled training samples. In addition, during the training process, most deep learning methods are time-consuming and sensitive to parameters. To make up for these deficiencies, a wide learning system (BLS) model that does not rely on the deep network structure has emerged. This model can also be well applied to hyperspectral imaging classification and can make full use of the spatial and spectral information of hyperspectral imaging. This model does not involve the coupling between deep layers, so it is simpler and more efficient, especially outstanding in terms of cost savings. In addition, the classification accuracy can be improved by expanding the network width, thereby enhancing the representation ability of features in the wide learning framework. However, the wide learning method still faces the following two problems when processing hyperspectral imaging data: on the one hand, random feature mapping easily leads to unstable classification performance and lack of interpretability; on the other hand, it is necessary to vectorize high-order tensor data, which will destroy its inherent three-dimensional data structure when processing hyperspectral imaging. Summary of the Invention
[0005] Aiming at the above deficiencies in the prior art, the hyperspectral image classification method based on the tensor width learning network provided by the present invention solves the problems that in the existing width learning method, the classification performance is easily unstable due to random feature mapping, and it is necessary to vectorize high-order tensor data, which will destroy its inherent three-dimensional data structure when processing hyperspectral imaging.
[0006] In order to achieve the above object of the invention, the technical solution adopted by the present invention is as follows: Provided is a hyperspectral image classification method based on a tensor width learning network, which includes the following steps: S1. Obtain a hyperspectral image and perform spatial-spectral feature extraction to construct a hyperspectral tensor feature representation based on feature extraction; S2. Input the hyperspectral tensor feature representation into the tensor width learning network, and perform CP decomposition by using the alternating least squares method to obtain a hyperspectral spatial-spectral double dictionary representation, which is used as the feature nodes of the tensor width learning network; S3. Randomly map the feature nodes of the tensor width learning network by using a standard orthogonal activation function to obtain the enhanced nodes of the tensor width learning network; S4. Perform decision-level fusion on the feature nodes and the enhanced nodes to obtain a hidden layer; calculate the weight of the fully connected layer between the hidden layer and the output layer of the tensor width learning network by calculating the matrix pseudo-inverse, and finally obtain the classification result.
[0007] Further, the formula for the hyperspectral tensor feature representation is:
[0008] Wherein, is the hyperspectral tensor feature representation, is the rank of the hyperspectral tensor feature representation is the outer product of vectors, is the summation function, is the summation function, , and are the first matrix, the second matrix and the third matrix of the hyperspectral tensor feature representation respectively, is the th weight parameter, is the th group of vectors of the first matrix is the th group of vectors of the second matrix is the th group of vectors of the third matrix; represents the , and Perform the outer product operation.
[0009] Furthermore, the alternating least squares algorithm in step S2 specifically includes: S2-1. Perform mode unfolding on the hyperspectral tensor feature representation to obtain the first-mode unfolding matrix of the hyperspectral tensor feature representation , the second-mode unfolding matrix and the third-mode unfolding matrix ; S2-2. Fix the second matrix of the hyperspectral tensor feature representation and the third matrix of the hyperspectral tensor feature representation , and calculate the first matrix of the hyperspectral tensor feature representation based on the first-mode unfolding matrix , that is, according to the following formula:
[0010] to obtain the updated first matrix , where is the transpose matrix, is the pseudo-inverse, is the Khatri-Rao product; S2-3. Fix the third matrix and the updated first matrix , and calculate the second matrix based on the second-mode unfolding matrix , that is, according to the following formula:
[0011] to obtain the updated second matrix ; where , and are the first matrix, the second matrix and the third matrix of the hyperspectral tensor feature representation respectively; S2-4. Fix the updated first matrix and the updated second matrix , and calculate the third matrix based on the third-mode unfolding matrix , that is, according to the following formula:
[0012] to obtain the feature node , that is, the hyperspectral spatial-spectral double dictionary representation.
[0013] Further, the formula in step S3 is as follows:
[0014] where, is the th enhanced node, is the orthonormal activation function, are the respective elements of the feature node respectively, and and are both random parameters.
[0015] Further, the formula for obtaining the classification result in step S4 is as follows:
[0016] where, is the classification result, is the weight of the fully connected layer between the hidden layer and the output layer of the tensor width learning network, is the feature node, is the enhanced node, is the hidden layer of the tensor width learning network obtained by performing decision-level fusion on the feature node and the enhanced node respectively.
[0017] The beneficial effects of the present invention are as follows: 1. It can directly process high-order tensor data in different scenarios, can well match the three-dimensional structure of hyperspectral imaging data, capture spatial and spectral structure information, and has high stability; 2. Feature nodes are generated from two dictionaries in space and spectrum, which can well avoid the uncertainty brought by the random feature mapping of the network and make the feature nodes more interpretable; 3. Considering that hyperspectral data has a non-linear structure, the feature nodes of the tensor subspace representation are enhanced through orthonormalized random mapping and activation functions, and enhanced nodes are generated to supplement the feature nodes, improving the non-linear data processing ability of the network and further reducing the training time of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is the flowchart of the invention; Figure 2 is the structural schematic diagram of the width learning network. DETAILED DESCRIPTION OF THE INVENTION
[0019] The specific embodiments of the present invention will be described below to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0020] As Figure 1 and Figure 2 shown, the hyperspectral image classification method based on the tensor width learning network includes the following steps: S1. Obtain a hyperspectral image and perform spatial-spectral feature extraction to construct a hyperspectral tensor feature representation based on feature extraction; generate hyperspectral spatial-spectral features, that is, hyperspectral features, by using the spectral of spatially adjacent pixels of the central pixel point through the window method.
[0021] The hyperspectral image feature tensor after spatial-spectral feature extraction is in the form of a third-order tensor (hereinafter referred to as third-order tensor data) (i.e., a three-dimensional array). The first dimension is the spatial feature direction, the second dimension is the spectral feature direction, and the third dimension represents the number of samples (i.e., the number of pixel points). Use tensor decomposition to construct a spatial-spectral double dictionary representation of the hyperspectral image. The dictionary and coefficient matrix of this representation are obtained by solving the following minimization objective function, that is .
[0022] S2. Input the hyperspectral tensor feature representation into the tensor width learning network, and perform CP decomposition using the alternating least squares method to obtain a hyperspectral spatial-spectral double dictionary representation, which is used as the feature node of the tensor width learning network; it can reduce the uncertainty caused by the randomness of the width learning network.
[0023] The formula for the hyperspectral tensor feature representation is:
[0024] where is the hyperspectral tensor feature representation, is the rank of the hyperspectral tensor feature representation , is the vector outer product, is the summation function, , and are the first matrix, the second matrix, and the third matrix of the hyperspectral tensor feature representation respectively, is the th weight parameter, is the th group of vectors of the first matrix is the th The group vector, is the third matrix of the group vector; denotes the outer product operation on , and .
[0025] The first matrix and the second matrix are the spatial dictionary and the spectral dictionary respectively. The spatial dictionary is a method based on energy level differentiation, which uses the inherent energy differences of molecular vibration and rotation to identify and separate the spectral features of ground objects; the spectral dictionary uses norm minimization problems such as nuclear norm and L2 norm to cluster the spectral features of ground objects; feature nodes are generated from the two dictionaries of space and spectrum , avoiding the uncertainty brought by the random feature mapping of the network and improving the classification effect of the model on hyperspectral images.
[0026] Unfolding the third-order tensor data along the first dimension into a matrix is denoted as , unfolding along the second dimension into a matrix is denoted as , and unfolding along the third dimension into a matrix is denoted as . respectively represent the 1-mode unfolding matrix, 2-mode unfolding matrix, and 3-mode unfolding matrix of the third-order tensor data . It should be noted that unfolding the tensor data is a prior art, such as Zhang Lefei's "Research on Tensor Expression and Manifold Learning Methods of Remote Sensing Images", and the specific unfolding process will not be elaborated here. Based on this, the alternating least squares algorithm in step S2 specifically includes: S2-1. Modal unfolding is performed on the hyperspectral tensor feature representation to obtain the 1-mode unfolding matrix , 2-mode unfolding matrix , and 3-mode unfolding matrix of the hyperspectral tensor feature representation respectively; S2-2. Fix the second matrix and the third matrix , then can be converted into matrix form as follows:
[0027] where is the norm F, is the spectral dictionary, that is, the second matrix. Then the optimal solution is obtained:
[0028] Since the pseudo-inverse of the Khatri-Rao product has the special form in the formula , the optimal solution is rewritten as:
[0029] The updated first matrix is obtained , where is the transpose matrix, is the pseudo-inverse, is the Khatri-Rao product; S2-3. Fix the third matrix and the updated first matrix , then can be converted into matrix form as follows:
[0030] Furthermore, we get:
[0031] Similarly, according to the following formula:
[0032] The updated second matrix is obtained; S2-4. Fix the updated first matrix and the updated second matrix , then can be converted into matrix form as follows:
[0033] Furthermore, according to the following formula:
[0034] The characteristic nodes are obtained, that is, the hyperspectral spatial-spectral double dictionary representation; where , is the number of pixels in the hyperspectral image, is the rank of the hyperspectral tensor feature representation .
[0035] S3. Randomly map the characteristic nodes of the tensor width learning network using an orthonormal activation function to obtain the enhanced nodes of the tensor width learning network; since the hyperspectral image data has a non-linear structure, the characteristic nodes are enhanced using an orthonormal random mapping and an activation function to obtain the enhanced nodes; the enhanced nodes supplement the characteristic nodes and improve the non-linear data processing ability of the network.
[0036] The formula for step S3 is:
[0037] Where, is the th enhanced node, is the orthonormal activation function, and the sigmoid activation function is specifically used during activation; are respectively the elements of the feature node , are both the j th feature node corresponding random parameters, that is and For each feature node a value will be randomly generated. When there are m feature nodes , there will be m groups of and , and thus corresponding to different enhanced nodes .
[0038] In this embodiment represents a random orthonormal mapping, is a bias term, represents performing an orthonormal mapping on the elements of the feature node, adding a bias after the orthonormal mapping, and then activating through the sigmoid activation function. Therefore as a whole represents the orthonormal activation function. The output result is used as the enhanced node.
[0039] S4. Perform decision-level fusion on the feature node and the enhanced node to obtain the output of the hidden layer; calculate the full connection layer weight between the hidden layer and the output layer of the tensor width learning network through matrix pseudoinversion, and finally obtain the classification result.
[0040] The formula for obtaining the classification result in step S4 is:
[0041] Where, is the classification result, is the weight (matrix) of the full connection layer between the hidden layer and the output layer of the tensor width learning network, is the feature node, is the enhanced node, is the feature node and the enhanced node The output of the hidden layer of the tensor width learning network obtained by decision-level fusion. Specifically, represents concatenating the feature nodes and the enhancement nodes into a matrix, which belongs to the output of the hidden layer. represents multiplying the matrix by the matrix W.
[0042] In this embodiment, the corresponding pseudocode during the training process of this method is:
[0043] The objective function during the above training process is , is the label matrix corresponding to the training samples. By using the Lagrange multiplier method to optimize this objective function, we can obtain . The calculation formula of this W can be calculated by the ridge regression theory. Using the trained W, we can perform actual classification according to the formula of the classification result. Among them is the regularization parameter, is the identity matrix, is the norm 2, is the variable function of the minimum value.
[0044] In an embodiment of the present invention, the generalized learning system (BLS), spectral space three-dimensional convolutional neural network (CNN-SS), three-dimensional deep learning method CNN (CNN-DL), multi-scale three-dimensional deep convolutional neural network (CNN-MS), context deep convolutional neural network (CNN-CD), SVM classifier, RPNET network are compared with this method, and the average accuracy (AA), overall accuracy (OA) and Kappa coefficient are used as evaluation indicators.
[0045] The experimental platform uses PyCharm, and the Indian pines dataset is selected. The Indian pines dataset contains data of 220 bands, the spectral range is from 0.2 to 2.5 meters, and the spatial resolution is 20 meters; it contains 16 ground truth classes, including corn, soybeans, wheat, grass, etc. There are 21025 pixels in the dataset image, among which 10249 are ground object pixels and 10776 are background pixels.
[0046] Randomly set 10% of the samples from each class of the dataset as training data. The network consists of a total of 17 feature nodes and 400 enhancement nodes. The classification performance is the best when the number of groups of feature nodes (tensor rank) is set close to the number of classes.
[0047] Table 1 Data comparison of various hyperspectral image classification methods
[0048] As can be seen from Table 1, in terms of the average accuracy (AA), the average accuracy of this method is 94%, which is greater than that of the Generalized Learning System (BLS), Spectral-Spatial 3D Convolutional Neural Network (CNN-SS), 3D Deep Learning Method CNN (CNN-DL), Multi-Scale 3D Deep Convolutional Neural Network (CNN-MS), Contextual Deep Convolutional Neural Network (CNN-CD), SVM classifier, and RPNET network; in terms of the overall accuracy, the overall accuracy (OA) of this method is as high as 96.23%, far exceeding the overall accuracy of other methods; in terms of the Kappa coefficient, the Kappa coefficient of this method is as high as 95.82%, far exceeding the Kappa coefficient of other methods; the calculation time of this method is 6.48 s, which is greater than that of the Generalized Learning System (BLS) but much less than that of the remaining methods; based on the above data, this method has better classification effect and classification accuracy, as well as less time consumption. In summary, the present invention can directly process high-order tensor data in different scenarios, can well match the three-dimensional structure of hyperspectral imaging data, capture spatial and spectral structure information, and has high stability; feature nodes can be generated from two dictionaries in the spatial and spectral dimensions, which can well avoid the uncertainty brought by the random feature mapping of the network and make the features of the feature nodes more interpretable; without calculating the entire connection weights, by adding an additional activation function calculation to complete the enhanced increment and supplementing the feature mapping nodes with the error obtained from the last decomposition of tensor decomposition, the time for training the network is reduced.
Claims
1. A hyperspectral image classification method based on a tensor width learning network, characterized in that: It includes the following steps: S1. Obtain a hyperspectral image and perform spatial-spectral feature extraction to construct a hyperspectral tensor feature representation based on feature extraction; S2. Input the hyperspectral tensor feature representation into a tensor wide learning network, and perform CP decomposition using the alternating least squares method to obtain a hyperspectral spatial-spectral double dictionary representation, which is used as the feature nodes of the tensor wide learning network; S3. Use an orthonormal activation function to perform random mapping on the feature nodes of the tensor wide learning network to obtain the enhanced nodes of the tensor wide learning network; S4. Perform decision-level fusion on the feature nodes and the enhanced nodes to obtain a hidden layer; calculate the weight of the fully connected layer between the hidden layer and the output layer of the tensor wide learning network by computing the Moore-Penrose pseudoinverse of the matrix, and finally obtain the classification result.
2. The hyperspectral image classification method based on a tensor width learning network according to claim 1, wherein The formula for the hyperspectral tensor feature representation is: Among them, is the hyperspectral tensor feature representation, is the hyperspectral tensor feature representation of the rank, is the outer product of vectors, is the summation function, and and are the first matrix, the second matrix and the third matrix of the hyperspectral tensor feature representation respectively, is the th weight parameter, is the th group of vectors of the first matrix is the th group of vectors of the second matrix is the th group of vectors of the third matrix; represents performing an outer product operation on and and .
3. The hyperspectral image classification method based on a tensor width learning network according to claim 1, characterized in that: The specific steps of the alternating least squares algorithm in step S2 include: S2-1. Perform modal expansion on the hyperspectral tensor feature representation to obtain the first-mode expansion matrix , the second-mode expansion matrix , and the third-mode expansion matrix of the hyperspectral tensor feature representation ; S2-2. Fixed Hyperspectral Tensor Feature Representation The second matrix and the hyperspectral tensor feature representation The third matrix , and based on the mode-1 unfolding matrix calculate the first matrix of the hyperspectral tensor feature representation according to the following formula: Obtain the updated first matrix , where is the transpose matrix, is the pseudo-inverse, is the Khatri-Rao product; S2-3. Fix the third matrix and the updated first matrix , and calculate the second matrix based on the two-mode expansion matrix using the following formula: Obtain the updated second matrix ; wherein , and are respectively the first matrix, the second matrix, and the third matrix of the hyperspectral tensor feature representation; S2-4. Fix the updated first matrix and the updated second matrix , and calculate the third matrix based on the triple-mode expansion matrix using the following formula: Obtain characteristic nodes , namely, hyperspectral spatial-spectral double dictionary representation.
4. The hyperspectral image classification method based on a tensor width learning network according to claim 3, characterized in that: The formula in step S3 is: Among them, is the th enhanced node, is the orthonormal activation function, are the respective elements of the feature node respectively, and are both random parameters.
5. The hyperspectral image classification method based on a tensor width learning network according to claim 4, characterized in that: The formula for obtaining the classification result in step S4 is: Among them, is the classification result, is the weight of the fully connected layer between the hidden layer and the output layer of the tensor width learning network, is the feature node, is the enhancement node, is the feature node and the enhancement node perform decision-level fusion to obtain the hidden layer of the tensor width learning network.