A Multilevel Feature Representation Method for Hyperspectral Images Based on Tensorized Autoencoder Network
Through the method based on tensorized autoencoder network, the problem of lack of multi-level feature representation of hyperspectral image feature representation in the prior art is solved, and the multi-level feature representation of hyperspectral images is realized, which improves the richness and accuracy of feature representation.
Patent Information
- Application Number
- CN202411443392.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-10-16
AI Technical Summary
The existing deep learning methods lack the ability to represent multi-level feature in hyperspectral image feature representation, ignore the role of shallow features of deep networks, and cannot fully combine shallow output and deep output.
A multi-level feature representation method for hyperspectral images based on tensorized autoencoder network is adopted. By inputting hyperspectral images into the autoencoder network, encoding and tensorization processing, all network layers output features are stacked, tensor forms are constructed, and the reverse transmission of the autoencoder network is realized through tensor decomposition terms.
The multi-level feature representation of hyperspectral images is realized, and the shallow and deep features can be used simultaneously, which improves the richness and accuracy of feature representation, and provides a more reliable basis for tasks such as classification.
Smart Images

Figure CN119229142B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of hyperspectral image processing, and particularly relates to a method for multi-level feature representation of hyperspectral images based on a tensorized autoencoder network. Background Art
[0002] Hyperspectral image feature representation is an important technology in the field of hyperspectral image processing. It aims to extract and express features from hyperspectral image data that can describe the image content and distinguish different ground objects or targets. These features usually combine spectral information and spatial information, and can comprehensively and accurately reflect the characteristics of hyperspectral images, providing strong data support for subsequent applications such as classification, recognition, and detection. In addition, through feature representation, the originally complex and redundant hyperspectral data can be made concise and efficient, thereby improving the efficiency and accuracy of data processing. At the same time, a good feature representation method can also enhance the discrimination between different category samples and improve the accuracy and robustness of classification algorithms.
[0003] Deep learning is considered to be the most advanced and effective means of hyperspectral image feature representation at present. It can mine and utilize the deep features hidden behind hyperspectral image data. Compared with traditional machine learning methods, the feature extraction method based on deep learning can obtain more complex structural information and has stronger robustness and feature perception ability. It should be noted that most of the existing deep learning methods use the output features of the deepest layer in the deep network as the feature representation form, but lack the ability of multi-level feature representation and ignore the role of the shallow features of the deep network. If the multi-level feature representation ability of deep learning can be improved, and the shallow output and deep output can be fully combined and utilized to construct a representation network that can output multi-level features, it will be very valuable and meaningful for analysis tasks related to hyperspectral images. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to design a multi-level feature representation method that can simultaneously utilize the shallow and deep features of hyperspectral images, so as to provide a reliable basis for related tasks such as classification.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A method for multi-level feature representation of hyperspectral images based on a tensorized autoencoder network, comprising the following steps:
[0006] Step 1, expand the hyperspectral image into a matrix representation X∈R Q×L , where Q = S×P represents the spatial resolution of the hyperspectral image, and L represents the spectral resolution of the hyperspectral image;
[0007] Step 2, input the hyperspectral image X into the autoencoder network to encode the hyperspectral image X, that is:
[0008]
[0009] Among them, λ represents the regularization parameter, and V j , U j , X j and Y j represent the encoding matrix, decoding matrix, input quantity, and output feature of the j-th layer in the network respectively; J represents the total number of layers of the autoencoder network;
[0010] X j =Y j-1 (j≥2) and X 1 =X;
[0011] Step 3, perform tensorization on Equation (1), stack all the output features Y j ∈R Q×L′ (1≤j≤J), and obtain the tensor form Equation (1) is transformed into:
[0012]
[0013] Among them, α is the coupling parameter, which is used to control the importance of the tensor decomposition terms; represents the tensor product; Cat(·) represents the operator used to stack all the layer output features Y j ; represents 's tensor kernel, which is obtained by performing tensor decomposition on , where Q″ << Q, K″ << L′, J″ << J; in addition, the factor matrix W 1 ∈R Q×Q″ , W 2 ∈R L′×L″ , W 3 ∈R J ×J″ ;
[0014] Since the factor matrices W 1 , W 2 and W 3 are the potential principal components of the extended matrix of in different modes respectively, the potential principal component corresponding to the factor matrix W 3 is constrained to the identity matrix. Therefore, Equation (2) is written as:
[0015]
[0016] Among them, represents the tensor kernel under the factor matrices W 1 ' and W 2 ';
[0017] Because are different Ys j connections along the network layer direction, so the multi-level feature representation model of the hyperspectral image in Equation (3) is rewritten in the following form:
[0018]
[0019] where, F j represents the eigenmatrix form of Y j ;
[0020] Step 4, solve the multi-level feature representation model of the hyperspectral image in Equation (4) to obtain the eigenmatrix F j of the output Y j for each layer and the factor matrix W′ 1 along the spectral direction;
[0021] Step 5, the multi-level feature representation form F of the final hyperspectral image, that is:
[0022]
[0023] It can be seen from Equation (20) that the final feature representation form is the product and accumulation of the eigenmatrix form of the output Y j for each layer and its factor matrix along the spectral direction. Compared with most traditional deep learning-based feature extraction methods, the features extracted by the present invention contain both the shallow features and the deep features of the hyperspectral image data, and have richer information.
[0024] Furthermore, in Step 4, solve the multi-level feature representation model of the hyperspectral image in Equation (4) to obtain the eigenmatrix F j of the output Y j for each layer and the factor matrix W′ 1 , and the specific steps are as follows:
[0025] First, rewrite Equation (4) with constraints into the following unconstrained form:
[0026]
[0027] where, γ is the coupling parameter;
[0028] Introduce the auxiliary variable A j = Y j , and Equation (5) is rewritten as:
[0029]
[0030] where, β is the coupling parameter;
[0031] Solve equation (6) using the alternating update optimization strategy:
[0032] 1) Update V j : Fix U j , Y j , F j , W′ 1 , W′ 2 and A j , then the variable V j is updated by the following formula:
[0033]
[0034] The closed - form solution of equation (7) is expressed as:
[0035]
[0036] 2) Update U j : Fix V j , Y j , F j , W′ 1 , W′ 2 and A j , then the variable U j is updated by the following formula:
[0037]
[0038] The solution of equation (9) is expressed as:
[0039] U j =(X j Y j T )(Y j Y j T ) -1 (8)
[0040] 3) Update Y j : Fix V j , U j , F j , W′ 1 , W′ 2 and A j , then the variable Y j is updated by the following formula:
[0041]
[0042] The solution of equation (10) is expressed as:
[0043]
[0044] Among them, I represents the identity matrix;
[0045] 4) Update F j : Fix V j , U j , Y j , W′ 1 , W′ 2 and A j , then the variable F j is updated by the following formula:
[0046]
[0047] The solution of formula (12) is expressed as:
[0048]
[0049] 5) Update W′ 1 : Fix V j , U j , Y j , F j , W′ 2 and A j , then the variable W′ 1 is updated by the following formula:
[0050]
[0051] The solution of formula (14) is expressed as:
[0052]
[0053] 6) Update W′ 2 : Fix V j , U j , Y j , F j , W′ 1 and A j , then the variable W′ 2 is updated by the following formula:
[0054]
[0055] The solution of formula (16) is expressed as:
[0056]
[0057] 7) Update A j : Fix V j , U j , Y j , F j , W′ 1 and W′ 2 , then the variable Aj Updated by the following formula:
[0058]
[0059] The solution of Equation (18) is expressed as:
[0060]
[0061] where soft(x, y) = sign(x)·max(|x| - y, 0) is the soft threshold operation function; here, sign(x) is the sign function, and max(|x| - y, 0) is the maximum value function used to compare the magnitudes of |x| - y and 0;
[0062] After iterative updates of 1) - 7), the optimal V j , U j , Y j , F j , W′ 1 , W′ 2 and A j are obtained.
[0063] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0064] The multi - level feature representation method for hyperspectral images based on the tensorized auto - encoder network proposed by the present invention stacks the outputs of each layer of the deep auto - encoder network to obtain a tensor form, and constructs and introduces a tensor decomposition term to realize the backpropagation of the auto - encoder network. Compared with existing deep learning methods, this method has significant multi - level feature representation ability, can unify and combine the outputs of each layer of the deep auto - encoder network, and covers both shallow - layer output characteristics and deep - layer output characteristics. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 is the flowchart of the steps of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0067] The present invention constructs a representation network capable of outputting multi-level features, which can simultaneously use the shallow output and the deep output as discriminative features for related tasks such as classification. It is not difficult to find that tensors can unify and combine the outputs of each layer of the deep network. Therefore, by performing tensorization on the existing deep network, it is expected to achieve multi-level feature representation of hyperspectral images. In this embodiment, a method for multi-level feature representation of hyperspectral images based on a tensorized autoencoder network is proposed. The outputs of each layer of the deep autoencoder network are stacked to obtain a tensor form, and a tensor decomposition term is constructed and introduced to realize the backpropagation of the autoencoder network. Compared with existing deep learning methods, this method has significant multi-level feature representation ability, can unify and combine the outputs of each layer of the deep autoencoder network, and simultaneously covers the characteristics of shallow output and deep output. The overall flowchart of the present invention is as shown in Figure 1 shown below, and the main steps involved are as follows:
[0068] Step 1: Input the hyperspectral image Its two-dimensional matrix form can be expressed as X ∈ R Q×L , where Q = S × P represents the spatial resolution of the hyperspectral image, and L represents the spectral resolution of the hyperspectral image.
[0069] Step 2: Introduce an autoencoder network to encode the hyperspectral image X, that is:
[0070]
[0071] where λ represents the regularization parameter, V j , U j , X j and Y j respectively represent the encoding matrix, decoding matrix, input quantity, and output feature of the j-th layer in the network. Here, X j = Y j-1 (j ≥ 2) and X 1 = X.
[0072] Step 3: Perform tensorization on Equation (1), that is, stack all the output features Y j ∈ R Q×L′ (1 ≤ j ≤ J) of all network layers to obtain the corresponding tensor form In this way, Equation (1) can be transformed in one step to:
[0073]
[0074] where α is a coupling parameter used to control the importance of the tensor decomposition term. represents the tensor product. Cat(·) represents stacking all the output features Y joperators. It should be noted that to ensure the smooth stacking of the output features Y of different layers j each Y j is set to have the same dimension. In Equation (2), denotes the tensor kernel, which is obtained by performing tensor decomposition on , where Q″ << Q, L″ << L′, J″ << J. In addition, the factor matrices W 1 ∈ R Q×Q″ , W 2 ∈ R L′×L″ , W 3 ∈ R J×J″ . It can be seen from Equation (2) that the output features of all network layers are stacked into a tensor and tensor decomposition operations are performed. At the same time, the tensor and its decomposition operations also affect the autoencoder network, thereby realizing multi-level feature representation of hyperspectral images.
[0075] Considering that Equation (2) contains both matrix operations and tensor operations, it is difficult to directly solve it. Therefore, Equation (2) needs to be transformed into a form that only contains matrix operations. Specifically: Since the factor matrices W 1 , W 2 , and W 3 are the latent principal components of the extended matrices of in different modes respectively, the latent principal components corresponding to the factor matrix W 3 are constrained to the identity matrix. Therefore, Equation (2) can be written as:
[0076]
[0077] where denotes the tensor kernel under the factor matrices W′ 1 and W′ 2 . Because is the connection of different Y j along the network layer direction, Equation (3) can be rewritten in the following form:
[0078]
[0079] where F j denotes the eigenmatrix form of Y j . It can be seen that Equation (4) only contains matrix operation forms.
[0080] Step 4: Design an alternating update optimization strategy to solve the multi-level feature representation model (4) of hyperspectral images based on the tensorized autoencoder network in Step 3. Specifically, first rewrite Equation (4) with constraint terms into the following unconstrained form:
[0081]
[0082] Among them, γ is the coupling parameter. Introduce the auxiliary variable A j = Y j , Equation (5) can be rewritten as:
[0083]
[0084] Among them, β is the coupling parameter. In Equation (6), the objective function is non-convex and contains multiple variables. However, for any single variable (fixing other variables), this objective function is convex. To sum up, the specific process of the alternating update optimization strategy is shown as follows:
[0085] 1) Update V j : Fix U j , Y j , F j , W′ 1 , W′ 2 , and A j , then the variable V j can be updated by the following formula:
[0086]
[0087] The closed-form solution of Equation (7) can be expressed as:
[0088]
[0089] 2) Update U j : Fix V j , Y j , F j , W′ 1 , W′ 2 , and A j , then the variable U j can be updated by the following formula:
[0090]
[0091] The solution of Equation (9) can be expressed as:
[0092] U j = (X j Y j T )(Y j Y j T ) -1 (10)
[0093] 3) Update Y j : Fix V j , U j , F j, W′ 1 , W′ 2 and A j , then the variable Y j can be updated by the following formula:
[0094]
[0095] The solution of Equation (10) can be expressed as:
[0096]
[0097] where I represents the identity matrix.
[0098] 4) Update F j : Fix V j , U j , Y j , W′ 1 , W′ 2 and A j , then the variable F j can be updated by the following formula:
[0099]
[0100] The solution of Equation (12) can be expressed as:
[0101]
[0102] 5) Update W′ 1 : Fix V j , U j , Y j , F j , W′ 2 and A j , then the variable W′ 1 can be updated by the following formula:
[0103]
[0104] The solution of Equation (14) can be expressed as:
[0105]
[0106] 6) Update W′ 2 : Fix V j , U j , Y j , F j , W′ 1 and A j , then the variable W′ 2 can be updated by the following formula:
[0107]
[0108] The solution of Equation (16) can be expressed as:
[0109]
[0110] 7) Update A j : Fix V j , U j , Y j , F j , W′ 1 and W′ 2 , then the variable A j can be updated by the following formula:
[0111]
[0112] The solution of Equation (18) can be expressed as:
[0113]
[0114] where soft(x, y) = sign(x)·max(|x| - y, 0) is the soft-threshold operation function. Here, sign(x) is the sign function, and max(|x| - y, 0) is the maximum function, which is used to compare the magnitudes of |x| - y and 0.
[0115] After the iterative updates of 1) to 7), the optimal V j , U j , Y j , F j , W′ 1 , W′ 2 and A j can be obtained.
[0116] Step Five. When the optimal solutions F j and W′ 1 are obtained by optimizing the multi-level feature representation model (4) of the hyperspectral image based on the tensorized autoencoder network proposed in Step Three through the alternating update optimization strategy in Step Four, the final multi-level feature representation form F of the hyperspectral image can be obtained, that is:
[0117]
[0118] It can be seen from Equation (20) that the final feature representation form is the product and accumulation of the eigenmatrix form of the output Y j of each layer and its factor matrix along the spectral direction. Compared with most traditional deep learning-based feature extraction methods, the features extracted by the present invention simultaneously contain the shallow features and deep features of the hyperspectral image data, and have richer information.
[0119] The technical means disclosed by the solution of the present invention are not limited to those disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A multi-level feature representation method for hyperspectral images, characterized in that: The steps include: Step 1: transform the hyperspectral image χ∈R S×P×L Expand into a matrix representation X∈R Q×L , where Q = S × P represents the spatial resolution of the hyperspectral image, and L represents the spectral resolution of the hyperspectral image; Step 2: Input the hyperspectral image X into the autoencoder network to encode the hyperspectral image X, that is: Among them, λ represents the regularization parameter, V j , U j , X j and Y j They represent the encoding matrix, decoding matrix, input quantity and output features of the jth layer in the autoencoder network respectively; J represents the total number of layers of the autoencoder network; The input quantity X1 of the first layer in the autoencoder network is written as: X1 = X; When j ≥ 2, X j =Y j-1 ; Step 3: Perform tensor quantization on equation (1) and stack all network layer output features Y j ∈R Q×L′ (1≤j≤J), get the tensor form Formula (1) is transformed into: Among them, α is a coupling parameter used to control the importance of tensor decomposition terms; Represents tensor product; Cat(·) represents the output feature Y for stacking all layers j The operator, express The tensor core of Tensor decomposition is performed to obtain, where Q″<<Q, L″<<L′, J″<<J; in addition, the factor matrix W1∈R Q×Q″ , W2∈R L′×L″ , W3∈R J×J″ ; Since the factor matrices W1, W2, and W3 are the potential principal components of the extended matrices of y in different modes, the potential principal components corresponding to the factor matrix W3 are constrained to be the unit matrix. Therefore, formula (2) is written as: s.t.Y j =V j X j ,j=1,2,…,J, in, represents the tensor kernel under the factor matrices W′1 and W′2; because It is different from Y j Along the connection in the direction of the network layer, the multi-level feature representation model of the hyperspectral image (3) is rewritten as follows: Among them, F j Represents Y j The eigenmatrix form of ; Step 4: Solve the multi-level feature representation model (4) of the hyperspectral image to obtain the output Y of each layer j The eigenvalue matrix F j The optimal solution and the factor matrix W′1 along the spectral direction; Step 5, the final multi-level feature representation F of the hyperspectral image is:
2. According to claim 1, a method for multi-level feature representation of hyperspectral images is characterized in that: In step 4, the multi-level feature representation model (4) of the hyperspectral image is solved to obtain the output Y of each layer j The eigenvalue matrix F j The optimal solution and the factor matrix W′1 along the spectral direction, the specific steps are as follows: First, rewrite the constraint-containing formula (4) into the following unconstrained form: Among them, γ is the coupling parameter; Introduce auxiliary variable A j =Y j , formula (5) is rewritten as: Among them, β is the coupling parameter; The alternating update optimization strategy is used to solve equation (6): 1) Update V j : Fixed U j , Y j 、F j , W′1, W′2 and A j , then the variable V j Update it with: The closed-form solution of formula (7) is expressed as: 2) Update U j : Fixed V j , Y j 、F j , W′1, W′2 and A j , then the variable U j Update it with: The solution of formula (9) is expressed as: OR j (X) j AND j T )(AND j AND j T ) -1 3) Update Y j : Fixed V j , U j 、F j , W′1, W′2 and A j , then the variable Y j Update it with: The solution of formula (10) is expressed as: Where I represents the identity matrix; 4) Update F j : Fixed V j , U j , Y j , W′1, W′2 and A j , then the variable F j Update it with: The solution of formula (12) is expressed as: 5) Update W′1: fix V j , U j , Y j 、F j , W′2 and A j , then the variable W′1 is updated by the following formula: The solution of formula (14) is expressed as: 6) Update W′2: fix V j , U j , Y j 、F j , W′1 and A j , then the variable W′2 is updated by the following formula: The solution of formula (16) is expressed as: 7) Update A j : Fixed V j , U j , Y j 、F j , W′1 and W′2, then variable A j Update it with: The solution of formula (18) is expressed as: Among them, soft(x,y)=sign(x)·max(|x|-y,0) is the soft threshold operation function; sign(x) is the sign function, and max(|x|-y,0) is the maximum value function used to compare the size of |x|-y and 0; After iterative updates from 1) to 7), the optimal V j , U j , Y j 、F j , W′1, W′2 and A j .
Citation Information
Patent Citations
Spectral super-division reconstruction method and system based on low-rank tensor network
CN115719309A
Dangerous chemical safety monitoring method based on big data
CN118506281A