Hyperspectral Image Feature Extraction Method Based on Multi-Level Variational Autoencoder
The spatial and spectral features of high-spectral images are extracted through a multi-level variational automatic encoder, and the homologousness is improved by using the spectral angular distance, which solves the problem of poor correlation of feature maps in the prior art, and achieves higher classification accuracy and noise resistance.
Patent Information
- Application Number
- CN202111627432.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-28
AI Technical Summary
The existing hyperspectral image feature extraction methods fail to fully consider the positional relationship between cells, ignore the homology between spatial features and spectral features, resulting in poor correlation of feature maps and affecting classification accuracy.
A multi-level variational automatic encoder is used to extract the spatial and spectral features of high-spectral images through deep neural networks, and use spectral angle distance to improve feature homology, combine long and short-term memory networks for feature fusion, and optimize the loss function to improve feature collaboration capabilities.
The classification accuracy and noise resistance of hyperspectral image features are improved, the classification accuracy problem is insufficient due to the distribution differences between spatial characteristics and spectral feature data, and better feature mapping correlation and classification performance are achieved.
Smart Images

Figure CN116416441B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of spectral imaging technology, and specifically relates to a hyperspectral image feature extraction method based on a multi-level variational autoencoder. Background Art
[0002] In the field of remote sensing, hyperspectral imaging technology is widely used in various studies. Hyperspectral images contain rich spatial and spectral features. The spatial features refer to the spatial position information of pixels at each wavelength, and the spectral features refer to the spectral curves composed of the spectral reflectance of a single pixel at each wavelength. By extracting features from hyperspectral images, low-dimensional embedded features containing rich discriminant information can be obtained, redundant information in the images can be reduced, and the recognition accuracy in subsequent classification studies can be improved. In the early stage, hyperspectral image feature extraction methods mainly extracted the spectral features of pixels and did not consider the position information between pixels, so it was difficult to obtain good results. With the improvement of computer computing power and the in-depth study of deep learning, some methods for extracting spatial-spectral features by training neural networks have been proposed one after another. These methods introduce the idea of multi-sensor data fusion and adopt the method of separately extracting and fusing spatial and spectral features, avoiding information loss and improving the performance of the algorithm.
[0003] Hyperspectral image feature extraction methods can be classified into spectral feature-based feature extraction methods and spatial-spectral feature-based feature extraction methods from the perspective of information sources.
[0004] The feature extraction method based on spectral features constructs a feature extractor using a single spectral curve in the hyperspectral image and ignores the position information of different pixels in the spatial dimension. Early widely used methods include Principle Component Analysis (PCA), Minimum Noise Fraction (MNF), Linear Discriminant Analysis (LDA), and so on. These methods generally consider the internal discriminant information of hyperspectral pixels to ensure the classification ability. With the continuous development of deep learning research, some deep network models have also been applied to the research of hyperspectral image feature extraction, including Auto-encoder (AE), Variational Auto-encoder (VAE), Long Short-term Memory (LSTM), and so on. However, these methods do not consider the positional relationship between different pixels, so they only describe the information in the image from the spectral level and do not fully utilize the advantage of "spectrum-image integration" of hyperspectral images, that is, there is unity and synergy between the spatial information and spectral information in the image. Currently, the mainstream method in this research field is the feature extraction method based on spatio-spectral features.
[0005] For the feature extraction method based on spatial-spectral features, hyperspectral images are a typical three-dimensional cube data. This data combines the spatial information of the target ground object with the spectral information at each wavelength and reflects them together in the complete data. Therefore, hyperspectral images have the characteristic of the unity of spectrum and image, that is, the spatial information and spectral information of hyperspectral images are consistent. At the same time, due to the influence of the variable shooting environment and external interference of hyperspectral images, there are still the phenomena of "one object with different spectra" and "one spectrum with different objects" in the images, which will also interfere with the results of hyperspectral analysis research. The spatial information of hyperspectral images can be understood as the local spatial neighborhood of a single pixel in the spatial dimension. This definition assumes that each pixel has a certain relationship with the pixels in its spatial neighborhood. Therefore, the position information of the pixel in the real ground object can be mastered by learning the distribution of its local spatial neighborhood. Currently, the more commonly used method is to use a convolutional neural network to extract the local information of the pixel, use a fully connected layer to extract the spectral features of the pixel, and finally use a concatenation layer to realize the combination of spatial-spectral features. When extracting spectral features, some studies also use long short-term memory networks to learn the continuous information in the spectral curve of the pixel. However, these methods have certain defects: ① These methods do not consider the continuous information in hyperspectral images from multiple aspects and only describe the continuous information from the perspective of the spectral curve, weakening the diversity of the data; ② When these methods extract spectral features and spatial features separately, the correlation and cooperation of the two mappings are poor. Only at the last step, a concatenation layer is used to realize feature fusion, and it is impossible to fully match the distributions of the two features; ③ These methods do not consider the homology between spatial features and spectral features. Although the two features describe the information of hyperspectral images from different levels, both describe the same hyperspectral pixel, so there must be homology between them.
[0006] In short, there are certain disadvantages in the current hyperspectral image feature extraction methods:
[0007] ① In the current hyperspectral image feature extraction methods based on spatial-spectral features, the continuous information of pixels is not considered from multiple levels. Most methods only describe the continuous information from the perspective of the spectral curve, weakening the diversity of the data;
[0008] ② In the current hyperspectral image feature extraction methods based on spatial-spectral features, the cooperation ability between the spatial feature mapping and the spectral feature mapping is poor. The two mappings are basically completely separated. After each extracts features separately, a concatenation layer is used for feature fusion. However, there are large differences in the data distributions of the spatial features and spectral features themselves, and it is difficult to achieve the expected purpose by simply splicing;
[0009] ③Currently, the hyperspectral image feature extraction methods based on spatial-spectral features do not consider the homology between spatial features and spectral features, ignoring the fact that both features are extracted from the same hyperspectral pixel. This further leads to an increase in the data distribution difference between the two features, which is not conducive to feature fusion expression and subsequent classification research. Summary of the Invention
[0010] To overcome the above defects, the purpose of this application is to propose a hyperspectral image feature extraction method based on a multi-level variational autoencoder (multi-level VAE) for hyperspectral images. This method uses the variational autoencoder as the basic framework of the method and adopts the finally corrected fusion feature as the spatial-spectral joint feature finally output after training.
[0011] To achieve the above purpose, this application adopts the following technical solutions.
[0012] A hyperspectral image feature extraction method based on a multi-level variational autoencoder, characterized in that the method comprises the following steps:
[0013] S1. Select a hyperspectral image, the size of the hyperspectral image is X×Y×B, where X and Y are the spatial sizes of the hyperspectral image at each wavelength, and B is the number of wavelengths of the hyperspectral image.
[0014] S2. Configure neighborhood information for each hyperspectral pixel in the hyperspectral image, that is, select the neighborhood pixels with a size of s×s around the hyperspectral pixel as the neighborhood information of the pixel, and the size of the neighborhood information is s×s×B, where the neighborhood information refers to a square area centered on the hyperspectral pixel, the side length is s, and s is an odd number.
[0015] S3. Transform the neighborhood information based on a deep network model to obtain a first sample with a size of 1×s 2 ×B, and the first sample is used as the input Input of the spatial feature extraction module p ,
[0016] A second sample of a 1×B hyperspectral pixel is used as the input Input of the spectral feature extraction module q , and the number of the first sample and the second sample is the same and they correspond one by one.
[0017] S4. Train a deep neural network.
[0018] S5. Feature splicing and calculation of the mean feature μ, that is, the input of the layer in the spatial feature extraction module is the output of the layer The output of the Input of the layer is the concatenation of the output of the -th layer and -th layer output, according to the calculation formula: where 1 < i < m, according to the calculation formula, the mean feature μ is obtained, and the size of μ is bs × s 2 × d,
[0019] S6. Pooling operation, that is, using the average pooling layer to perform pooling operations on the mean feature μ and the standard deviation feature δ to obtain the pooled mean feature and the pooled standard deviation feature both with the size of bs × d,
[0020] S7. Obtain the fused feature O based on the feature fusion module and input the fused feature O into the decoder module, and the decoder is used to reconstruct the data of the fused feature O
[0021] S8. Network optimization, that is, construct the loss function during the training of the network model according to the formula Г = Γ R + Γ KL + Γ Homo where ∑(·) adds up all the contents in the parentheses. Among them, Γ
[0022]
[0023]
[0024]
[0025] wherein, ∑(·) sums up all the content in the brackets. Among them, Γ R uses the Euclidean distance to calculate the similarity between the input of the spectral feature extraction module and the output of the encoder, and Γ KL is the loss function in the variational autoencoder VAE, which uses the KL divergence to calculate the similarity between the Gaussian distribution and the embedded feature distribution, and Γ Homo uses the spectral angle distance to calculate the similarity between the output of all spatial feature extraction modules and the output of the corresponding spectral feature extraction modules.
[0026] Preferably, in step S1, it further includes:
[0027] Perform normalization preprocessing on the hyperspectral image, set the neighborhood size s, the number of network layers m in the spatial feature extraction module and the spectral feature extraction module, the number of network layers n in the decoder module, and the embedded feature dimension d, where d is an even number greater than 0.
[0028] Preferably, in step S4, it includes: from X × Y ones with the size of 1 × s 2A small batch of samples is randomly selected from the first sample of ×B and the second samples of X×Y with a size of 1×B and input into the deep neural network for training the deep neural network. The number of small batch pixels of the first sample and the second sample is both bs, and the activation function in the network is the Tanh activation function.
[0029] Preferably, the hyperspectral image feature extraction method based on the multi-level variational autoencoder is characterized in that it further includes:
[0030] Normalize all X×Y hyperspectral pixels so that the value range is between -1 and 1. The normalization formula is as follows:
[0031]
[0032] where x min represents the minimum value in the pixel data, and x max is the maximum value.
[0033] Preferably, the calculation formula of the Tanh activation function is:
[0034] Preferably, in step S8, it further includes: using the loss function, selecting the Adam optimizer with a step size of 10 -3 to optimize the deep network model. After the model reaches stability, use the pooled mean feature as the output.
[0035] Preferably, step S3 includes:
[0036] Transform the hyperspectral pixel x with a size of 1 to obtain a second sample with a size of . The second sample is used as the input Input of the spectral feature extraction module q ,
[0037] If B is not an integer multiple of s 2 , then remove ε wavelengths so that B - ε is an integer multiple of s 2 , where q is used to refer to the relevant variables in the spectral feature extraction module.
[0038] Preferably, step S6 further includes:
[0039] Use the long short-term memory network layer L to obtain the standard deviation feature δ. The input of the long short-term memory network layer L is with the number of nodes being d, and the size of δ is bs×s 2 ×d.
[0040] Preferably, in step S7, the decoder module includes:
[0041] n fully-connected network layers, where each network layer is respectively {d1, d2…d n}, and the inputs of each network layer are respectively {ind1, ind2…ind n}, and the outputs of each network layer are respectively {outd1, outd2…outd n}. The number of nodes in the last layer is B, and the number of nodes in other network layers is d.
[0042] Preferably, step S7 further includes:
[0043] According to the formula the fused feature O is obtained, where γ is a randomly generated noise matrix, which conforms to the Gaussian distribution and has a size of bs×d.
[0044] Beneficial effects
[0045] Compared with the prior art, the multi-continuous feature integration method for hyperspectral images in the embodiments of the present application is used to extract the spatial features and spectral features of hyperspectral images. This method takes into account two types of continuous information contained in hyperspectral images, and through designing a deep network model, the purpose of multi-angle description of information is achieved. In addition, in this method, through the multi-level spatial feature correction method, the spatial features at different stages are used to correct the spectral features at the corresponding stages in turn, so as to improve the cooperation ability of the two features in the extraction process. In this method, based on the spectral angle distance, the multi-level spatial-spectral feature homology improvement method is used to gradually improve the homology between the spatial-spectral features at each level, and to enhance the correlation between the spatial features and the spectral features in the feature mapping stage, so as to solve the problem that it is difficult to improve the classification accuracy due to the large difference in the data distributions of the two features. Description of the drawings
[0046] Figure 1 is a schematic flowchart of the feature extraction method in the embodiments of the present application,
[0047] Figure 2 is a schematic flowchart of the overall differential demodulation process in the embodiments of the present application. Detailed implementation manners
[0048] The above solution will be further described below with specific embodiments. It should be understood that these embodiments are used to illustrate the present application and are not limited to restricting the scope of the present application. The implementation conditions adopted in the embodiments can be further adjusted according to the conditions of specific manufacturers, and the implementation conditions not specified are usually the conditions in conventional experiments.
[0049] The present application provides a hyperspectral image feature extraction method based on a multi-level variational autoencoder, which uses a long short-term memory network layer to extract the continuous features of pixels from the spatial level and the spectral level, and uses a splicing layer to fuse the two continuous features, solving the problem of single continuous information in traditional feature extraction algorithms. The multi-level spatial-spectral feature homology enhancement method based on spectral angular distance uses spectral angular distance to calculate and increase the homology between spatial-spectral features at each stage, solving the problem of difficulty in improving subsequent classification accuracy due to the large difference in data distribution between spatial features and spectral features, and improving the correlation between the two features in the feature mapping stage.
[0050] The hyperspectral image feature extraction method based on a multi-level variational autoencoder proposed in this application is described below with reference to the accompanying drawings.
[0051] like Figure 1 FIG. 1 is a flow chart of a feature extraction method, which includes:
[0052] S1. Select / acquire a hyperspectral image with a size of X×Y×B, where X and Y are the spatial sizes of the hyperspectral image at each wavelength, and B is the number of wavelengths of the hyperspectral image. Perform normalization preprocessing on the hyperspectral image, set the neighborhood size s, the number of network layers in the spatial feature extraction module and the spectral feature extraction module m, the number of network layers in the decoder module n, and the embedding feature dimension d, where d needs to be an even number greater than 0.
[0053] S2. Set neighborhood information for each hyperspectral pixel (there are X×Y hyperspectral pixels in total), select the surrounding neighborhood pixels with a size of s×s as the neighborhood information of the pixel, and the size of the neighborhood information is s×s×B.
[0054] S3. Build a deep network model and transform the neighborhood information of the pixel to obtain a size of 1×s 2 The first sample of ×B is used as the input of the spatial feature extraction module p . Transform the hyperspectral pixel x of size 1×B to obtain The second sample is used as the input of the spectral feature extraction module q , if B is not s 2 If ε is an integer multiple of , then remove ε wavelengths so that B-ε is s 2 An integer multiple of . The superscript p is used to refer to the relevant variables in the spatial feature extraction module, and the superscript q is used to refer to the relevant variables in the spectral feature extraction module. The number of the first sample and the second sample is the same and one-to-one corresponding. The deep network model includes a spatial feature extraction module, a spectral feature extraction module, a feature fusion module and a decoder module.
[0055] S4. Train the deep neural network. Randomly select a small batch of samples from all X×Y hyperspectral pixels and input them into the deep neural network. The number of pixels in the small batch is bs. The activation function in the network is the Tanh activation function. Except for the last layer of the decoder module, a batch normalization layer (Batch Normalization Layer) is connected after all other network layers.
[0056] S5. Feature concatenation operation
[0057] That is, take Input p as the input of the layer in the spatial feature extraction module, and obtain the output
[0058] Take Inputq as the input of the layer in the spectral feature extraction module, and obtain the output where and both have the size of bs×s 2 ×d. Take as the input of the layer in the spatial feature extraction module According to the following formula
[0059]
[0060] where Concat(·) is the concatenation operation, concatenate the two in the third dimension to obtain an output with the size of bs×s 2 ×2d and use it as the input of the layer in the spectral feature extraction module
[0061] The input of the layer in the spatial feature extraction module is the output of the layer The input of the layer in the spectral feature extraction module is the concatenation of the outputs of the layer and
[0062]
[0063] where 1 < i < m. According to the following formula:
[0064]
[0065] Obtain the mean feature μ, and the size of μ is bs×s2 ×d.
[0066] S6. Pooling operation, that is, using the average pooling layer (Average Pooling Layer) to perform pooling operations on the mean feature μ and the standard deviation feature δ to obtain the pooled mean feature and the pooled standard deviation feature both with the size of bs × d. This step also includes: using a long short-term memory network layer L to obtain the standard deviation feature δ, and the input of this network layer is with the number of nodes being d and the size of δ being bs × s 2 ×d.
[0067] S7. Obtain the fused feature O based on the feature fusion module, and use the fused feature O as the input of the decoder module; in this step, according to the formula obtain the fused feature O in the feature fusion module, where γ is a randomly generated noise matrix, which conforms to the Gaussian distribution and has the size of bs × d.
[0068] S8. According to the formula Г = Γ R + Γ KL + Γ Homo construct the loss function during the training of the network model, where
[0069]
[0070]
[0071]
[0072] where ∑(·) adds up all the contents inside the brackets. Among them, Γ R uses the Euclidean distance to calculate the similarity between the input of the spectral feature extraction module and the output of the encoder, Γ HL is the loss function in the variational autoencoder VAE, which uses the KL divergence to calculate the similarity between the Gaussian distribution and the embedded feature distribution, Γ Homo uses the spectral angle distance to calculate the similarity between the output of the entire spatial feature extraction module and the output of the corresponding spectral feature extraction module. In this embodiment, the spatial feature extraction module includes: m long short-term memory network layers (Long Short-term Memory Layer), and each network layer is the input of each network layer is respectively the output of each network layer is respectively The number of nodes in the last layer is d / 2, and the number of nodes in other network layers is d. The spectral feature extraction module consists of: m long short-term memory network layers, and each network layer is The inputs of each network layer are respectively The outputs of each network layer are respectively The number of nodes in the last layer is d / 2, and the number of nodes in other network layers is d. The decoder module consists of n fully connected network layers, and each network layer is {d1, d2... d n}, and the inputs of each network layer are {ind1, ind2... ind n}, and the outputs of each network layer are {outd1, outd2... outd n}, and the number of nodes in the last layer is B, and the number of nodes in other network layers is d. The decoder is used to reconstruct the data of the fused feature O, forming a structure similar to an autoencoder, which helps to ensure the consistency of sample information.
[0073] In one embodiment, it further includes: using the above loss function, selecting an Adam optimizer with a step size of 10 -3 to optimize the deep network model. After the model reaches stability, the pooled mean feature is used as the output, and all the first samples and second samples are used as test samples to obtain the expected embedded features.
[0074] Preferably, in step S4, all X×Y hyperspectral pixels need to be divided into a training set and a test set according to a certain ratio, and normalized so that the value range is between -1 and 1. The normalization formula is as follows:
[0075]
[0076] where x min represents the minimum value in the pixel data, and x max is the maximum value. Then the training set pixels are randomly sorted and packed, that is, divided into multiple batches of sample packs. Each sample pack contains bs pixels. Only one sample pack is selected and input into the neural network for each iterative optimization, and the selected sample packs are different each time. The calculation formula of the used Tanh activation function is:
[0077]
[0078] Next, the above method will be verified in combination with specific embodiments.
[0079] The hyperspectral image feature extraction method based on the multi-level variational autoencoder is used to extract the spatial-spectral features of hyperspectral images and for subsequent classification research. Taking the Indiana Forest Dataset as an example, the image size is 145×145×200, with a total of 21,025 pixels. Each pixel contains 200 spectral wavelengths, and the entire dataset contains 16 valid classes and a background noise class. After removing the pixels belonging to the background noise class, a total of 10,366 valid pixels remain. The deep network structure is as Figure 2 shown:
[0080] Input: The input hyperspectral image is an image with a size of 145×145×200.
[0081] Parameter setting: The neighborhood size is 5, the number of network layers in the spatial feature extraction module and the spectral feature extraction module is 3, the number of network layers in the decoder module is 3, and the embedding feature dimension is 40.
[0082] Neighborhood information is selected. For each pixel, neighborhood information of size 5×5×200 is obtained, and the pixel and neighborhood information are input into the deep network for training.
[0083] Train this CNN
[0084] Among the 10,366 training set data, 40% of the samples are selected for training the deep network model. These samples are randomly sorted and packed, and the number of pixels in a small batch is 512. Only one sample pack is used for each training. After training, all 10,366 training set data are input into the deep model for testing, and embedding features of size 10,366×40 are obtained. Finally, an SVM classifier is used for classification. 10% of the samples are randomly selected to train the SVM classifier, and the remaining 90% of the samples are used for testing. The classification results are finally obtained, and the overall classification accuracy and average classification accuracy are selected to evaluate the classification results. The overall classification result refers to the ratio of the number of correctly classified samples to the total number of all samples. The average classification accuracy is first the ratio of the number of correctly classified samples in each class to the number of samples in that class, and then the average value of the ratios of each class is calculated.
[0085] The classification results obtained by using the hyperspectral image feature extraction method based on the multi-level variational autoencoder proposed in this application and the ordinary variational autoencoder (the ordinary variational autoencoder, including an encoder, a feature fusion module, and a decoder, where the encoder consists of 3 fully connected layers, the decoder consists of 3 fully connected layers, and the number of nodes in the network layer and the structure of the feature fusion module are the same as those of the method implemented in this application) are shown in the following table.
[0086] Overall classification accuracy Average classification accuracy Method implemented in this application 85.3% 79.1% Adding random Gaussian noise 81.4% 72.3% Ordinary variational autoencoder 76.7% 66.3%
[0087] As can be seen from the table, the method of the present application can better improve the classification performance of the embedded features and has fewer misclassified samples. In addition, by adding a certain amount of random Gaussian noise to the original hyperspectral image and repeating the above experiments, the overall classification accuracy obtained is 81.4% (the overall classification accuracy reaches 85.3% without adding random Gaussian noise). It can be seen that the method of the present application has strong anti-noise interference ability. Therefore, the method of the present application can effectively improve the classifiable ability and classification accuracy of the embedded features, and can also improve the anti-noise interference ability of the model.
[0088] The above embodiments are only used to illustrate the technical concept and features of the present application, and the purpose is to enable those who are familiar with this technology to understand the content of the present application and implement it accordingly, and cannot be used to limit the protection scope of the present application. Any equivalent transformation or modification made in the spirit of the present application should be covered within the protection scope of the present application.
Claims
1. A hyperspectral image feature extraction method based on a multi-level variational autoencoder, characterized in that The method includes the following steps: S1. Select a hyperspectral image, the size of the hyperspectral image is X×Y×B, where X and Y are the spatial sizes of the hyperspectral image at each wavelength, B is the number of wavelengths of the hyperspectral image, the number of network layers m in the spatial feature extraction module and the spectral feature extraction module, the number of network layers n in the decoder module, and the embedded feature dimension d, where d is an even number greater than 0. S2. Configure neighborhood information for each hyperspectral pixel in the hyperspectral image, that is, select the neighborhood pixels with a size of s×s around the hyperspectral pixel as the neighborhood information of the pixel. The size of the neighborhood information is s×s×B, where the neighborhood information refers to a square area centered on the hyperspectral pixel, with a side length of s, and s is an odd number. S3. Build a deep network model, and transform the neighborhood information based on the deep network model to obtain 1×s 2 ×B’s first sample, which is used as the input of the spatial feature extraction module p , The second sample with a size of 1×B hyperspectral pixels, which is used as the input Input of the spectral feature extraction module q , where the number of the first sample and the second sample is the same and they correspond one by one, S4. Train a deep neural network by randomly selecting mini-batch samples from X×Y first samples of size 1×s 2 ×B and X×Y second samples of size 1×B and inputting them into the deep neural network for training the deep neural network. The number of mini-batch pixels of the first samples and the second samples is both bs, and the activation function in the network is the Tanh activation function. S5. Feature splicing and calculation of the mean feature μ, that is, the input of the layer in the spatial feature extraction module is the output of the layer The input of the layer in the spectral feature extraction module is the splicing of the outputs of the layer and the layer, according to the calculation formula: where 1 < i < m, according to the calculation formula, the mean feature μ is obtained, and the size of μ is bs × s 2 × d, S6. Pooling operation, that is, using the average pooling layer to perform pooling operations on the mean feature μ and the standard deviation feature δ to obtain the pooled mean feature and the pooled standard deviation feature both with the size of bs×d S7. Obtain the fused feature O based on the feature fusion module, and input the fused feature O into the decoder module. The decoder is used to reconstruct the data of the fused feature O. The decoder module includes: n fully-connected network layers, where each network layer is respectively {d1, d2... d n}, and the inputs of each network layer are respectively {ind1, ind2... ind n}, and the outputs of each network layer are respectively {outd1, outd2... outd n}, the number of nodes in the last layer is B, and the number of nodes in other network layers is d, S8. Network optimization, that is, according to the formula Γ = Γ R + Γ KL + Γ Homo Construct the loss function during the training of the network model, where where ∑(·) adds up all the contents within the parentheses.
2. The hyperspectral image feature extraction method based on a multi-level variational autoencoder according to claim 1, wherein In step S1, it further includes: Perform normalization preprocessing on the hyperspectral image and set the neighborhood size s.
3. The hyperspectral image feature extraction method based on a multi-level variational autoencoder according to claim 1, wherein It further includes: Normalize all X×Y hyperspectral pixels so that the value range is between -1 and 1. The normalization formula is as follows: Among them, x min represents the minimum value in the pixel data, and x max is the maximum value.
4. The hyperspectral image feature extraction method based on a multi-level variational autoencoder according to claim 1, characterized in that The calculation formula of the Tanh activation function is as follows:
5. The hyperspectral image feature extraction method based on a multi-level variational autoencoder according to claim 1, characterized in that Step S8 further includes: using a loss function, selecting an Adam optimizer with a step size of 10 -3 to optimize the constructed network model. After the model reaches stability, the pooled mean features are used as the output, and the first sample and the second sample are used as test samples to obtain the expected embedded features.
6. The hyperspectral image feature extraction method based on a multi-level variational autoencoder according to claim 1, characterized in that In step S3, it includes: Transform the hyperspectral pixel x with a size of 1×B to obtain a second sample with a size of The second sample is used as the input of the spectral feature extraction module q , If B is not an integer multiple of s 2 then remove ε wavelengths so that B - ε is an integer multiple of s 2 where q is used to refer to relevant variables in the spectral feature extraction module.
7. The hyperspectral image feature extraction method based on a multi-level variational autoencoder according to claim 1, wherein In step S6, it further includes: The standard deviation feature δ is obtained by using the long short-term memory network layer L, and the input of the long short-term memory network layer L is The number of nodes is d, and the size of δ is bs×s 2 ×d.
8. The hyperspectral image feature extraction method based on a multi-level variational autoencoder according to claim 1, characterized in that, Step S7 further includes: According to the formula the fused feature O is obtained, where γ is a randomly generated noise matrix that follows a Gaussian distribution and has a size of bs×d.
Citation Information
Patent Citations
A densely connected three-dimensional space-spectrum separation convolution depth network and construction method
CN109376753A
Image classification method and related device
WO2021082480A1