A Hyperspectral Unmixing Method Based on a Cascade Autoencoder with Spatial Constraints

Through the hyperspectral demix method of a cascaded autoencoder based on spatial constraints, problems such as spatial resolution limiting and spectral variability in hyperspectral image demix are solved, and the accurate identification and classification of landform components are achieved, improving the understanding of mixing effect and accuracy.

CN119808838BActive Publication Date: 2025-06-24NANCHANG INST OF TECH +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411870472.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-06-24
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

Hyperspectral image demixing faces problems such as spatial resolution limitations, complex and diverse landform distribution, low data signal-to-noise ratio, and spectral variability, which affects the accuracy and robustness of landform component recognition and classification.

Method used

The hyperspectral demix method of a cascaded autoencoder based on spatial constraints is adopted to construct a cascaded autoencoder demix model of spatial constraints through layer-by-layer training, and a spatial weighting factor and noise sparse regular terms are introduced to fully explore the spatial information and multi-scale features in hyperspectral images.

Benefits of technology

It realizes accurate identification and classification of land objects, improves the mixing effect and accuracy, effectively resists noise interference, and improves the effect of classification and identification of land objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119808838B_ABST
    Figure CN119808838B_ABST
Patent Text Reader

Abstract

The present invention discloses a hyperspectral unmixing method based on a spatially constrained cascaded autoencoder, comprising the following steps: S1, obtaining a ground object hyperspectral image; S2, constructing a spatially constrained cascaded autoencoder unmixing model; S3, using the constructed spatially constrained cascaded autoencoder unmixing model to unmix the obtained ground object hyperspectral image to obtain the abundance of each ground object; the cascaded autoencoder unmixing model is trained layer by layer; its autoencoder includes at least one encoder, a fusion layer and a decoder; the encoder includes a multi-scale feature extraction module, a BN layer and an activation function layer; the multi-scale feature extraction module includes a small convolutional kernel for focusing on local feature extraction and a large convolutional kernel for capturing global context information. The present invention realizes the accurate identification and classification of ground object components.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hyperspectral image processing in the field of remote sensing, and particularly relates to a hyperspectral unmixing method based on a spatially constrained cascaded autoencoder. Background Art

[0002] Imaging spectroscopy technology organically combines the ground object spectrum determined by the material composition with the spatial distribution image reflecting the existence pattern of the ground object, providing an innovative method. Each pixel on the remote sensing image not only has spectral information characterizing its physical properties but also contains spatial information reflecting its geometric properties. In the process of ground object recognition and classification, hyperspectral imaging technology is particularly important. By acquiring hyperspectral images of surface coverings, researchers can deeply analyze the spectral characteristics of various ground objects. After interpreting these spectral sample data, the composition, chemical composition, and structural characteristics of the ground objects existing in the detection data can be accurately determined, thereby identifying potential ground object resources. Furthermore, by comparing with the spectral database of known ground object resources, rapid recognition and quantitative analysis of potential ground object resources can be achieved. This process not only improves the efficiency of data processing but also provides important experimental data support for the development and utilization of ground object resources. Such data support is of great significance for fields such as environmental monitoring, land management, and resource assessment. Therefore, hyperspectral technology shows broad application prospects in ground object analysis. With the continuous progress of technology and the improvement of data processing capabilities, it is expected that hyperspectral imaging technology will play a more crucial role in the exploration and development of ground object resources in the future, contributing to sustainable development and ecological protection.

[0003] The difficulties in ground object recognition are mainly reflected in the following aspects. First, the spectral characteristics of different ground objects may be similar, resulting in confusion and misrecognition. Second, the spectral information of ground objects is affected by environmental factors such as climate, season, and soil conditions, which challenges the stability and reliability of spectral data. In addition, the large amount and complexity of data require efficient algorithms and models for processing and analysis, which pose higher requirements for computing power and technical level. Finally, the lack of a comprehensive spectral database also limits the accurate recognition and analysis of potential resources.

[0004] Spectral variability is caused by the interaction of various factors such as the surface material of the ground object, lighting conditions, and atmospheric effects, resulting in differences in the spectral responses of the same ground object under different conditions. This characteristic makes it difficult for the linear mixing model to accurately capture the true spectral characteristics during the unmixing process, thus affecting the accuracy and robustness of the results.

[0005] Therefore, it is necessary to design a new hyperspectral unmixing method. Summary of the Invention

[0006] The object of the present invention is to provide a hyperspectral unmixing method based on a spatially constrained cascaded autoencoder to solve problems such as spatial resolution limitation, complex and diverse ground object distributions, low data signal-to-noise ratio, and spectral variability faced in hyperspectral image unmixing.

[0007] To achieve the above object, the present invention provides a hyperspectral unmixing method based on a spatially constrained cascaded autoencoder, including the following steps: S1. Obtain a ground object hyperspectral image;

[0008] S2. Construct a spatially constrained cascaded autoencoder unmixing model;

[0009] S3. Use the constructed spatially constrained cascaded autoencoder unmixing model to unmix the obtained ground object hyperspectral image to obtain the ground object abundance corresponding to each ground object;

[0010] The spatially constrained cascaded autoencoder unmixing model is trained layer by layer, that is, first train the first-layer autoencoder, and then use the output of the trained first-layer autoencoder as the input of the second-layer autoencoder, and so on for iteration until the termination condition of the iteration is reached;

[0011] The autoencoder specifically includes a preprocessing stage module and a spectral unmixing stage module; the preprocessing stage module includes at least one encoder, and the spectral unmixing stage module includes a fusion layer and a decoder;

[0012] The encoder includes a multi-scale feature extraction module, a BN layer, and an activation function layer. The multi-scale feature extraction module is a Conv layer, and the activation function layer uses a Leaky ReLU activation function; the multi-scale feature extraction module includes a small convolutional kernel for focusing on local feature extraction and a large convolutional kernel for capturing global context information.

[0013] In a specific implementation manner, a spatial constraint term is introduced into the autoencoder. The spatial constraint term is used to take into account the information of adjacent pixels to optimize the parameters of the autoencoder.

[0014] In a specific implementation manner, when the spatially constrained cascaded autoencoder unmixing model uses a deep learning-based spatial regularization unmixing algorithm to process the hyperspectral image of a ground object:

[0015] First, quantify the proximity between adjacent pixels; use the Euclidean distance as an index to measure the similarity of adjacent pixels, and calculate the Euclidean distance between pixels to describe the proximity between pixels;

[0016] Subsequently, based on the Euclidean distance information between pixels, further define the neighborhood weight , and the neighborhood weight is expressed as:

[0017]

[0018] Among them, im (·) represents a pixel x i and x j The Euclidean distance between them. When the positions of pixel x i and pixel x j are closer, the Euclidean distance is smaller, and the weight is larger;

[0019] Using (a, b) and (c, d) as the spatial coordinates of pixel x i and pixel x j respectively, the Euclidean distance is expressed as ;

[0020] Utilize neighborhood weights to mine spatial correlation, the method is as follows:

[0021]

[0022] In the formula, represents the neighborhood set of abundance elements;

[0023] Construct a spatial weighting factor:

[0024] Introduce the constructed spatial weight into the model:

[0025] .

[0026] In a specific implementation manner, the cascaded autoencoder unmixing model with spatial constraints sets a loss function, which is used to evaluate the accuracy of spectral unmixing and the effective utilization degree of spatial information; set an optimization algorithm to minimize the loss function, and iteratively update the model parameters of the cascaded autoencoder unmixing model with spatial constraints until convergence to the optimal solution, so as to ensure that the cascaded autoencoder unmixing model with spatial constraints can best fuse spectral and spatial information.

[0027] In a specific implementation manner, the construction of the loss function is:

[0028] Among them, represents the average value of the sum of squares of the differences between the true pixel vector Y and the predicted pixel vectors and respectively. Y represents the true pixel vector, and both represent the predicted pixel vectors, and , , where N represents the total number of pixel points.

[0029] In a specific embodiment, the autoencoder further introduces a structured sparsity loss function:

[0030] .

[0031] In a specific embodiment, the fusion layer is a fusion layer.

[0032] In a specific embodiment, the training process of the spatially constrained cascaded autoencoder unmixing model uses six encoders working together. When the encoders perform multi-scale feature extraction, two encoders are selected from three encoders with convolutional kernels of different scale sizes. One of the encoders uses a small 1x1 convolutional kernel to focus on local feature extraction, and the other encoder uses a large 5x5 or 9x9 convolutional kernel to capture global context information.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] The present invention solves the problems of spatial resolution limitation, complex and diverse ground object distributions, low data signal-to-noise ratio, and spectral variability faced by hyperspectral image unmixing, thereby achieving accurate identification and classification of ground object components.

[0035] The present invention innovatively adds a spatial weighting factor in the framework of deep learning. By constructing a spatially constrained cascaded autoencoder that fuses multiple encoders, a new multi-scale feature extraction mechanism is constructed to utilize the spatial weighting factor, fully mining the spatial information in the hyperspectral image. At the same time, a noise sparse regularization term is introduced to resist noise interference to the greatest extent, effectively improving the effect of ground object classification and recognition. Experimental results on simulated hyperspectral data and real hyperspectral data sets show that the present invention has achieved good recognition and classification effects.

[0036] In addition to the objectives, features, and advantages described above, the present invention has other objectives, features, and advantages. The present invention will be further described in detail below. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0038] Figure 1 is a schematic flowchart of an embodiment of the present invention;

[0039] Figure 2 is the true abundance map of each endmember in an embodiment of the simulated data of the present invention;

[0040] Figure 3 is the unmixed abundance map of endmember 2 in an embodiment of the simulated data of the present invention;

[0041] Figure 4 is the color image of the endmember simulated by the model in an embodiment of the present invention. Detailed implementation manners

[0042] The following describes the embodiments of the present invention in detail. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0043] The present invention provides a hyperspectral unmixing method based on a spatially constrained cascaded autoencoder, including the following steps: S1. Obtain a hyperspectral image of a ground object;

[0044] S2. Construct a spatially constrained cascaded autoencoder unmixing model;

[0045] S3. Use the constructed spatially constrained cascaded autoencoder unmixing model to unmix the obtained hyperspectral image of the ground object to obtain the ground object abundance corresponding to each ground object;

[0046] The spatially constrained cascaded autoencoder unmixing model adopts a layer-by-layer training method, that is, first train the first-layer autoencoder, and then use the output of the trained first-layer autoencoder as the input of the second-layer autoencoder, and so on for iteration until the termination condition of the iteration is reached;

[0047] The autoencoder specifically includes a preprocessing stage module and a spectral unmixing stage module; the preprocessing stage module includes at least one encoder, and the spectral unmixing stage module includes a fusion layer and a decoder;

[0048] The encoder includes a multi-scale feature extraction module, a BN layer, and an activation function layer. The multi-scale feature extraction module is a Conv layer, and the activation function layer uses a Leaky ReLU activation function; the multi-scale feature extraction module includes a small convolutional kernel for focusing on local feature extraction and a large convolutional kernel for capturing global context information.

[0049] The Conv layer performs a linear transformation to lay the foundation for data flow. The conv layer adopts a layer-by-layer training method.

[0050] The BN layer improves the model stability and accelerates convergence. The fusion layer can integrate features from different encoders or different modalities to form a more comprehensive feature representation. Using the Leaky ReLU activation function solves the gradient vanishing problem of ReLU and further enhances the model's non-linear processing ability. The decoder reconstructs the high-dimensional output from the fused low-dimensional features and compares it with the original pixel data to calculate the reconstruction error, ensuring the stability and accuracy of the output.

[0051] The present invention utilizes a spatially constrained cascaded autoencoder unmixing model to accurately estimate each ground object component contained in a mixed pixel; according to the solved ground object abundances, the proportion corresponding to each ground object is quantified to achieve accurate identification and classification.

[0052] A spatial constraint term is introduced into the autoencoder, and the spatial constraint term is used to take into account the information of adjacent pixels to optimize the parameters of the autoencoder.

[0053] When the spatially constrained cascaded autoencoder unmixing model processes the hyperspectral image of a ground object using a spatially regularized unmixing algorithm based on deep learning:

[0054] To deeply exploit the rich spatial information contained in the image, the proximity degree between adjacent pixels is first quantified; the Euclidean distance is used as an index to measure the similarity between adjacent pixels, and the Euclidean distance between pixels is calculated to describe the proximity degree between pixels;

[0055] Subsequently, based on the Euclidean distance information between pixels, the neighborhood weight is further defined so as to make full use of these spatial correlations in subsequent processing, thereby improving the unmixing effect and accuracy. The neighborhood weight is expressed as:

[0056]

[0057] where im (·) represents the Euclidean distance between pixel x i and x j When the positions of pixel x i and pixel x j are closer, the Euclidean distance is smaller, and the weight is larger;

[0058] Using (a, b) and (c, d) as the spatial coordinates of pixel x i and pixel x j , the Euclidean distance is expressed as ;

[0059] Spatial correlation is mined using neighborhood weights as follows:

[0060]

[0061] In the formula, represents the neighborhood of the abundance element set;

[0062] Construct a spatial weighting factor:

[0063] Introduce the constructed spatial weight into the model:

[0064] .

[0065] In the process of mining the spatial information contained in the image, two encoders in a dual-path parallel manner perform multi-scale feature extraction to mine the spatial information in the image, and the outputs of the two encoders are fused through the fusion layer.

[0066] The cascaded autoencoder unmixing model with spatial constraints sets a loss function, which is used to evaluate the accuracy of spectral unmixing and the effective utilization degree of spatial information; an optimization algorithm is set to minimize the loss function, and the model parameters of the cascaded autoencoder unmixing model with spatial constraints are iteratively updated until convergence to the optimal solution, so as to ensure that the cascaded autoencoder unmixing model with spatial constraints can best fuse spectral and spatial information.

[0067] The loss function is used to evaluate the model performance and provides a comprehensive quantitative standard for the model performance.

[0068] Preferably, the optimization algorithm is the gradient descent method.

[0069] The construction of the loss function is as follows:

[0070] Wherein, represents the average value of the sum of squares of the differences between the true pixel vector Y and the predicted pixel vectors and respectively. Y represents the true pixel vector, and both represent the predicted pixel vectors, and , , and N represents the total number of pixel points.

[0071] In order to improve the unmixing accuracy, reduce the computational complexity or optimize the model performance, which can promote sparsity, prevent overfitting and utilize the data group structure, and help improve the performance and stability of the model, the following constraint terms are also introduced, which can also be called the structured sparsity loss function:

[0072] .

[0073] In the formula, represents the norm, The norm plays a role in promoting sparsity, preventing overfitting, and utilizing the data group structure in hyperspectral image processing, which helps to improve the performance and stability of the model.

[0074] The added structured sparsity loss function also sparsifies the reconstruction result. The structured sparsity takes into account the context information of the image and the correlation between pixels. By introducing the structured sparsity loss, the sparsity and detail retention ability of the image can be further improved.

[0075] In addition, the SAD constraint is introduced. Introducing the SAD constraint can improve the accuracy and reliability of spectral data correction, and at the same time ensure that the corrected spectral vector is consistent with the true spectrum in shape and direction. The SAD constraint is defined as:

[0076] .

[0077] Based on the above introduction, the overall constraint can be written as:

[0078]

[0079] The fusion layer is the fusion layer.

[0080] In the training process of the cascade autoencoder unmixing model with spatial constraints, six encoders work together. When the encoders perform multi-scale feature extraction, two encoders are selected from three encoders with different scale sizes of convolutional kernels. One of the encoders uses a small 1x1 convolutional kernel to focus on local feature extraction, and the other encoder uses a large 5x5 or 9x9 convolutional kernel to capture global context information.

[0081] The multi-scale feature extraction mechanism is introduced to capture local and global features through different convolutional kernels, such as 1x1, 5x5, 9x9, to enhance the model's understanding of complex input data.

[0082] The decoder reconstructs the high-dimensional output data. At the output layer of the decoder, according to the specific task requirements, the model can flexibly select the Sigmoid or Softmax function for probability output or normalization processing to meet the needs of different classification scenarios.

[0083] Based on deep learning, hyperspectral image unmixing is completed to obtain the abundances of various ground objects. Identification is carried out according to the obtained abundances of ground objects. As is well known, the abundances of ground objects are represented by a matrix composed of a set of numbers in the model. The method will automatically match the final unmixed abundances to complete the identification and classification of abundances.

[0084] Example 1

[0085] A hyperspectral unmixing method based on a spatially constrained cascaded autoencoder, the specific steps are as follows:

[0086] Put in the hyperspectral image that needs to identify ground objects. The algorithm starts a multi-scale feature extraction mechanism using different convolutional kernels. This mechanism works collaboratively through three independent encoders. The entire spatially constrained cascaded autoencoder unmixing model is set up with six encoders working together to obtain the trained spatial features, achieving a comprehensive capture of the input data features. One of the encoders uses a small convolutional kernel (1x1) and focuses on extracting fine local features. One of the other two encoders uses a large convolutional kernel (such as 9x9 or 5x5) to capture more extensive and global spatial context information. Then, the outputs of the two encoders are fused through a fusion layer to construct a spatial weighting factor:

[0087] W spa represents the t +1-th iteration of the i row and the j column elements.

[0088] Among them, represents spatial correlation, and the specific expression is:

[0089]

[0090] Among them represents the neighborhood set of the element and

[0091] represents a function for mining spatial correlation through the neighborhood system. Further define the neighborhood weight so as to make full use of these spatial correlations in subsequent processing, thereby improving the unmixing effect and accuracy. The neighborhood weight

[0092] Among them, im (·) represents the Euclidean distance between the pixel x i and x j When the pixel x iand pixels x j The closer the position of j is to the pixel, the smaller the Euclidean distance, and thus the weight is larger;

[0093] Using (a, b) and (c, d) as the spatial coordinates of the pixel i x i and the pixel x j respectively, the Euclidean distance is expressed as ;

[0094] The model integrates 6 encoders, 1 fusion layer, and 1 decoder to construct an efficient spectral data processing system. Among them, the Conv layer, as a fully connected (linear) layer, is responsible for the linear transformation of data, laying the foundation for data flow. The BN layer (batch normalization layer) effectively improves the stability and convergence speed of the model through its unique batch normalization operation, ensuring a smoother training process.

[0095] Using LReLU (Leaky Rectified Linear Unit) as the activation function cleverly solves the problem of gradient disappearance that ReLU may encounter in the negative value region, further enhancing the model's non-linear processing ability.

[0096] To accurately evaluate and quantify the difference between the network output and the target output, the mean squared error is used as the core measurement index, and the REC constraint is as follows:

[0097]

[0098] At the same time, norm constraint is added, and the constraint is as follows:

[0099]

[0100] In addition, introducing the SAD constraint can improve the accuracy and reliability of spectral data correction, while ensuring that the corrected spectral vector is consistent with the true spectrum in shape and direction. The SAD constraint is defined as:

[0101]

[0102] Based on the above introduction, the overall constraint can be written as:

[0103]

[0104] The decoder reconstructs the high-dimensional output data. At the output layer of the decoder, according to the specific task requirements, the model can flexibly select the Sigmoid or Softmax function for probability output or normalization processing to meet the needs of different classification scenarios.

[0105] Based on deep learning, hyperspectral image unmixing is completed to obtain the abundances of various ground objects.

[0106] Example 2

[0107] In this example, a set of simulated data is constructed, which consists of 5 endmembers. To comprehensively cover the frequency range of spectral data, 162 bands are set in this example. Each band is regarded as a unique data dimension to capture the spectral responses of substances at different frequencies. Particularly crucial is that within each band, a 162×162 pixel matrix is constructed. This matrix structure not only reflects the fine distribution of spectral data in the spatial dimension but also ensures the integrity and analyzability of the dataset. To simulate the complexity and uncertainty that spectral data may encounter in real scenarios, 30 dB of noise is specifically introduced into this set of data to enhance the impact of spectral variability.

[0108] In this dataset, the learning rate of the encoder of the model is set to 4.5×(10)^(-3), and the learning rate of the decoder is 1×(10)^(-3); the number of iterations is 300; for the loss function, its parameters are beta1, beta2, beta3, beta4, beta5 respectively, where beta1 is set to 8×(10)^(-7); beta2 is set to 4.4×(10)^(-6); beta3 is set to 9.5×(10)^(-2); beta4 is set to 3×(10)^(-1); beta5 is set to 5×(10)^(-2).

[0109] The Fully Constrained Least Squares (FCLS), the Sparse Unmixing via Variable Splitting and Augmented Lagrangian (SUnSAL), the Augmented Linear Mixed Model (ALMM), and the Convolutional Neural Network for Automatic Endmember Extraction and Unmixing (CNNAEU) are respectively used for comparative experiments to verify the effectiveness of the model.

[0110] SAD, aSAD, and aRMSE are used to quantitatively measure the quality of the measurement results, and their definitions are as follows:

[0111]

[0112] SAD is the abbreviation of Sum of Absolute Differences, which is used to measure the difference in pixel values between two images. It is obtained by calculating the sum of the absolute differences of each corresponding pixel between two image patches of the same size. Mathematically, for two image patches A and B with size h×w.

[0113] Where A(i,j) and B(i,j) represent the pixel values of image patches A and B at position (i,j) respectively.

[0114]

[0115] aSAD is a variant of SAD, which is obtained by dividing the SAD value by the total number of pixels in the image patch, so as to obtain an average difference value.

[0116] aRMSE is the abbreviation of "Average Root Mean Square Error", which measures the global error between two images. First, calculate the sum of the squares of the differences of each corresponding pixel between two image patches, and then take the square root of the average value. Mathematically, for two image patches A and B with size h×w, RMSE is defined as:

[0117]

[0118] The smaller the SAD, the better the obtained result. aSAD and aRMSE can reflect the overall performance of the algorithm.

[0119] Table 1. Comparison results of different algorithms (×10 -2 )

[0120]

[0121] In order to more intuitively display the unmixing effect, the abundances of the endmembers estimated by each algorithm and the true abundance map are shown. Compared with other comparison methods, the abundance image obtained in this embodiment better fits the real situation, which proves the effectiveness of the abundance estimation in this embodiment.

[0122] Example 3

[0123] This embodiment uses the Samson real dataset to further verify the beneficial effects of the present invention.

[0124] The Samson dataset occupies a prominent position in the field of hyperspectral image unmixing and is one of the widely used benchmark datasets. The original image of this dataset consists of 952×952 pixel points, covers a wide wavelength range from 401 nm to 889 nm, and contains a total of 156 bands, with a fine spectral resolution of 3.13 nm between each band.

[0125] The sub-image region used in this embodiment is a 95×95 pixel area cropped from this complete image.

[0126] In this dataset, this embodiment focuses on three basic and representative substances, which are often regarded as pure endmembers in remote sensing and spectral analysis. They are trees, soil, and water respectively.

[0127] The selection of these three pure endmembers is not only based on their wide distribution on the earth's surface, but also because of their significant differences in spectral characteristics, which provides natural discrimination for the unmixing and analysis of spectral data.

[0128] In this dataset, the learning rate of the encoder of the model is set to 9.5×(10)^(-2), and the learning rate of the decoder is 4.5×(10)^(-2); the number of iterations is 300 times; for the loss function, beta1 is set to 7×(10)^(-7); beta2 is set to 1×(10)^(-5); beta3 is set to 1×(10)^(-1); beta4 is set to 2×(10)^(-1); beta5 is set to 1×(10)^(-5); the spatial weight factor incorporates the information within the 7x7 neighborhood around each pixel during evaluation and imposes corresponding constraint conditions accordingly.

[0129] Table 2. Comparison results of different algorithms on the Samson dataset (×10 -2 )

[0130]

[0131] Through experimental verification, the present invention has shown excellent advantages in two key evaluation indicators, aSAD (average spectral angle distance) and aRMSE (average root mean square error).

[0132] This achievement not only demonstrates the superiority of the algorithm itself, but also strongly proves the key role played by the spatial weighting factor in improving the unmixing accuracy.

[0133] As can be seen from Table 2, in the extraction of the three endmembers, the present invention has achieved the optimal value among all methods. Especially in the extraction of the soil endmember, the value is far less than the sub-optimal value.

[0134] aSAD and aRMSE also reach the best state among these methods, and there is an obvious gap in numerical values from the sub-optimal methods.

[0135] The predicted pixel data reconstructed and generated is analyzed by this method to complete identification and classification.

[0136] In summary, it can be proved that the present invention is accurate in unmixing, and it can also show that the spatial weighting factor has an obvious effect on endmember extraction, greatly improving the accuracy.

[0137] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions and substitutions can still be made, which should all be regarded as belonging to the protection scope of the present invention.

Claims

1. A hyperspectral unmixing method based on a cascaded autoencoder with spatial constraints, characterized in that: The following steps are involved: S1. Obtaining hyperspectral images of ground objects; S2, construct a spatially constrained cascaded autoencoder unmixing model; S3, using the constructed spatially constrained cascade autoencoder unmixing model to unmix the acquired hyperspectral image of the ground objects, and obtain the ground object abundance corresponding to each ground object; The spatially constrained cascaded autoencoder unmixing model adopts a layer-by-layer training method, that is, the first layer of autoencoder is trained first, and then the output of the trained first layer of autoencoder is used as the input of the second layer of autoencoder, and so on, iteratively performed until the termination condition of the iteration is reached; The autoencoder specifically includes a preprocessing stage module and a spectral unmixing stage module; The preprocessing stage module includes at least one encoder, and the spectral unmixing stage module includes a fusion layer and a decoder; The encoder includes a multi-scale feature extraction module, a BN layer and an activation function layer. The multi-scale feature extraction module is a Conv layer, and the activation function layer adopts a Leaky ReLU activation function. The multi-scale feature extraction module includes a small convolution kernel for focusing on local feature extraction and a large convolution kernel for capturing global context information. A spatial constraint term is introduced into the autoencoder, and the spatial constraint term is used to take the information of adjacent pixels into consideration, and the spatial constraint term is used as a part of the loss function to optimize the parameters of the autoencoder; The training process of the spatially constrained cascaded autoencoder unmixing model uses six encoders to work together. When the encoders perform multi-scale feature extraction, two encoders are selected from three encoders with convolution kernels of different scales. One of the encoders uses a small 1x1 convolution kernel to focus on local feature extraction, and the other encoder uses a large 5x5 or 9x9 convolution kernel to capture global context information.

2. The hyperspectral unmixing method based on the spatially constrained cascaded autoencoder according to claim 1, characterized in that: When the spatially constrained cascaded autoencoder unmixing model uses a deep learning-based spatial regularization unmixing algorithm to process hyperspectral images of ground objects: First, quantify the similarity between adjacent pixels; use Euclidean distance as an indicator to measure the similarity of adjacent pixels, and describe the similarity between pixels by calculating the Euclidean distance between pixels; Then, based on the Euclidean distance information between pixels, the neighborhood weight is further defined , neighborhood weight It is expressed as: in, im (·) indicates pixel x i and x j The Euclidean distance between pixels x i and pixel x j The closer the position, the smaller the Euclidean distance, and the weight The bigger; (a, b) and (c, d) are pixels respectively. x i and pixel x j The spatial coordinates of , then the Euclidean distance is expressed as ; Use neighborhood weights to mine spatial correlations as follows: In the formula, Represents the neighborhood of abundance elements gather; Construct spatial weighting factors: Introduce the constructed spatial weights into the model: 。 3. The hyperspectral unmixing method based on the spatially constrained cascaded autoencoder according to claim 1, characterized in that: The spatially constrained cascaded autoencoder unmixing model sets a loss function, which is used to evaluate the accuracy of spectral unmixing and the effective utilization of spatial information; an optimization algorithm is set to minimize the loss function, and the model parameters of the spatially constrained cascaded autoencoder unmixing model are iteratively updated until convergence to an optimal solution, thereby ensuring that the spatially constrained cascaded autoencoder unmixing model can optimally fuse spectral and spatial information.

4. The hyperspectral unmixing method based on the spatially constrained cascaded autoencoder according to claim 3 is characterized in that: The loss function is constructed as: in, Represents the true pixel vector Y and the predicted pixel vector and The average value of the sum of squares of the differences between , Y represents the true pixel vector, and denotes the predicted pixel vector, and , , N represents the total number of pixels.

5. The hyperspectral unmixing method based on the spatially constrained cascaded autoencoder according to claim 1, characterized in that: The autoencoder also introduces a structured sparsity loss function: 。 6. The hyperspectral unmixing method based on spatially constrained cascaded autoencoders according to claim 1, characterized in that: The fusion layer is a fusion layer.

Citation Information

Patent Citations

  • Mineral identification and classification method based on hyperspectral unmixing technology

    CN118072186A

  • Bridge cable apparent defect segmentation method based on deep learning

    CN118351316A