Hyperspectral image unmixing method and device based on feedback information guided spectral variable attention
By using a spectral variable attention neural network guided by feedback information, the problems of spectral variability and noise interference in hyperspectral unmixing are solved, achieving high-precision unmixing results.
Patent Information
- Application Number
- CN202410827733.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-06-25
AI Technical Summary
Existing hyperspectral unmixing techniques have limited unmixing accuracy when dealing with spectral variability and noise interference. Traditional methods have low solution efficiency, and deep learning methods lack targeted feature extraction and modeling of spectral variability factors.
A spectrally variable attention neural network guided by feedback information is adopted, including a moving window Transformer encoder, an abundance encoding layer, a feedback enhancement module, a spectrally variable attention module, and a generative decoder. The feedback enhancement strategy suppresses spectrally variable interference and extracts and models abundance, endmembers, and spectrally variable information.
It improves the precision of hyperspectral image unmixing, effectively suppresses the influence of spectral variability on the unmixing effect, and achieves higher unmixing precision and accuracy.
Smart Images

Figure CN118864869B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image processing, and in particular to a hyperspectral image unmixing method and device based on feedback information guided spectral variable attention. BACKGROUND
[0002] Hyperspectral unmixing refers to a technology of separating and identifying different substances or components mixed in the same pixel by analyzing and processing hyperspectral data, i.e., extracting endmembers and corresponding abundances in the data. Hyperspectral images can provide rich spectral information and can finely depict the absorption and reflection characteristics of each substance to the spectrum, making it possible to analyze the material at the sub-pixel level. Hyperspectral unmixing has been widely used in crop growth estimation, water quality assessment, and mineral exploration.
[0003] Due to the influence of factors such as light and terrain, the same material may exhibit different spectral curves, i.e., the phenomenon of "same object different spectrum". In view of such problems, spectral variability unmixing algorithms have been widely studied. Current spectral variability unmixing methods can be mainly divided into traditional methods and deep learning methods. Traditional methods can be divided into methods based on endmember beams, methods based on probability distribution, and algorithms based on improved linear mixing models, but these algorithms are limited by solving efficiency and modeling ability of the model, and it is difficult to achieve high-precision unmixing. Deep learning-based methods usually use deep neural networks to fit the distribution of data or use convolution, Transformer, etc. to fully exploit and utilize the spatial and spectral features of hyperspectral data, and realize the extraction of material information. However, the random distribution of interference factors such as spectral variability and noise in hyperspectral images is complex, and the current algorithm pays more attention to the mining of material characteristics, ignoring the modeling and suppression of interference factors, and the unmixing accuracy is limited. In order to suppress the influence of spectral variability on unmixing, some algorithms start from the perspective of its influence on the spectrum, and model the interference factors based on scaling and perturbation terms. This kind of algorithm improves the unmixing performance to some extent, however, the current algorithm is more based on convex optimization method for solving, which cannot fully utilize the spatial information, and the network based on deep learning lacks targeted feature extraction strategy, and lacks research on modeling of spectral variability factors.
[0004] Feedback mechanism is the most commonly used strategy in automatic control loop, which can feedback the output signal of the system as posteriori information to the input part of the loop, so as to adjust the function of the system to achieve better performance. Attention mechanism is a commonly used technique in the field of artificial intelligence, which is used to simulate the attention allocation process of human in information processing. It assigns different weights or attention to different parts of the input, so that the model can more targetedly process different parts of the input data. By introducing the attention mechanism combined with feedback, the network can adjust the feature extraction strategy based on posteriori information and assign appropriate weights to different features, and be used for subsequent extraction and modeling of spectral variable information to improve the accuracy of unmixing. SUMMARY
[0005] In view of the shortcomings of the existing hyperspectral unmixing technology, the present application provides a hyperspectral image unmixing method and device based on spectral variable attention guided by feedback information.
[0006] To solve the technical problems of the present application, the technical solutions of the present application are as follows,
[0007] The purpose of the present application is to provide a hyperspectral image unmixing method based on spectral variable attention guided by feedback information, which comprises the following steps:
[0008] Step 1): selecting effective spectral information from the original hyperspectral image with high-dimensional waveband number and projecting the image into a feature map with lower spectral dimension;
[0009] Step 2): constructing a neural network based on spectral variable attention guided by feedback information, the neural network comprising a mobile window-based Transformer encoder layer, an abundance encoding layer, a linear reconstruction layer, a feedback enhancement module, a spectral variable attention module and a generative decoder module; using random parameters as the initial weights of the network; wherein the mobile window-based Transformer module is used to mine deep features from the feature map; the abundance encoding layer extracts abundance information from the deep features; the linear reconstruction layer linearly reconstructs the image based on the abundance information; the feedback enhancement module obtains a complementary feature map based on the difference of the deep features of the reconstructed image; the spectral variable attention module mines spectral variable information in the deep features under the guidance of the complementary feature map; the generative decoder module is used to model the spectral variable information and generate scaling items and perturbation items;
[0010] Step 3): training the neural network based on spectral variable attention guided by feedback information using training samples, adjusting the network parameter weights, and obtaining a trained unmixing neural network model;
[0011] Step 4): input the hyperspectral image to be unblended into the trained unblending neural network model, the abundance encoding layer of the model outputs an abundance map as an abundance estimation result, and the weight of the linear reconstruction layer is taken as an endmember estimation result, so that unblending of the hyperspectral image is realized.
[0012] Another object of the present application is to provide a hyperspectral image unblending device based on feedback information guided spectral variable attention, comprising:
[0013] An image acquisition module is configured to acquire a hyperspectral image of a region of interest.
[0014] An image preprocessing module is configured to filter effective spectral information from the original hyperspectral image with a high number of wavebands and project the image into a feature map with a lower spectral dimension.
[0015] A neural network module based on feedback information guided spectral variable attention is configured to mine effective information from the feature map after dimension reduction, so as to model the abundance, endmember information and spectral variable information at the same time.
[0016] An unblending result prediction module is configured to extract an abundance map and an endmember estimation result estimated by the network from the neural network weight based on feedback information guided spectral variable attention after training.
[0017] An unblending result output module is configured to output the unblended abundance map and endmember estimation result.
[0018] The present application has the following advantages:
[0019] 1) The present application maps the difference between the linear reconstruction image and the original image into a picture attention mask through a feedback enhancement strategy, and emphasizes the areas in the image that are seriously affected by spectral variable interference. The spectral variable attention and the generated decoder considering scaling and disturbance are extracted and modeled, which can effectively suppress the influence of spectral variable on unblending and improve the unblending precision.
[0020] 2) The present application proposes a feedback enhancement module, which first generates attention weights based on the error between the linear reconstruction image and the input in the field of hyperspectral unblending, and guides the modeling of spectral variable.
[0021] 3) The present application proposes a spectral variable attention module, which designs a variable sampling strategy based on feedback enhanced features, which can adaptively sample from pixels seriously affected by spectral variable, and realize global attention calculation and mine spectral variable information.
[0022] 4) The present application designs a decoder module considering scaling and disturbance, which first generates a disturbance term based on a probabilistic generative model, and can generate more rich disturbance terms to fit the spectral variable interference in the image. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 Basic step flow chart of hyperspectral unmixing method embodiment of the present application;
[0024] Figure 2 Structure schematic diagram of hyperspectral unmixing device of the present application;
[0025] Figure 3 Spectral variable attention mechanism schematic diagram of the present application
[0026] Figure 4 Schematic diagram of generative decoder of the present application
[0027] Figure 5 Jasper Ridge hyperspectral image dataset for experiment;
[0028] Figure 6 Unmixing abundance result diagram after unmixing of Jasper Ridge hyperspectral image dataset by the embodiment of the present application and different methods. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be described in detail below in combination with specific embodiments and drawings, and specific embodiments are described to simplify the present application. However, it should be recognized that the present application is not limited to the described embodiments, and various modifications of the present application are possible without departing from the basic principles, and these equivalent forms also fall within the scope defined by the appended claims of the present application.
[0030] As shown in Figure 1 , the basic step flow chart of the hyperspectral change detection method of the present application in the present embodiment mainly includes:
[0031] Step 1) Image preprocessing: for the input hyperspectral image with high channel dimension, the present embodiment uses 32 two-dimensional convolutions to project the image into a feature map with channel number , let represent the input hyperspectral image, and the calculation formula is as follows:
[0032]
[0033] , wherein represents the data feature map after dimension reduction, and represent the width and height of the image respectively, represents the channel number of the feature map after dimension reduction.
[0034] Step 2): constructing a neural network for guiding spectrum-variable attention based on feedback information, the neural network comprising a moving window-based Transformer encoder layer, an abundance encoding layer, a linear reconstruction layer, a feedback enhancement module, a spectrum-variable attention module, and a generative decoder module; using random parameters as initial weights of the network; wherein the moving window-based Transformer module is used to mine deep features from the feature map; the abundance encoding layer extracts abundance information from the deep features; the linear reconstruction layer linearly reconstructs the image based on the abundance information; the feedback enhancement module obtains a complementary feature map based on the difference in deep features of the reconstructed image; the spectrum-variable attention module mines spectrum-variable information in the deep features under the guidance of the complementary feature map; and the generative decoder module is used to model the spectrum-variable information and generate a scaling term and a perturbation term. In this embodiment, random parameters are used as initial weights of the network; the specific working processes of the modules in the neural network are as follows:
[0035] Step 21): the structure of the moving window-based Transformer module is as shown in Figure 1 This module divides the data into multiple windows and performs self-attention calculation on each window. The self-attention mechanism projects the feature map into query , key and value , respectively, and calculates the attention weight based on the correlation.
[0036] The attention calculation process can be represented by the following formula:
[0037]
[0038]
[0039] wherein, , , are projection matrices of , and matrices, respectively. represents the self-attention mechanism. is a nonlinear function that can convert the input into a weight with a sum of 1. Based on the above attention calculation, the moving window-based Transformer module can be represented as:
[0040]
[0041]
[0042]
[0043]
[0044] where, and denote the local attention computation based on moving window and the inter-window attention computation, respectively. denotes a multi-layer perceptron consisting of three linear fully connected networks. denotes the deep feature map of the image, is the intermediate feature map of the moving window based Transformer module.
[0045] Step 22): After passing through the moving window based Transformer module, the deep features of the image are processed via the abundance encoding layer, which consists of one linear layer and one nonlinear activation function, and can represent the deep features as low-dimensional abundance representation. The formula can be expressed as:
[0046]
[0047] where and denote the weight and bias of the linear layer, respectively, and the output of the abundance encoding layer is the abundance estimated by the network for the current input image, and P is the number of endmembers in the image.
[0048] Step 23): The linear reconstruction layer consists of one linear layer, which can convert the low-dimensional abundance to the original spectral dimension image through linear mapping. The formula is expressed as:
[0049]
[0050] where is the weight of the linear layer, which is regarded as the endmember matrix extracted by the network in this invention. denotes the linearly reconstructed image.
[0051] Step 24): The feedback enhancement module structure used in this embodiment is shown in Figure 1 First, the module projects the linearly reconstructed image through the two-dimensional convolution of to the same dimension as , and the formula is expressed as:
[0052]
[0053] where is the feature map of the linearly reconstructed image after dimension reduction. Then the module generates an image mask based on the difference between the feature map and The calculation formula of the mask is as follows:
[0054]
[0055] where The mask represents the image, and the value of each pixel represents the degree of influence of the spectral variability in the pixel. is a nonlinear function that can map the value to the range of .
[0056] Based on the generated mask, the feedback enhancement module weights the features of the feature map to obtain complementary features , which can be expressed by the formula:
[0057]
[0058] where represents the complementary feature map, and the endmember-related information in the complementary feature map is suppressed, and the spectral variability information is highlighted. represents the Hadamard product, which means multiplying the corresponding positions of the matrices, and the output is a matrix of the same size.
[0059] Step 25): The complementary features obtained through feedback enhancement contain rich spectral variability information. For this feature, the spectral variability attention module is designed to adaptively sample from the complementary feature map and extract effective features based on the attention mechanism, and its structure is shown in Figure 3 .
[0060] First, according to the size of the deep features of the image, the image is divided into multiple windows according to the size of , and the reference sampling point coordinates in each window are obtained based on uniform sampling . Based on the complementary features , the query matrix is projected, which can be expressed by the formula:
[0061] where is the complementary feature query matrix; is the corresponding projection matrix.
[0062] The spectral variability attention module generates an adjustment amount for the reference sampling point position based on the offset generation network. The offset generation network adopts a lightweight design, which is composed of one layer of depth separable convolution, GELU nonlinear activation function layer, two-dimensional convolution layer and Tanh function, denoted by . In addition, in order to make the model training stable, the spectral variability attention module uses the coefficient to control the maximum value of the offset. The corresponding formula is as follows:
[0063]
[0064] where is a tensor of the offset amount, and the value in it represents the offset amount of each generated reference sampling point coordinate. Based on this, the appropriate sampling point can be determined according to the complementary feature map. Since the sampling coordinates may be decimal numbers, the module determines the pixel value of the sampling point by using bilinear interpolation, and the formula can be expressed as:
[0065]
[0066]
[0067] wherein represents a bilinear interpolation function, and . represents the position of a pixel point in the deep feature map X of the image, represents the position of the sampling point. represents the feature map after sampling, and based on this, the projection into the key and the value two matrices are calculated, and the effective features are mined, and the formula is expressed as follows:
[0068]
[0069]
[0070] wherein is the attention weighted spectral variable feature map; and respectively represent the projection matrices of and matrices.
[0071] Step 26): the spectral variable feature is used to generate the scaling term and the disturbance term in the image caused by the spectral variable effect by using a generative decoder, and the structure of the generative decoder is as shown in Figure 4 . The generative decoder includes a scaling term decoder, a disturbance term encoder and a disturbance term decoder;
[0072] wherein, for the scaling term branch, the scaling term decoder is composed of a linear fully connected layer, a ReLU activation function, a linear fully connected layer and a Tanh function, and is denoted by ; and for the disturbance term branch, the disturbance term encoder composed of a linear fully connected layer, a ReLU activation function and a linear fully connected layer encodes the spectral variable feature map into the mean and variance of the probability distribution, then samples from the standard normal distribution, and then passes through the disturbance term decoder composed of a linear fully connected layer, a ReLU activation function, a linear fully connected layer and a Tanh function to generate the disturbance term in the image caused by the spectral variable effect based on the probability. The whole process can be expressed by the formula:
[0073]
[0074]
[0075]
[0076]
[0077] in This indicates the scaling effect present at each pixel location in the image; and These represent the mean and standard deviation after encoding, respectively. To sample from the standard normal distribution, These are the latent variables after transformation; The perturbation matrix represents the irregular perturbations in the data.
[0078] Step 27): Based on the linearly reconstructed image, the final reconstructed image is obtained by comprehensively considering the linearly reconstructed image, the scaling term, and the perturbation term, as expressed by the following formula:
[0079]
[0080] in This indicates a reconstructed image.
[0081] Step 3: Neural Network Model Training: The entire hyperspectral image is used as the training sample. The network adopts a self-supervised training method. The sample is input into the network model. The loss function is the reconstruction error, the sparsity error of the abundance and perturbation terms, and the KL divergence of the probability distribution. The network weights are updated based on the Adam gradient descent method with adaptive adjustment of the learning rate. In this embodiment, the training is iterated 300 times.
[0082] Step 4: Based on the trained neural network model, the output of the abundance encoding layer is used as the abundance estimated by the network, and the weights of the linear reconstruction layer are extracted as the endmembers extracted by the network to obtain the demixing result.
[0083] Corresponding to the aforementioned embodiment of a hyperspectral image demixing method based on spectral variable attention guided by feedback information, the present invention also provides an embodiment of a hyperspectral image demixing device based on spectral variable attention guided by feedback information.
[0084] Figure 2 This is a block diagram illustrating a hyperspectral image demixing device based on spectral variable attention guided by feedback information, according to an embodiment. Figure 2 As shown, the device includes:
[0085] An image acquisition module is configured to acquire a hyperspectral image of a region of interest.
[0086] An image preprocessing module is configured to filter effective spectral information from the original hyperspectral image with a high number of bands and project the image into a feature map with a lower spectral dimension.
[0087] A neural network module with feedback information guided spectral variable attention is configured to mine effective information from the feature map after dimension reduction, so as to model the abundance, endmember information and spectral variable information at the same time.
[0088] A de-mixing result prediction module is configured to extract an estimated abundance map and an estimated endmember result from the neural network weight after training.
[0089] A de-mixing result output module is configured to output the de-mixed abundance map and the estimated endmember result.
[0090] As to the device in the above embodiment, the specific manner in which the various modules perform operations has been described in detail in the embodiment of the method, and will not be described in detail here.
[0091] For the device embodiment, since it basically corresponds to the method embodiment, the relevant part can be referred to the part of the method embodiment. The device embodiment described above is only illustrative, for example, the image preprocessing module can be a logical function division, and in actual implementation, there can be another division manner, for example, multiple modules can be combined or integrated into another unit. In addition, the connection between the displayed or discussed modules can be a communication connection through some interface, which can be electrical or other forms. According to actual needs, part or all of the modules can be selected to achieve the purpose of the scheme of the present application. Those skilled in the art can understand and implement without creative labor. The following takes a real hyperspectral image as an example to illustrate the specific implementation manner, so as to embody the technical effect of the present application, and the specific steps in the embodiment will not be described.
[0092] Embodiment
[0093] Next, the Jasper Ridge hyperspectral dataset is taken as the research object to verify the de-mixing method of the present application. In order to realize visualization and quantification of the comparison of de-mixing performance from two angles, the estimated abundance map and the evaluation indexes: root mean square error (aRMSE) of abundance and spectral angle distance (eSAD) of endmember are used to evaluate the proposed de-mixing method; the formulas are as follows:
[0094]
[0095]
[0096] wherein, represents the true abundance value of the endmember in the i-th pixel; represents the true abundance value of the endmember in the i-th pixel; represents the true abundance value of the endmember in the i-th pixel; represents the estimated abundance value of the endmember in the i-th pixel; represents the estimated abundance value of the endmember in the i-th pixel; represents the estimated abundance value of the endmember in the i-th pixel; represents the true value of the endmember; represents the true value of the endmember; represents the true value of the endmember.
[0097] The Jasper Ridge dataset contains a total of 1, 000, 000 pixels, each pixel contains 198 wavebands, the wavelength range is 0.38 microns to 2.5 microns, and the spectral resolution is 9.46 nanometers. The dataset contains a total of four endmembers, namely trees, water, soil and roads. are the spectral curves of the hyperspectral image and the corresponding endmember. Figure 5
[0098] Figure 6 are the unmixing abundance maps of the Jasper Ridge dataset obtained by the embodiments of the present application and different unmixing algorithms.
[0099] Table 1 Jasper Ridge hyperspectral image dataset unmixing result evaluation index
[0100]
[0101] The comparative method VCA-FCLS is from: D. C. Heinz, “Fully constrained least squares linear spectral mixture analysis method for material quantification in hyperspectral imagery,” IEEE Transactions on Geoscience and Remote Sensing, vol. 39, no. 3, pp. 529–545, 2001.
[0102] The comparative method PLMM comes from: P.-A. Thouvenin, N. Dobigeon, and J.-Y. Tourneret, “Hyperspectral unmixing with spectral variability using a perturbed linear mixing model,” IEEE Transactions on Signal Processing, vol. 64, no. 2, pp. 525-538, 2016.
[0103] The comparative method ELMM comes from: L. Drumetz, M.-A. Veganzones, S. Henrot, R. Phlypo, J. Chanussot, and C. Jutten, “Blind hyperspectral unmixing using an extended linear mixing model to address spectral variability,” IEEE Transactions on Signal Processing, vol. 25, no. 8, pp. 3890-3905, 2016.
[0104] The comparative method ALMM comes from: D. Hong, N. Yokoya, J. Chanussot, and X. X. Zhu, “An augmented linear mixing model to address spectral variability for hyperspectral unmixing,” IEEE Transactions on Signal Processing, vol. 28, no. 4, pp. 1923-1938, 2019.
[0105] The comparative method PGMSU comes from: S. Shi, M. Zhao, L. Zhang, Y. Altmann, and J. Chen,“Probabilistic generative model for hyperspectral unmixing accounting for endmember variability,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–15, 2022.
[0106] The comparative method DeepTrans comes from: P. Ghosh, S. K. Roy, B. Koirala, B. Rasti, and P. Scheunders,“Deep hyperspectral unmixing using transformer network,” IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1–16, 2022.
[0107] The comparative method UST-Net comes from: Z. Yang, M. Xu, S. Liu, H. Sheng, and J. Wan,“UST-net: A u-shaped transformer network using shifted windows for hyperspectral unmixing,” IEEE Transactions on Geoscience and Remote Sensing, pp. 1–1, 2023.
[0108] The comparative method ULA-Net comes from: S. Xiang, X. Li, J. Ding, S. Chen, and Z. Hua,“Unidirectional local-attention autoencoder network for spectral variability unmixing,” IEEE Transactions on Geoscience and Remote Sensing, vol. 62, pp. 1–15, 2024.
[0109] The unmixing results of the Jasper Ridge dataset are as follows: Figure 6 As shown in Table 1, the quantization results demonstrate that the method presented in this invention achieves the best performance, particularly for endmembers with complex distributions such as trees and soil, where it can more accurately estimate their abundance. Compared to the local attention mechanism of ULA-Net, the global attention mechanism in UST-Net and the proposed method is better at capturing useful information in complex data. Furthermore, the introduction of a feedback mechanism and spectral variability modeling further improves the unmixing accuracy. Moreover, based on the unmixed abundance map and endmember curves, this algorithm more accurately estimates the abundance of trees and soil. It also effectively distinguishes between soil and roads, avoiding confusion between them.
[0110] The above examples are merely specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments and many variations are possible. All variations that can be directly derived or conceived by those skilled in the art from the disclosure of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A method for unmixing hyperspectral images based on feedback information guided spectral variable attention, characterized in that The method comprises the following steps: Step 1): screening effective spectral information from the original hyperspectral image with a high number of spectral bands and projecting the image into a feature image with a lower spectral dimension; Step 2): constructing a neural network based on spectral variable attention guided by feedback information, the neural network comprising a mobile window-based Transformer encoder layer, an abundance encoding layer, a linear reconstruction layer, a feedback enhancement module, a spectral variable attention module and a generative decoder module; using random parameters as the initial weights of the network; wherein the mobile window-based Transformer module is used to mine deep features from the feature image; The abundance encoding layer extracts abundance information from the deep features; the linear reconstruction layer linearly reconstructs the image based on the abundance information; the feedback enhancement module obtains a complementary feature image based on the difference in deep features of the reconstructed image; the spectral variable attention module mines spectral variable information in the deep features under the guidance of the complementary feature image; and the generative decoder module is used to model the spectral variable information and generate a scaling term and a perturbation term; Step 3): training the neural network based on spectral variable attention guided by feedback information using training samples, adjusting the network parameter weights, and obtaining a trained unmixing neural network model; Step 4): inputting a hyperspectral image to be unmixed into the trained unmixing neural network model, outputting an abundance map from the abundance encoding layer of the model as an abundance estimation result, and outputting the weight of the linear reconstruction layer as an endmember estimation result, thereby realizing unmixing of the hyperspectral image.
2. The method of claim 1, wherein the feedback information is obtained by a method comprising: In step 1), effective spectral information is screened and the image is projected into a feature image with a lower dimension, specifically: For the input high-dimensional waveband number of hyperspectral image, use One Two-dimensional convolution projects the image into a feature map, let Denote the input hyperspectral image, the calculation formula is as follows: ; wherein represents the feature map of the data after dimension reduction, and respectively represent the width and height of the image, represents the number of channels of the feature map after dimension reduction.
3. The hyperspectral image demixing method based on spectral variable attention guided by feedback information according to claim 2, characterized in that... In step 2), the mobile window-based Transformer module The data is divided into multiple windows, and self-attention calculation is performed respectively, wherein the self-attention mechanism respectively projects the feature maps into queries , keys and values , and calculates attention weighting based on correlation; The attention calculation process is represented by the following formula: ; ; in, , , They represent , and The projection matrix of a matrix. This indicates the self-attention mechanism. It is a non-linear function that can transform the input into weights that sum to 1; Based on the above attention calculation, the mobile window-based Transformer module is represented by: ; ; ; ; wherein, and denote the local attention computation based on moving window and the inter-window attention computation, respectively; denotes a multi-layer perceptron consisting of three linear fully connected networks; denotes a deep feature map of the image, is an intermediate feature map of the moving window based Transformer module.
4. The hyperspectral image unmixing method based on spectral variable attention guided by feedback information according to claim 1, characterized in that The abundance encoding layer is composed of a linear layer and a nonlinear activation function, and can represent the deep feature image as a low-dimensional abundance representation, which is represented by the following formula: ; where and denote the weight and bias of the linear layer, respectively, and the output of the abundance encoding layer is the estimated abundance for the current input image, and P is the number of endmembers in the image.
5. The hyperspectral image unmixing method based on spectral variable attention guided by feedback information according to claim 4, characterized in that The linear reconstruction layer is composed of a linear layer, which can convert the low-dimensional abundance into an image with the original spectral dimension through linear mapping, which is represented by the following formula: ; wherein is the weight of the linear layer, is considered as the endmember matrix extracted by the network, represents the spectral dimension of the original hyperspectral data; represents the linearly reconstructed image.
6. The hyperspectral image demixing method based on spectral variable attention guided by feedback information according to claim 5, characterized in that, The feedback enhancement module projects the linearly reconstructed image through a two-dimensional convolution to the same dimension as the feature map , which is formulated as: ; wherein is the feature map after dimensionality reduction of the linearly reconstructed image; The module then generates an image mask based on the difference map with the difference map, the mask being calculated according to the following formula: ; wherein a mask representing an image, the value of each pixel representing the degree to which the pixel is affected by the spectral variability; is a non-linear function capable of mapping a numerical value to the range of 0 to 1. Based on the generated mask, the feedback enhancement module weights the features of the feature map to obtain a complementary feature map which is expressed by the formula: ; wherein represents a complementary feature map, endmember-related information in the complementary feature map is suppressed, and spectral variable information is highlighted; represents a Hadamard product, indicates that the corresponding positions of the matrices are multiplied, and an output matrix result of the same scale is obtained.
7. The method of claim 1, wherein, The The spectral variable attention module adaptively samples from the complementary feature image and extracts effective features based on the attention mechanism, and the specific process is as follows: First, according to the size of the image deep feature map X, the image is divided into multiple windows, and the reference sampling point coordinates in each window are obtained based on uniform sampling ; based on complementary features , projection to obtain a query matrix, the formula is as follows: ; wherein, is a complementary feature query matrix; is a corresponding projection matrix; The spectral variable attention module generates an adjustment amount for the reference sampling point position based on an offset generation network; the offset generation network uses to represent; in addition, the spectral variable attention module controls the maximum value of the offset amount by using a coefficient ; the corresponding formula is as follows: ; wherein is a tensor of offsets, wherein the values in it represent the offset of each generated reference sampling point coordinate; based on this, a suitable sampling point is determined according to the complementary feature map; since the sampling coordinates can be decimal numbers, the module determines the pixel value of the sampling point by using bilinear interpolation, and the formula is expressed as: ; ; wherein represents a double-thread interpolation function, and ; represents the position of a pixel point in the image deep feature map X, represents the position of a sampling point; indicates a feature map after sampling, based on which a projection into a key and a value Two matrices are calculated, effective features are mined, and the formula is represented as follows: ; ; wherein is an attention-weighted spectral variable feature map; and respectively represent and a projection matrix of a matrix.
8. The hyperspectral image unmixing method based on spectral variable attention guided by feedback information according to claim 7, characterized in that The generative decoder decodes the scaling term and the perturbation term in the image caused by the spectral variable effect based on the spectral variable features; and the generative decoder comprises a scaling term decoder, a perturbation term encoder and a perturbation term decoder; where the scaling term decoder consists of a linear fully connected layer, a ReLU activation function, a linear fully connected layer, and a Tanh function, denoted as for the scaling term branch, while the perturbation term encoder consists of a linear fully connected layer, a ReLU activation function, and a linear fully connected layer, denoted as for the perturbation term branch. The spectral variable feature map is encoded into the mean and variance of a probability distribution, then sampled from a standard normal distribution, and passed through a perturbation term decoder consisting of a linear fully connected layer, a ReLU activation function, a linear fully connected layer, and a Tanh function to generate the perturbation term in the image due to the spectral variability effect, which is formulated as ; ; ; ; where denotes the zoom effect present at each pixel position in the image; and denote the encoded mean and standard deviation, respectively, is a sample from the standard normal distribution, is the transformed latent variable; is a perturbation matrix, representing irregular perturbations in the data.
9. The method of claim 1, wherein, The step 2) is specifically as follows: On the basis of the linearly reconstructed image, the final reconstructed image is obtained by comprehensively considering the linearly reconstructed image, a scaling term and a perturbation term, and is expressed as follows: ; wherein represents the reconstructed image.
10. The method of claim 1, wherein, The step 3) is specifically: Taking the whole hyperspectral image as a training sample, the network adopts a self-supervised training manner, the training sample is input into the neural network, the loss function is taken as the reconstruction error, the sparsity error of the abundance and the perturbation term and the KL divergence of the probability distribution, and the network weight is updated based on the gradient descent method Adam of adaptive adjustment learning rate.
11. A hyperspectral image unmixing device based on feedback information guided spectral variable attention for implementing the method of claim 1, characterized by, Comprise: An image acquisition module configured to acquire a hyperspectral image of a region of interest; An image preprocessing module configured to filter effective spectral information from the original hyperspectral image with a high number of spectral bands, and project the image into a feature map with a lower spectral dimension; A neural network module based on spectral variable attention guided by feedback information, configured to mine effective information from the feature map after dimension reduction, so as to model the abundance, endmember information and spectral variable information at the same time; An unmixing result prediction module configured to extract the abundance map and the endmember estimation result estimated by the neural network from the neural network weight based on spectral variable attention guided by feedback information after training; An unmixing result output module configured to output the abundance map and the endmember estimation result after unmixing.
Citation Information
Patent Citations
Hyperspectral image sparse unmixing method based on random projection
CN102314685A
Method for hyperspectral imagery exploitation and pixel spectral unmixing
US6665438B1