Hyperspectral Image Unmixing Method and System Based on Endmember-Oriented Transformer and Endmember Bundle

By using the neural network method of end-element-oriented Transformer and end-element beam in hyperspectral image demix, the problem of unstable performance of hyperspectral image demix algorithm in the prior art when dealing with spectral variability is solved, and more accurate feature extraction and end-element generation are achieved to improve the understanding of mixing performance.

CN119516376BActive Publication Date: 2025-06-17ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411586852.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-08
Publication Date
2025-06-17
Estimated Expiration
2044-11-08

AI Technical Summary

Technical Problem

The existing hyperspectral image demixing algorithms have unstable performance when dealing with spectral variability, which may generate end elements with unclear physical significance, resulting in inaccurate abundance estimation.

Method used

Using a neural network method based on end-element-oriented Transformer and end-element beam, abundance estimation and end-element extraction of hyperspectral images are realized through end-element-oriented Transformer encoder module, end-element generator module based on end-element beam, heterogeneous information fusion module and abundance decoder module.

Benefits of technology

This method can provide more accurate feature extraction results, improve network fitting efficiency, generate more stable and physically meaningful end element results, and improve the demixing performance of hyperspectral images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516376B_ABST
    Figure CN119516376B_ABST
Patent Text Reader

Abstract

The present invention discloses a hyperspectral image unmixing method and system based on endmember-guided Transformer and endmember bundle, which relates to the field of hyperspectral unmixing. The method includes: image preprocessing; constructing a neural network based on endmember-guided Transformer and endmember bundle; training the neural network to obtain the trained neural network; and using the trained neural network to obtain the unmixing result of the hyperspectral image to be unmixed. Based on the endmember-guided Transformer, the present invention focuses on mining the features of each endmember category and weakening the influence of spectral variability on the unmixing performance, and specifically designs an endmember-guided Transformer encoder module, a heterogeneous information fusion module, and an endmember generator module based on endmember bundle, which can better suppress the influence of spectral variability on unmixing and improve the accuracy of endmember extraction and abundance estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and in particular relates to a hyperspectral image unmixing method and system based on an endmember-guided Transformer and an endmember bundle. Background Art

[0002] Hyperspectral unmixing refers to using the advantages of hyperspectral images in the spectral dimension to analyze the internal composition of the pixels of the image with limited spatial resolution, solve the basic components and corresponding proportions of the hyperspectral image, and obtain the end members and abundance of the image. With the advantage of its spectral resolution, hyperspectral images can depict spectral curves with richer details, providing an effective way to achieve material analysis of imaging targets.

[0003] Due to the differences in lighting, environment and other factors in real imaging scenes, or changes in the imaging target, the phenomenon of "same object, different spectrum" is often presented in hyperspectral images, which brings challenges to hyperspectral unmixing. In order to reduce the impact of the above phenomenon on the performance of hyperspectral unmixing, unmixing algorithms that consider spectral variability have been widely studied. At present, unmixing algorithms that consider spectral variability can be generally divided into algorithms based on endmember bundles, algorithms based on mixture models and algorithms based on deep learning. Traditional unmixing algorithms are usually based on fixed prior assumptions and cannot provide a suitable general solution for many complex data, so their unmixing performance is limited. Algorithms based on deep learning use modules such as fully connected layers and convolutional layers to mine information about spectral variability and fit its impact on the spectrum, and based on this, model the variable information to weaken its impact on abundance estimation. When considering spectral variability, current deep learning-based algorithms often directly output the fitted spectral variable terms based on generative models, and urge network training based on various regularization terms. However, the performance of existing algorithms is unstable, and endmembers with unclear physical meanings may be generated, resulting in inaccurate abundance estimates.

[0004] The endmember beam algorithm can maintain the physical meaning of solving the endmember by extracting multiple spectra to represent the same endmember. However, its optimization speed based on traditional heuristic algorithms is slow and inefficient. Currently, there are few deep learning-based algorithms that can bring out the advantages of the endmember beam algorithm. Summary of the invention

[0005] In order to solve the problems in the prior art, the present invention proposes a hyperspectral image unmixing method and system based on endmember-guided Transformer and endmember bundle.

[0006] The technical solution adopted by the present invention is as follows:

[0007] In a first aspect, the present invention discloses a hyperspectral image unmixing method based on an endmember-guided Transformer and an endmember bundle, comprising the following steps:

[0008] Step 1): Obtain the initial endmembers of the hyperspectral image, and obtain the data feature map and endmember tokens;

[0009] Step 2): Construct a neural network based on the endmember-guided Transformer and endmember bundle, which is used to obtain the abundance estimation result and endmember extraction result of the hyperspectral image based on the data feature map and endmember tokens. The neural network includes an endmember-guided Transformer encoder module, an endmember generator module based on the endmember bundle, a heterogeneous information fusion module, and an abundance decoder module;

[0010] Step 3): Train the neural network based on the endmember-guided Transformer and endmember bundle to obtain a trained neural network;

[0011] Step 4): Use the trained neural network to obtain the abundance estimation result and endmember extraction result of the hyperspectral image to be unmixed, and realize the unmixing of the hyperspectral image.

[0012] In a second aspect, the present invention discloses a hyperspectral image unmixing system based on the endmember-guided Transformer and endmember bundle for the method, including:

[0013] An image preprocessing module, which is used to obtain the initial endmembers of the hyperspectral image, and obtain the data feature map and endmember tokens;

[0014] A neural network construction module, which is used to construct a neural network based on the endmember-guided Transformer and endmember bundle, which is used to obtain the abundance estimation result and endmember extraction result of the hyperspectral image based on the data feature map and endmember tokens. The neural network includes an endmember-guided Transformer encoder module, an endmember generator module based on the endmember bundle, a heterogeneous information fusion module, and an abundance decoder module;

[0015] A neural network training module, which is used to train the neural network based on the endmember-guided Transformer and endmember bundle to obtain a trained neural network;

[0016] A hyperspectral image unmixing module, which is used to use the trained neural network to obtain the abundance estimation result and endmember extraction result of the hyperspectral image to be unmixed, and realize the unmixing of the hyperspectral image.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0018] 1) The present invention proposes an end-member guided Transformer encoder module, which realizes feature extraction based on directional projection of data feature maps, and for the first time proposes a low-redundancy attention mechanism based on a mixture of experts system, which can effectively constrain the solution space of the model, provide more accurate feature extraction results, and improve the network fitting efficiency.

[0019] 2) The present invention proposes an end-member generator module based on end-member bundles, and for the first time adopts the method of converting end-member guided features into within-class weights and combining end-member bundles to generate the estimation results of end-members.

[0020] 3) The present invention proposes a heterogeneous information fusion module, which can introduce the prior of complementary spatial distribution of end-members during information fusion to achieve full fusion of different end-member guided features. Description of the Drawings

[0021] Figure 1 It is a basic step flow chart of an embodiment of the hyperspectral unmixing method of the present invention;

[0022] Figure 2 It is a schematic structural diagram of the hyperspectral unmixing system of the present invention;

[0023] Figure 3 It is a schematic diagram of the end-member guided Transformer module of the present invention;

[0024] Figure 4 It is a schematic diagram of the heterogeneous information fusion module of the present invention;

[0025] Figure 5 It is the Urban hyperspectral image dataset for experiments;

[0026] Figure 6 It is the unmixing abundance result diagram of the Urban hyperspectral image dataset unmixed by the embodiment of the present invention and different methods. Detailed Embodiments

[0027] The following further elaborates and explains the present invention in combination with specific embodiments. The embodiments are only demonstrations of the present disclosure content and do not delimit the scope of limitation. Without conflict, the technical features of each embodiment in the present invention can be combined accordingly.

[0028] As Figure 1 shown, it is a basic step flow chart of the hyperspectral image unmixing method of the present invention, mainly including:

[0029] Step 1): (Image preprocessing) For the hyperspectral image to be processed, in this embodiment, the vertex component analysis algorithm VCA is used to extract the initial end-members, and let Y ∈ R N×LDenote the input hyperspectral image as \(I\), \(N\) represents the total number of pixels in the image, and \(L\) represents the number of bands of the hyperspectral image. The calculation formula is as follows:

[0030] M VCA =VCA(Y)

[0031] Where \(M\) VCA \(\in\mathbb{R}\) P×L is the initial endmember, representing the number of endmembers in the image; \(P\) represents the number of endmembers of the hyperspectral image.

[0032] Then, for the hyperspectral image and the initial endmember, using a projection layer composed of 1×1 two-dimensional convolution, project the high-dimensional original data into a lower dimension. The calculation formula is as follows:

[0033] \(X = Proj(Y)\)

[0034] \(E = Proj(M\) VCA )

[0035] Where \(X\in\mathbb{R}\) N×D represents the data feature map obtained by projection, and \(Proj(*)\) is the projection function; \(E\in\mathbb{R}\) P×D represents the endmember token obtained by projection. \(D\) represents the number of channels of the data feature map and the endmember token. In this embodiment, \(D = 32\).

[0036] Step 2): Construct a neural network based on endmember-guided Transformer and endmember bundle. The neural network includes an endmember-guided Transformer encoder module, an endmember generator module based on endmember bundle, a heterogeneous information fusion module, and an abundance decoder module. Among them, the endmember-guided Transformer encoder module is used to mine endmember-guided features highly correlated with the endmember token from the data feature map; the endmember generator module based on endmember bundle converts the endmember-guided features into intra-class weights and weights the endmember bundle based on the intra-class weights to obtain the final endmember curve; the heterogeneous information fusion module fuses multiple endmember-guided features from the perspectives of spatial complementarity and channel interaction to obtain heterogeneous information fusion features; the abundance decoder module processes the heterogeneous information fusion features and outputs the abundance map. In this embodiment, random parameters are used as the initial weights of the network.

[0037] The specific working processes of each module in the neural network based on endmember-guided Transformer and endmember bundle are as follows:

[0038] Step 21): The endmember-guided Transformer encoder module includes \(P\) parallel endmember-guided Transformer modules; the structure of the endmember-guided Transformer module is as Figure 3As shown, the end - member - oriented Transformer module first constructs an end - member - oriented projection operator \(P\) based on the input end - member token \(E\). oriented , and the calculation formula is as follows:

[0039] P oriented =(E i ) T ·(E i (E i ) T ) -1 ·E i

[0040] Where \(E\ i is the \(i\) - th vector in the end - member token, \(() T represents the transpose of the matrix, and \(() -1 represents the inverse of the matrix.

[0041] Then, the end - member - oriented Transformer module applies projection to the end - member token and the data feature map and performs layer normalization. The calculation formula is as follows:

[0042] [E′ i , X'] = LN(P oriented ·[E i , X])

[0043] Where \(E′ i , X′ represent the end - member token and the data feature map after end - member - oriented projection respectively, and [] represents the matrix concatenation operation; LN(*) is the function of layer normalization processing.

[0044] Then, the end - member - oriented Transformer module performs a low - redundancy attention mechanism calculation on the end - member token and the data feature map after end - member - oriented projection. The process can be expressed as:

[0045] Q = E′ i W q , K = X'W k , V = X'W v

[0046] head1,..., head n = Split([Q, K, V])

[0047]

[0048] Where \(W q , W k , W v represent the projection matrices of the query matrix Q, the key matrix K, and the value matrix V respectively; head iIndicates the i-th attention head after the data in the data feature map projected by the end-member orientation is split. In this embodiment, there are a total of h attention heads; Split([Q, K, V]) is the attention head splitting function; Softmax is the activation function that can convert the input into weights with a sum of 1 and weight them on the value matrix V; head′ i Represents the i-th attention head after attention weighting; Q i Is the query matrix of the i-th attention head; Is the transpose of the key matrix of the i-th attention head; V i Is the value matrix of the i-th attention head.

[0049] Then, the redundancy removal module based on the mixture of experts system measures the similarity between the attention heads after attention weighting using the spectral information divergence. The calculation formula is expressed as:

[0050]

[0051] Among them, Sim(i, j) is the similarity between head′ i And head′ j ; head′ i Is the i-th attention head after attention weighting; head′ j Is the j-th attention head after attention weighting.

[0052] Based on the calculation of the similarity between all the attention heads after attention weighting, the similarity metric matrix Can be obtained, where h is the number of attention heads.

[0053] The redundancy removal module based on the mixture of experts system determines the importance of each attention head after attention weighting according to the similarity metric matrix and calculates the weighted sum as the output of the low-redundancy attention mechanism, that is, the fused attention-weighted end-member token is obtained. The calculation process can be expressed as:

[0054] g = 1 - Softmax(FC(Sim))

[0055]

[0056] Among them, Represents the gating weight for each attention head after attention weighting, g i Is the gating weight of the i-th attention head after attention weighting; FC represents the fully connected network layer, Represents the fused attention-weighted end-member token.

[0057] To achieve feature fusion between channels, the endmember-oriented Transformer module uses layer normalization and a multi-layer perceptron to process the data after attention weighting, which is expressed by the following formula:

[0058] [E″′ i , I] = [E″ i , X] + MLP(LN([E″ i , X]))

[0059] where E″′ i is the final endmember token, is the endmember-oriented feature, and the endmember-oriented feature is used as the final output of the endmember-oriented Transformer encoder module. Therefore, for P parallel endmember-oriented Transformer modules, the output of the endmember-oriented Transformer encoder module can be expressed as:

[0060] I i = EOT(E i , X)

[0061] where I i represents the endmember-oriented feature of the i-th endmember; EOT(*) is the function of the endmember-oriented Transformer encoder module. Therefore, the encoder outputs a total of P features: I1, …, I P .

[0062] Step 22): The endmember-oriented feature contains exclusive features belonging to each endmember category. The present invention designs an endmember generator module based on endmember bundles for generating a set of endmembers for each pixel, and its structure is as Figure 1 shown. The endmember generator module based on endmember bundles first uses a two-dimensional convolutional layer and a Softmax activation function to convert the endmember-oriented feature into the within-class weight of the endmember bundle, which is expressed by the following formula:

[0063] I weight,i = Softmax(Conv(I i ))

[0064] where represents the within-class weight of the i-th endmember bundle corresponding to each pixel in the image, H is the height of the hyperspectral image, W is the width of the hyperspectral image, and K is the number of spectral bands within the endmember bundle. Thus, the endmember generator module based on endmember bundles weights each endmember bundle based on the within-class weight to obtain the finally extracted endmember curve, which is expressed by the following formula:

[0065] M i = I weight,i ·EM i

[0066] M = Concat(M1,..., M P )

[0067] where EM i is the i-th endmember bundle among P endmember bundles, represents P endmember bundles, represents the i-th endmember corresponding to each position in the hyperspectral image; Concat(*) is a concatenation function; thus, the final endmember result extracted by the model can be obtained by concatenating each endmember tensor

[0068] Step 23): The heterogeneous information fusion module aims to fully fuse the endmember-oriented features belonging to different endmembers, and its structure is as Figure 4 shown. Taking the processing of the i-th endmember-oriented feature as an example, first, the heterogeneous information fusion module rearranges the endmember-oriented feature into a three-dimensional tensor, that is

[0069] To extract the meaningful part of each endmember-oriented feature, the heterogeneous information fusion module generates a spatial mask based on two-dimensional convolution and the Sigmoid activation function, and the formula can be expressed as:

[0070] m i = Sigmoid(Conv(I i ′))

[0071] m j→i = Sigmoid(Conv(I j ′))

[0072] where is the mask of the valuable spatial information in I i ′, while is the mask of the valuable spatial information of I j ′ for fusing I i ′.

[0073] Based on the prior information that the endmembers are spatially heterogeneous, the present invention designs a complementary feature fusion mechanism to enhance the endmember-oriented feature of the current category based on the complement set of the endmember-oriented features of other categories, and the specific formula is expressed as:

[0074]

[0075] where is the spatial fusion feature of the i-th endmember, and ⊙ represents the Hadamard product, which can achieve element-wise multiplication. After performing the above spatial fusion enhancement on each endmember-directed feature, the heterogeneous information fusion module generates P spatial fusion features I″1,…, I″ P .

[0076] Before implementing channel - dimension fusion, the heterogeneous information fusion module uses a concatenation layer to perform concatenation along the channel dimension on all spatially fused features. The formula can be expressed as:

[0077] I spa =Concat(I″1,…,I″ P )

[0078] Where, is the spatially fused feature after concatenation.

[0079] To achieve full interaction in the channel dimension and maximize the retention of valuable information, the heterogeneous information fusion module implements channel - dimension fusion based on channel rearrangement and grouped convolution layers to obtain the features after heterogeneous information fusion. The formula is expressed as follows:

[0080] I fuse =GConv(CS(I spa ))

[0081] Where, represents the heterogeneous information fusion feature, CS represents the channel rearrangement layer, and GConv represents the grouped convolution layer.

[0082] Step 24): The abundance decoder module aims to extract effective abundance information from the heterogeneous information fusion feature and output the final abundance map. Its structure is as Figure 1 shown. The whole process is implemented based on a multi - layer neural network and is expressed by the formula:

[0083] A=Softmax(Conv(GELU(Conv(I fuse ))))

[0084] Where, is the abundance map estimated by the model. GELU and Softmax respectively represent two non - linear activation functions, and Conv represents the two - dimensional convolution layer.

[0085] Step 3: (Neural network model training) Using the entire hyperspectral image as the training sample, the network adopts a self - supervised training method. The sample is input into the network model, and the root mean square error of reconstruction and the spectral angle distance of reconstruction are used as the loss function. The network weights are updated based on the Adam gradient descent method with adaptive learning rate adjustment. In this embodiment, a total of 200 iterations of training are performed.

[0086] Step 4: Based on the trained neural network model, taking the output of the endmember generator module based on the endmember bundle as the endmembers extracted by the network, and taking the output of the abundance decoder as the abundance map estimated by the network, the unmixing result is obtained.

[0087] Figure 2A block diagram of a hyperspectral image unmixing system based on an endmember-guided Transformer and an endmember bundle shown according to an embodiment is as follows Figure 2 As shown, the system includes:

[0088] An image preprocessing module (jointly composed of the image acquisition module and the image preprocessing module in Figure 2 ), which is used to obtain the initial endmembers of the hyperspectral image and obtain the data feature map and endmember tokens;

[0089] A neural network construction module, which is used to construct a neural network based on an endmember-guided Transformer and an endmember bundle. The neural network is used to obtain the abundance estimation result and endmember extraction result of the hyperspectral image based on the data feature map and endmember tokens. The neural network includes an endmember-guided Transformer encoder module, an endmember generator module based on an endmember bundle, a heterogeneous information fusion module, and an abundance decoder module;

[0090] A neural network training module, which is used to train the neural network based on an endmember-guided Transformer and an endmember bundle to obtain a trained neural network; (The neural network construction module and the neural network training module jointly form the neural network module based on an endmember-guided Transformer and an endmember bundle in Figure 2 )

[0091] A hyperspectral image unmixing module (jointly composed of the unmixing result prediction module and the unmixing result output module in Figure 2 ), which is used to use the trained neural network to obtain the abundance estimation result and endmember extraction result of the hyperspectral image to be unmixed, and realize the unmixing of the hyperspectral image.

[0092] Regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.

[0093] For system embodiments, since they basically correspond to method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The system embodiments described above are only illustrative. For example, the image preprocessing module can be a logical function division, and there can be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another unit. Another point is that the connections between the displayed or discussed modules can be communication connections through some interfaces, which can be electrical or other forms. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts. The following takes real hyperspectral images as an example to illustrate the specific implementation manners to reflect the technical effects of the present invention, and the specific steps in the embodiments will not be elaborated.

[0094] Embodiment

[0095] Next, taking the Urban hyperspectral dataset as the research object, the unmixing method of the present invention is verified. In order to comprehensively compare the unmixing performance and display the unmixing results from the perspectives of visualization and quantification, the estimated abundance maps and evaluation indicators: root mean square error of abundance (aRMSE) and spectral angle distance of endmembers (eSAD) are respectively used to evaluate the proposed unmixing method; the formulas are expressed as follows:

[0096]

[0097]

[0098] where a ji represents the true abundance value of the j-th endmember in the i-th pixel; represents the estimated abundance value of the j-th endmember in the i-th pixel; m j represents the true value of the j-th endmember; represents the estimated result of the j-th endmember.

[0099] The Urban dataset contains 307×307 pixels in total. Each pixel contains 162 bands, the wavelength range is from 0.4 to 2.5 microns, and the spectral resolution is 10 nanometers. There are five endmembers in this dataset, namely road, grassland, tree, roof, and dirt. Figure 5 Is the false color map of the hyperspectral image.

[0100] Figure 6 Is the unmixing abundance map obtained from the embodiments of the present invention and different unmixing algorithms for the Urban dataset.

[0101] Table 1 Evaluation indicators of the unmixing results of the Urban hyperspectral image dataset

[0102]

[0103] The comparison method FCLS is from: D.C. Heinz, “Fully constrained least squares linear spectral mixture analysis method for material quantification in hyperspectral imagery,” IEEE Trans. Geosci. Remote Sensing, vol. 39, no. 3, pp. 529–545, 2001.

[0104] The comparison method MiSiCNet is from: B. Rasti, B. Koirala, P. Scheunders, and J. Chanussot, “Misicnet: Minimum simplex convolutional network for deep hyperspectral unmixing,” IEEE Trans. Geosci. Remote Sensing, vol. 60, pp. 1–15, 2022.

[0105] The comparison method IMPSO is from: R. Liu, P. Wang, B. Du, and B. Qu, “Endmember bundle extraction based on improved multiobjective particle swarm optimization,” IEEE Geosci. Remote Sensing Lett., vol. 20, pp. 1–5, 2023.

[0106] The comparison method DOEBE is from: R. Liu, C. Lei, L. Xie, and X. Qin, “A novel endmember bundle extraction framework for capturing endmember variability by dynamic optimization,” IEEE Trans. Geosci. Remote Sensing, vol. 62, pp. 1–17, 2024.

[0107] The comparison method PGMSU is from: S. Shi, M. Zhao, L. Zhang, Y. Altmann, and J. Chen, “Probabilistic generative model for hyperspectral unmixing accounting for endmember variability,” IEEE Trans. Geosci. Remote Sensing, vol. 60, pp. 1–15, 2022.

[0108] The comparison method PPM is from: W. Gao, J. Yang, and J. Chen, “Proportional perturbation model for hyperspectral unmixing accounting for endmember variability,” IEEE Geosci. Remote Sensing Lett., vol. 21, pp. 1–5, 2024.

[0109] The comparison method DGMSSU is from: S. Shi, L. Zhang, Y. Altmann, and J. Chen, “Deep generative model for spatial–spectral unmixing with multiple endmember priors,” IEEE Trans. Geosci. Remote Sensing, vol. 60, pp. 1–14, 2022.

[0110] The unmixing results of the Urban dataset are as Figure 6 shown, and the quantization results are shown in Table 1. In the quantization results, the present invention achieved the best performance, and better unmixing effects were obtained for both the roof and soil endmembers with complex distributions and small occupied areas. Based on the heterogeneous information fusion module, this algorithm can introduce the prior of the heterogeneous distribution of endmembers in space and achieve more accurate abundance estimation when estimating abundances; in addition, the endmember generator based on the endmember bundle can provide more stable and physically meaningful endmember results.

[0111] The above-described embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the present invention. For those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A hyperspectral image unmixing method based on endmember-guided Transformer and endmember bundle, characterized in that: The following steps are involved: Step 1): Obtain the initial endmembers of the hyperspectral image and obtain the data feature map and endmember tokens; Step 2): constructing a neural network based on endmember-guided Transformer and endmember bundles, wherein the neural network is used to obtain abundance estimation results and endmember extraction results of hyperspectral images based on data feature maps and endmember tokens, and the neural network includes an endmember-guided Transformer encoder module, an endmember generator module based on endmember bundles, a heterogeneous information fusion module, and an abundance decoder module; Step 3): Train the neural network based on the endmember-guided Transformer and the endmember bundle to obtain a trained neural network; Step 4): using the trained neural network to obtain the abundance estimation result and endmember extraction result of the hyperspectral image to be unmixed, so as to realize the unmixing of the hyperspectral image; In step 2), the endmember-guided Transformer encoder module mines endmember-guided features related to endmember tokens from the data feature graph; the endmember generator module based on endmember bundles converts the endmember-guided features into intra-class weights of endmember bundles, and then weights the endmember bundles based on the intra-class weights to obtain the final endmember curve, and uses the endmember curve as the endmember extraction result; the heterogeneous information fusion module fuses the endmember-guided features from the perspective of spatial complementarity and channel interaction to obtain heterogeneous information fusion features; The abundance decoder module obtains an estimated abundance map based on the heterogeneous information fusion feature, and uses the abundance map as the abundance estimation result; The endmember-guided Transformer encoder module mines endmember-guided features related to endmember tokens from the data feature graph; including: The endmember-guided Transformer encoder module includes P parallel endmember-guided Transformer modules; First, the endmember-guided Transformer module constructs the endmember-guided projection operator P based on the input endmember tokens. oriented , the formula is as follows: P oriented =(And i ) T ·(AND i (AND i ) T ) -1 ·AND i Among them, E i is the i-th vector in the end-element token; () T represents the transpose of a matrix, () -1 represents the inverse of a matrix; Then, the endmember tokens and data feature maps are subjected to endmember-guided projection and layer normalization to obtain the endmember tokens and data feature maps after endmember-guided projection; the calculation formula is as follows: [HAVE BEEN' i ,X′]=LN(P oriented ·[HAVE BEEN i ,X]) Where E is the end-member token, E∈R P×D , D represents the number of channels of data feature graph and endmember token, P is the number of endmembers of hyperspectral image; E i ′ is the endmember token after endmember-guided projection; X is the data feature map, X∈R N×D , N represents the total number of pixels in the image; X′ is the data feature map after end-member guided projection; [] represents the concatenation operation of the matrix; LN(*) is the function of layer normalization processing; Perform low-redundancy attention mechanism calculation on the endmember tokens and data feature maps after endmember-guided projection to obtain fused attention-weighted endmember tokens; Based on the layer normalization method and multi-layer perceptron processing fused attention weighted endmember tokens, the endmember-oriented features are obtained; the calculation formula is as follows: [E″′ i ,I]=[E″ i ,X]+MLP(LN([E″ i ,X])) Among them, E″′ i is the final end member token; I is the end member guide feature, E″ i is the fused attention weighted endmember token, Finally, the output of the endmember-guided Transformer encoder module is represented as: I i =EOT(E i ,X) Among them, I i is the endmember-guided feature of the i-th endmember; EOT(*) is the function of the endmember-guided Transformer encoder module; The endmember-guided Transformer encoder module outputs a total of P endmember-guided features, namely: I1,…,I P .

2. The hyperspectral image unmixing method based on endmember-guided Transformer and endmember bundle according to claim 1 is characterized in that: The step 1) comprises: The initial endmembers of the hyperspectral image are extracted based on the vertex component analysis algorithm; the calculation formula is as follows: M VCA =VCA(Y) Where Y is the input hyperspectral image, N is the total number of pixels in the image, L is the number of bands in the hyperspectral image; M VCA represents the initial end member, M VCA ∈R P×L , P represents the number of endmembers of the hyperspectral image; The hyperspectral image and the initial endmember are projected using a projection layer consisting of a 1×1 two-dimensional convolution to obtain the data feature map and endmember token; the calculation formula is as follows: X=Proj(Y) E=Proj(M VCA ) Among them, X represents the data feature map obtained by projection, X∈r N×D ; E∈R P×D represents the end-member token obtained by projection, E∈R P ×D ; D represents the number of channels of the data feature map and the end-member token; Proj(*) is the projection function.

3. The hyperspectral image unmixing method based on endmember-guided Transformer and endmember bundle according to claim 1 is characterized in that: The endmember tokens after the endmember-guided projection and the data feature map are subjected to low-redundancy attention mechanism calculation to obtain fused attention-weighted endmember tokens; include: The low-redundancy attention mechanism is calculated for the endmember tokens and data feature maps after the endmember-guided projection, and the formula is: Q=E i ′W q ,K=X′W k ,V=X′W v head1,…,head h =Split([Q,K,V]) Among them, W q , W k , W v are the projection matrices of the query matrix Q, key matrix K, and value matrix V respectively; head i It represents the i-th attention head after the data in the data feature map after the end-member guided projection is split, and there are h attention heads in total; Split([Q,K,V]) is the attention head splitting function; Softmax is the activation function; head′ i is the i-th attention head after attention weighting; Q i is the query matrix of the i-th attention head; is the transpose of the key matrix of the i-th attention head; V i is the value matrix of the i-th attention head; Then, the spectral information divergence is used to measure the similarity between the attention heads after attention weighting. The calculation formula is expressed as: Among them, Sim(i,j) is head′ i and head′ j Similarity between i is the i-th attention head after attention weighting; head′ j is the jth attention head after attention weighting; Based on the calculation of the similarity between all attention heads after weighting, the similarity measurement matrix is ​​obtained Among them, h is the number of attention heads; The importance of each attention-weighted attention head is measured according to the similarity measurement matrix, and the weighted sum is calculated to obtain the fused attention-weighted end-member token. The calculation formula is: g = 1-Softmax(FC(Sim)) Among them, g represents the gating weight of the attention head after attention weighting, FC(*) represents the function of the fully connected network layer, E″ i is the fused attention weighted endmember token, g i is the gating weight of the i-th attention head after attention weighting.

4. The hyperspectral image unmixing method based on endmember-guided Transformer and endmember bundle according to claim 1, characterized in that: The endmember generator module based on endmember bundles converts the endmember guidance features into intra-class weights of the endmember bundles, and then weights the endmember bundles based on the intra-class weights to obtain the final endmember curve, including: The two-dimensional convolutional layer and the Softmax activation function are used to transform the endmember-guided features into the intra-class weights of the endmember bundles. The calculation formula is: I weight,i =Softmax(Conv(I i )) Among them, I weight,i is the intra-class weight of the i-th endmember bundle corresponding to each pixel in the image; H is the height of the hyperspectral image, W is the width of the hyperspectral image, and K is the number of spectra in the endmember beam; Then, each endmember bundle is weighted based on the intra-class weight to obtain the final endmember curve. The calculation formula is: M i =I weight,i ·EM i M=Concat(M1,...,M P ) Among them, EM i is the i-th endmember bundle among the P endmember bundles, M i is the i-th endmember corresponding to each position in the hyperspectral image, M is the final end member curve, Concat(*) is a concatenation function.

5. The hyperspectral image unmixing method based on endmember-guided Transformer and endmember bundle according to claim 1, characterized in that: The heterogeneous information fusion module fuses the end-member guidance features from the perspective of spatial complementarity and channel interaction to obtain heterogeneous information fusion features; including: First, the end member is directed to feature I i Rearrange into a three-dimensional tensor I i ′, Generate a spatial mask based on two-dimensional convolution and Sigmoid activation function, the formula is: m i =Sigmoid(Conv(I i ′)) m j→i =Sigmoid(Conv(I j ′)) Among them, m i For I i ′ contains a mask with valuable spatial information, m j→i isI j 'Fusion I i ′, a mask with valuable spatial information; Then, based on the complement of the endmember-guided features of other categories, the endmember-guided features of the current category are enhanced to obtain the spatial fusion features of the current category. The specific formula is: Among them, I″ i is the spatial fusion feature of the i-th endmember, ⊙ represents Hadamard; Perform spatial fusion enhancement on each end-member guided feature to obtain P spatial fusion features I″1,…,I″ P ; Then all the spatial fusion features are spliced ​​along the channel dimension to obtain the spliced ​​spatial fusion features; the formula is expressed as: I spa =Concat(I″1,…,I″ P ) Among them, I spa is the spatial fusion feature after splicing, Finally, based on the channel rearrangement layer and the grouped convolution layer, the spliced ​​spatial fusion features are fused in the channel dimension to obtain the heterogeneous information fusion features. The formula is expressed as: I fuse =GConv(CS(I spa )) Among them, I fuse is the heterogeneous information fusion feature, CS(*) is the function of the channel reordering layer, and GConv(*) is the function of the grouped convolutional layer.

6. The hyperspectral image unmixing method based on endmember-guided Transformer and endmember bundle according to claim 1, characterized in that: The abundance decoder module obtains an abundance map based on heterogeneous information fusion features, including: The abundance decoder module obtains the abundance map through a multi-layer neural network and based on heterogeneous information fusion features, and the formula is expressed as: A=Softmax(Conv(GELU(Conv(I fuse )))) Where A is the abundance map estimated by the abundance decoder model, GELU and Softmax represent two nonlinear activation functions, respectively, and Conv(*) represents the function of the two-dimensional convolutional layer.

7. The hyperspectral image unmixing method based on endmember-guided Transformer and endmember bundle according to claim 1, characterized in that: In the step 3), The reconstructed root mean square error and the reconstructed spectral angle distance are used as loss functions, and the weights of the neural network are updated based on a gradient descent method with adaptively adjusted learning rate.

8. A hyperspectral image unmixing system based on endmember-guided Transformer and endmember bundle for implementing the method of claim 1, characterized in that: include: An image preprocessing module, which is used to obtain the initial endmembers of the hyperspectral image and obtain the data feature map and endmember tokens; A neural network building module, which is used to construct a neural network based on an endmember-guided Transformer and an endmember bundle, wherein the neural network is used to obtain an abundance estimation result and an endmember extraction result of a hyperspectral image based on a data feature map and an endmember token, and the neural network includes an endmember-guided Transformer encoder module, an endmember generator module based on an endmember bundle, a heterogeneous information fusion module, and an abundance decoder module; A neural network training module, which is used to train a neural network based on an endmember-guided Transformer and an endmember bundle to obtain a trained neural network; The hyperspectral image unmixing module is used to obtain the abundance estimation result and endmember extraction result of the hyperspectral image to be unmixed by using the trained neural network, so as to realize the unmixing of the hyperspectral image.