Hyperspectral Image Unmixing Method and Device Based on Prior Knowledge Correction and Swin Transformer
By using a Swin Transformer dual-branch network corrected by prior knowledge in hyperspectral image demix, the end element prior information and weight sharing strategy is used to solve the problems of low demix accuracy and poor reliability in the existing methods, and the demix effect of high precision and low complexity is achieved.
Patent Information
- Application Number
- CN202410505185.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-25
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-04-25
AI Technical Summary
The existing hyperspectral image demixing methods have low demix accuracy and insufficient reliability and good generalization. Traditional convolutional neural networks are difficult to capture global context features and have high computational complexity, and deep learning models are prone to generate meaningless end elements.
The Swin Transformer dual-branch hyperspectral image demix network model based on prior knowledge correction is adopted, and the pre-extracted end elements provide pure pixel prior information. Feature extraction and decomposition are performed through weight sharing strategy combined with the Swin Transformer structure to improve the demix accuracy and reliability.
It improves the accuracy and reliability of hyperspectral image demix, reduces the computational complexity, and achieves better demix performance and accuracy of abundance map.
Smart Images

Figure CN118736428B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hyperspectral image unmixing, and particularly to a hyperspectral image unmixing method and device based on prior knowledge correction and Swin Transformer. Background Art
[0002] With the development of science and technology, remote sensing earth observation technology has become increasingly mature, and refinement, practicality, and integration have become the future development direction of high-resolution earth observation systems. As an advanced technology in aerospace science and technology, hyperspectral remote sensing technology has data with rich spectral information and high spectral resolution. It has broad application requirements in civilian fields such as environmental monitoring and disaster assessment, fine classification of crops and vegetation, marine resource census, detection and identification of rock minerals, investigation of illegal cultivation, and military fields such as military target reconnaissance. However, the complex diversity of ground objects leads to the existence of mixed pixels in hyperspectral images. Coupled with the generally low resolution of hyperspectral imaging systems, a single pixel usually covers a range of dozens or even hundreds of meters. This greatly restricts the development of quantitative applications of hyperspectral data. Therefore, how to reduce the influence of mixed pixels, study hyperspectral image unmixing methods, and improve the unmixing accuracy of hyperspectral images is of great practical significance for improving the quantitative application accuracy of hyperspectral image data, enhancing the target detection ability, and realizing the fine classification of ground object targets.
[0003] Currently, spectral unmixing mainly uses two methods: constructing a non-linear unmixing model through deep learning and a physical-based linear model to fit the mixture. However, due to the scale factors caused by illumination and terrain changes, complex noise generated by environmental conditions or instrument sensors, physical and chemical atmospheric effects, and non-linear mixing of materials in spectral imaging, simple linear models can hardly fully fit complex spectral mixing phenomena and are difficult to accurately decompose mixed pixels. Therefore, the research on non-linear models for spectral decomposition is particularly important. The powerful non-linear fitting ability of deep learning methods represented by convolutional neural networks performs excellently in the field of spectral unmixing.
[0004] Existing methods at the present stage mainly focus on convolutional neural networks, mainly learning features in local intervals and lacking the ability to capture global context features. In addition, operations such as downsampling and pooling in them result in the loss of details during the unmixing process. Moreover, the unmixing network based on the Transformer structure has a relatively high computational complexity and has problems with poor ability to capture multi-scale features. In addition, when most deep learning models perform spectral decomposition, they are prone to generating meaningless or non-existent endmembers and do not have sufficient reliability and good generalization. Summary of the Invention
[0005] To solve the technical problems that existing hyperspectral image unmixing methods based on deep learning have relatively low unmixing accuracy and the unmixing methods do not have sufficient reliability and good generalization, the embodiments of the present invention provide a hyperspectral image unmixing method and device based on prior knowledge correction and Swin Transformer. The technical solutions are as follows:
[0006] On the one hand, a hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer is provided. This method is implemented by a hyperspectral image unmixing device, and the method includes:
[0007] S1. Obtain a hyperspectral image to be unmixed.
[0008] S2. Input the hyperspectral image into a trained hyperspectral image unmixing network model.
[0009] Among them, the hyperspectral image unmixing network model includes a first hyperspectral image unmixing network model branch and a second hyperspectral image unmixing network model branch.
[0010] The first hyperspectral image unmixing network model branch includes a first branch encoder, a first branch Swin Transformer module, and a first branch decoder.
[0011] The second hyperspectral image unmixing network model branch includes a second branch encoder, a second branch Swin Transformer module, and a second branch decoder.
[0012] A weight sharing strategy is used between the first hyperspectral image unmixing network model branch and the second hyperspectral image unmixing network model branch.
[0013] S3. Obtain a hyperspectral image unmixing result according to the hyperspectral image and the hyperspectral image unmixing network model.
[0014] Optionally, the training process of the hyperspectral image unmixing network model in S2 includes:
[0015] S21. Obtain a hyperspectral dataset, extract pure pixel points from the original hyperspectral images in the hyperspectral dataset, and obtain the extracted images.
[0016] S22. Obtain the weights of the first branch according to the extracted images and the first hyperspectral image unmixing network model branch.
[0017] S23. Obtain the second hyperspectral image unmixing network model branch with weights according to the weights of the first branch and the second hyperspectral image unmixing network model branch.
[0018] S24. Input the original hyperspectral image into the second hyperspectral image unmixing network model branch with weights assigned to obtain the reconstructed hyperspectral image.
[0019] S25. Design a loss function and train the hyperspectral image unmixing network model according to the loss function to obtain the trained hyperspectral image unmixing network model.
[0020] Optionally, the hyperspectral image unmixing network model in S2 is as shown in the following formula (1):
[0021] Y = EA + N, s.t. 0 ≤ E ≤ 1, A ≥ 0 (1)
[0022] In the formula, Y represents the observed original hyperspectral image, E represents the endmember matrix, A represents the abundance matrix, and N represents the noise matrix.
[0023] Optionally, the loss function in S25 is as shown in the following formulas (2)-(4):
[0024] L = ηL downstreamRE + σL downstreamSAD (2)
[0025]
[0026]
[0027] In the formula, L represents the total loss function, η represents the regularization parameter weight, L downstreamRE represents the downstream network reconstruction error loss, σ represents the regularization parameter weight, L downstreamSAD represents the downstream network spectral angle distance loss, e represents the true endmember, represents the estimated endmember, H represents the height of the original image, W represents the width of the original image, represents the j-th band of the i-th endmember estimated by the network, e ij represents the j-th band of the i-th endmember of the image, M represents the number of endmembers of the original hyperspectral image, e i represents the i-th endmember of the image, represents the i-th endmember estimated by the network.
[0028] Optionally, the first branch encoder is used to extract features from the input image to obtain the output of the first branch encoder.
[0029] The first branch Swin Transformer module is used to divide the output of the first branch encoder into non-overlapping tokens, obtain the context global information, and get the output of the first branch Swin Transformer module.
[0030] The first branch decoder is used to reconstruct the hyperspectral image from the output of the first branch Swin Transformer module.
[0031] Optionally, the first branch encoder includes a first convolutional layer, a second convolutional layer, and a third convolutional layer.
[0032] Among them, the first convolutional layer includes a first Conv 2D layer, a first batch normalization (BN) layer, a first Dropout layer, and a Leaky ReLU function.
[0033] The second convolutional layer includes a second Conv 2D layer, a second BN layer, and a second Dropout layer.
[0034] The third convolutional layer includes a third Conv 2D layer and a third BN layer.
[0035] Optionally, the first branch Swin Transformer module includes one or more Swin Transformer encoders.
[0036] Among them, each Swin Transformer encoder includes one or more shifted window multi-head self-attention networks.
[0037] On the other hand, a hyperspectral image unmixing device based on prior knowledge correction and Swin Transformer is provided. The device is applied to a hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer. The device includes:
[0038] An acquisition module for acquiring the hyperspectral image to be unmixed.
[0039] An input module for inputting the hyperspectral image into the trained hyperspectral image unmixing network model;
[0040] Among them, the hyperspectral image unmixing network model includes a first hyperspectral image unmixing network model branch and a second hyperspectral image unmixing network model branch.
[0041] The first hyperspectral image unmixing network model branch includes a first branch encoder, a first branch Swin Transformer module, and a first branch decoder.
[0042] The second hyperspectral image unmixing network model branch includes a second branch encoder, a second branch Swin Transformer module, and a second branch decoder.
[0043] A weight sharing strategy is used between the first hyperspectral image unmixing network model branch and the second hyperspectral image unmixing network model branch.
[0044] An output module, configured to obtain a hyperspectral unmixing result according to a hyperspectral image and a hyperspectral image unmixing network model.
[0045] Optionally, the input module is further configured to:
[0046] S21. Obtain a hyperspectral dataset, extract pure pixel points from the original hyperspectral image in the hyperspectral dataset, and obtain the extracted image.
[0047] S22. According to the extracted image and the first branch of the hyperspectral image unmixing network model, obtain the weights of the first branch.
[0048] S23. According to the weights of the first branch and the second branch of the hyperspectral image unmixing network model, obtain the second branch of the hyperspectral image unmixing network model with weights assigned.
[0049] S24. Input the original hyperspectral image into the second branch of the hyperspectral image unmixing network model with weights assigned, and obtain the reconstructed hyperspectral image.
[0050] S25. Design a loss function, and train the hyperspectral image unmixing network model according to the loss function to obtain a trained hyperspectral image unmixing network model.
[0051] Optionally, the hyperspectral image unmixing network model is shown as the following formula (1):
[0052] Y = EA + N, s.t. 0 ≤ E ≤ 1, A ≥ 0 (1)
[0053] In the formula, Y represents the observed original hyperspectral image, E represents the endmember matrix, A represents the abundance matrix, and N represents the noise matrix.
[0054] Optionally, the loss function is shown as the following formulas (2)-(4):
[0055] L = ηL downstreamRE + σL downstreamSAD (2)
[0056]
[0057]
[0058] In the formula, L represents the total loss function, η represents the regularization parameter weight, L downstreamRE represents the downstream network reconstruction error loss, σ represents the regularization parameter weight, L downstreamSAD represents the downstream network spectral angle distance loss, e represents the true endmember, represents the estimated endmember, H represents the height of the original image, W represents the width of the original image, Denote the j-th band of the i-th endmember estimated by the network, e ij Denote the j-th band of the i-th endmember of the image, M represents the number of endmembers of the original hyperspectral image, e i Denote the i-th endmember of the image, Denote the i-th endmember estimated by the network.
[0059] Optionally, a first-branch encoder is used to extract features from the input image to obtain the output of the first-branch encoder.
[0060] The first-branch Swin Transformer module is used to divide the output of the first-branch encoder into non-overlapping tokens to obtain context global information, and obtain the output of the first-branch Swin Transformer module.
[0061] The first-branch decoder is used to reconstruct the hyperspectral image from the output of the first-branch Swin Transformer module.
[0062] Optionally, the first-branch encoder includes a first convolutional layer, a second convolutional layer, and a third convolutional layer.
[0063] Among them, the first convolutional layer includes a first Conv 2D layer, a first batch normalization BN layer, a first Dropout layer, and a Leaky ReLU function.
[0064] The second convolutional layer includes a second Conv 2D layer, a second BN layer, and a second Dropout layer.
[0065] The third convolutional layer includes a third Conv 2D layer and a third BN layer.
[0066] Optionally, the first-branch Swin Transformer module includes one or more Swin Transformer encoders.
[0067] Among them, each Swin Transformer encoder includes one or more shifted window multi-head self-attention networks.
[0068] On the other hand, a hyperspectral image unmixing device is provided, and the hyperspectral image unmixing device includes: a processor; a memory, and computer-readable instructions are stored on the memory. When the computer-readable instructions are executed by the processor, any one of the methods in the above-mentioned hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer is implemented.
[0069] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned hyperspectral image unmixing methods based on prior knowledge correction and Swin Transformer.
[0070] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0071] In the embodiments of the present invention, a Swin Transformer double-branch hyperspectral image unmixing network model with prior information correction is constructed. The upper branch uses pre-extracted endmembers to provide pure pixel prior information. The lower branch adopts the Swin Transformer structure for feature extraction and decomposition processing. A weight sharing strategy is adopted between the two branches to introduce prior knowledge to improve the reliability of the network model, focus on global information, and improve the unmixing accuracy with relatively low computational complexity. Description of the Drawings
[0072] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0073] Figure 1 is a flowchart of a hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer provided by the embodiments of the present invention;
[0074] Figure 2 is a schematic structural diagram of a hyperspectral image unmixing network model provided by the embodiments of the present invention Figure 1 ;
[0075] Figure 3 is a schematic structural diagram of a hyperspectral image unmixing network model provided by the embodiments of the present invention Figure 2 ;
[0076] Figure 4 is a schematic comparison diagram of the endmember spectral curves (Extracted curves) extracted by using the method of the present invention on the Samson dataset and the true endmember spectral curves (GT curves) provided by the embodiments of the present invention;
[0077] Figure 5It is a comparison schematic diagram of the endmember spectral curves (Extracted curves) extracted by using the existing comparison method Deep Hyperspectral Unmixing using Transformer Network on the Samson dataset and the true endmember spectral curves (GT curves) provided by the embodiments of the present invention;
[0078] Figure 6 It is a comparison schematic diagram of the abundance maps extracted by using the method of the present invention on the Samson dataset and the true abundance maps provided by the embodiments of the present invention;
[0079] Figure 7 It is a comparison schematic diagram of the abundance maps extracted by using the existing comparison method Deep Hyperspectral Unmixing using Transformer Network on the Samson dataset and the true abundance maps provided by the embodiments of the present invention;
[0080] Figure 8 It is a block diagram of a hyperspectral image unmixing device based on prior knowledge correction and Swin Transformer provided by the embodiments of the present invention;
[0081] Figure 9 It is a structural schematic diagram of a hyperspectral image unmixing device provided by the embodiments of the present invention. Detailed implementation manners
[0082] Next, the technical solutions in the present invention will be described with reference to the accompanying drawings.
[0083] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or more advantageous than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.
[0084] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when not emphasizing their differences, the meanings they express are the same. "(of)", "corresponding" and "corresponding" can sometimes be used interchangeably. It should be noted that when not emphasizing their differences, the meanings they express are the same.
[0085] In the embodiments of the present invention, sometimes subscripts such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, their intended meanings are the same.
[0086] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0087] The embodiments of the present invention provide a hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer. This method can be implemented by a hyperspectral image unmixing device, which can be a terminal or a server. As Figure 1 shown in the flowchart of the hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer, the processing flow of this method can include the following steps:
[0088] S1. Obtain the hyperspectral image to be unmixed.
[0089] In a feasible implementation, spectral decomposition is an important preprocessing method before spectral data analysis. Its purpose is to simultaneously obtain each component (endmember) in the spectral map and the proportion (abundance) of each component. The powerful non-linear fitting ability of deep learning performs excellently in the field of spectral unmixing. However, since the local spatial information utilized by traditional convolutional models is easily lost in the reconstruction process, resulting in certain errors between the extracted endmembers and their abundances and the true values. Although the model based on the Transformer network considers global information, its computational complexity is high and it is not conducive to the learning of multi-scale features. Also, since most deep learning models lack the guidance of reliable prior knowledge, when spectral pixels are highly mixed, it is easy to lead to unsatisfactory network unmixing results and an abundance map with poor effects cannot be obtained.
[0090] To address the above problems, the present invention designs a dual-branch network. The upper branch network obtains reliable network weights by learning pure pixel blocks, corrects the unmixing results of the lower branch Swin Transformer network for hyperspectral images through a shared weight strategy, and uses the window multi-head self-attention and shifted window multi-head self-attention of the Swin Transformer module to calculate the non-local spatial correlation between hyperspectral pixels to improve the quality of the estimated abundances and enhance the unmixing performance.
[0091] S2. Input the hyperspectral image into the trained hyperspectral image unmixing network model.
[0092] As Figure 2 、 Figure 3 shown, the hyperspectral image unmixing network model includes a first hyperspectral image unmixing network model branch and a second hyperspectral image unmixing network model branch.
[0093] A weight sharing strategy is used between the first hyperspectral image unmixing network model branch and the second hyperspectral image unmixing network model branch.
[0094] In a feasible implementation, the upper branch (the first branch) uses pre-extracted endmembers to provide pure pixel prior information and obtain reliable network weight parameters; the lower branch (the second branch) adopts the Swin Transformer structure for feature extraction and decomposition processing; a weight sharing strategy is adopted between the two branches, and the query parameters generated in the Swin Transformer encoder of the upper branch are shared to the lower branch.
[0095] The first hyperspectral image unmixing network model branch includes a first branch encoder, a first branch Swin Transformer module, and a first branch decoder.
[0096] Optionally, the first branch encoder is used to extract features from the input image to obtain the output of the first branch encoder.
[0097] The first branch Swin Transformer module is used to divide the output of the first branch encoder into non-overlapping tokens, obtain context global information, and obtain the output of the first branch Swin Transformer module.
[0098] The first branch decoder is used to reconstruct the hyperspectral image from the output of the first branch Swin Transformer module.
[0099] The second hyperspectral image unmixing network model branch includes a second branch encoder, a second branch Swin Transformer module, and a second branch decoder.
[0100] In a feasible implementation, each branch includes an encoder, a Swin Transformer module, and a decoder; among them, the encoder is used to extract the main features of the input image and reduce the data dimension; the Swin Transformer module is used to divide the downsampled image into non-overlapping tokens, extract information between different windows, and obtain context global information; the decoder is used to reconstruct the hyperspectral image from the output of the Swin Transformer module.
[0101] Optionally, the first branch encoder includes a first convolutional layer, a second convolutional layer, and a third convolutional layer.
[0102] Among them, the first convolutional layer includes a first convolutional Conv 2D layer, a first batch normalization BN layer, a first Dropout layer, and a Leaky ReLU function.
[0103] The second convolutional layer includes a second Conv2D layer, a second BN layer, and a second Dropout layer.
[0104] The third convolutional layer includes a third Conv2D layer and a third BN layer.
[0105] In a feasible implementation, the encoder includes 3 convolutional layers: the first convolutional layer is successively composed of a Conv2D layer, a BN (Batch Normalization) layer, a Dropout layer, and a Leaky ReLU function; the second convolutional layer is successively composed of a Conv2D layer, a BN layer, and a Dropout layer; the third convolutional layer is successively composed of a Conv2D layer and a BN layer. The specific parameters are shown in Table 1:
[0106] Table 1
[0107]
[0108] Optionally, the first branch Swin Transformer module includes one or more Swin Transformer encoders.
[0109] Among them, each Swin Transformer encoder includes one or more moving window multi-head self-attention networks.
[0110] Optionally, the training process of the hyperspectral image unmixing network model in S2 includes:
[0111] S21. Obtain a hyperspectral dataset, extract pure pixel points from the original hyperspectral images in the hyperspectral dataset to obtain the extracted images.
[0112] S22. According to the extracted images and the first hyperspectral image unmixing network model branch, obtain the weights of the first branch.
[0113] S23. According to the weights of the first branch and the second hyperspectral image unmixing network model branch, obtain the second hyperspectral image unmixing network model branch with weights assigned.
[0114] S24. Input the original hyperspectral images into the second hyperspectral image unmixing network model branch with weights assigned to obtain the reconstructed hyperspectral images.
[0115] S25. Design a loss function, and train the hyperspectral image unmixing network model according to the loss function to obtain the trained hyperspectral image unmixing network model.
[0116] Optionally, the hyperspectral image unmixing network model in S2 is as shown in the following formula (1):
[0117] Y = EA + N, s.t. 0 ≤ E ≤ 1, A ≥ 0 (1)
[0118] Wherein, Y represents the observed original hyperspectral image, E represents the endmember matrix, including r endmembers; A represents the corresponding abundance matrix, where the sum of abundances of each endmember is 1, and N represents the noise matrix.
[0119] Optionally, the loss function in S25 is shown in the following formulas (2)-(4):
[0120] L = ηL downstreamRE + σL downstreamSAD (2)
[0121]
[0122]
[0123] Wherein, L represents the total loss function, η represents the regularization parameter weight, L downstreamRE represents the downstream network reconstruction error loss, σ represents the regularization parameter weight, L downstreamSAD represents the downstream network spectral angle distance loss, e represents the true endmember, represents the estimated endmember, H represents the height of the original image, W represents the width of the original image, represents the j-th band of the i-th endmember estimated by the network, e ij represents the j-th band of the i-th endmember of the image, M represents the number of endmembers of the original hyperspectral image, e i represents the i-th endmember of the image, represents the i-th endmember estimated by the network.
[0124] Regarding the values of the weights η and σ, the selection of the parameter values can be determined by comparing the effects of different values of the parameters on the experimental results. The parameter selection is different on different data sets, Figure 4 is a comparison schematic diagram of the endmember spectral curves (Extracted curves) extracted by using the method of the present invention on the Samson data set and the true endmember spectral curves (GT curves); among them, the left, middle and right three figures respectively represent the comparison schematic diagrams of the spectral curves of the soil, tree and water three endmembers in the hyperspectral image.
[0125] S3. According to the hyperspectral image and the hyperspectral image unmixing network model, the hyperspectral image unmixing result is obtained.
[0126] In this embodiment, the Samson data set is taken as an example to further elaborate the hyperspectral image unmixing method of the present invention.
[0127] The Samson dataset is adopted. This dataset is obtained by Samson sensors and is one of the most widely used hyperspectral datasets in hyperspectral applications. The original image contains 95×95 pixels and 156 bands, covering the wavelength range [401–889] nm. In this research scenario, three endmembers including #1 soil, #2 tree, and #3 water are included.
[0128] According to the method of the present invention, pure endmembers are extracted using an endmember extraction algorithm, and pure pixel blocks are obtained by replicating and deforming the matrix results through MATLAB, which are used as the input of the upper branch of the network, and the original image is input to the lower branch network.
[0129] Both the upper and lower branches input their respective inputs into the encoder. The encoder extracts a basic feature map of the input data, specifically: the input hyperspectral image is represented by 24 channels through three convolutional layers in sequence to discriminate features, greatly reducing the spectral bands of the hyperspectral image, and extracting and forming the high-level features required for the next Transformer module. At this time, the feature output by the encoder is a discriminant feature map with a size of 95×95×24.
[0130] Then it enters the Swin Transformer module. The Swin Transformer module first normalizes the features through a normalization layer (LN), divides the hyperspectral image into non-overlapping local windows of size 5×5, and performs MSA (Multi-head Self-Attention) calculations within each window. To overcome the lack of information exchange between windows, which limits the model's ability to capture long-range dependencies, the former window is offset by [M / 2,M / 2] pixels to achieve cross-window information interaction. Then there is another LN, MLP (Multilayer Perceptron), and residual operations to obtain the output of this layer.
[0131] The output of the Swin Transformer module passes through a multilayer perceptron to obtain an output of 3×95×95. To meet the constraints of the sum of abundances in the abundance map being one and abundances being non-negative, a softmax function is introduced, thereby obtaining the final abundance matrix of 3×95×95.
[0132] Finally, in the decoder, a 1×1 convolution operation and the ReLU (Rectified Linear Unit) function are used to increase the dimension of the abundance map, with the weight being the endmember matrix, resulting in a reconstructed hyperspectral image of 156×95×95. The quantitative results are provided by the RMSE (Root Mean Squared Error) between the estimated abundance fractions and the actual abundance fractions, as well as the SAD between the estimated endmembers and the ground truth endmembers. The lower the values of these metrics, the better the unmixing effect and the closer it is to the true value. Figure 5 (The three figures on the left, in the middle, and on the right respectively show a comparison schematic diagram of the spectral curves of the three endmembers of soil, tree, and water in the hyperspectral image), Figure 6 (The three figures on the left, in the middle, and on the right on the upper side respectively show the true abundance maps of the three endmembers of soil, tree, and water in the hyperspectral image; the three figures on the left, in the middle, and on the right on the lower side respectively show the abundance maps of the three endmembers of soil, tree, and water extracted by the method of the present invention), Figure 7 (The three figures on the left, in the middle, and on the right on the upper side respectively show the true abundance maps of the three endmembers of soil, tree, and water in the hyperspectral image; the three figures on the left, in the middle, and on the right on the lower side respectively show the abundance maps of the three endmembers of soil, tree, and water extracted by the comparative method DeepHyperspectral Unmixing using Transformer Network) and Table 2 (RMSE comparison on the Samson dataset), Table 3 (SAD comparison on the Samson dataset) respectively give the endmember maps, abundance maps, and corresponding quantitative results of the Samson dataset. The present invention is superior to other technologies in general results, with an average root mean squared error of 0.0686, showing the best performance among all methods; the average SAD value is 0.0596, and its result is second only to the method of the AE combine Transformer model in the comparative methods, but shows the best performance in the results of soil and trees.
[0133] Table 2
[0134]
[0135] Table 3
[0136]
[0137] In the embodiment of the present invention, a Swin Transformer double-branch hyperspectral image unmixing network model with prior information correction is constructed. The upper branch uses the pre-extracted endmembers to provide pure pixel prior information. The lower branch adopts the Swin Transformer structure for feature extraction and decomposition processing. A weight sharing strategy is adopted between the two branches to introduce prior knowledge to improve the reliability of the network model, focus on global information, and improve the unmixing accuracy with a relatively small computational complexity.
[0138] Figure 8 It is a block diagram of a hyperspectral image unmixing device based on prior knowledge correction and Swin Transformer shown according to an exemplary embodiment. This device is used for a hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer. Refer to Figure 8 This device includes an acquisition module 310, an input module 320, and an output module 330. Among them:
[0139] The acquisition module 310 is used to acquire the hyperspectral image to be unmixed.
[0140] The input module 320 is used to input the hyperspectral image into the trained hyperspectral image unmixing network model;
[0141] Among them, the hyperspectral image unmixing network model includes a first hyperspectral image unmixing network model branch and a second hyperspectral image unmixing network model branch.
[0142] The first hyperspectral image unmixing network model branch includes a first branch encoder, a first branch Swin Transformer module, and a first branch decoder.
[0143] The second hyperspectral image unmixing network model branch includes a second branch encoder, a second branch Swin Transformer module, and a second branch decoder.
[0144] A weight sharing strategy is used between the first hyperspectral image unmixing network model branch and the second hyperspectral image unmixing network model branch.
[0145] The output module 330 is used to obtain the hyperspectral image unmixing result according to the hyperspectral image and the hyperspectral image unmixing network model.
[0146] Optionally, the input module 320 is further used for:
[0147] S21. Acquire a hyperspectral dataset, extract pure pixel points from the original hyperspectral images in the hyperspectral dataset, and obtain the extracted image.
[0148] S22. According to the extracted image and the first hyperspectral image unmixing network model branch, obtain the weights of the first branch.
[0149] S23. According to the weights of the first branch and the second hyperspectral image unmixing network model branch, obtain the second hyperspectral image unmixing network model branch with weights assigned.
[0150] S24. Input the original hyperspectral image into the second hyperspectral image unmixing network model branch with weights assigned to obtain the reconstructed hyperspectral image.
[0151] S25. Design a loss function and train the hyperspectral image unmixing network model according to the loss function to obtain the trained hyperspectral image unmixing network model.
[0152] Optionally, the hyperspectral image unmixing network model is as shown in the following formula (1):
[0153] Y = EA + N, s.t. 0 ≤ E ≤ 1, A ≥ 0 (1)
[0154] In the formula, Y represents the observed original hyperspectral image, E represents the endmember matrix, A represents the abundance matrix, and N represents the noise matrix.
[0155] Optionally, the loss function is as shown in the following formulas (2)-(4):
[0156] L = ηL downstreamRE + σL downstreamSAD (2)
[0157]
[0158]
[0159] In the formula, L represents the total loss function, η represents the regularization parameter weight, L downstreamRE represents the downstream network reconstruction error loss, σ represents the regularization parameter weight, L downstreamSAD represents the downstream network spectral angle distance loss, e represents the true endmember, represents the estimated endmember, H represents the height of the original image, W represents the width of the original image, represents the j-th band of the i-th endmember estimated by the network, e ij represents the j-th band of the i-th endmember of the image, M represents the number of endmembers of the original hyperspectral image, e i represents the i-th endmember of the image, represents the i-th endmember estimated by the network.
[0160] Optionally, the first branch encoder is used to extract features from the input image to obtain the output of the first branch encoder.
[0161] The first branch Swin Transformer module is used to divide the output of the first branch encoder into non-overlapping tokens, obtain the context global information, and obtain the output of the first branch Swin Transformer module.
[0162] The first branch decoder is used to reconstruct the hyperspectral image from the output of the first branch Swin Transformer module.
[0163] Optionally, the first branch encoder includes a first convolutional layer, a second convolutional layer, and a third convolutional layer.
[0164] Among them, the first convolutional layer includes a first convolutional Conv 2D layer, a first batch normalization BN layer, a first Dropout layer, and a Leaky ReLU function.
[0165] The second convolutional layer includes a second Conv 2D layer, a second BN layer, and a second Dropout layer.
[0166] The third convolutional layer includes a third Conv 2D layer and a third BN layer.
[0167] Optionally, the first branch Swin Transformer module includes one or more Swin Transformer encoders.
[0168] Among them, each Swin Transformer encoder includes one or more moving window multi-head self-attention networks.
[0169] In the embodiments of the present invention, a Swin Transformer dual-branch hyperspectral image unmixing network model with prior information correction is constructed. The upper branch uses pre-extracted endmembers to provide pure pixel prior information. The lower branch uses the Swin Transformer structure for feature extraction and decomposition processing. A weight sharing strategy is adopted between the two branches to introduce prior knowledge to improve the reliability of the network model, focus on global information, and improve the unmixing accuracy with a relatively small computational complexity.
[0170] Figure 9 It is a schematic structural diagram of a hyperspectral image unmixing device provided by the embodiments of the present invention. As Figure 9 shown, the hyperspectral image unmixing device may include the above-mentioned Figure 8 hyperspectral image unmixing device based on prior knowledge correction and Swin Transformer. Optionally, the hyperspectral image unmixing device 410 may include a first processor 2001.
[0171] Optionally, the hyperspectral image unmixing device 410 may further include a memory 2002 and a transceiver 2003.
[0172] Among them, the first processor 2001 is connected to the memory 2002 and the transceiver 2003, such as through a communication bus.
[0173] Next, in combination with Figure 9Specifically introduce each component of the hyperspectral image unmixing device 410:
[0174] Among them, the first processor 2001 is the control center of the hyperspectral image unmixing device 410, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0175] Optionally, the first processor 2001 can execute various functions of the hyperspectral image unmixing device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0176] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 9 the CPU0 and CPU1 shown in
[0177] In a specific implementation, as an embodiment, the hyperspectral image unmixing device 410 may also include multiple processors, such as Figure 9 the first processor 2001 and the second processor 2004 shown in
[0178] Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0178] Among them, the memory 2002 is used to store the software program for implementing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation method can refer to the above method embodiments and will not be elaborated here.
[0179] Optionally, the memory 2002 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 can be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through the interface circuit of the hyperspectral image unmixing device 410 ( Figure 9 not shown in the figure), and the embodiments of the present invention do not make specific limitations in this regard.
[0180] The transceiver 2003 is used to communicate with network devices or communicate with terminal devices.
[0181] Optionally, the transceiver 2003 can include a receiver and a transmitter ( Figure 9 not shown separately in the figure). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0182] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through the interface circuit of the hyperspectral image unmixing device 410 ( Figure 9 not shown in the figure), and the embodiments of the present invention do not make specific limitations in this regard.
[0183] It should be noted that Figure 9 the structure of the hyperspectral image unmixing device 410 shown in the figure does not constitute a limitation on the router. The actual knowledge structure recognition device can include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0184] In addition, the technical effects of the hyperspectral image unmixing device 410 can refer to the technical effects of the hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer described in the above method embodiments, and will not be elaborated here.
[0185] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0186] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM) and direct rambus RAM (DR RAM).
[0187] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0188] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context.
[0189] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0190] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0191] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0192] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0193] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings, direct couplings, or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0194] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0195] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0196] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0197] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer, characterized in that, The method includes: S1. Obtain a hyperspectral image to be unmixed; S2. Input the hyperspectral image into a trained hyperspectral image unmixing network model; Among them, the hyperspectral image unmixing network model includes a first branch and a second branch, and the query parameters generated by the Swin Transformer encoder of the first branch Swin Transformer module in the first branch are shared using a weight sharing strategy; each branch includes an encoder, a Swin Transformer module, and a decoder; the first branch encoder includes a first convolutional layer, a second convolutional layer, and a third convolutional layer; The Swin Transformer module includes one or more Swin Transformer encoders; each Swin Transformer encoder includes one or more shifted window multi-head self-attention networks; the Swin Transformer module uses window shifting by [M / 2, M / 2] pixels to achieve cross-window information interaction; M represents the number of heads of multi-head attention; the first branch Swin Transformer module is used to divide the output of the first branch encoder into non-overlapping tokens, obtain context global information, and obtain the output of the first branch Swin Transformer module; S3. Obtain a hyperspectral image unmixing result according to the hyperspectral image and the hyperspectral image unmixing network model; The training process of the hyperspectral image unmixing network model in S2 includes: S21. Extract pure pixel points from the original hyperspectral images in the obtained hyperspectral dataset to obtain an extracted image; S22. Obtain the weights of the first branch according to the extracted image and the first branch; S23. Obtain the weighted second branch according to the query parameters in the weights of the first branch and the second branch; S24. Input the original hyperspectral image into the weighted second branch to obtain a reconstructed hyperspectral image; S25. Design a loss function to train the hyperspectral image unmixing network model to obtain a trained hyperspectral image unmixing network model.
2. The hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer according to claim 1, wherein The hyperspectral image unmixing network model in S2 is shown as the following formula (1): Y = EA + N, s.t. 0 ≤ E ≤ 1, A ≥ 0 (1) In the formula, Y represents the observed original hyperspectral image, E represents the endmember matrix, A represents the abundance matrix, and N represents the noise matrix.
3. The hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer according to claim 1, wherein, The loss function in S25 is shown as the following formulas (2)-(4): L = ηL downstreamRE + ηL downstreamSAD (2) where \(L\) represents the total loss function, \(\eta\) represents the regularization parameter weight, \(L\) downstreamRE represents the downstream network reconstruction error loss, \(\sigma\) represents the regularization parameter weight, \(L\) downstreamSAD represents the downstream network spectral angle distance loss, \(e\) represents the true endmember, represents the estimated endmember, \(H\) represents the height of the original image, \(W\) represents the width of the original image, represents the \(j\)-th band of the \(i\)-th endmember estimated by the network, \(e\) ij represents the \(j\)-th band of the \(i\)-th endmember of the image, \(M\) represents the number of endmembers of the original hyperspectral image, \(e\) i represents the \(i\)-th endmember of the image, represents the \(i\)-th endmember estimated by the network.
4. The hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer according to claim 1, characterized in that, The first branch encoder is used to extract features from the input image to obtain the output of the first branch encoder; The first branch decoder is used to reconstruct the hyperspectral image from the output of the first branch Swin Transformer module.
5. The hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer according to claim 1, characterized in that, The first convolutional layer includes a first convolutional Conv 2D layer, a first batch normalization BN layer, a first Dropout layer, and a Leaky ReLU function; The second convolutional layer includes a second Conv 2D layer, a second BN layer, and a second Dropout layer; The third convolutional layer includes a third Conv 2D layer and a third BN layer.
6. A hyperspectral image unmixing device based on prior knowledge correction and Swin Transformer, the hyperspectral image unmixing device based on prior knowledge correction and Swin Transformer is used to implement the hyperspectral image unmixing method based on prior knowledge correction and Swin Transformer according to any one of claims 1-5, characterized in that, The device includes: An acquisition module, configured to acquire a hyperspectral image to be unmixed; An input module, configured to input the hyperspectral image into a trained hyperspectral image unmixing network model; Wherein, the hyperspectral image unmixing network model includes a first hyperspectral image unmixing network model branch and a second hyperspectral image unmixing network model branch; The first hyperspectral image unmixing network model branch includes a first branch encoder, a first branch Swin Transformer module, and a first branch decoder; The second hyperspectral image unmixing network model branch includes a second branch encoder, a second branch Swin Transformer module, and a second branch decoder; A weight sharing strategy is used between the first hyperspectral image unmixing network model branch and the second hyperspectral image unmixing network model branch; An output module, configured to obtain a hyperspectral image unmixing result according to the hyperspectral image and the hyperspectral image unmixing network model.
7. A hyperspectral image unmixing device, characterized in that, The hyperspectral image unmixing device includes: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 5 is implemented.
8. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Hyperspectral image super-resolution method based on spectral unmixing convolutional neural network
CN113744134A
Method and system for training a neural network model to perform image classification and image localization and / or segmentation
WO2024080929A1