Hyperspectral de-mixing method based on double-branch cross-weighted network
By adopting a dual-branch cross-weighted network in hyperspectral image processing, combining adaptive feature selection module and spatial channel reconstruction block, the problem of insufficient processing capability of mixed pixels in hyperspectral images is solved, and a more efficient and accurate hyperspectral solution-mixing effect is achieved.
Patent Information
- Application Number
- CN202510221265.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-17
AI Technical Summary
The existence of mixed pixels in hyperspectral images leads to limited data processing capabilities. The existing nonlinear mixed models require prior knowledge and are difficult to effectively apply in actual scenarios.
Using a hyperspectral demix method based on a dual-branch cross-weighted network, hyperspectral images and lidar data are processed through a spectral encoder and a spatial encoder, and feature fusion and reconstruction are performed in combination with an adaptive feature selection module and a spatial channel reconstruction block.
It improves the accuracy and efficiency of hyperspectral image demixing, and can more accurately estimate the abundance coefficient and reconstruct the hyperspectral image on multimodal data sets, with strong stability and effectiveness.
Smart Images

Figure CN120162736A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hyperspectral image processing, and particularly to a hyperspectral unmixing method based on a dual-branch cross-weighted network. Background Art
[0002] Hyperspectral images are obtained by imaging spectrometers on different space platforms for target areas, and are valuable resources in the field of remote sensing. Hyperspectral images have been widely applied in various industries. However, the spectral resolution and spatial complexity of imagers have led to the existence of mixed pixels in hyperspectral images, which inevitably affects the data processing ability. Therefore, the purpose of hyperspectral unmixing is to estimate the spectral features and their corresponding abundance coefficients contained in the mixed pixels, so as to better utilize hyperspectral images in practical applications.
[0003] For the hyperspectral unmixing task, there are usually two spectral mixture model assumptions: the linear mixture model and the non-linear mixture model. The non-linear mixture model assumes that the observed pixel points can be approximated as a linear combination of pure spectral features, and then weighted according to their corresponding proportional factor abundances. However, the interaction between light reflection and spectral changes makes the basic assumption of the linear mixture model difficult to hold. Therefore, in order to solve more complex non-linear interaction problems and apply non-linear measurement models to practical scenarios, many non-linear measurement models have been proposed. However, using non-linear mixture models usually requires prior knowledge of non-linear features to help establish the unmixing model. Therefore, we need to seek a more general and simpler method to help us unmix hyperspectral images. Summary of the Invention
[0005] Aiming at the defects of the prior art, the purpose of the present invention is to provide a hyperspectral unmixing method based on a dual-branch cross-weighted network.
[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is: a hyperspectral unmixing method based on a dual-branch cross-weighted network, including the following steps: S1. Obtain a hyperspectral multimodal dataset, which contains hyperspectral image data and lidar data, and input the hyperspectral image data and lidar data into the dual-branch cross-weighted network; S2. Process the hyperspectral image data through a spectral encoder to obtain primary spectral features; S3. Process the lidar image data through a spatial encoder to obtain primary spatial features; S4. Construct an adaptive feature selection module. The adaptive feature selection module includes a first branch and a second branch. Input the primary spectral features and primary spatial features into the first branch for processing to obtain precise spectral features. S5. Input the primary spectral features and primary spatial features into the second branch for processing to obtain precise spatial features. S6. Feature fusion: Add the obtained precise spectral features and precise spatial features together so that the precise spectral features and precise spatial features are fused. S7. Pass the result after fusion in step S6 through the Softmax activation function to obtain the estimated abundance coefficient map and the corresponding endmembers. Use the decoder to obtain the reconstructed hyperspectral image by integrating the abundance coefficient map and the corresponding endmembers.
[0007] Furthermore, in step S4, it includes the following sub-steps: S41. Construct a spatial channel reconstruction block. S42. When processing the primary spectral features, a 1×1 convolutional layer is used. When processing the primary spatial features, a spatial channel reconstruction block, a 1×1 convolutional layer, a rectified linear unit, and squeeze-and-excitation attention are used.
[0008] Furthermore, in step S5, it includes the following sub-steps: S51. When processing the primary spectral features, a 1×1 convolutional layer is used. When processing the primary spatial features, a spatial channel reconstruction block, a 1×1 convolutional layer, a rectified linear unit, and squeeze-and-excitation attention are used.
[0009] The spatial channel reconstruction block consists of a horizontal branch, a vertical branch, and a channel branch. Denote the size of the primary spatial features as H×W×C, where H is the height, W is the width, and C is the number of channels. In the horizontal branch, the data tensor is redefined as (HC)×W, and a weight W of size W×W is applied to each HC row. W ; In the vertical branch, the data tensor is redefined as (WC)×H, and a weight W of size H×H is applied to each WC row. H ; In the channel branch, the data tensor is redefined as (HW)×C, and a weight W of size C×C is applied to each HW row. C . Finally, the outputs of the three paths are fused together by adding them, thus obtaining an output feature tensor with the same dimension as the input features.
[0010] Furthermore, in step S7, the decoder includes a 1×1 convolutional layer and a rectified linear unit.
[0011] The beneficial effects of the present application are as follows: In the double-branch cross-weighted network of the present application, a spatial channel reconstruction block and an adaptive feature selection module are introduced. The spatial channel reconstruction block can help the network obtain spatial feature information more accurately from three aspects: horizontal, vertical, and channel. The adaptive feature selection module can effectively utilize spectral feature information and spatial feature information, and better focus on the discrimination of materials in the scene through weighted fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is the overall structure diagram of the present invention;
[0013] Figure 2 is the structure diagram of the spatial channel reconstruction block of the present invention;
[0014] Figure 3 is the selected area diagram in the Muffle dataset;
[0015] Figure 4 is the selected area diagram in the Houston dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0017] An embodiment of the present invention provides a hyperspectral unmixing method based on a double-branch cross-weighted network, which includes the following steps: S1. Obtain a hyperspectral multimodal dataset, which contains hyperspectral image data and lidar data, and input the hyperspectral image data and lidar data into the double-branch cross-weighted network; S2. Process the hyperspectral image data through a spectral encoder to obtain primary spectral features; S3. Process the lidar image data through a spatial encoder to obtain primary spatial features; S4. Construct an adaptive feature selection module, which includes a first branch and a second branch. Input the primary spectral features and primary spatial features into the first branch for processing to obtain precise spectral features, S5. Input the primary spectral features and primary spatial features into the second branch for processing to obtain precise spatial features; S6. Feature fusion: Add the obtained precise spectral features and precise spatial features to fuse the precise spectral features and precise spatial features; S7. Pass the result after fusion in step S6 through a Softmax activation function to obtain an estimated abundance coefficient map and corresponding endmembers, and use a decoder to obtain a reconstructed hyperspectral image by integrating the abundance coefficient map and the corresponding endmembers.
[0018] In step S4, the following sub-steps are included: S41. Construct a spatial channel reconstruction block; S42. When processing the primary spectral features, a 1×1 convolutional layer is used. When processing the primary spatial features, a spatial channel reconstruction block, a 1×1 convolutional layer, a rectified linear unit, and a squeeze-and-excitation attention are used.
[0019] In step S5, when processing the primary spectral features, a 1×1 convolutional layer is used. When processing the primary spatial features, a spatial channel reconstruction block, a 1×1 convolutional layer, a rectified linear unit, and a squeeze-and-excitation attention are used. The adaptive feature selection module can extract precise spectral features and precise spatial features from the primary spectral features and the primary spatial features respectively, analyze the uniqueness of each feature and its contribution to the task objective. Then the weights are adjusted to strengthen the influence of the key features while weakening the interference of the irrelevant features or the noise features, so as to ensure that the model focuses more on the truly important information during the training process.
[0020] The spatial channel reconstruction block consists of a horizontal branch, a vertical branch, and a channel branch. Denote the size of the primary spatial features as H×W×C, where H is the height, W is the width, and C is the number of channels. In the horizontal branch, the data tensor is redefined as (HC)×W, and a weight W of size W×W is applied to each HC row W ; in the vertical branch, the data tensor is redefined as (WC)×H, and a weight W of size H×H is applied to each WC row H ; in the channel branch, the data tensor is redefined as (HW)×C, and a weight W of size C×C is applied to each HW row C . Finally, the outputs of the three paths are fused together by addition to obtain an output feature tensor with the same dimension as the input features.
[0021] In step S7, the decoder contains a 1×1 convolutional layer and a rectified linear unit.
[0022] Experiments are conducted on two real multi-modal datasets and are competitive compared with seven other representative methods.
[0023] The following is the processing of two real hyperspectral multi-modal datasets:
[0024] Muffle dataset: The original image has 325×220 pixels and 64 bands in the spectrum with a range of 0.375~1.050μm. We selected a popular region of interest with a size of 130×90 pixels, as Figure 3 shown. Five dominant materials in this scene are studied: #1 roof, #2 grassland, #3 tree, #4 shadow, and #5 asphalt. The experimental results are shown in Table 1 below.
[0025] Table 1 Comparison of the proposed application and seven other unmixing methods in the real Muffle multimodal dataset
[0026] Houston dataset: It is the Houston data acquired by the ITRES CASI-1500 sensor on the campus of the University of Houston, USA in June 2012. This dataset was initially released by the 2013 IEEE GRSS Data Fusion Contest and has been widely used to evaluate the performance of land cover classification. The original image records 349×1905 pixels in 144 bands, with a range of 0.364 - 1.046 μm. We studied 170×170 pixel sub-images cropped from the original image, as Figure 4 shown. The four endmembers in this scene are: #1 Parking Lot 1, #2 Parking Lot 2, #3 Runway, and #4 Healthy Grass. The experimental results are shown in Table 2.
[0027] Table 2 Comparison of the proposed application and seven other unmixing methods in the real Houston multimodal dataset
[0028] Tables 1 and 2 respectively show the quantitative evaluation metrics of the proposed application and seven representative methods on the Muffle multimodal dataset and the Houston multimodal dataset. Although the application did not obtain all the best eSAD results for individual endmembers, it performed best in terms of the average eSAD of all endmembers. More importantly, all the endmember results extracted by the application are close to the best endmember results on the Muffle multimodal dataset and the Houston multimodal dataset, which also fully demonstrates the stability and effectiveness of the proposed application. After being tested by two evaluation metrics, aRMSE and Mean eSAD, on multiple datasets, compared with other methods, the application can obtain more accurate endmembers and abundance results, which fully proves its superior performance in actual unmixing tasks.
[0029] To prove the effectiveness of the method proposed in this patent, we also conducted ablation experiments on the real Houston multimodal dataset. The ablation experiment results are shown in Tables 3 and 4.
[0030] Table 3 Ablation analysis of the application on the Houston multimodal dataset
[0031] As can be seen from Table 3, when we introduce the spatial channel reconstruction block and weight it, the unmixing performance of the present application is significantly improved. To some extent, this indicates that the present application with cross-weighted integration has certain improvements in estimating endmembers and abundances. The spatial channel reconstruction block obtains spatial information from three aspects, and after introducing weights, it can pay more attention to important information. This can further improve the aRMSE and eSAD of different materials. Therefore, the cross-weight design is effective in the unmixing of the dual-branch network.
[0032] Table 4 explores the influence of the number of stacked blocks on the network of the present application in the Houston multimodal dataset
[0033] To explore the influence of the number of stacked blocks on the unmixing performance, we set up six experiments. M and N represent the number of spatial channel reconstruction blocks in stacked blocks A and B respectively. To ensure the fairness of comparison, we set the same hyperparameters for the networks with different numbers of stacked blocks and selected the network with the best unmixing performance for comparative analysis. As shown in Table 4, as the number of stacked blocks increases, the unmixing performance also improves. By introducing the spatial channel reconstruction block in the adaptive feature selection module, the present application after integration has certain improvements in endmember and abundance estimation. Therefore, the design of the spatial channel reconstruction block plays an important role in the field of multimodal unmixing.
[0034] The present application is mainly achieved by minimizing the reconstruction error. Therefore, we use the SAD measurement method, which can be expressed as:
[0035] where N is the number of pixels in the hyperspectral image data, y i and represent the i-th pixel in the input hyperspectral image Y and the i-th pixel in the reconstructed hyperspectral image respectively.
[0036] To obtain sparse abundance results, we also introduce L 1 / 2 sparse regularization, which can be expressed as: where P and N represent the inclusions and the number of pixels in the hyperspectral image spectral library respectively, represents the abundance element in the i-th row and j-th column of the abundance matrix .
[0037] In addition, to effectively handle the endmember extraction problem, we adopt the Minimum Volume Constraint (MVC) to help find the compact simplex surrounded by endmembers. In the measurement based on endmember distance, the decoder part uses MVC regularization to obtain robust endmember results, so it can be expressed as: where L is the number of spectral bands, P is the number of endmembers in the hyperspectral image spectral library, and are the endmember vectors used to estimate the i-th and j-th ones respectively.
[0038] Therefore, the total loss function of this application is: L = L A + ηL S + ξL MVC
[0039] where η and ξ play the role of hyperparameters when balancing the three objective functions.
[0040] The present invention can separately process the primary spectral features and primary spatial features in two branches. By using the spatial channel reconstruction block, more representative spatial attributes in the data are obtained in terms of horizontal, vertical, and channel aspects, thereby improving the unmixing performance and accuracy.
[0041] The above are only the preferred embodiments of the present invention, and do not impose any formal limitations on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the equivalent embodiments by using the above-disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any simple modification, equivalent change, and modification made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A hyperspectral demixing method based on a double-branch cross-weighted network, characterized in that: The following steps are involved: S1, obtaining a hyperspectral multimodal data set, the hyperspectral multimodal data set including hyperspectral image data and laser radar data, and inputting the hyperspectral image data and the laser radar data into a dual-branch cross-weighted network; S2, processing the hyperspectral image data through a spectral encoder to obtain primary spectral features; S3, processing the lidar image data through a spatial encoder to obtain primary spatial features; S4, constructing an adaptive feature selection module, the adaptive feature selection module comprising a first branch and a second branch, inputting the primary spectral feature and the primary spatial feature into the first branch for processing to obtain accurate spectral features, S5, inputting the primary spectral features and the primary spatial features into the second branch for processing to obtain accurate spatial features; S6, feature fusion: adding the obtained precise spectral features and precise spatial features to fuse the precise spectral features and the precise spatial features; S7, the result after fusion in step S6 is subjected to a Softmax activation function to obtain an estimated abundance coefficient map and corresponding end members, and a decoder is used to integrate the abundance coefficient map and the corresponding end members to obtain a reconstructed hyperspectral image.
2. The hyperspectral unmixing method based on a double-branch cross-weighted network according to claim 1 is characterized in that: Step S4 includes the following sub-steps: S41, constructing a spatial channel reconstruction block; S42. When processing primary spectral features, a 1×1 convolutional layer was used, and when processing primary spatial features, a spatial channel reconstruction block, a 1×1 convolutional layer, a rectified linear unit, and squeeze-excited attention were used.
3. The hyperspectral unmixing method based on a double-branch cross-weighted network according to claim 2 is characterized in that: Step S5 includes the following sub-steps: S51. When processing primary spectral features, a 1×1 convolutional layer was used, and when processing primary spatial features, a spatial channel reconstruction block, a 1×1 convolutional layer, a rectified linear unit, and squeeze-excited attention were used.
4. The hyperspectral unmixing method based on a double-branch cross-weighted network according to claim 3 is characterized in that: The spatial channel reconstruction block consists of a horizontal branch, a vertical branch, and a channel branch. The primary spatial feature size is denoted as H×W×C, where H is the height, W is the width, and C is the number of channels. In the horizontal branch, the data tensor is redefined as (HC)×W, and a weight W of size W×W is applied to each HC row. W ; In the vertical branch, the data tensor is redefined as (WC)×H, and a weight W of size H×W is applied to each WC row H ; In the channel branch, redefine the data tensor as (HW)×C and apply a weight W of size C×C to each HW row C Finally, the outputs of the three paths are fused together by adding them together to obtain an output feature tensor with the same dimension as the input feature.
5. The hyperspectral unmixing method based on a dual-branch cross-weighted network according to claim 1, characterized in that: In step S7, the decoder includes a 1×1 convolutional layer and a rectified linear unit.
Citation Information
Cited By
Hyperspectral image super-resolution method based on unmixing
CN121481848A
Hyperspectral image super-resolution method based on unmixing
CN121481848B