A broadband snapshot-type hyperspectral fusion imaging method, system, and medium
By fusing panchromatic images and aliased spectral encoded mosaic images using a multi-scale preserved spatial spectral self-attention mechanism module (MRSSA), the limitations of light loss and spectral range in existing technologies are overcome, enabling hyperspectral imaging with high spatial resolution and wide spectral range.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUNAN UNIV
- Filing Date
- 2026-06-17
- Publication Date
- 2026-07-17
AI Technical Summary
Existing hyperspectral imaging technologies struggle to achieve a balance of high spatial resolution, high temporal resolution, and wide spectral range in a single exposure, and are limited by high light loss, hardware complexity, and cost.
A broadband snapshot-type hyperspectral fusion imaging method is adopted. By combining panchromatic images and aliased spectral encoded mosaic images through a multi-scale preserved spatial spectral self-attention mechanism module (MRSSA), feature extraction and fusion are performed to generate broadband hyperspectral images.
Achieving high-quality spatial-spectral fusion and spectral super-resolution reconstruction under a single exposure overcomes limitations in light loss, spectral range, and hardware complexity, resulting in hyperspectral imaging with high spatial resolution and a wide spectral range.
Smart Images

Figure CN122415334A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of hyperspectral fusion imaging technology, specifically to a broadband snapshot-type hyperspectral fusion imaging method, system, and medium. Background Technology
[0002] Existing hyperspectral acquisition technologies are inherently limited by trade-offs between temporal resolution, spatial resolution, and spectral resolution. Scanning imaging schemes (such as point scanning, line scanning, and area scanning) can achieve high spectral resolution, but require multiple exposures, resulting in low temporal resolution and making them unsuitable for dynamic imaging scenarios such as handheld, airborne, or endoscopic imaging. Snapshot imaging schemes rely on narrowband mosaic spectral arrays, but are constrained by micron-level coating processes, typically resulting in a limited number of channels, a spectral range generally around 200 nm, and a fixed bandwidth per channel, making it difficult to balance a wide spectral range with high-precision spectral reconstruction. Computational imaging methods, such as CASSI, can acquire hyperspectral data cubes in a single exposure, but their encoding elements introduce over 70% optical loss. Furthermore, the complex optical path structure not only easily produces artifacts and spectral distortion under undersampling conditions but also limits the miniaturization and low-cost implementation of the system. While RGB image-based super-resolution (SSR) methods have a simple structure, the RGB three channels only cover the visible light region, resulting in significant spectral information loss, making it difficult to guarantee spectral fidelity and reliable substance identification during reconstruction. Existing hyperspectral fusion imaging methods typically combine multi-source image information through spatial-spectral fusion strategies to improve reconstruction results. However, these methods often rely on high-quality hyperspectral images, thus limiting their application in snapshot imaging systems. Current technologies, limited by low scanning temporal resolution, insufficient narrowband mosaic spectral width, high optical loss, and missing spectral information, still cannot simultaneously achieve the imaging requirements of a ≥400nm wide spectral range, high spatial resolution, high signal-to-noise ratio, and system-on-chip integration under single-exposure conditions. Summary of the Invention
[0003] The technical problem to be solved by this invention is to provide a broadband snapshot-type hyperspectral fusion imaging method, system, and medium to address the above-mentioned problems in the prior art. This invention aims to overcome the limitations of existing imaging methods in terms of light loss, spectral range, hardware complexity, and cost, and to achieve hyperspectral imaging with high spatial resolution, high temporal resolution, and wide spectral range coverage in a single exposure.
[0004] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A broadband snapshot-type hyperspectral fusion imaging method includes the following steps: inputting a hyperspectral image An aliased spectral-coded mosaic image is generated through the combined effects of spectral aliasing and mosaic sampling. aliased spectral encoding mosaic image Compared with the input panchromatic image Feature extraction and fusion are performed separately to obtain the initial fused features. ; Initial fusion features The input is a K-level cascaded multi-scale preserving spatial-spectral self-attention mechanism module (MRSSA). Each level of the MRSSA module then integrates the input features with the panchromatic image. The spatial error map is combined to extract enhancement features, and the final K-level enhancement features are concatenated along the channel dimension and refined to obtain refined features. ; refine features With initial fusion features After summing the residuals, the data is input into the reconstruction module to generate a broadband hyperspectral image. The spatial error map is obtained by processing the input features through a linear layer and then comparing them with a panchromatic image. Subtraction yields that the Multi-Scale Preservative Spatial Spectrum Self-Attention Mechanism (MRSSA) module includes a first Preservative Spatial Spectrum Self-Attention Module (RSSA), a downsampling module, a second Preservative Spatial Spectrum Self-Attention Module (RSSA), an upsampling module, a concatenation module, and a 1×1 convolutional layer. The concatenation module concatenates the output features of the first Preservative Spatial Spectrum Self-Attention Module (RSSA) and the output features of the upsampling module, and then processes them through a 1×1 convolutional layer to obtain the output features.
[0005] Optionally, the input hyperspectral image An aliased spectral-coded mosaic image is generated through the combined effects of spectral aliasing and mosaic sampling. The function expression is: ; in, For the number of bands, For the input hyperspectral image, This represents the Hadamard product. and The first The spectral response function and mosaic sampling mask corresponding to each band For the number of channels, The height and width of the hyperspectral image; the aliased spectral encoding mosaic image With panchromatic image Feature extraction and fusion are performed separately to obtain the initial fused features. Includes: encoding aliased spectra into mosaic images With panchromatic image Features are extracted using independent 3×3 convolutions, and the extracted features are concatenated along the channel dimension. Then, a 1×1 convolution is used to fuse the features to obtain the initial fused features. .
[0006] Optionally, the Retaining Spatial Spectrum Self-Attention Module (RSSA) includes: The attention module is used to perform attention calculations on the input features to obtain attention features. ; The Multi-Source Cueing State Space Module (MPSS) combines the Mamba architecture with the Transformer mechanism to extract features from input features, obtaining features that fuse long-range spatial dependencies and multi-source cueing information. ; The multiplication module is used to incorporate attention features. Features that integrate long-range spatial dependencies and multi-source cue information Multiply to obtain the output features.
[0007] Optionally, the attention calculation is performed on the input features to obtain attention features. The function expression is: ; ; ; in, Input features The pooling results For average pooling, For pooling size, , and The query matrix, key matrix, and value matrix are obtained by performing average pooling and linear post-processing on the input features, respectively. , and These are the weight matrices corresponding to the query matrix, key matrix, and value matrix, respectively. The softmax activation function is used. for transpose, and These are the weighting coefficients. This represents the Hadamard product. As a spectral distance prior, Represents the exponential decay coefficient. It is a non-causal mask matrix. ,in and These represent different spectral channel indices. The element in the i-th row and j-th column of the noncausal mask matrix. This is a coarse similarity prior obtained based on the spectral separation mosaic image and the spectral response matrix.
[0008] Optionally, the calculation function expression for the coarse similarity prior obtained based on the spectral separation mosaic image and the spectral response matrix is as follows: ; ; ; in, To separate mosaic images and spectral response matrices based on spectral density The obtained coarse similarity prior, for transpose, This is a low-resolution hyperspectral image obtained by inverse matrix mapping. Encoding mosaic images from aliased spectra The spectral separation mosaic image obtained from it, For spectral separation operation, Spectral response matrix The inverse matrix, for Pooling characteristics, For linear layers, For average pooling, Pooling size; spectral response matrix It is an n x m matrix, where n represents different spectral channels, m represents different wavelengths, and each element represents the sensor's response intensity to incident light at a specific wavelength.
[0009] Optionally, the step of combining the Mamba structure with the Transformer mechanism to extract features from the input features obtains features that fuse long-range spatial dependencies and multi-source cue information. include: The spatial error map is used to generate a memory selection matrix using an error-guided selector. The error-guided selector includes multiple convolutional layers, and the output of each convolutional layer is equipped with a LeakyReLU activation function. Stage-shared memory features Stage-specific memory characteristics Multiply to obtain segmented memory features Among them, the characteristics of stage-shared memory Stage-specific memory features are learnable hyperparameters shared by all multi-scale preserved spatial spectral self-attention mechanism modules (MRSSA) at each level. Each level of the Multi-Scale Preservative Spatial Spectrum Self-Attention Mechanism (MRSSA) module has its own learnable hyperparameters, and the stages share memory features. Stage-specific memory characteristics During training, all results are initialized with one and updated via backpropagation to obtain the final result. Memory selection matrix and segmented memory features Multiplication yields specific error indications ; Specific error prompts The input features are fed into the Mamba state-space model to obtain features that fuse long-range spatial dependencies and multi-source cue information. The Mamba state-space model will provide specific error hints. Add to output mapping model parameters To generate the output features at time k from the input features at time k, the functional expression of the Mamba state-space model is: ; in, The output feature at time k is... These are the output mapping model parameters, used to map the hidden states to the output feature space; This is the hidden state vector of the system at the current time k, used to represent the dynamic representation of the input features in the state space; These are skip connection terms, used to introduce input features or bias terms to enhance the expressiveness and stability of the model; The input features are for the current time k; ; These are the parameters of the state transition model, used to characterize the evolution of the hidden states between adjacent time steps. For the previous time k The system hidden state vector corresponding to 1, These are the parameters of the input mapping model, used to map the current input features to the state space and inject the hidden state update process.
[0010] Optionally, the final K-level enhanced features are concatenated along the channel dimension and then refined to obtain refined features. In this context, the feature refinement process refers to achieving feature refinement through a 1×1 convolution operation; the refinement of features... With initial fusion features After summing the residuals, the data is input into the reconstruction module to generate a broadband hyperspectral image. At that time, the reconstruction module consists of a 1×1 convolutional layer, which is used to integrate the channel information of the fused features and restore the corresponding spectral distribution to generate a broadband hyperspectral image. .
[0011] The present invention also provides a broadband snapshot hyperspectral fusion imaging system, including a microprocessor and a memory interconnected thereto, wherein the microprocessor is programmed or configured to execute the broadband snapshot hyperspectral fusion imaging method.
[0012] The present invention also provides a computer-readable storage medium storing a computer program or instructions that are programmed or configured to execute the broadband snapshot hyperspectral fusion imaging method by a processor.
[0013] The present invention also provides a computer program product, including a computer program or instructions that are programmed or configured to execute the broadband snapshot hyperspectral fusion imaging method via a processor.
[0014] Compared with existing technologies, the present invention mainly achieves the following beneficial effects: The method of the present invention includes generating an aliased spectral encoded mosaic image from an input hyperspectral image under the combined action of spectral aliasing and mosaic sampling, extracting features from the input panchromatic image and fusing them to obtain initial fusion features; inputting the initial fusion features into a K-level cascaded multi-scale preserving spatial-spectral self-attention mechanism module to obtain refined features; adding the residuals of the refined features and the initial fusion features and inputting them into a reconstruction module to generate a broadband hyperspectral image. The present invention can achieve high-quality spatial-spectral fusion and spectral super-resolution reconstruction, breaking through the limitations of existing imaging methods in terms of light loss, spectral range, hardware complexity and cost, and achieving hyperspectral imaging with high spatial resolution, high temporal resolution and broadband spectral range in a single exposure. Attached Figure Description
[0015] Figure 1 This is a schematic diagram illustrating the basic principle of the method in an embodiment of the present invention.
[0016] Figure 2 This is a schematic diagram of the network structure of the Multiscale Retaining Spatial Spectrum Self-Attention Mechanism (MRSSA) module in an embodiment of the present invention.
[0017] Figure 3 This is a schematic diagram of the network structure of the Retained Spatial Spectrum Self-Attention Module (RSSA) in an embodiment of the present invention.
[0018] Figure 4 This is a schematic diagram illustrating the generation principle of coarse similarity priors in an embodiment of the present invention.
[0019] Figure 5This is a feature in this embodiment of the invention that integrates long-range spatial dependency and multi-source cue information. A schematic diagram illustrating the generation principle. Detailed Implementation
[0020] This invention aims to achieve high-quality spatial-spectral fusion and spectral super-resolution reconstruction by utilizing the complementary spatial-spectral information of a single-frame panchromatic image (PAN) and an aliased coded mosaic image (ASM). This overcomes the limitations of existing imaging methods in terms of optical loss, spectral range, hardware complexity, and cost, achieving hyperspectral imaging with high spatial resolution, high temporal resolution, and a wide spectral range in a single exposure. To enable those skilled in the art to better understand the technical solution of this invention, the following detailed description, in conjunction with the accompanying drawings of the embodiments, will further illustrate the technical solution of this invention.
[0021] like Figure 1 As shown, the broadband snapshot hyperspectral fusion imaging method in this embodiment includes the following steps: inputting a hyperspectral image... An aliased spectral-coded mosaic image is generated through the combined effects of spectral aliasing and mosaic sampling. aliased spectral encoding mosaic image Compared with the input panchromatic image Feature extraction and fusion are performed separately to obtain the initial fused features. ; Initial fusion features The input is a K-level cascaded multi-scale preserving spatial-spectral self-attention mechanism module (MRSSA). Each level of the MRSSA module then integrates the input features with the panchromatic image. The spatial error map is combined to extract enhancement features, and the final K-level enhancement features are concatenated along the channel dimension and refined to obtain refined features. ; refine features With initial fusion features After summing the residuals, the data is input into the reconstruction module to generate a broadband hyperspectral image. .
[0022] In this embodiment, the input hyperspectral image An aliased spectral-coded mosaic image is generated through the combined effects of spectral aliasing and mosaic sampling. The function expression is: ; in, For the number of bands, For the input hyperspectral image, and The first The spectral response function and mosaic sampling mask corresponding to each band For the number of channels, represents the height and width of the hyperspectral image. The mosaic sampling mask takes values of 0 or 1, with a value of 0 indicating that the first digit is not a scalar pixel. If a band is not sampled, then sampling is indicated.
[0023] like Figure 1 As shown, in this embodiment, the aliased spectral encoding mosaic image is used. With panchromatic image Feature extraction and fusion are performed separately to obtain the initial fused features. Includes: encoding aliased spectra into mosaic images With panchromatic image Features are extracted using independent 3×3 convolutions, and the extracted features are concatenated along the channel dimension. Then, a 1×1 convolution is used to fuse the features to obtain the initial fused features. The initial fusion features Input a K-level cascaded multi-scale preserved spatial spectrum self-attention mechanism module (MRSSA), and each level outputs enhanced features. .
[0024] The input features of the Kth stage of the K-level cascaded multiscale preserved spatial spectrum self-attention mechanism module (MRSSA) are: The K-level processing flow can then be represented as: ,in This represents the output feature of the Kth stage. For example... Figure 2 As shown, the enhanced features output by each Multiscale Preservative Spatial Spectrum Self-Attention Mechanism (MRSSA) module are... This will serve as input for the next level module, enabling phased feature enhancement and multi-scale augmentation. For example... Figure 2 As shown, the Multi-Scale Preservative Spatial Spectrum Self-Attention (MRSSA) mechanism module includes a first Preservative Spatial Spectrum Self-Attention (RSSA) module, a downsampling module, a second Preservative Spatial Spectrum Self-Attention (RSSA) module, an upsampling module, a concatenation module, and a 1×1 convolutional layer. The concatenation module concatenates the output features of the first Preservative Spatial Spectrum Self-Attention (RSSA) module and the output features of the upsampling module, and then processes them through a 1×1 convolutional layer to obtain the output features. The MRSSA mechanism module adopts a cascaded dual-scale structure, including K cascaded Preservative Spatial Spectrum Self-Attention (RSSA) modules. The larger-scale branch is used to capture image details and spatial texture information; the smaller-scale branch extracts global contextual dependencies through feature compression and attention modeling. Its overall processing flow can be represented as follows: ; ; ; ; ; ; in, and These represent the outputs of the two RSSA modules, and This corresponds to the features after downsampling and upsampling. This indicates that it is achieved through strided convolution. The operation of double-space downsampling is specifically adopted. Convolutional layer (stride 4) followed by The convolutional layer (stride 2) ultimately yields an overall downsampling factor of 8. Upsampling is then performed using the PixelShuffle (PS) operation. This process requires a 1×1 convolutional layer to expand the channel dimension of the low-resolution features. The K-stage MRRA process ultimately generates K intermediate feature maps. These feature maps are then fed into the reconstruction module for processing.
[0025] like Figure 3 As shown, the Retaining Spatial Spectrum Self-Attention Module (RSSA) includes: The attention module is used to perform attention calculations on the input features to obtain attention features. ; The Multi-Source Cueing State Space Module (MPSS) combines the Mamba architecture with the Transformer mechanism to extract features from input features, obtaining features that fuse long-range spatial dependencies and multi-source cueing information. ; The multiplication module is used to incorporate attention features. Features that integrate long-range spatial dependencies and multi-source cue information Multiply to obtain the output features.
[0026] like Figure 3 As shown, in the Feature-Preserving Spatial-Spectral Self-Attention (RSSA) module, the attention module uses average pooling to reduce spatial resolution, guiding the attention calculation to focus more on the overall structure of each spectral channel, thereby effectively suppressing the influence of local texture noise. Finally, based on the pooled features... Recalculate attention features The step involves performing attention calculations on the input features to obtain attention features. The function expression is: ; ; ; in, Input features The pooling results For average pooling, For pooling size, , and The query matrix, key matrix, and value matrix are obtained by performing average pooling and linear post-processing on the input features, respectively. , and These are the weight matrices corresponding to the query matrix, key matrix, and value matrix, respectively. The softmax activation function is used. for transpose, and These are the weighting coefficients. This represents the Hadamard product, which is element-wise multiplication. As a spectral distance prior, Represents the exponential decay coefficient. It is a non-causal mask matrix. ,in and These represent different spectral channel indices. The element in the i-th row and j-th column of the noncausal mask matrix. This is a coarse similarity prior obtained based on the spectral separation mosaic image and the spectral response matrix. By introducing a spectral distance prior to gradually reduce the focus on long-range information, smaller spectral intervals typically exhibit stronger inter-band correlations, thus achieving a spectral attention mechanism with preservative properties. This represents the exponential decay coefficient, used to control the degree to which attention weights decrease with positional (or temporal) distance in the sequence. This represents the Hadamard product. ,in and These represent different spectral channel indices. This is transformed into a non-causal mask that captures the prior correlation between spectral bands based on their relative positions; the smaller the spectral spacing, the stronger the correlation between the bands. As an optional implementation, different... A priori mask is used to encourage each attention head to focus on the correlation at different spectral distances.
[0027] The Retaining Spatial Spectrum Self-Attention Module (RSSA) introduces a coarse similarity prior based on the spectral separation mosaic image and the spectral response matrix. Coarse similarity prior based on inverse spectral response mapping This can further enhance the representation of overall spectral characteristics. For example... Figure 4 As shown, the computational function expression for the coarse similarity prior obtained based on the spectral separation mosaic image and the spectral response matrix is as follows: ; ; ; in, To separate mosaic images and spectral response matrices based on spectral density The obtained coarse similarity prior, for transpose, This is a low-resolution hyperspectral image obtained by inverse matrix mapping. Encoding mosaic images from aliased spectra The spectral separation mosaic image obtained from it, For spectral separation operation, Spectral response matrix The inverse matrix, for Pooling characteristics, For linear layers, For average pooling, Pooling size; spectral response matrix This is an n x m matrix, where n represents different spectral channels, m represents different wavelengths, and each element represents the sensor's response intensity to incident light at a specific wavelength. Referring to the formula above, it can be seen that the aliased spectral encoding mosaic image is first performed. Spectral separation processing is performed, ignoring the influence of mosaic sampling and treating it as a fully sampled spectral image, by applying the spectral response matrix. The inverse matrix can be used to obtain coarse, low-resolution hyperspectral images (i.e., From this coarse, low-resolution hyperspectral image, spectral priors are derived. .
[0028] The Reservative Spatial Spectrum Self-Attention Module (RSSA) combines the Mamba structure with the Transformer mechanism, introduces a Multi-Prompt State Space (MPSS) module, and integrates it into the attention mechanism to enhance the network model's ability to perceive and represent spatial structure and details. Figure 5 This embodiment incorporates features that integrate long-range spatial dependencies and multi-source cue information. The diagram illustrates the generation principle, which involves combining the Mamba structure with the Transformer mechanism to extract features from the input features, thereby obtaining features that fuse long-range spatial dependencies and multi-source cue information. include: The spatial error map is used to generate a memory selection matrix using an error-guided selector. The error-guided selector includes multiple convolutional layers, and the output of each convolutional layer is equipped with a LeakyReLU activation function; for example... Figure 5 As shown, the spatial error map is generated by processing the input features through a linear layer and then comparing them with the panchromatic image. Subtraction yields the result, which can be expressed as: ; ; in, For the output features of the linear layer, As input features, For the output weights of the linear layer, This is a spatial error map. The input is a panchromatic image; based on this, a memory selection matrix can be generated from the spatial error map using an error-guided selector. It is worth noting that the spatial error map The memory selection matrix is shared across different scales within the same Multi-Scale Preservative Spatial Spectrum Self-Attention Mechanism (MRSSA) module and downsampled using bilinear interpolation to ensure consistent spatial resolution across the corresponding scales. The spatial error map is then used to generate a memory selection matrix. It can be represented as: ; in, For error-guided selectors, such as Figure 5 As shown, the error-guided selector consists of a series of convolutional layers, activation function layers (LeakyReLU activation function), and another convolutional layer followed by an activation function layer (LeakyReLU activation function). This utilizes stage-shared memory features. Stage-specific memory characteristics Multiply to obtain segmented memory features Among them, the characteristics of stage-shared memory Stage-specific memory features are learnable hyperparameters shared by all multi-scale preserved spatial spectral self-attention mechanism modules (MRSSA) at each level. Each level of the Multi-Scale Preservative Spatial Spectrum Self-Attention Mechanism (MRSSA) module has its own learnable hyperparameters, and the stages share memory features. Stage-specific memory characteristics During training, all results are initialized with one-to-one values and updated via backpropagation to obtain the final result; segmented memory features. This can be viewed as a form of statistical information, memorizing common features of different images during training and providing prior knowledge of data distribution during inference; features are memorized in segments. It can be decomposed into stage-shared memory features Stage-specific memory characteristics : ,in This represents stage-specific memory features, which are maintained independently in different modules; and The representative stages share memory features, which are used by all modules. This decomposition facilitates interaction between different modules, reduces parameter redundancy, and enhances the expressiveness and flexibility of the memory mechanism; Memory selection matrix and segmented memory features Multiplication yields specific error indications , can be represented as: ; This error-aware mechanism intelligently filters out suitable prompts from the memory feature library.
[0029] Specific error prompts The input features are fed into the Mamba state-space model to obtain features that fuse long-range spatial dependencies and multi-source cue information. The functional expression of the traditional Mamba state-space model is: ; This embodiment improves upon the traditional Mamba state-space model, which includes specific error indicators. Add to output mapping model parameters To generate the output features at time k from the input features at time k, the functional expression of the Mamba state-space model is: ; in, The output feature at time k is... These are the output mapping model parameters, used to map the hidden states to the output feature space; This is the hidden state vector of the system at the current time k, used to represent the dynamic representation of the input features in the state space; These are skip connection terms, used to introduce input features or bias terms to enhance the expressiveness and stability of the model; The input features are for the current time k; ; These are the parameters of the state transition model, used to characterize the evolution of the hidden states between adjacent time steps. For the previous time k The system hidden state vector corresponding to 1, These are the parameters of the input mapping model, used to map the current input features to the state space and inject the hidden state update process.
[0030] In this embodiment, the final K-level enhanced features are concatenated along the channel dimension and then refined to obtain the refined features. In this context, the feature refinement process refers to the refinement of features through a 1×1 convolution operation, thereby compressing and fusing channel information, effectively reducing redundant features and enhancing the expressive power of key information. The refined features can be used as input to the subsequent feature fusion module or reconstruction module to improve the consistency of the overall feature representation and the reconstruction quality.
[0031] In this embodiment, the refined features will be... With initial fusion features After summing the residuals, the data is input into the reconstruction module to generate a broadband hyperspectral image. At that time, the reconstruction module consists of a 1×1 convolutional layer, which is used to integrate the channel information of the fused features and restore the corresponding spectral distribution to generate a broadband hyperspectral image. .
[0032] The method in this embodiment is based on the complementary spatial-spectral information of a single-frame panchromatic (PAN) image and an aliased coded mosaic (ASM) image to achieve high-quality spatial-spectral fusion and spectral super-resolution reconstruction. It has the following advantages: (1) Significantly enhanced spectral feature expression capability: The method in this embodiment combines spectral distance prior with coarse-grained similarity prior by designing a multi-scale preserved spatial-spectral self-attention mechanism module (MRSSA), which realizes effective modeling of fine spectral dependencies. It can enhance local spectral discrimination capability while maintaining global consistency, thereby significantly improving the expression and reconstruction accuracy of spectral features; (2) More efficient and accurate spatial structure modeling: The multi-source cueing state space module (MPSS) proposed in this embodiment introduces spatial prior and multi-source cueing mechanism in the Mamba structure. It achieves adaptive optimization of spatial features through an error-guided selector, which can achieve more accurate spatial information reconstruction for complex scenes and structural regions, and improve the accuracy and robustness of overall fusion imaging. In order to verify the effectiveness of the method in this embodiment, this embodiment conducted experimental verification on three mainstream public datasets: CAVE, KAIST and ICVL. The CAVE dataset contains 32 indoor hyperspectral images, each consisting of 31 spectral bands ranging from 400 nm to 700 nm in 10 nm intervals, with a spatial resolution of 512 × 512 pixels. The first 20 images in this dataset were used as the training set, and the remaining 12 were used as the test set. The KAIST dataset contains 30 hyperspectral images, each providing 31 spectral channels ranging from 420 nm to 720 nm in 10 nm intervals, with a spatial resolution of 2704 × 3376 pixels. The first 20 images in this dataset were used for training, and the remaining 10 were used for testing. The ICVL dataset contains 200 hyperspectral images, each with a spatial resolution of 1392 × 1300 pixels. The 31 spectral bands are uniformly sampled within the range of 400 nm to 700 nm, with 10 nm intervals. Forty images were randomly selected from the ICVL dataset as the training set, and 20 as the test set. To evaluate the quality of the reconstructed broadband high-resolution hyperspectral images, six widely used quantitative metrics were employed: Peak Signal-to-Noise Ratio (PSNR), Root Mean Square Error (RMSE), Synthetic Relative Dimensionless Global Error (ERGAS), Spectral Angle Error (SAM), Structural Similarity Index (SSIM), and Distortion (DD). PSNR measures the ratio between the maximum possible power of the signal and the noise power affecting accuracy. It reflects reconstruction quality by evaluating the ratio between the original signal (i.e., the reference image) and the noise introduced during reconstruction; a higher PSNR value indicates better fidelity. RMSE is used to calculate the average level of prediction error; a lower value indicates smaller reconstruction error. ERGAS is a comprehensive metric used to evaluate the overall error level of the reconstruction results.This index quantifies the global consistency between the reconstructed image and the reference image by calculating the relative errors of each spectral band and considering spatial resolution. A lower ERGAS value indicates that the reconstructed image is closer to the reference image in both overall spectral and spatial dimensions, with smaller errors. SAM measures spectral similarity by calculating the angle between the corresponding spectral vectors of the reference and reconstructed images; a smaller value indicates higher spectral consistency. SSIM is a perceptual evaluation index based on the human visual system, used to measure the similarity between two images in terms of brightness, contrast, and structural information. This index ranges from 0 to 1; a value closer to 1 indicates higher image fidelity and perceptual quality. DD directly reflects the degree of image distortion; a lower value indicates less distortion. Spectral distortion between the reconstructed and reference images is assessed by calculating the spectral change based on the difference in average gray values between the two images; a larger value indicates greater spectral distortion. Several representative hyperspectral fusion imaging methods were selected for comparative experiments, including OcTree implicit adaptive sampling OTIAS (2025), dual spatial-spectral pyramid network DSP (2025), similarity-guided graph attention and variational autoencoder-transformer network CSGAV (2025), LRTN (2025), multi-input multi-output spatial-spectral transformer method MIMO (2024), and mutually guided spatial-spectral expansion network SMGU-Net (2025). Tables 1, 2, and 3 show the results of the comparative experiments on the CAVE, KAIST, and ICVL datasets, respectively.
[0033] Table 1: Experimental results on the CAVE dataset
[0034] Table 2: Experimental results on the KAIST dataset
[0035] Table 3: Experimental results on the ICVL dataset
[0036] The experimental results in Tables 1, 2, and 3 demonstrate that this embodiment consistently outperforms the comparative methods on most evaluation metrics of the CAVE, KAIST, and ICVL datasets. Specifically, the proposed method achieves superior reconstruction performance compared to existing techniques on the CAVE, KAIST, and ICVL datasets. On the CAVE dataset, this method leads in all six evaluation metrics: compared to the second-best performing method, it improves PSNR by approximately 1.11 dB, reduces RMSE and SAM by approximately 0.25 and 0.92 respectively, while achieving lower ERGAS and DD and higher SSIM, indicating superior performance in spatial detail preservation and spectral reconstruction accuracy. On the KAIST dataset, this method outperforms the comparative methods in all six metrics, with improved overall spatial fidelity and spectral recovery accuracy, exhibiting the highest PSNR and SSIM and the lowest SAM, RMSE, ERGAS, and DD. On the ICVL dataset, the method proposed in this embodiment achieves the best performance in most metrics, with a PSNR of 53.49 dB, which is about 0.62 dB higher than the second best method. The SSIM is consistent with the comparison method, while SAM, RMSE, ERGAS and DD all achieve lower values. The results demonstrate that the method proposed in this embodiment has high imaging quality and reliability in broadband high-resolution hyperspectral fusion imaging tasks.
[0037] To further evaluate the model size of the method proposed in this embodiment, the number of parameters for each comparative method was statistically analyzed. The results show that the method proposed in this embodiment achieves the best reconstruction results despite a lower number of parameters. This model contains approximately 585,700 parameters, which is about 90.3% less than the highest parameter count of DSP (approximately 6,052,500 parameters) and about 22.8% less than the second lowest parameter count of SMGU-Net (approximately 758,700 parameters). These results demonstrate that the method proposed in this embodiment maintains high reconstruction performance while significantly reducing model size. It is evident that the method in this embodiment obtains a high-resolution hyperspectral image covering a wide spectral range by fusing a panchromatic image with an aliased spectral mosaic image. The key to this method is the use of a multi-scale preserving spatial-spectral self-attention mechanism module (MRSSA). The MRSSA module combines the Mamba structure with the Transformer mechanism through a multi-source cueing state space module (MPSS) to extract features from the input features, obtaining features that fuse long-range spatial dependencies and multi-source cueing information. The method combines spectral distance priors to capture fine-grained correlations between adjacent bands and utilizes spectral response similarity priors to enhance global spectral feature representation. The Multi-Source Cueing State Space Module (MPSS) introduces multi-source cueing information into the state space modeling process to maintain spatial structure constraints during the inference phase and optimizes the global spatial representation through a spatial error guidance mechanism. Through this design, the method in this embodiment can achieve high-precision, high-fidelity broadband high-resolution hyperspectral image reconstruction in a single exposure, based on fully utilizing the spatial-spectral complementary information of multi-source images.
[0038] Those skilled in the art will understand that the technical solutions provided by this invention can take the form of methods, systems, or computer program products. For example, this application can provide a broadband snapshot hyperspectral fusion imaging system, including a microprocessor and a memory interconnected, wherein the microprocessor is programmed or configured to execute the broadband snapshot hyperspectral fusion imaging method. This application can provide a computer-readable storage medium storing a computer program or instructions programmed or configured to execute the broadband snapshot hyperspectral fusion imaging method via a processor. This invention can provide a computer program product including a computer program or instructions programmed or configured to execute the broadband snapshot hyperspectral fusion imaging method via a processor. Furthermore, this invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this invention can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of a flowchart and / or block diagram, and combinations of blocks in a flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0039] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A broadband snapshot-type hyperspectral fusion imaging method, characterized in that, The process includes the following steps: inputting a hyperspectral image An aliased spectral-coded mosaic image is generated through the combined effects of spectral aliasing and mosaic sampling. aliased spectral encoding mosaic image Compared with the input panchromatic image Feature extraction and fusion are performed separately to obtain the initial fused features. ; Initial fusion features The input is a K-level cascaded multi-scale preserving spatial-spectral self-attention mechanism module (MRSSA). Each level of the MRSSA module then integrates the input features with the panchromatic image. The spatial error map is combined to extract enhancement features, and the final K-level enhancement features are concatenated along the channel dimension and refined to obtain refined features. ; refine features With initial fusion features After summing the residuals, the data is input into the reconstruction module to generate a broadband hyperspectral image. The spatial error map is obtained by processing the input features through a linear layer and then comparing them with a panchromatic image. Subtraction yields that the Multi-Scale Preservative Spatial Spectrum Self-Attention Mechanism (MRSSA) module includes a first Preservative Spatial Spectrum Self-Attention Module (RSSA), a downsampling module, a second Preservative Spatial Spectrum Self-Attention Module (RSSA), an upsampling module, a concatenation module, and a 1×1 convolutional layer. The concatenation module concatenates the output features of the first Preservative Spatial Spectrum Self-Attention Module (RSSA) and the output features of the upsampling module, and then processes them through a 1×1 convolutional layer to obtain the output features.
2. The broadband snapshot hyperspectral fusion imaging method according to claim 1, characterized in that, The input hyperspectral image An aliased spectral-coded mosaic image is generated through the combined effects of spectral aliasing and mosaic sampling. The function expression is: ; in, For the number of bands, For the input hyperspectral image, This represents the Hadamard product. and The first The spectral response function and mosaic sampling mask corresponding to each band For the number of channels, The height and width of the hyperspectral image; the aliased spectral encoding mosaic image With panchromatic image Feature extraction and fusion are performed separately to obtain the initial fused features. Includes: encoding aliased spectra into mosaic images With panchromatic image Features are extracted using independent 3×3 convolutions, and the extracted features are concatenated along the channel dimension. Then, a 1×1 convolution is used to fuse the features to obtain the initial fused features. .
3. The broadband snapshot hyperspectral fusion imaging method according to claim 1, characterized in that, The Retaining Spatial Spectrum Self-Attention Module (RSSA) includes: The attention module is used to perform attention calculations on the input features to obtain attention features. ; The Multi-Source Cueing State Space Module (MPSS) combines the Mamba architecture with the Transformer mechanism to extract features from input features, obtaining features that fuse long-range spatial dependencies and multi-source cueing information. ; The multiplication module is used to incorporate attention features. Features that integrate long-range spatial dependencies and multi-source cue information Multiply to obtain the output features.
4. The broadband snapshot hyperspectral fusion imaging method according to claim 3, characterized in that, Attention features are obtained by performing attention calculation on the input features. The function expression is: ; ; ; in, Input features The pooling results For average pooling, For pooling size, , and The query matrix, key matrix, and value matrix are obtained by performing average pooling and linear post-processing on the input features, respectively. , and These are the weight matrices corresponding to the query matrix, key matrix, and value matrix, respectively. The softmax activation function is used. for transpose, and These are the weighting coefficients. This represents the Hadamard product. As a spectral distance prior, Represents the exponential decay coefficient. It is a non-causal mask matrix. ,in and These represent different spectral channel indices. The element in the i-th row and j-th column of the noncausal mask matrix. This is a coarse similarity prior obtained based on the spectral separation mosaic image and the spectral response matrix.
5. The broadband snapshot hyperspectral fusion imaging method according to claim 4, characterized in that, The calculation function expression for the coarse similarity prior obtained based on the spectral separation mosaic image and the spectral response matrix is as follows: ; ; ; in, To separate mosaic images and spectral response matrices based on spectral density The obtained coarse similarity prior, for transpose, This is a low-resolution hyperspectral image obtained by inverse matrix mapping. Encoding mosaic images from aliased spectra The spectral separation mosaic image obtained from it, For spectral separation operation, Spectral response matrix The inverse matrix, for Pooling characteristics, For linear layers, For average pooling, Pooling size; spectral response matrix It is an n x m matrix, where n represents different spectral channels, m represents different wavelengths, and each element represents the sensor's response intensity to incident light at a specific wavelength.
6. The broadband snapshot hyperspectral fusion imaging method according to claim 3, characterized in that, The method combines the Mamba structure with the Transformer mechanism to extract features from the input features, thereby obtaining features that integrate long-range spatial dependencies and multi-source cue information. include: The spatial error map is used to generate a memory selection matrix using an error-guided selector. The error-guided selector includes multiple convolutional layers, and the output of each convolutional layer is equipped with a LeakyReLU activation function. Stage-shared memory features Stage-specific memory characteristics Multiply to obtain segmented memory features Among them, the characteristics of stage-shared memory Stage-specific memory features are learnable hyperparameters shared by all multi-scale preserved spatial spectral self-attention mechanism modules (MRSSA) at each level. Each level of the Multi-Scale Preservative Spatial Spectrum Self-Attention Mechanism (MRSSA) module has its own learnable hyperparameters, and the stages share memory features. Stage-specific memory characteristics During training, all results are initialized with one and updated via backpropagation to obtain the final result. Memory selection matrix and segmented memory features Multiplication yields specific error indications ; Specific error prompts The input features are fed into the Mamba state-space model to obtain features that fuse long-range spatial dependencies and multi-source cue information. The Mamba state-space model will provide specific error hints. Add to output mapping model parameters To generate the output features at time k from the input features at time k, the functional expression of the Mamba state-space model is: ; in, The output feature at time k is... These are the output mapping model parameters, used to map the hidden states to the output feature space; This is the hidden state vector of the system at the current time k, used to represent the dynamic representation of the input features in the state space; These are skip connection terms, used to introduce input features or bias terms to enhance the expressiveness and stability of the model; The input features are for the current time k; ; These are the parameters of the state transition model, used to characterize the evolution of the hidden states between adjacent time steps. For the previous time k The system hidden state vector corresponding to 1, These are the parameters of the input mapping model, used to map the current input features to the state space and inject the hidden state update process.
7. The broadband snapshot hyperspectral fusion imaging method according to claim 1, characterized in that, The final K-level enhanced features are then concatenated along the channel dimension and refined to obtain refined features. In this context, the feature refinement process refers to achieving feature refinement through a 1×1 convolution operation; the refinement of features... With initial fusion features After summing the residuals, the data is input into the reconstruction module to generate a broadband hyperspectral image. At that time, the reconstruction module consists of a 1×1 convolutional layer, which is used to integrate and fuse the channel information of the features and restore the corresponding spectral distribution to generate a broadband hyperspectral image. .
8. A broadband snapshot-type hyperspectral fusion imaging system, comprising a microprocessor and a memory interconnected, characterized in that, The microprocessor is programmed or configured to perform the broadband snapshot hyperspectral fusion imaging method according to any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program or instructions, characterized in that, The computer program or instructions are programmed or configured to execute the broadband snapshot hyperspectral fusion imaging method of any one of claims 1 to 7 via a processor.
10. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instructions are programmed or configured to execute the broadband snapshot hyperspectral fusion imaging method of any one of claims 1 to 7 via a processor.