Hyperspectral image reconstruction method and application thereof

CN122473004BActive Publication Date: 2026-09-04WESTLAKE UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610916434.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-04
Estimated Expiration
2046-06-24

AI Technical Summary

Technical Problem

[0004]本发明实施例提供了一种高光谱图像重建方法及其应用,针对现有快照式高光谱重建方案仅依赖缺乏先验信息的单模态观测进行反演,且在直接展开的序列化建模过程中破坏了高维数据的空间结构邻接性与谱间响应连续性,导致重建结果易出现边缘模糊、纹理缺失及局部光谱失真等问题

Benefits of technology

[0017]本发明的主要贡献和创新点如下:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122473004B_ABST
    Figure CN122473004B_ABST
Patent Text Reader

Abstract

The application provides a hyperspectral image reconstruction method and application thereof, and belongs to the technical field of hyperspectral compressive sensing imaging and image reconstruction. In order to solve the problems of fuzzy spatial details and low spectral fidelity in the existing method, the two-dimensional compressive measurement image and the auxiliary RGB image are synchronously acquired by a double-path snapshot compressive imaging device, the initial hyperspectral estimation is fused with the RGB image to construct a multi-modal input, and the input is input into a reconstruction network comprising an encoder and a decoder; a multi-view two-dimensional slice scanning module in the network recombines three-dimensional features into a two-dimensional sequence retaining spatial and spectral continuity and inputs the two-dimensional sequence into a state space sequence model for global modeling; a three-dimensional local patch slice scanning module divides a feature map into local cubic subblocks and repeatedly scans the local cubic subblocks to enhance local details, so that a global and local collaborative reconstruction mechanism is formed. The application is used for high-quality reconstruction of hyperspectral images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hyperspectral image compression sensing and computational reconstruction technology, and in particular to a hyperspectral image reconstruction method and its application. Background Technology

[0002] Hyperspectral imaging, by acquiring spatial and spectral information of a scene across dozens to hundreds of consecutive spectral bands, is widely used in remote sensing, medical diagnosis, agricultural monitoring, and industrial sorting. Traditional hyperspectral image acquisition relies on point-by-point scanning or push-broom methods, which suffer from long acquisition times and difficulty in adapting to dynamic scene changes. To improve imaging speed, snapshot hyperspectral imaging technology has become a research hotspot, with a typical implementation being the coded aperture snapshot spectral imaging system. This system introduces a coded mask and dispersive elements into the optical imaging link, compressing and encoding a three-dimensional hyperspectral data cube into a two-dimensional measurement image through a single exposure, and then using computational reconstruction algorithms to reconstruct the complete hyperspectral image.

[0003] Currently, mainstream reconstruction methods include optimization iterative algorithms based on sparse priors and end-to-end reconstruction networks based on deep learning. Methods based on Transformers or convolutional neural networks improve reconstruction accuracy by establishing a joint spatial-spectral mapping, but they still have limitations in restoring spatial texture details and spectral continuity under high compression ratios due to limitations in receptive field size or computational complexity. Other methods attempt to introduce sequence modeling structures to process hyperspectral data, unfolding 3D images into one-dimensional sequences before feeding them into state-space models or recurrent neural networks to capture long-range dependencies. However, simple sequence unfolding operations can disrupt the inherent spatial adjacency and spectral dimensional continuity of the original hyperspectral data, leading to problems such as blurred edges and distorted spectral curves in the reconstruction results. Summary of the Invention

[0004] This invention provides a hyperspectral image reconstruction method and its application. It addresses the problem that existing snapshot-based hyperspectral reconstruction schemes rely solely on single-modal observations lacking prior information for inversion. Furthermore, the direct unfolding of the sequential modeling process disrupts the spatial structure adjacency and spectral response continuity of high-dimensional data, leading to problems such as blurred edges, missing textures, and local spectral distortion in the reconstruction results.

[0005] The core technology of this invention is to construct a multimodal fusion input through dual-channel synchronous acquisition, and to adopt a progressive sequence modeling of multi-view two-dimensional slice scanning and three-dimensional local block slice scanning in the reconstruction network to achieve the synergy of global spatial spectral dependency modeling and local fine feature refinement enhancement.

[0006] In a first aspect, the present invention provides a hyperspectral image reconstruction method, the method comprising the following steps:

[0007] Acquire a two-dimensional compressed measurement image of the target scene and an auxiliary RGB image corresponding to the two-dimensional compressed measurement image space; A three-dimensional initial estimate hyperspectral image is generated from the two-dimensional compressed measurement image, and the three-dimensional initial estimate hyperspectral image is fused with the auxiliary RGB image to construct a multimodal input feature map; The multimodal input feature map is fed into a trained hyperspectral reconstruction network, where the following operations are performed: The multi-view 2D slice scanning module reassembles the input 3D feature map into a 2D sequence that includes at least a spatial main view slice and a spatial dimension coupled with a spectral dimension view slice. The state space sequence model is then used to model the global dependency relationship of the reassembled sequence. The feature map processed by the multi-view two-dimensional slice scanning module is divided into multiple local cubic sub-blocks through the three-dimensional local block slicing scanning module. The multi-view two-dimensional slice scanning operation is repeated in each local cubic sub-block to refine and enhance the local spatial and spectral features. Spatial resolution is restored based on the modeled features, and the reconstructed hyperspectral image is output.

[0008] Furthermore, the state-space sequence model is a Mamba module; the multi-view two-dimensional slice scanning module is specifically used for: The input feature map is sliced ​​along the channel dimension, width dimension, and height dimension respectively, generating a frontal slice representing the structural distribution in the spatial plane, a hyperspectral slice representing the joint correlation between the height dimension and the spectral dimension, and a broadband slice representing the joint correlation between the width dimension and the spectral dimension. Each slice is expanded into a one-dimensional sequence and input into the Mamba module for sequence modeling. The sequence features output by the Mamba module are reshaped into the original shape of the slices and then spliced ​​and fused.

[0009] Furthermore, the 3D local block slicing scanning module is specifically used for: The input feature map is divided into local cubic sub-blocks of multiple scales according to the spatial and spectral dimensions. The multiple scales include the basic block size and its proportionally scaled block size. Perform a multi-view 2D slice scan operation on each local cube sub-block; The processing results of all local cubic sub-blocks are recombined into an enhanced feature map with the same size as the input feature map, and residual connections are made with the input feature map.

[0010] Furthermore, the basic block size of the local cubic sub-block is determined by the height, width, and number of channels of the input feature map according to a preset block coefficient.

[0011] Furthermore, the hyperspectral reconstruction network adopts an encoder-decoder structure, with multi-view two-dimensional slice scanning modules and three-dimensional local block slice scanning modules respectively in the multiple downsampling stages of the encoder and the multiple upsampling stages of the decoder; and the modules in different stages process feature maps of corresponding scales to form a progressive global and local joint modeling mechanism.

[0012] Further, acquiring a two-dimensional compressed measurement image of the target scene and an auxiliary RGB image corresponding to the two-dimensional compressed measurement image space includes: Two-dimensional compressed measurement images and auxiliary RGB images are acquired simultaneously using a dual-path snapshot-type compressed imaging device. The dual-path snapshot-type compressed imaging device includes a compressed measurement acquisition optical path and an auxiliary imaging acquisition optical path. The compressed measurement acquisition optical path includes an encoding mask and a dispersive element, which are used to compress and encode the three-dimensional hyperspectral image of the target scene into a two-dimensional observation image. The auxiliary imaging acquisition optical path is used to record the spatial texture and edge structure information of the scene.

[0013] Furthermore, a three-dimensional initial estimated hyperspectral image is generated from the two-dimensional compressed measurement image, including: By combining the calibration parameters of the imaging system, the dispersion offset, and the mask response information, the two-dimensional compressed measurement image is converted into a three-dimensional initial estimated hyperspectral image through back projection, shift inverse mapping, or physical prior initialization methods.

[0014] In a second aspect, the present invention provides a hyperspectral image reconstruction apparatus, comprising: The dual-channel acquisition module is used to acquire a two-dimensional compressed measurement image of the target scene and an auxiliary RGB image corresponding to the two-dimensional compressed measurement image space. The initialization and fusion module is used to generate a three-dimensional initial estimate hyperspectral image based on the two-dimensional compressed measurement image, and to fuse the three-dimensional initial estimate hyperspectral image with the auxiliary RGB image to construct a multimodal input feature map; The hyperspectral reconstruction network module includes: The multi-view 2D slice scanning submodule is used to reassemble the input 3D feature map into a 2D sequence that includes at least a spatial main view slice and a spatial dimension coupled with a spectral dimension view slice, and to model the global dependency relationship of the reassembled sequence through a state space sequence model. The 3D local block slicing scanning submodule is used to divide the feature map processed by the multi-view 2D slicing scanning submodule into multiple local cube sub-blocks, and repeatedly perform the multi-view 2D slicing scanning operation within each local cube sub-block to refine and enhance local spatial and spectral features. The resolution restoration submodule is used to restore spatial resolution based on the modeled features and output a reconstructed hyperspectral image.

[0015] Thirdly, the present invention provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the above-described hyperspectral image reconstruction method.

[0016] Fourthly, the present invention provides a readable storage medium storing a computer program, the computer program including program code for controlling a process to execute the process, the process including the hyperspectral image reconstruction method described above.

[0017] The main contributions and innovations of this invention are as follows: 1. By employing a dual-path acquisition and multimodal feature fusion mechanism, this invention breaks through the information bottleneck of traditional snapshot imaging, which relies solely on single two-dimensional spectral coding observations. Utilizing the high-fidelity spatial texture, edge contours, and structural prior information provided by auxiliary RGB images, it effectively compensates for the high-frequency spatial details lost during projection in spectral compression measurements, greatly improving the information completeness and modal complementarity of the reconstruction network input.

[0018] 2. By introducing a multi-view 2D slice scanning mechanism, this invention avoids the physical structure fragmentation caused by the unidirectional unfolding of 3D hyperspectral features in traditional sequential modeling. By extracting slice sequences from the main spatial viewpoint and two types of spatial-spectral coupled viewpoints respectively, the network can explicitly model long-distance spatial dependence and correlation between co-located continuous spectra while maintaining the same spatial positional correspondence, thereby achieving spatial and spectral consistency representation of high-dimensional data within the global receptive field.

[0019] 3. By constructing a three-dimensional local block slicing scanning mechanism, this invention further sinks global features to local three-dimensional neighborhoods at multiple scales. Repeatedly performing recombination modeling in local spatial-spectral sub-regions at different network levels can significantly enhance the network's perception of the adjacency relationships between pixels in local heterogeneous regions and subtle spectral line change trends, effectively improving the ability to restore fine-grained textures and local spectral abrupt changes in complex scenes.

[0020] 4. By forming a hierarchical and progressive collaborative architecture of "first global reconstruction and modeling, then local refinement and enhancement," this invention balances overall consistency across regions and bands with refined characterization of local multi-scale features. This collaborative mechanism enables the invention to maintain excellent spatial texture fidelity and spectral fidelity even when facing high compression ratio challenges, effectively improving the overall imaging accuracy and robustness of the snapshot-type hyperspectral compression reconstruction system.

[0021] Details of one or more embodiments of the present invention are set forth in the following drawings and description, so that other features, objects and advantages of the invention will be more readily understood. Attached Figure Description

[0022] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a flowchart of a multimodal hyperspectral reconstruction method based on sliced ​​and segmented Mamba according to an embodiment of the present invention; Figure 2 A flowchart illustrating the specific scanning process of the Mamba sequence after slicing and segmentation according to an embodiment of the present invention; Figure 3 This is an apparatus for a multimodal hyperspectral reconstruction method based on sliced ​​and segmented Mamba according to an embodiment of the present invention; Figure 4 MST simulation was used to reconstruct the synthetic image and spectral curves; Figure 5 To simulate and reconstruct synthetic images and spectral curves according to embodiments of the present invention; Figure 6 Reconstruct synthetic images and spectral curves for real-world scenes using SSABNet with 100 channels; Figure 7 To reconstruct synthetic images and spectral curves of a 100-channel real scene according to embodiments of the present invention; Figure 8 Synthesized images and spectral curves for real-world scene reconstruction using a 100-channel hyperspectral line scan.

[0023] Figure 9 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0024] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0025] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in this specification in other embodiments. In some other embodiments, the methods may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; and multiple steps described in this specification may be combined into a single step in other embodiments.

[0026] This invention provides a multimodal hyperspectral reconstruction method based on sliced ​​and block-based Mamba and its application. The method simultaneously acquires two-dimensional compressed measurement images and auxiliary RGB images using a dual-path snapshot-type compressed imaging device. The initial hyperspectral estimate is fused with the RGB images to construct a multimodal input, which is then input into a hyperspectral reconstruction network containing an encoder and a decoder. Within the network, a multi-view two-dimensional slice scanning module reconstructs three-dimensional features into a two-dimensional sequence that preserves spatial neighborhood continuity and isotopic spectral continuity, and inputs this sequence into a state-space sequence model for global modeling. A three-dimensional local block-based slice scanning module divides the feature map into local cubic sub-blocks and repeatedly slices and scans to enhance local details, forming a global and local collaborative reconstruction mechanism.

[0027] Example 1 like Figure 1 As shown, the multimodal hyperspectral reconstruction method based on sliced ​​and segmented Mamba provided in this embodiment includes the following steps: Step S1: Simultaneously acquire the target scene using a dual-channel snapshot compression imaging device to obtain a two-dimensional compressed measurement image Meas and an auxiliary RGB image I.

[0028] Figure 3 The optical path structure and working principle of the dual-path snapshot-type compressed imaging device of the present invention are illustrated. Light emitted from the target scene first enters the system through the camera lens 30, and is then split into two beams by the beam splitter 31. The first beam enters the RGB camera 32, forming an auxiliary imaging acquisition optical path for real-time recording of visible light image information of the scene and acquiring an auxiliary RGB image. The second beam enters the compressed measurement acquisition optical path composed of an encoding aperture 33, a first relay lens 34, a prism 35, a second relay lens 36, and a grayscale camera 37. The encoding aperture 34 spatially modulates the scene's light field, and the prism 35 performs dispersion shifting on different wavelength components, ultimately forming a two-dimensional compressed measurement image on the target surface of the grayscale camera 37. Since the two light signals originate from incident light at the same time and in the same field of view, they are synchronously imaged after beam splitting, allowing simultaneous acquisition of the spatially corresponding auxiliary RGB image and the two-dimensional compressed measurement image. The auxiliary RGB image is used to provide prior information on spatial texture, edge contours, and structure for hyperspectral reconstruction. Since the two optical signals originate from the incident light at the same time and in the same field of view, they are imaged synchronously after beam splitting, allowing for the simultaneous acquisition of spatially corresponding RGB images and compressed measurement images.

[0029] The imaging process of the compressed measurement acquisition optical path can be further described as follows: the hyperspectral image is regarded as a tensor form X∈ H×W×C Where H and W are spatial dimensions, and C is the number of spectral channels. First, an encoding mask M∈ H×WSpatial modulation is applied to each band to obtain a mask-modulated image X'=X·M. After passing through a dispersive element, each band image generates a channel-dependent displacement d along the row direction. c The modulated images of each band are superimposed in a staggered manner to generate a two-dimensional observation image Y∈ H×W Spatial location A point can be represented as:

[0030] Where c is the channel index of the current spectral band, and its value ranges from 1 to C. To reduce sensor noise. The auxiliary imaging acquisition optical path synchronously acquires the corresponding RGB image. .

[0031] Step S2: Obtain the initial three-dimensional hyperspectral image based on the two-dimensional compressed measurement image Meas, and stitch the initial three-dimensional hyperspectral image with the auxiliary RGB image I along the channel dimension to construct a multimodal input tensor.

[0032] When using simulated data, existing hyperspectral image data can be used to generate corresponding compressed observation samples through virtual coding mask modulation and dispersion shifting processes. Simultaneously, corresponding auxiliary RGB images can be generated or matched from the hyperspectral images, and training and testing data pairs can be constructed accordingly. When using real-world data, it is necessary to combine system calibration parameters, dispersion shift, mask response information, and necessary initial demixing processes to generate a three-dimensional initial estimated hyperspectral image for subsequent network reconstruction. The initial 3D estimate can be obtained from the 2D observation map through back projection, shift inverse mapping, channel expansion, or physical prior initialization methods, thereby converting the 2D observation domain information into a 3D tensor form suitable for network processing.

[0033] Furthermore, the three-dimensional initial estimated hyperspectral image The network input tensor is constructed by concatenating the auxiliary RGB image I along the channel dimension: Z (0)= Concat ( , ) Here, Concat(·) represents the channel stitching operation. In this way, the coded observation information from the two-dimensional compressed measurement and the high-resolution spatial structure prior provided by the RGB image are jointly introduced into the reconstruction network.

[0034] Step S3: Convert the multimodal input tensor Z (0) Input to a pre-built hyperspectral image reconstruction network, which includes an encoder and a decoder.

[0035] The hyperspectral image reconstruction network employs an encoder-decoder structure. In the encoder's multiple downsampling stages and the decoder's multiple upsampling stages, there are multi-view 2D slice scanning modules and 3D local block slice scanning modules, respectively. Modules at different stages process feature maps of corresponding scales to form a progressive global and local joint modeling mechanism. In each stage of the encoder, the input features first undergo convolution and normalization operations to adjust the number of channels and extract local features, and then spatial downsampling is completed through convolution or pooling operations with a stride of 2. Specifically, the output feature sizes of the first to fourth encoding stages are as follows: Stage 1: Stage 2: Stage 3: Stage 4: .

[0036] In each stage of the encoder, the downsampled features sequentially pass through a multi-view 2D slice scanning module and a 3D local block slice scanning module to achieve progressive joint modeling of spatial information, spectral information, and their coupling relationships at different levels. The multi-view 2D slice scanning module and the 3D local block slice scanning module are not independent of each other or simply connected in series, but form a hierarchical collaborative mechanism of "first global reorganization modeling, then local refinement and enhancement": the former module is used to establish a consistent overall representation across regions and bands, and the latter module is used to supplement local fine structures in local spatial and spectral sub-regions.

[0037] The following provides a detailed explanation of the specific operations of the two core modules mentioned above.

[0038] 1. Multi-view 2D slice scanning module Let the multimodal feature map input to this module be... This module is used to extract three types of two-dimensional slices: Frontal slice (primary spatial view): sliced ​​according to channel dimension C There are C sheets in total, used to characterize the structural distribution features within a spatial plane.

[0039] Hyperspectral slices (a coupled perspective of spatial and spectral dimensions): extraction by width dimension A total of W sheets are used to characterize the joint correlation between the height dimension and the spectral dimension.

[0040] Broadband slicing (a coupled perspective of spatial and spectral dimensions): extraction by height dimension There are H sheets in total, used to characterize the joint correlation between the width dimension and the spectral dimension.

[0041] Where k is the current index in the channel dimension, j is the current index in the width dimension, and i is the current index in the height dimension.

[0042] The three types of two-dimensional slices mentioned above do not only expand the spatial plane directionally, but also reorganize the multimodal fusion features from the main spatial perspective and the coupling perspective of spatial and spectral dimensions, respectively. This enables the network to not only capture long-distance dependencies within the spatial neighborhood during sequence modeling, but also to model the interspectral correlations between different bands while maintaining the same spatial positional correspondence.

[0043] like Figure 2 (a) Figure 2 (c) and Figure 2 As shown in (e), each slice is unfolded into a one-dimensional sequence according to a preset order (such as a serpentine scanning order, i.e., alternating reverse scanning from left to right and from top to bottom), and input into a state-space sequence model for sequence modeling to obtain enhanced sequence features. These enhanced features are then reshaped back into the original corresponding two-dimensional slice shape. Here, the state-space sequence model refers to a type of neural network structure that models sequence data based on state-space equations. It captures long-distance dependencies in the sequence through recursive updates of hidden states. Typical implementations of this type of model include, but are not limited to, the Mamba module and structured state-space sequence models. In this embodiment, the Mamba module is preferably used as the specific implementation of the state-space sequence model, but the use of other state-space models with similar sequence modeling capabilities is not excluded.

[0044] The Mamba module is built on a selective state-space model. Its calculation process includes: the input slice of one-dimensional sequence is linearly projected, then sequentially processed by 1D convolution to extract local features and SiLU activation function to introduce nonlinearity; then the sequence propagation is modeled by discretizing the state-space equation, where the state transition matrix A, input matrix B, output matrix C and direct connection matrix D are obtained through structured parameterization, and the input dependency weights are dynamically adjusted by a selective scanning mechanism; finally, the enhanced sequence features are output after linear projection.

[0045] Subsequently, the slice features output from the three directions are concatenated and adaptively fused using 1×1 convolution to obtain a unified output feature. Preferably, the fusion process can also incorporate residual connections to superimpose the fused features with the input features, thereby improving information transmission stability and network training robustness.

[0046] 2. 3D Local Block Slicing Scanning Module like Figure 2 (b) Figure 2 (d) and Figure 2As shown in (f), based on the completion of the above global spatial and spectral coupling modeling, the three-dimensional local block slicing scanning module further divides the feature map into multiple local cubic sub-blocks, and repeatedly performs the multi-view two-dimensional slicing scanning operation within each sub-block, so that the slicing scanning unit is further sinking from the global feature level to the local spatial and spectral sub-regions.

[0047] Specifically, the current feature map is divided into several cubic blocks of fixed size at multiple scales according to spatial and spectral dimensions, including... , , Where h, w, and c represent the basic block sizes of the input feature map in the height, width, and spectral dimensions, respectively, and are determined proportionally by the overall dimensions H, W, and C of the input feature map. = / n , = / n , = / n , where n is the partitioning coefficient. Preferably, the partitioning coefficient n is 2, so that the local sub-blocks maintain an appropriate scale in both spatial and spectral ranges, balancing the ability to represent local details with network computational efficiency.

[0048] Several local sub-blocks were obtained ( , and Afterwards, the multi-view two-dimensional slicing scan operation is repeatedly performed on each local sub-block to extract its frontal slice, hyperspectral slice, and broadband slice, respectively. The slices are then unfolded into one-dimensional sequences and input into the state space sequence model in parallel. Local information is extracted using the state space modeling mechanism, and the output sequences are then reshaped into a cube shape identical to the input. Subsequently, all sub-blocks are reassembled back into their original spatial layout to obtain an enhanced feature map with the same size as the input. This enhanced feature map is then residually fused with the original features as the output of this stage. .

[0049] By employing the aforementioned three-dimensional block scanning method, the original hyperspectral feature map can be divided into multiple local spatial-spectral sub-regions. This allows the network to establish more compact pixel relationships and more stable inter-spectral connections within a local area, thereby enhancing its ability to represent fine-grained textures, edge contours, and local spectral variations. Since the cubic blocks correspond to input features of different scales at different network levels, and the input features at each level contain multimodal information formed by fusing the two-dimensional compressed measurement image Meas with the RGB image, differentiated receptive fields can be formed at different stages of the encoder and decoder. This enables collaborative adaptation to multi-scale targets, complex texture structures, local heterogeneous regions, and non-stationary inter-spectral variations.

[0050] Step S4: The decoder structure is symmetrical to the encoder. Each stage first performs upsampling (e.g., transposed convolution), then fuses it with the high-resolution feature map output from the corresponding encoder stage through skip connections, and finally performs convolution to enhance semantic expressiveness. Subsequently, in each decoding stage, after upsampling, the same multi-view 2D slice scanning module and 3D local block slice scanning module are used for spatial-spectral information modeling. Finally, the decoder outputs a reconstructed hyperspectral image P∈ with the same size as the original input. H×W×C .

[0051] The beneficial effects of the method of the present invention will be explained below with reference to a specific experimental example.

[0052] This experiment uses a single-exposure hyperspectral compression imaging device based on prisms and lenses, with the optical structure as follows: Figure 3 As shown in the figure, the hyperspectral reconstruction method based on slice and block Mamba and the snapshot compression device proposed in this invention are used to achieve rapid compression acquisition and accurate high-dimensional information reconstruction of 100-channel hyperspectral data. The specific steps are as follows: Step 1: Load the CAVE hyperspectral dataset and the KAIST hyperspectral dataset. The CAVE dataset contains 32 sets of hyperspectral images, each with a spatial size of 512×512; the KAIST dataset contains 30 sets of hyperspectral images, each with a spatial size of 2704×3376. The CAVE dataset is used as the training data source, and 10 scenes are selected from the KAIST dataset as the test set to verify the reconstruction performance of the method of this invention under cross-dataset conditions.

[0053] Step 2: Preprocess the loaded hyperspectral images. For the training images in the CAVE dataset, several fixed-size hyperspectral image patches are cropped and used as training samples for the network. Based on the snapshot-style compressed imaging physical model, each hyperspectral image patch undergoes coded mask modulation and dispersion shift compression to generate a corresponding two-dimensional compressed measurement image (Meas). Simultaneously, an auxiliary RGB image corresponding to the hyperspectral image's spatial representation is generated as auxiliary modality input. For the test images in the KAIST dataset, the two-dimensional compressed measurement image (Meas) and auxiliary RGB image are generated in the same manner to construct the test samples.

[0054] Step 3: Generate a three-dimensional initial hyperspectral estimate from the two-dimensional compressed measurement image Meas obtained in Step 2 according to the inversion process of snapshot compressed imaging; and stitch the auxiliary RGB image with the initial hyperspectral estimate in the channel dimension to form multimodal input data.

[0055] Step 4: Feed the multimodal input constructed in Step 3 into the hyperspectral reconstruction network. In each stage of the encoder, shallow features are extracted through convolution and normalization operations, and spatial downsampling is gradually completed through convolution or pooling operations with a stride of 2. In each stage of the decoder, spatial resolution is gradually restored through upsampling and skip connections. In different stages of the encoder and decoder, features at the corresponding scales are input into the multi-view 2D slice scanning module and the 3D local block slice scanning module for modeling, respectively.

[0056] Step 5: Train the multimodal hyperspectral reconstruction network using training set samples. Using the original hyperspectral image as a supervision signal, compare the reconstructed hyperspectral image output by the network with the real hyperspectral image. Constrain the error between the network reconstruction result and the real data using a loss function to optimize the model parameters. During training, stochastic gradient descent or Adam optimization algorithms can be used to update the network parameters.

[0057] Preferably, during the model training phase, the Adam optimizer is used, with an initial learning rate set to 1×10⁻⁶. -4 The hyperspectral image was attenuated using a cosine annealing strategy. The batch size was set to 8, and the training run was 300 epochs. The loss function was defined as the weighted sum of the mean square error loss between the reconstructed hyperspectral image and the real hyperspectral image, and the spectral angle mapping loss, to simultaneously constrain spatial reconstruction accuracy and spectral fidelity. All experiments were implemented using the PyTorch framework and completed on a single NVIDIA RTX 4090 or higher performance GPU.

[0058] Step 6: Load the trained model parameters, perform reconstruction tests on the 2D compressed measurement images (Meas) and auxiliary RGB images in the KAIST test set, output the corresponding hyperspectral reconstruction results, and compare the reconstruction results with the real hyperspectral images in the test set. Calculate the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM) evaluation metrics. The reconstruction evaluation results are shown in Table 1. Table 1

[0059] The results show that the method of this invention achieves PSNR and SSIM of 40.19 and 0.974 respectively in the simulation reconstruction experiment, which are significantly better than the existing MST method (33.56 and 0.921) and the SSABNet method (35.43 and 0.955). Combined with... Figure 4-8It can be further seen that in the simulation scene, the synthetic image reconstructed by the method of the present invention has a clearer contour and a more natural edge transition, and the corresponding spectral curve is closer to the real curve. In the real scene experiment of 100 channels, the reconstruction result of the method of the present invention is closer to the line scan hyperspectral acquisition result than SSABNet. It can not only better restore the texture level and structural details of ground objects, but also fit the reflectance change trend more accurately in the visible light to near-infrared band.

[0060] Example 2 Based on the same concept, the present invention also proposes a hyperspectral image reconstruction apparatus, comprising: The device comprises an auxiliary imaging acquisition optical path, a compressed measurement acquisition optical path, and a reconstruction processor. The auxiliary imaging acquisition optical path is used to acquire an auxiliary RGB image of the target scene; the compressed measurement acquisition optical path includes coded optical components for acquiring a two-dimensional compressed measurement image of the target scene, wherein the two-dimensional compressed measurement image spatially corresponds to the auxiliary RGB image; the reconstruction processor is used to execute the multimodal hyperspectral reconstruction method based on slice and block Mamba as described in any of the above embodiments to reconstruct a hyperspectral image based on the auxiliary RGB image and the two-dimensional compressed measurement image. The specific structure and working principle of this device have been described in detail in the above method embodiments and will not be repeated here.

[0061] Example 3 This embodiment also provides an electronic device, see reference. Figure 9 It includes a memory 404 and a processor 402, wherein the memory 404 stores a computer program and the processor 402 is configured to run the computer program to perform the steps in any of the above method embodiments.

[0062] Specifically, the processor 402 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement embodiments of the present invention.

[0063] Memory 404 may include a mass storage device for data or instructions. For example, and not limitingly, memory 404 may include a hard disk drive (HDD), a floppy disk drive, a solid-state drive (SSD), flash memory, an optical disk drive, a magneto-optical disk drive, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 404 may include removable or non-removable (or fixed) media. Where appropriate, memory 404 may be internal or external to a data processing device. In a particular embodiment, memory 404 is non-volatile memory. In a particular embodiment, memory 404 includes read-only memory (ROM) and random access memory (RAM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable read-only memory (PROM), an erasable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), an electrically alterable read-only memory (EAROM), or flash memory, or a combination of two or more of these. Where appropriate, the RAM can be Static Random-Access Memory (SRAM) or Dynamic Random-Access Memory (DRAM). DRAM can be Fast Page Mode Dynamic Random-Access Memory (FPMDRAM), Extended Data Out Dynamic Random-Access Memory (EDODRAM), Synchronous Dynamic Random-Access Memory (SDRAM), etc.

[0064] The memory 404 can be used to store or cache various data files that need to be processed and / or communicated, as well as possible computer program instructions executed by the processor 402.

[0065] The processor 402 reads and executes computer program instructions stored in the memory 404 to implement any of the hyperspectral image reconstruction methods in the above embodiments.

[0066] Optionally, the electronic device may further include a transmission device 406 and an input / output device 408, wherein the transmission device 406 is connected to the processor 402, and the input / output device 408 is connected to the processor 402.

[0067] The transmission device 406 can be used to receive or send data via a network. Specific examples of the network described above may include wired or wireless networks provided by the communication provider of the electronic device. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 406 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0068] Input / output device 408 is used to input or output information.

[0069] Example 4 This embodiment also provides a readable storage medium storing a computer program, the computer program including program code for controlling a process to execute the process, the process including the hyperspectral image reconstruction method according to Embodiment 1.

[0070] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.

[0071] Generally, various embodiments can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. Some aspects of the invention can be implemented in hardware, while others can be implemented by firmware or software executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Although various aspects of the invention may be shown and described as block diagrams, flowcharts, or using some other graphical representation, it should be understood that, by way of non-limiting example, these blocks, apparatuses, systems, techniques, or methods described herein can be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.

[0072] Embodiments of the present invention can be implemented by computer software, which may be executable by a data processor of a mobile device, such as a processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs (also referred to as program products) including software routines, applets, and / or macros can be stored in any device-readable data storage medium, and they include program instructions for performing specific tasks. The computer program product may include one or more computer-executable components configured to perform the embodiments when the program is run. The one or more computer-executable components may be at least one piece of software code or a portion thereof. Additionally, it should be noted in this respect that, as Figure 1 Any box in the logical flow can represent a program step, or interconnected logic circuits, boxes and functions, or a combination of program steps and logic circuits, boxes and functions. Software can be stored on physical media such as memory chips or blocks of storage implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants, CDs, etc. The physical medium is a non-transient medium.

[0073] Those skilled in the art should understand that the technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments have been described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0074] The above embodiments are merely illustrative of several implementations of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention should be determined by the appended claims.

Claims

1. A hyperspectral image reconstruction method, characterized in that, Includes the following steps: Acquire a two-dimensional compressed measurement image of the target scene and an auxiliary RGB image corresponding to the two-dimensional compressed measurement image space; A three-dimensional initial estimate hyperspectral image is generated from the two-dimensional compressed measurement image, and the three-dimensional initial estimate hyperspectral image is fused with the auxiliary RGB image to construct a multimodal input feature map; The multimodal input feature map is fed into a trained hyperspectral reconstruction network, where the following operations are performed: The multi-view 2D slice scanning module reassembles the input 3D feature map into a 2D sequence that includes at least a spatial main view slice and a spatial dimension coupled with a spectral dimension view slice. The state space sequence model is then used to model the global dependency relationship of the reassembled sequence. The feature map processed by the multi-view two-dimensional slice scanning module is divided into multiple local cubic sub-blocks through the three-dimensional local block slicing scanning module. The multi-view two-dimensional slice scanning operation is repeated in each local cubic sub-block to refine and enhance the local spatial and spectral features. Spatial resolution is restored based on the modeled features, and the reconstructed hyperspectral image is output.

2. The hyperspectral image reconstruction method as described in claim 1, characterized in that, The state-space sequence model is a Mamba module; the multi-view 2D slice scanning module is specifically used for: The input feature map is sliced ​​along the channel dimension, width dimension, and height dimension respectively, generating a frontal slice representing the structural distribution in the spatial plane, a hyperspectral slice representing the joint correlation between the height dimension and the spectral dimension, and a broadband slice representing the joint correlation between the width dimension and the spectral dimension. Each slice is expanded into a one-dimensional sequence and input into the Mamba module for sequence modeling. The sequence features output by the Mamba module are reshaped into the original shape of the slices, and then spliced ​​and fused.

3. The hyperspectral image reconstruction method as described in claim 1 or 2, characterized in that, The 3D local block slicing scanning module is specifically used for: The input feature map is divided into local cubic sub-blocks of multiple scales according to the spatial and spectral dimensions. The multiple scales include the basic block size and its proportionally scaled block size. Perform a multi-view 2D slice scan operation on each local cube sub-block; The processing results of all local cubic sub-blocks are recombined into an enhanced feature map with the same size as the input feature map, and residual connections are made with the input feature map.

4. The hyperspectral image reconstruction method as described in claim 3, characterized in that, The basic block size of the local cube sub-block is determined by the height, width and number of channels of the input feature map according to the preset block coefficient.

5. The hyperspectral image reconstruction method as described in claim 1, characterized in that, The hyperspectral reconstruction network adopts an encoder-decoder structure. In the multiple downsampling stages of the encoder and the multiple upsampling stages of the decoder, there are multi-view two-dimensional slice scanning modules and three-dimensional local block slice scanning modules, respectively. The modules at different stages process feature maps of corresponding scales to form a progressive global and local joint modeling mechanism.

6. The hyperspectral image reconstruction method as described in claim 1, characterized in that, Acquire a two-dimensional compressed measurement image of the target scene and an auxiliary RGB image corresponding to the two-dimensional compressed measurement image space, including: Two-dimensional compressed measurement images and auxiliary RGB images are acquired simultaneously using a dual-path snapshot-type compressed imaging device. The dual-path snapshot-type compressed imaging device includes a compressed measurement acquisition optical path and an auxiliary imaging acquisition optical path. The compressed measurement acquisition optical path includes an encoding mask and a dispersive element, which are used to compress and encode the three-dimensional hyperspectral image of the target scene into a two-dimensional observation image. The auxiliary imaging acquisition optical path is used to record the spatial texture and edge structure information of the scene.

7. The hyperspectral image reconstruction method as described in claim 1, characterized in that, Generate an initial three-dimensional estimated hyperspectral image from the two-dimensional compressed measurement image, including: By combining the calibration parameters of the imaging system, the dispersion offset, and the mask response information, the two-dimensional compressed measurement image is converted into a three-dimensional initial estimated hyperspectral image through back projection, shift inverse mapping, or physical prior initialization methods.

8. A hyperspectral image reconstruction device, characterized in that, include: The dual-channel acquisition module is used to acquire a two-dimensional compressed measurement image of the target scene and an auxiliary RGB image corresponding to the two-dimensional compressed measurement image space. The initialization and fusion module is used to generate a three-dimensional initial estimate hyperspectral image based on the two-dimensional compressed measurement image, and to fuse the three-dimensional initial estimate hyperspectral image with the auxiliary RGB image to construct a multimodal input feature map; The hyperspectral reconstruction network module includes: The multi-view 2D slice scanning submodule is used to reassemble the input 3D feature map into a 2D sequence that includes at least a spatial main view slice and a spatial dimension coupled with a spectral dimension view slice, and to model the global dependency relationship of the reassembled sequence through a state space sequence model. The 3D local block slicing scanning submodule is used to divide the feature map processed by the multi-view 2D slicing scanning submodule into multiple local cube sub-blocks, and repeatedly perform the multi-view 2D slicing scanning operation within each local cube sub-block to refine and enhance local spatial and spectral features. The resolution restoration submodule is used to restore spatial resolution based on the modeled features and output a reconstructed hyperspectral image.

9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the hyperspectral image reconstruction method according to any one of claims 1 to 7.

10. A readable storage medium, characterized in that, The readable storage medium stores a computer program, the computer program including program code for controlling a process to execute the process, the process including the hyperspectral image reconstruction method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Hyperspectral image reconstruction method and device based on high compression ratio snapshot compression and readable storage medium thereof

    CN120355808A

  • Hyperspectral image reconstruction method and device, equipment and storage medium

    CN121544741A