High-resolution snapshot hyperspectral imaging system and method based on deep learning
Through a deep learning-based high-resolution snapshot hyperspectral imaging system, combined with broadband multispectral filtering arrays and metasurface technology, HSITNet is used to achieve joint optimization of encoding and decoding, solving the conflict between spectral resolution and spatial resolution in spectral imaging, and realizing the reconstruction of high-quality spectral cubes.
Patent Information
- Application Number
- CN202510569332.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-01
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art cannot take into account both spectral resolution and spatial resolution during spectral imaging, resulting in a decrease in the mass of the spectral cube, the density limitation of the detector cell leads to a large loss of information entropy, and the design efficiency of filter devices is inefficient.
A high-resolution snapshot hyperspectral imaging system based on deep learning is adopted, including an encoding reception module and an algorithm decoding reconstruction module, using a broadband multispectral filtering array and detector, combining metasurface technology and deep learning network, deep spectral features are extracted through the encoding module, the decoding module restores spatial resolution, and uses HSITNet to achieve joint optimization of encoding and decoding.
High-quality spectral reconstruction of 151 spectral channels is achieved, with lossless spatial resolution, improving spectral imaging quality, breaking through the limitations of information entropy transfer, and enhancing the upper limit of spectral reconstruction.
Smart Images

Figure CN120403862A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of spectral imaging, and particularly relates to a high-resolution snapshot hyperspectral imaging system and method based on deep learning. Background Art
[0002] Spectral imaging technology is a new imaging technology that combines imaging technology and spectral technology, and can observe a target simultaneously in the spatial dimension and the spectral dimension. By recording the pixel values of each pixel point in space at different spectral bands, a three-dimensional data cube of the observed target is generated. Due to having both spatial information and spectral information, it has a wide range of applications in many fields, such as remote sensing monitoring, biomedicine, and food safety fields.
[0003] In recent years, snapshot spectral imaging technology has developed rapidly. Compared with scanning spectrometers, snapshot spectral imaging systems have a simple structure and do not require complex mechanical structures and special optical systems to extract the spectral cube of dynamic targets. However, during the imaging process, the spectral resolution and spatial resolution restrict and conflict with each other, making it difficult to obtain a high-quality spectral cube. In terms of improving the spectral resolution, in 2022, Yang et al. from Tsinghua University achieved an ultra-high spectral resolution of 601 channels in the 460nm - 740nm band through a free-form meta-atom metasurface and by using a point-by-point spectral reconstruction method, but the spatial resolution of spectral imaging decreased by 25 times. In terms of improving the comprehensive quality of the spectral cube, in 2024, Bian et al. achieved hyperspectral imaging with 61 wavelength channels in the range of 400nm - 1000nm and 96 wavelength channels in the range of 400nm - 1700nm through broadband multi-spectral filtering array (BMSFA) manufacturing technology combined with a high-performance neural network, while ensuring the non-deterioration of the spatial resolution. In terms of improving the spectral reconstruction speed, in 2025, the Wen team used a multi-layer thin film structure to achieve 33 frames of high-quality spectral reconstruction through a Massively Parallel Neural Network, but there are still problems of insufficient utilization.
[0004] Currently, due to the limitation of the detector pixel density, the existing technical methods cannot simultaneously take into account the spectral resolution and spatial resolution during the spectral imaging process, resulting in a significant decline in the quality of the final spectral cube obtained. The main factor restricting the spectral imaging quality at present is the encoding loss of the information entropy of the target spectral cube. The core steps of snapshot spectral imaging are the spectral response encoding (SRE) of the incident spectral cube by BMSFA and the decoding of the information received by the detector by the backend neural network. Since the detector receives information in a two-dimensional plane, a large amount of information entropy is lost during the process of encoding and compressing the three-dimensional information of the spectral cube into two-dimensional information by BMSFA, which makes it impossible for any decoding method to reconstruct the spectral cube of the scene losslessly. In addition, the design of the filter device is only based on the correlation coefficient, and the design efficiency is low. Summary of the Invention
[0005] To solve the above problems, the present invention provides a high-resolution snapshot hyperspectral imaging system and method based on deep learning.
[0006] The first object of the present invention is to provide a high-resolution snapshot hyperspectral imaging system based on deep learning, which is characterized in that it includes an encoding and receiving module and an algorithm decoding and reconstruction module; The encoding and receiving module includes an imaging unit, a broadband multispectral filter array and a detector; the target spectral cube is imaged on the surface of the broadband multispectral filter array through the imaging unit; subsequently, the broadband multispectral filter array performs spectral response encoding on the incident light; the encoded spectral information is received on the surface of the detector, realizing efficient encoding and compression from three-dimensional spectral information to two-dimensional plane information; The algorithm decoding and reconstruction module includes a spectral reconstruction network; the encoding and receiving module combines the spectral channel encoding curve of the broadband multispectral filter array and the data received by the detector, and inputs the combined data into the spectral reconstruction network to obtain a high-quality spectral cube through the spectral reconstruction network.
[0007] Preferably, the imaging unit includes a lens group for imaging the light field of the target scene on the surface of the broadband multispectral filter array; The broadband multispectral filter array is composed of a plurality of identical filter arrays, the filter arrays match the size of the detector, and the size of each filter unit in the filter array is the same as that of the detector pixel; the filter arrays in the broadband multispectral filter array are prepared by using the metasurface technology.
[0008] Preferably, the filter array is composed of a plurality of metasurface units, and the structure of the metasurface unit includes a SiO2 substrate, a Si3N4 substrate and a top TiO2 layer stacked tightly in sequence.
[0009] Preferably, the overall framework of the spectral reconstruction network uses a U-shaped connection, including an encoding module, a decoding module, and a normalization layer; the deep spectral features are extracted by the encoding module, the spatial resolution is gradually restored by the decoding module, and feature fusion is enhanced, and finally the spectral cube is reconstructed by the normalization layer.
[0010] Preferably, the encoding module includes multiple downsampling units and a spectral channel attention mechanism unit. The data size is reduced by continuous downsampling, and the deep features are further extracted by the spectral channel attention mechanism unit; the multiple downsampling units implement downsampling through convolutional layers and pooling layers. After each downsampling, the spatial resolution is halved and the number of channels is doubled to extract deep spectral features; the spectral channel attention mechanism unit weights different spectral channels to enhance the correlation between spectral channels and further extract spectral features; The decoding module is connected to the encoding module to form a U-shaped network, upsamples the obtained deep features, and decodes them with a structure symmetric to the encoding module to obtain hyperspectral data with a spatial resolution reduced by half; the decoding module includes a transposed convolution upsampling unit and a NiN unit. The transposed convolution upsampling unit realizes upsampling through transposed convolutional layers. After each upsampling, the spatial resolution is doubled and the number of channels is halved to gradually restore the spatial resolution of the data; the NiN unit is used to process the data with equal number of channels in the encoding module and the decoding module and perform skip connections. The feature maps of the skip connections are processed through convolutional layers to reduce the number of channels and enhance the non-linear expression of the features at the same time; The normalization layer uses a shifted window-based vision transformers structure, processes the feature maps using the self-attention mechanism, improves the perception field of view, reduces the joint distribution information entropy of the spectrum, completes the spectral detail filling, maps various spectral features and spatial details into a complete spectral cube, and finally outputs a high-quality spectral cube.
[0011] Preferably, the spectral reconstruction network further includes an optimization module. The optimization module includes a broadband multispectral filter array structure parameter optimization unit. The structure parameters of the broadband multispectral filter array are trained as part of the spectral reconstruction network. The mapping relationship between the structure parameters of the broadband multispectral filter array and the spectral response curve is fitted through a multi-layer perceptron, and the structure parameters of the broadband multispectral filter array are automatically optimized through the neural network gradient backpropagation.
[0012] The second object of the present invention is to provide a high-resolution snapshot hyperspectral imaging method based on deep learning, which uses a high-resolution snapshot hyperspectral imaging system based on deep learning for imaging. The specific steps are as follows: S1. The target spectrum is imaged on the surface of the broadband multispectral filter array through an optical system. The broadband multispectral filter array encodes the spectral response of the incident light and encodes the three-dimensional spectral information into two-dimensional information; S2. Measure the spectral response encoding characteristics at different positions in the broadband multispectral filter array, form a three-dimensional mask cube, and record the mask information as prior information; S3. The detector receives the spectral information encoded by the broadband multispectral filter array; S4. Fuse the encoded data received by the detector with the measured mask information and input it into the spectral reconstruction network; S5. The spectral reconstruction network extracts features from the input fused information and finally realizes the high-quality reconstruction of the target spectrum; S6. Establish a mapping relationship between the structural parameters of the broadband multispectral filter array and the spectral response curve through simulation software and a multi-layer perceptron, optimize the structural parameters of the broadband multispectral filter array, fabricate an actual broadband multispectral filter array and conduct measurements, and correct the model to adapt to the characteristics of the actual device; S7. Retrain and optimize the decoding module of the spectral reconstruction network based on the actually measured spectral response curve to adapt to the characteristics of the actual device and ensure the high-quality spectral cube reconstruction; S8. Through the optimization module, use the neural network gradient backpropagation to automatically optimize the structural parameters of the broadband multispectral filter array to ensure the joint optimization of the encoding and decoding processes.
[0013] Preferably, step S5 specifically includes the following steps: S501. Perform high-dimensional feature processing on the input fused information through the embedding layer of the spectral reconstruction network, compress the high-dimensional features into a low-dimensional space through dimensionality reduction technology, generate a low-dimensional feature representation, and reduce the number of model parameters and computational complexity; S502. Perform deep spectral feature extraction on the input fused information through the encoding module of the spectral reconstruction network, and then gradually restore the spatial resolution and enhance the feature fusion through the decoding module; S503. Use the normalization layer to complete the reconstruction of the spectral cube and output a high-quality spectral cube.
[0014] Preferably, step S6 specifically includes the following steps: S601. Select the initial structural parameters of the broadband multispectral filter array and input them into the simulation software for nested parameter scanning to obtain the spectral response encoding characteristics corresponding to the structural parameters; S602. Use the multi-layer perceptron model, take the scanned data as the training data set; train the multi-layer perceptron model to obtain the mapping relationship between the structural parameters of the broadband multispectral filter array and the spectral response encoding characteristics; S603. Input the selected initial structural parameters into the established mapping relationship, output the spectral response characteristic curve corresponding to the initial structure, open the gradient corresponding to the spectral response characteristic curve and optimize it, and output the optimized spectral response characteristic curve; then input the optimized spectral response characteristic curve into the mapping relationship to obtain the optimized broadband multispectral filter array structural parameters; S604. Perform photolithography processing according to the optimized structural parameters to fabricate a broadband multispectral filter array; measure the actual spectral response characteristic curve of each metasurface unit in the broadband multispectral filter array, and record the deviation between the actual and the designed spectral response characteristic curves; S605. Input the actually measured spectral response characteristic curve into the model, and turn off the gradient corresponding to the spectral response characteristic curve; perform model optimization again, correct the spectral reconstruction strategy, and finally obtain the spectral reconstruction strategy corresponding to the spectral response characteristic curve of the actual broadband multispectral filter array.
[0015] Preferably, step S8 specifically includes the following steps: S801. Input the optimized broadband multispectral filter array structural parameters as initial parameters into the spectral reconstruction network; S802. In the spectral reconstruction network, calculate the gradient of the structural parameters through gradient backpropagation; adjust the structural parameters according to the gradient information to optimize the encoding process; S803. During the optimization process, consider the performance of both the encoding and decoding processes simultaneously to ensure collaborative optimization of encoding and decoding; through multiple iterations, improve the overall performance of the system; S804. Evaluate the performance of the optimized system and determine the final broadband multispectral filter array structural parameters.
[0016] Compared with the prior art, the present invention can achieve the following beneficial effects: (1) The present invention adopts a method of jointly optimizing the encoding strategy and the decoding strategy, realizes the automatic design of the BMSFA device during the training process, and ensures the adaptation of its encoding characteristics to the decoding strategy, which theoretically improves the upper limit of the reconstructed spectral quality; (2) The displacement window type vision transformer structure is selected, which achieves a balance between the computational field of view size and the required computational resources, realizes the consistency between the detector resolution and the spatial resolution of the reconnected spectrum, and has higher spectral imaging quality; (3) The self-attention mechanism is added in the spectral dimension to extract deep spectral features, combined with the features of spatial distribution, and high-quality spectral reconstruction of 151 spectral channels is achieved with high precision.
[0017] In summary, the present invention jointly designs the encoding and decoding methods. By fusing the forward design neural network of the metasurface and the spectral reconstruction neural network, the hyperspectral imaging transformers network (HSITNet) is derived. It has extremely strong feature extraction and prior modeling capabilities, and can directly obtain the structural parameters of the broadband multispectral filter array (BMSFA) during the decoding strategy optimization process, making the characteristics of the BMSFA match the decoding method, achieving joint optimization, ultimately improving the upper limit of spectral imaging resolution, and also avoiding the step of finding the structural parameters of the BMSFA based on the correlation coefficient. The experimental results prove that this method can achieve snapshot hyperspectral imaging with 151 wavelength channels in the 400-1000nm band and no loss of spatial resolution. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 FIG. is a structural diagram of a high-resolution snapshot hyperspectral imaging system based on deep learning according to an embodiment of the present invention.
[0019] Figure 2 FIG. is a schematic structural diagram of a filter array in a broadband multispectral filter array according to an embodiment of the present invention.
[0020] Figure 3 FIG. is an overall structural diagram of a spectral reconstruction network according to an embodiment of the present invention.
[0021] Figure 4 FIG. is a processing flow chart of a high-resolution snapshot hyperspectral imaging method based on deep learning according to an embodiment of the present invention.
[0022] Figure 5 FIG. is a practical operation flow chart for establishing the mapping relationship between the structural parameters of the broadband multispectral filter array and the spectral response encoding characteristics and establishing the broadband multispectral filter array design and spectral reconstruction strategy; in the figure, (a) represents the establishment of the mapping relationship between the structural parameters of the broadband multispectral filter array and the spectral response encoding characteristics, and (b) represents the practical operation flow for establishing the broadband multispectral filter array design and spectral reconstruction strategy.
[0023] REFERENCE SIGNS 1. Target spectral cube; 2. Imaging unit; - 3. Broadband multispectral filter array; 31. Filter array; 301. Metasurface unit; 3011. SiO2 substrate; 3012. Si3N4 substrate; 3013. Top TiO2 layer; 4. Detector; 5. Spectral reconstruction network; DETAILED DESCRIPTION OF THE EMBODIMENTS In the following, embodiments of the present invention will be described with reference to the accompanying drawings. In the following description, the same modules are denoted by the same reference numerals. In the case of the same reference numerals, their names and functions are also the same. Therefore, their detailed descriptions will not be repeated.
[0024] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.
[0025] See Figure 1 , the present invention provides a high-resolution snapshot hyperspectral imaging system based on deep learning, which includes two parts: an encoding and receiving module at the front end and an algorithm decoding and reconstruction module at the back end; The encoding and receiving module includes an imaging unit 2, a broadband multispectral filter array 3 (BMSFA), and a detector 4; the target spectral cube 1 is imaged on the surface of the broadband multispectral filter array 3 (BMSFA) through the imaging unit 2; subsequently, the broadband multispectral filter array 3 (BMSFA) performs spectral response encoding (SRE) on the incident target spectral cube 1, compressing the three-dimensional spectral information into two-dimensional information; the encoded spectral information is received on the surface of the detector 4 behind the broadband multispectral filter array (BMSFA), achieving efficient encoding and compression from three-dimensional spectral information to two-dimensional plane information; Specifically, as Figure 1 shown, the target spectral cube 1 to be measured, which includes two-dimensional spatial information and one-dimensional spectral information; the imaging unit 2 includes a lens group for imaging the light field of the target scene with high quality on the surface of the broadband multispectral filter array 3; the broadband multispectral filter array 3, which is the core component of the system, is used to perform spectral response encoding on the one-dimensional spectral information of the incident light; the detector 4 is closely attached to the broadband multispectral filter array 3 and receives the optical signal processed by the broadband multispectral filter array 3 in the form of light intensity.
[0026] The filter array in the BMSFA matches the size of the detector, and the size of each filter unit in the filter array is the same as that of the detector pixel, ensuring that the light intensity received by each pixel of the detector is subjected to SRE by the periodically arranged filters; The filter array in the BMSFA is fabricated using metasurface technology (sub-wavelength structure technology); see Figure 2, the BMSFA is composed of multiple identical filter array 31, and each filter array 31 is composed of multiple metasurface units 301 with different SREs; in the initial structure selection of the metasurface unit 301, in order to obtain the freedom degree of the SRE curve, a metasurface is selected to fabricate the BMSFA. The structure of the metasurface unit includes a SiO2 substrate 3011, a Si3N4 substrate 3012, and a top TiO2 layer 3013 that are tightly stacked in sequence, and a square structure is etched out of it. The structure of the metasurface unit and the overall structure of the BMSFA are as Figure 2 shown. The parameters that can be changed in the BMSFA design are the thickness and period of the substrate, and the thickness and width of the top titanium dioxide (TiO2) layer. Four by four metasurface units with different structural parameters form a filter array, and several filter arrays with the same arrangement form the BMSFA device.
[0027] Compared with the traditional imaging system, the encoding and receiving module at the front end of the system adds a BMSFA in front of the detector surface, which is one of the core devices of the whole system. The size of the filter array in the BMSFA matches the size of the detector, and the size of each filter unit is the same as the size of the detector pixel, ensuring that the light intensity received by each pixel of the detector is subjected to SRE by the periodically arranged filters. The performance of the system mainly depends on the SRE method of the BMSFA. It is necessary to ensure that the spectral response curve of the filter array has a high degree of freedom during design, and an array composed of multiple filters with different corresponding curves is required. This makes it difficult for the filters fabricated by the traditional coating process to meet the requirements, so a metasurface is used to fabricate the filter array. A metasurface is a two-dimensional artificial element composed of sub-wavelength structural units, which can modulate the phase, amplitude, polarization and other characteristics of light waves or other electromagnetic waves through a precisely designed structure. Therefore, using a metasurface to manufacture the BMSFA can complete a more refined spectral response encoding operation.
[0028] The algorithm decoding and reconstruction module at the back end includes a spectral reconstruction network; the encoding and receiving module combines the spectral channel encoding curve of the broadband multi-spectral filter array and the data received by the detector, inputs the combined data into the spectral reconstruction network, and obtains the spectral cube reconstruction result through the spectral reconstruction network; combines the spectral dimension information obtained by the spectral reconstruction network with the two-dimensional spatial information to output a high-quality spectral cube.
[0029] The core of the algorithm decoding and reconstruction module is to decode two-dimensional information through deep learning methods to complete the reconstruction of the target spectral cube. To achieve higher spectral resolution, that is, the number of spectral dimension channels of the spectral cube to be decoded is larger. Due to the limitation of information entropy transmission, it is difficult for traditional regularization and compressive sensing methods to achieve this process. Deep learning methods can induce the prior information in the spectral imaging process from the data in the dataset, make full use of the correlation in the data, extract deep features, act on the design of BMSFA in reverse, and at the same time realize the joint optimization of the encoding and decoding processes, reduce the joint distribution entropy of the spectral cube, break through the system information entropy transmission limitation, and reconstruct a higher-quality spectral cube.
[0030] Through the research on the information entropy transmission of the spectral cube, it is proved that the correlation between adjacent pixels is important for the spectral imaging quality. Therefore, when designing the architecture of HSITNet, its encoding and decoding modules need to take into account a larger sampling field of view to reduce the overall joint distribution information entropy. And the BMSFA design module is integrated into the spectral reconstruction network to realize the role of joint optimization of the encoding and decoding processes, improve the upper limit of information entropy transmission in the system, and thus achieve spectral imaging with higher spectral resolution and lossless spatial resolution.
[0031] Specifically, the overall framework of the spectral reconstruction network uses a U-shaped connection, including an encoding module, a decoding module, and a normalization layer; the deep spectral features are extracted through the encoding module, the spatial resolution is gradually restored and the feature fusion is enhanced through the decoding module, and finally the reconstruction of the spectral cube is completed through the normalization layer; Among them, the encoding module includes multiple downsampling units and a spectral channel attention mechanism unit. The data size is reduced by continuous downsampling, and the spectral channel attention mechanism unit further extracts deep features. At the same time, information is transmitted to the decoding layer through one-dimensional convolution during this process; the multiple downsampling units realize downsampling through convolutional layers and pooling layers (or strided convolutions). After each downsampling, the spatial resolution is halved and the number of channels is doubled. The multiple downsampling units are used to gradually reduce the spatial resolution of the input data while increasing the number of channels for deep spectral feature extraction; the spectral channel attention mechanism unit weights different spectral channels (highlighting important channels and suppressing unimportant channels), enhances the correlation between spectral channels, extracts spectral features, and improves the expression ability of spectral information; The decoding module and the encoding module form a U-shaped network connection, which upsamples the obtained deep features and decodes them in a structure symmetric to the encoding module to obtain hyperspectral data with a halved spatial resolution. It includes a transposed convolution upsampling unit and a NiN unit. The transposed convolution upsampling unit realizes upsampling through a transposed convolution layer. After each upsampling, the spatial resolution is doubled and the number of channels is halved. The transposed convolution upsampling unit is used to gradually restore the spatial resolution of the data and reduce the number of channels to reconstruct a high-resolution spectral cube. The NiN unit is used to process the data with equal numbers of channels in the encoding module and the decoding module and perform skip connections. The feature maps of the skip connections are processed through a convolutional layer (1×1) to reduce the number of channels and enhance the non-linear expression of the features. Through the skip connections, low-level spatial detail information (such as edges and textures) can be combined with high-level spectral feature information, enabling the spectral reconstruction network to better restore the details of the spectral cube during feature reconstruction. At the same time, the low-level features can be reused in deeper layers of the spectral reconstruction network through the skip connections, improving the utilization rate of the features. The normalization layer uses a shifted window-based vision transformer (SwinViT) structure to process the feature maps using the self-attention mechanism, improve the perception field of view, reduce the joint distribution information entropy of the spectrum, complete the spectral detail filling in a large computational field of view range, map various spectral features and spatial details into a complete spectral cube, complete the spectral reconstruction network, and finally output a spectral cube with lossless spatial resolution. The overall structure of the spectral reconstruction network is as Figure 3 shown.
[0032] The spectral reconstruction network also includes an optimization module, which realizes the optimization of the broadband multispectral filter array structure parameters through the optimization module, ensures the joint optimization of the encoding and decoding processes, breaks through the limitations of information entropy transmission, and realizes high-quality spectral cube reconstruction. Specifically, the optimization module includes a broadband multispectral filter array structure parameter optimization unit, which trains the structure parameters of the broadband multispectral filter array as part of the spectral reconstruction network, fits the mapping relationship between the structure parameters of the broadband multispectral filter array and the spectral response curve through a multi-layer perceptron (MLP), and automatically optimizes the structure parameters of the broadband multispectral filter array through the neural network gradient backpropagation to ensure the joint optimization of the encoding and decoding processes.
[0033] The U-shaped connection makes the loss surface of the spectral reconstruction network smoother, helps the network converge faster, and due to the improvement of gradient propagation and the direct transmission of information, the training process of the network becomes more stable and efficient. It can still learn effective spatial distribution patterns and spectral features even when the dataset is small, has stronger robustness to the spectral data after encoding compression, and can better handle noisy and incomplete spectral data.
[0034] Based on the above imaging system, the present invention further provides a high-resolution snapshot hyperspectral imaging method based on deep learning (for the processing flow, see Figure 4 ), which specifically includes the following steps: S1. The target spectrum is imaged on the surface of a broadband multispectral filter array (BMSFA) through an optical system. The broadband multispectral filter array performs spectral response encoding (SRE) on the incident light, encoding three-dimensional spectral information into two-dimensional information; S2. Measure the spectral response encoding characteristics at different positions in the broadband multispectral filter array to form a three-dimensional mask cube, and record the mask information as prior information; S3. The detector receives the spectral information encoded by the broadband multispectral filter array, achieving efficient encoding compression from three-dimensional spectral information to two-dimensional plane information; S4. Fuse the encoded data received by the detector with the measured mask information and input it into the spectral reconstruction network; S5. The spectral reconstruction network extracts features from the input fused information (three-dimensional tensor information), and finally realizes high-quality reconstruction of the target spectrum; specifically, it includes the following steps: S501. Perform high-dimensional feature processing on the input fused information through the embedding layer of the spectral reconstruction network, and compress the high-dimensional features into a low-dimensional space through dimensionality reduction techniques (such as PCA, t-SNE, or autoencoders) to generate low-dimensional feature representations, thereby reducing the number of model parameters and computational complexity; S502. Perform deep spectral feature extraction on the input fused information through the encoding module of the spectral reconstruction network, and then gradually restore the spatial resolution and enhance feature fusion through the decoding module; The encoding module includes multiple layers of downsampling units and a spectral channel attention mechanism unit. The multiple layers of downsampling units achieve downsampling through convolutional layers and pooling layers (or strided convolutions). After each downsampling, the spatial resolution is halved and the number of channels is doubled for deep spectral feature extraction; the spectral channel attention mechanism unit weights different spectral channels to enhance the correlation between spectral channels, extracts spectral features, and improves the expression ability of spectral information; The decoding module is connected to the encoding module to form a U-shaped network, including a transposed convolution upsampling unit and a NiN unit; the transposed convolution upsampling unit realizes upsampling through transposed convolution layers. After each upsampling, the spatial resolution is doubled and the number of channels is halved to reconstruct a high-resolution spectral cube; the NiN unit is used to process the data with equal numbers of channels in the encoding module and the decoding module and perform skip connections, and processes the feature maps of the skip connections through convolutional layers (1×1) to reduce the number of channels and enhance the non-linear expression of features at the same time; S503. Use a normalization layer (such as SwinViT) to complete the reconstruction of the spectral cube and output a high-quality spectral cube; The normalization layer uses a shifted window vision transformer (SwinViT) structure to process the feature map using the self-attention mechanism, improve the perception field of view, reduce the joint distribution information entropy of the spectrum, complete the spectral detail filling in a large computational field of view, map various spectral features and spatial details into a complete spectral cube, and finally output a high-quality spectral cube with lossless spatial resolution. [[ID={1}]] [[ID={2}]]
[0035] In the present invention, the algorithm decoding and reconstruction module is composed of the HSITNet algorithm, which has two modes. In the training mode, HSITNet trains suitable encoding and decoding strategies according to the target scene and jointly optimizes them to improve the upper limit of the spectral cube reconstruction quality. In the working mode, first, the corresponding BMSFA is processed according to the encoding strategy designed in the training mode, and the physical process encoding is realized by it. HSITNet only needs to run the decoding part to finally complete the spectral cube reconstruction of the scene, and respectively realize the automatic optimization of the BMSFA structure parameters and the spectral decoding and reconstruction function from the two-dimensional information received at the front end to the three-dimensional information.
[0036] The key points of the core algorithm HSITNet mainly include the following aspects: (1) Establish the mapping relationship between the structural parameters of BMSFA and its spectral response curve, connect this relationship to the spectral reconstruction network, and realize the adaptation of the BMSFA encoding characteristics and decoding strategy during the training process, while realizing the automatic design of the BMSFA device. (2) Add a shifted window vision transformer structure during the spectral reconstruction process to improve the perception field of view of the spectral reconstruction network, better reduce the joint distribution information entropy of the spectrum, improve the spatial bandwidth product of the detector, and realize the function of lossless spatial resolution during the spectral imaging process. (3) Add a self-attention mechanism in the spectral dimension to improve the interaction efficiency between spectral information, obtain the deep features of the spectral dimension data distribution, store them as prior information in the parameters of the spectral operation module, and finally realize the high-quality spectral reconstruction of 151 spectral channels.
[0037] S6. Establish the mapping relationship between the structural parameters of the broadband multispectral filter array and the spectral response curve through simulation software and a multi-layer perceptron (MLP), optimize the structural parameters of the broadband multispectral filter array, fabricate an actual device and conduct measurements, and correct the model to adapt to the characteristics of the actual device; specifically including the following sub-steps: S601. Structural parameter scanning: Select the initial structural parameters of the broadband multispectral filter array and input them into the simulation software for nested parameter scanning to obtain the spectral response encoding characteristics corresponding to the structural parameters; S602. Establishing a mapping relationship: Using a multi-layer perceptron (MLP) model, the scanned data is used as a training data set; training the multi-layer perceptron model to obtain a mapping relationship between the structural parameters of the broadband multispectral filter array and the spectral response encoding characteristics; S603. Optimizing structural parameters: Inputting the selected initial structural parameters into the established mapping relationship, outputting the spectral response characteristic curve (SRcurves) corresponding to the initial structure, opening the gradient corresponding to the spectral response characteristic curve and optimizing it, and outputting the optimized spectral response characteristic curve; then inputting the optimized spectral response characteristic curve into the mapping relationship to obtain the optimized broadband multispectral filter array structural parameters; S604. Actual device fabrication and measurement: Perform photolithography according to the optimized structural parameters to fabricate a broadband multispectral filter array; measure the actual spectral response characteristic curve of each metasurface unit in the broadband multispectral filter array, and record the deviation between the actual and designed spectral response characteristic curves; S605. Model modification and optimization: Input the spectral response characteristic curve obtained from actual measurement into the model, close the gradient corresponding to the spectral response characteristic curve; optimize the model again, modify the spectral reconstruction strategy, and finally obtain the spectral reconstruction strategy corresponding to the spectral response characteristic curve of the actual broadband multi-spectral filter array. The establishment of the mapping relationship and the actual operation process of BMSFA design and spectral reconstruction strategy establishment are shown in Figure 5 shown.
[0038] To raise the upper limit of spectral reconstruction resolution, it is necessary to fully exploit the encoding performance of the BMSFA and overcome the bottleneck of system information entropy transfer. The structural design of the traditional BMSFA relies on a correlation matrix, minimizing the correlation coefficient between BMSFAs and improving encoding capabilities. However, this approach fails to consider the algorithmic strategy of the decoding process, resulting in a mismatch between the two and an inability to fully compress the information entropy of the target spectral cube, leading to significant loss of spectral information during system transmission. In this system, the SRE characteristics of the BMSFA are integrated into the spectral reconstruction process, and a closed loop is formed through neural network gradient backpropagation to automatically optimize the BMSFA's structural parameters.
[0039] A mapping relationship between the metasurface structural parameters and the corresponding transmittance curve is established, and a multi-layer perceptron (MLP) is used to fit the functional relationship between the structural parameters and the transmittance curve. This structure is then connected to the back-end spectral reconstruction network. The MLP gradient is fixed to ensure that the fitted functional relationship remains unchanged. An initial set of structural parameters is input to the MLP, and the entire network is trained for spectral reconstruction. Under the action of the optimizer, the optimal solution for the structural parameters is automatically found, achieving simultaneous optimization of the encoding and decoding processes.
[0040] Considering the mapping error between the MLP for device parameters and the transmittance curve, as well as the process error during device processing, there is a certain deviation between the actual performance and transmittance curve of the final device and the transmittance curve used in the model network. For this situation, it is necessary to adjust the network parameter weights and use a better spectral decoding strategy for the transmittance curve of the real device. The complete process of BMSFA structure design is as Figure 5 shown in (b). The initial BMSFA structure obtains the corresponding spectral response curves (SRcurves) through the mapping relationship and inputs them into HSITNet. The improved BMSFA structure is obtained through the overall learning and optimization of the neural network. The actual SRcurves of BMSFA are fabricated and measured according to the structure. Based on the actual SRcurves, the decoding network is retrained and optimized, and finally the actual decoding module is obtained. The closed-loop process of system encoding and decoding joint optimization is as Figure 5 .
[0041] S7. Retraining and optimization of the decoding module of the spectral reconstruction network: Retrain and optimize the decoding module of the spectral reconstruction network based on the actually measured spectral response curves (SRcurves) to adapt to the characteristics of the actual device and ensure high-quality spectral cube reconstruction; specifically, it includes the following sub-steps: S701. Data preparation: Integrate the actually measured spectral response curves (SRcurves) with the encoded data received by the detector to form a new training dataset; S702. Decoding module retraining: Retrain the decoding module of the spectral reconstruction network using the new training dataset; Adjust the parameters of the decoding module through an optimization algorithm (such as gradient descent) to minimize the reconstruction error; S703. Optimization verification: Evaluate the performance of the decoding module using the validation dataset to ensure that the spectral cube can be accurately reconstructed; According to the verification results, further adjust the parameters of the decoding module as the actually available decoding module.
[0042] S8. Automatically optimize the BMSFA structure parameters using the optimization module: Through the optimization module, use the neural network gradient backpropagation to automatically optimize the structure parameters of the broadband multispectral filter array to ensure the joint optimization of the encoding and decoding processes; specifically, it includes the following sub-steps: S801. Structure parameter initialization: Input the optimized BMSFA structure parameters as the initial parameters into the spectral reconstruction network; S802. Gradient backpropagation: In the spectral reconstruction network, calculate the gradient of the structure parameters through gradient backpropagation; Adjust the structure parameters of BMSFA according to the gradient information to optimize the encoding process; Joint Optimization: During the optimization process, consider the performance of both the encoding and decoding processes simultaneously to ensure collaborative optimization of encoding and decoding. Through multiple iterations, gradually improve the overall performance of the system. Determination of Optimal Parameters: Evaluate the performance of the optimized system and determine the final structural parameters of BMSFA.
[0043] Since the output parameters and input parameters set by the spectral reconstruction network are the same, and the number of downsampling times and upsampling times of data in the neural network are the same, spectral imaging with lossless spatial resolution is ultimately achieved. The HSITNet structure determines that the spatial resolution is not affected by the size of the filter array. Setting a larger filter array can reduce the overall correlation between SRcurves, and finally, an array arrangement of 4×4 is selected.
[0044] In practical applications, dynamically adjust the parameters of BMSFA according to environmental changes and user requirements to ensure the high performance and adaptability of the system. This includes the following aspects: (1) Environmental Monitoring: Continuously monitor changes in the working environment, such as lighting conditions and target object characteristics; (2) Parameter Adjustment: Dynamically adjust the parameters of BMSFA based on the monitoring results to adapt to the current working conditions; use the optimization module to optimize the system parameters in real time to ensure the best performance; (3) Performance Evaluation: Regularly evaluate the performance of the system to ensure its stability and reliability under different conditions; further adjust the system parameters based on the evaluation results to improve performance; (4) Comprehensive Testing: Conduct comprehensive tests on the system, including key indicators such as spectral resolution, spatial resolution, and signal-to-noise ratio. Use standard spectral samples and actual application scenarios for testing to evaluate the performance of the system.
[0045] The key technical points and advantages of the present invention are summarized as follows: The present invention proposes a novel spatially resolved non-destructive snapshot hyperspectral imaging system and method. First, from the three-dimensional spectral cube information entropy theory, it can be known that there is a correlation between adjacent pixel points in both the spectral dimension and the spatial dimension. When this correlation is retained, the joint entropy between pixels is much smaller than the information entropy calculated separately for each pixel, that is, more information is retained during compression coding in the former case. Based on this, when designing the neural network for decoding, a transformer structure with a larger field of view is selected. In addition, the traditional BMSFA design method is too complicated, and the evaluation criterion is the correlation coefficient between them. In this study, the encoding method and the decoding method are jointly designed. By fusing the forward design neural network of the metasurface and the spectral reconstruction neural network, the hyperspectral imaging transformers network (HSITNet) is derived, which has extremely strong feature extraction and prior modeling capabilities. It can directly obtain the structural parameters of BMSFA during the optimization process of the decoding strategy, making the characteristics of BMSFA match the decoding method, achieving joint optimization, ultimately improving the upper limit of the spectral imaging resolution, and also avoiding the step of finding the BMSFA structural parameters based on the correlation coefficient. Experimental results prove that this method can achieve snapshot hyperspectral imaging with 151 wavelength channels in the wavelength range of 400 nm to 1000 nm and without loss of spatial resolution.
[0046] To improve the quality of the reconstructed spectral cube, first, the optimal encoding method is selected according to the data characteristics of the imaging target to design BMSFA, and a joint optimization function is added to HSITNet to directly obtain the structural parameters of BMSFA during the optimization process of the decoding strategy, making the characteristics of BMSFA match the decoding method, and also avoiding the step of finding the BMSFA structural parameters based on the correlation coefficient, thereby improving the design efficiency of the BMSFA device. Second, a displacement window-based vision transformers architecture is used in the spectral reconstruction part to improve the perceptual field of view of the spectral reconstruction network and achieve non-destructive spectral imaging with spatial resolution at a relatively high accuracy. Finally, to improve the spectral resolution, the self-attention mechanism is applied to the spectral channels to increase the interaction efficiency between spectral information, obtain the deep features of the spectral dimension data distribution, and store them as prior information in the parameters of the spectral operation module.
[0047] It should be understood that the various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recorded in the disclosure of the present invention can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in the present invention can be achieved, and no limitations are imposed herein.
[0048] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A high-resolution snapshot hyperspectral imaging system based on deep learning, characterized in that: It includes an encoding reception module and an algorithm decoding and reconstruction module; The encoding reception module includes an imaging unit, a broadband multispectral filter array, and a detector; The target spectral cube is imaged on the surface of the broadband multispectral filter array by the imaging unit; subsequently, the broadband multispectral filter array performs spectral response encoding on the incident light; The encoded spectral information is received on the surface of the detector, achieving efficient encoding and compression from three-dimensional spectral information to two-dimensional plane information; The algorithm decoding and reconstruction module includes a spectral reconstruction network; the encoding reception module combines the spectral channel encoding curves of the broadband multispectral filter array and the data received by the detector, and inputs the combined data into the spectral reconstruction network, and a high-quality spectral cube is obtained through the spectral reconstruction network.
2. The high-resolution snapshot hyperspectral imaging system based on deep learning according to claim 1, characterized in that: The imaging unit includes a lens group for imaging the light field of the target scene on the surface of the broadband multispectral filter array; The broadband multispectral filter array is composed of multiple identical filter arrays, the filter arrays match the size of the detector, and the size of each filter unit in the filter array is the same as that of the detector pixel; the filter arrays in the broadband multispectral filter array are prepared by metasurface technology.
3. The high-resolution snapshot hyperspectral imaging system based on deep learning according to claim 2, characterized in that: The filter array is composed of multiple metasurface units, and the structure of the metasurface unit includes a SiO2 substrate, a Si3N4 substrate, and a top TiO2 layer stacked tightly in sequence.
4. A high-resolution snapshot hyperspectral imaging system based on deep learning according to claim 1, characterized in that: The overall framework of the spectral reconstruction network uses a U-shaped connection, including an encoding module, a decoding module, and a normalization layer; deep spectral features are extracted through the encoding module, the decoding module gradually restores the spatial resolution and enhances feature fusion, and finally the spectral cube is reconstructed through the normalization layer.
5. A high-resolution snapshot hyperspectral imaging system based on deep learning according to claim 4, characterized in that: The encoding module includes multiple downsampling units and a spectral channel attention mechanism unit, which reduces the data size by continuous downsampling and further extracts deep features through the spectral channel attention mechanism unit; the multiple downsampling units achieve downsampling through convolutional layers and pooling layers. After each downsampling, the spatial resolution is halved and the number of channels is doubled for deep spectral feature extraction; the spectral channel attention mechanism unit weights different spectral channels to enhance the correlation between spectral channels and further extracts spectral features; The decoding module is connected to the encoding module to form a U-shaped network, upsamples the obtained deep features, and decodes them with a structure symmetric to that of the encoding module to obtain hyperspectral data with a spatial resolution reduced by half; the decoding module includes a transposed convolution upsampling unit and a NiN unit. The transposed convolution upsampling unit realizes upsampling through a transposed convolution layer. After each upsampling, the spatial resolution is doubled and the number of channels is halved, gradually restoring the spatial resolution of the data; the NiN unit is used to process the data with equal numbers of channels in the encoding module and the decoding module and perform skip connections, and processes the feature maps of the skip connections through convolutional layers to reduce the number of channels and enhance the non-linear expression of features at the same time; The normalization layer uses a displacement window-based vision transformers structure, processes the feature map using the self-attention mechanism, improves the perception field of view, reduces the joint distribution information entropy of the spectrum, completes the spectral detail filling, maps various spectral features and spatial details into a complete spectral cube, and finally outputs a high-quality spectral cube.
6. The high-resolution snapshot hyperspectral imaging system based on deep learning according to claim 4, wherein: The spectral reconstruction network further includes an optimization module. The optimization module includes a broadband multispectral filter array structure parameter optimization unit, which trains the structure parameters of the broadband multispectral filter array as part of the spectral reconstruction network, fits the mapping relationship between the structure parameters of the broadband multispectral filter array and the spectral response curve through a multi-layer perceptron, and automatically optimizes the structure parameters of the broadband multispectral filter array through the backpropagation of the neural network gradient.
7. A high-resolution snapshot hyperspectral imaging method based on deep learning, which uses a high-resolution snapshot hyperspectral imaging system according to claim 1 for imaging, and is characterized in that: Specifically, it includes the following steps: S1. The target spectrum is imaged on the surface of the broadband multispectral filter array by the optical system. The broadband multispectral filter array encodes the spectral response of the incident light and encodes the three-dimensional spectral information into two-dimensional information. S2. Measure the spectral response encoding characteristics at different positions in the broadband multispectral filter array, form a three-dimensional mask cube, and record the mask information as prior information. S3. The detector receives the spectral information encoded by the broadband multispectral filter array. S4. Fuse the encoded data received by the detector with the measured mask information and input it into the spectral reconstruction network. S5. The spectral reconstruction network extracts features from the input fused information and finally realizes the high-quality reconstruction of the target spectrum. S6. Establish the mapping relationship between the structure parameters of the broadband multispectral filter array and the spectral response curve through simulation software and a multi-layer perceptron, optimize the structure parameters of the broadband multispectral filter array, manufacture the actual broadband multispectral filter array and measure it, and correct the model to adapt to the actual device characteristics. S7. Retrain and optimize the decoding module of the spectral reconstruction network based on the actually measured spectral response curve to adapt to the actual device characteristics and ensure the high-quality spectral cube reconstruction. S8. Through the optimization module, use the backpropagation of the neural network gradient to automatically optimize the structure parameters of the broadband multispectral filter array to ensure the joint optimization of the encoding and decoding processes.
8. A high-resolution snapshot hyperspectral imaging method based on deep learning according to claim 7, characterized in that: The specific steps of step S5 include the following steps: S501. Perform high-dimensional feature processing on the input fused information through the embedding layer of the spectral reconstruction network, compress the high-dimensional features into a low-dimensional space through dimensionality reduction technology, generate low-dimensional feature representations, and reduce the number of model parameters and computational complexity. S502. Perform deep spectral feature extraction on the input fused information through the encoding module of the spectral reconstruction network, and then gradually restore the spatial resolution and enhance feature fusion through the decoding module. S503. Use the normalization layer to complete the reconstruction of the spectral cube and output a high-quality spectral cube.
9. A high-resolution snapshot hyperspectral imaging method based on deep learning according to claim 7, characterized in that: The specific steps of step S6 include the following steps: S601. Select the initial structure parameters of the broadband multispectral filter array and input them into the simulation software for nested parameter scanning to obtain the spectral response encoding characteristics corresponding to the structure parameters. S602. Use a multi-layer perceptron model and use the scanned data as the training data set. Train a multi-layer perceptron model to obtain the mapping relationship between the broadband multi-spectral filter array structure parameters and the spectral response coding characteristics; S603. Input the selected initial structure parameters into the established mapping relationship, output the spectral response characteristic curve corresponding to the initial structure, open the gradient corresponding to the spectral response characteristic curve and optimize it, and output the optimized spectral response characteristic curve; then input the optimized spectral response characteristic curve into the mapping relationship to obtain the optimized broadband multi-spectral filter array structure parameters; S604. Perform photolithography processing according to the optimized structure parameters to fabricate a broadband multi-spectral filter array; measure the actual spectral response characteristic curve of each metasurface unit in the broadband multi-spectral filter array, and record the deviation between the actual and the designed spectral response characteristic curves; S605. Input the actually measured spectral response characteristic curve into the model and turn off the gradient corresponding to the spectral response characteristic curve; perform model optimization again, correct the spectral reconstruction strategy, and finally obtain the spectral reconstruction strategy corresponding to the spectral response characteristic curve of the actual broadband multi-spectral filter array.
10. A high-resolution snapshot hyperspectral imaging method based on deep learning according to claim 7, characterized in that: The step S8 specifically includes the following steps: S801. Input the optimized broadband multi-spectral filter array structure parameters as initial parameters into the spectral reconstruction network; S802. In the spectral reconstruction network, calculate the gradient of the structure parameters through gradient backpropagation; adjust the structure parameters according to the gradient information to optimize the coding process; S803. During the optimization process, consider the performance of both the coding and decoding processes simultaneously to ensure collaborative optimization of coding and decoding; Improve the overall system performance through multiple iterations; S804. Evaluate the optimized system performance and determine the final broadband multi-spectral filter array structure parameters.
Citation Information
Cited By
Infrared spectrum detection chip-oriented adjustable coding and deep learning reconstruction inversion method and system
CN120974440A
Silicon carbide substrate-based metasurface waveguide array and spectrum sensing system
CN121252958A