Scene-adaptive tunable calculation hyperspectral imaging system
By introducing LCoS-SLM and hardware-guided spectral multi-headed self-attention modules in hyperspectral imaging systems, combined with the CNN-Transformer architecture, the existing system has solved the problem of insufficient adaptability to real-time capture and complex scenes in dynamic scenarios, and achieved high-precision spectral reconstruction and system adaptive optimization.
Patent Information
- Application Number
- CN202510244239.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
The existing hyperspectral imaging systems are difficult to achieve real-time capture in dynamic scenarios, and fail to effectively handle complex and changeable real-world scenarios, resulting in insufficient spectral reconstruction accuracy and difficult to adaptively adjust system parameters.
A tunable computing hyperspectral imaging system with scene adaptability is proposed, and phase encoding is achieved using liquid crystal silicon spatial light modulator (LCoS-SLM) without the need to manufacture diffraction optical components. Combined with hardware-guided spectral multi-head self-attention module and CNN-Transformer architecture, it realizes efficient spectral reconstruction and system adaptive optimization.
It realizes high-precision spectral information acquisition and reconstruction in complex dynamic scenarios, and improves the system's adaptability and accuracy and dynamic adaptability in complex scenarios.
Smart Images

Figure CN120176842A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of hyperspectral imaging, and particularly relates to a scene-adaptive tunable computational hyperspectral imaging system. Background Art
[0002] Hyperspectral imaging technology has evolved from traditional color imaging technology and enables more precise identification and analysis by acquiring the spectral information of objects. Due to its ability to simultaneously obtain spatial and high-resolution spectral information, hyperspectral imaging technology has played an important role in fields such as agriculture, healthcare, mineral exploration, cultural relics protection, and national defense security. However, traditional hyperspectral imaging systems (such as pushbroom, filter-based, and interferometric systems) are difficult to meet the requirements of real-time capture in dynamic scenarios due to limitations such as large equipment volume, high cost, and slow acquisition speed. To address these issues, snapshot hyperspectral imaging technology has emerged.
[0003] A snapshot hyperspectral imaging (SHI) system is an advanced imaging technology that can capture the reflected or radiated data of an object within a continuous wavelength range in a single exposure. It compresses spectral data through the physical encoding of optical elements and combines relevant algorithms for spectral reconstruction, achieving the miniaturization of the device and real-time imaging. According to different encoding methods, SHI systems are mainly divided into three categories: amplitude-based, wavelength-based, and phase-based. Among them, the encoding system based on amplitude modulation has a simple structure and mature technology, but the system structure is complex and vulnerable to external factors such as changes in the external environment and system random noise; the encoding system based on wavelength modulation has portable equipment and low cost, but the spectral reconstruction accuracy is limited. In contrast, the diffraction spectral imaging system based on phase modulation utilizes the refractive index difference of light in different bands by diffraction optical elements to introduce different phase delays, achieving phase encoding and spectral separation. Combined with a spectral reconstruction algorithm based on deep learning, it fully explores the correlation between spatial and spectral features to achieve a high-performance snapshot hyperspectral imaging system.
[0004] Snapshot hyperspectral imaging technology based on traditional RGB images uses a traditional camera or a dedicated array filter to achieve spectral reconstruction only through RGB images. Based on the spectral response function (SRF) and sensor characteristics, the system establishes a robust mapping relationship between the captured image and spectral data, exploring the correlation of spatial spectral features.
[0005] The snapshot spectral imaging technology based on coded aperture mainly adopts a Coded Aperture Snapshot Spectral Imaging (CASSI) system. Through the combination of optical elements such as physical masks, prisms, filters, and lenses, compression and capture are achieved in the spectral dimension, and spectral reconstruction is realized using image reconstruction algorithms.
[0006] The snapshot hyperspectral imaging technology based on diffractive optical elements uses diffractive optical elements (DOEs) to perform phase modulation on optical signals, realizes spectral encoding through a pre-designed Point Spread Function (PSF), and maps the encoded spectral information onto the sensor plane in the form of an RGB image. Subsequently, using a spectral reconstruction algorithm, the captured encoded RGB image is spectrally reconstructed.
[0007] It can be seen that the training of most existing technologies only relies on open-source and ordinary datasets, so it cannot handle the high complexity and variability in real-world scenarios; existing systems fail to simultaneously obtain spectral ground truth data and encoded data in actual scenarios, making it difficult to perform adaptive adjustment and correction of system parameters. Although mature spectral devices can be used to separately collect the ground truth of real scenarios, during the separate collection process, changes in the viewing angle and environmental noise may cause large deviations in the spatial and spectral characteristics between the target scenario and the dataset scenario, thus greatly reducing the accuracy of the reconstructed spectrum; existing technologies fail to fully utilize the real spectral data of the target scenario to optimize the model, so it is difficult to correct the spectral deviation of the system in actual scenarios. At the same time, the current reconstruction algorithms based on diffractive optical imaging mainly use shallow convolutional neural networks, and their image decoding projects directly from the encoded low-dimensional image to the high-dimensional hyperspectral space, ignoring the role of spatial-spectral correlation and spectral characteristics between channels, which affects the reconstruction quality. Existing systems usually cannot simultaneously obtain the real spectral data (Ground truth) and encoded reconstruction data of the scenario, so they lack a reference benchmark for actual scenarios, restricting the adaptive optimization and correction of the system. Summary of the Invention
[0008] To solve the adaptability problem of hyperspectral imaging in various application scenarios, the present invention provides a scene-adaptive tunable computational hyperspectral imaging system that can achieve phase encoding without manufacturing diffractive optical elements and realize hyperspectral imaging of the target scenario.
[0009] A scene-adaptive tunable computational hyperspectral imaging system includes a light-receiving module, a filter wheel, a spatial modulation module, a color CMOS camera, and a reconstruction model; wherein, the filter wheel includes 1 polarizer and 31 narrowband filters with different central wavelengths.
[0010] Let the visible light emitted by the target scene pass through the light receiving module and then enter the polarizer of the filter wheel, and the light emitted from the polarizer forms an initial image; the initial image enters the spatial modulation module for phase modulation to obtain a modulated image; the modulated image enters the color CMOS camera for encoding to obtain an RGB image;
[0011] Let the visible light emitted by the target scene pass through the light receiving module and then enter the narrowband filter of each channel of the filter wheel in turn, and the light emitted from each narrowband filter forms a true value image respectively;
[0012] Input the RGB image and the true value images of 31 channels into the reconstruction model, and the reconstruction model outputs the hyperspectral images of 31 channels corresponding to the target scene.
[0013] Furthermore, the reconstruction model includes a first spatial multi-head self-attention module SAMⅠ, a second spatial multi-head self-attention module SAMⅡ, a first hardware-guided spectral multi-head self-attention module HS-MSA Ⅰ, a second hardware-guided spectral multi-head self-attention module HS-MSA Ⅱ, a first feed-forward network module FNNⅠ to a fourth feed-forward network module FNNⅣ, and 1 convolutional module Conv;
[0014] The first feed-forward network module FNNⅠ extracts the local features of the RGB image to obtain a first local feature map; the second feed-forward network module FNNⅡ extracts the local features of the first local fusion feature map obtained by superimposing the RGB image and the first local feature map to obtain a second local feature map; the third feed-forward network module FNNⅢ extracts the local features of the second local fusion feature map obtained by superimposing the first local fusion feature map and the second local feature map to obtain a third local feature map; the fourth feed-forward network module FNNⅣ extracts the local features of the third local fusion feature map obtained by superimposing the second local fusion feature map and the third local feature map to obtain a fourth local feature map;
[0015] The first spatial multi-head self-attention module SAMⅠ extracts the spatial features of the true value images of each channel to obtain a first spatial feature map; the first hardware-guided spectral multi-head self-attention module HS-MSA Ⅰ extracts the spectral channel features of the first spatial feature map to obtain a first spectral channel feature map; the second spatial multi-head self-attention module SAMⅡ extracts the spatial features of the first spectral fusion feature map obtained by superimposing the true value images of each channel and the first spectral channel feature map to obtain a second spatial feature map; the second hardware-guided spectral multi-head self-attention module HS-MSA Ⅱ extracts the spectral channel features of the second spatial feature map to obtain a second spectral channel feature map;
[0016] The convolution module Conv performs feature aggregation on the global feature map obtained by superimposing the first spectral fusion feature map, the second spectral channel feature map, and the fourth local feature map, and finally outputs the hyperspectral image of each channel.
[0017] Further, any hardware-guided spectral multi-head self-attention module HS-MSA includes a feature embedding unit, a three-way feature mapping unit, a first product unit, a second product unit, a third product unit, and a superimposing unit;
[0018] The feature embedding unit is used to convert the feature map received by itself from a high-dimensional sparse state to a low-dimensional dense state, obtaining a low-dimensional dense feature map;
[0019] The three-way feature mapping unit is used to project the low-dimensional dense feature map through a K transformation matrix, a Q transformation matrix, and a V transformation matrix to obtain a K-transformed feature map, a Q-transformed feature map, and a V-transformed feature map respectively;
[0020] The first product unit is used to multiply the K-transformed feature map and the Q-transformed feature map to obtain a first product feature map;
[0021] The second product unit is used to multiply the V-transformed feature map element-wise with the point spread function to obtain a second product feature map; among them, the phase modulation pattern loaded on the spatial modulation module is different, and the corresponding point spread function is different;
[0022] The third product unit is used to multiply the first product feature map and the second product feature map to obtain a third product feature map;
[0023] The superimposing unit is used to fuse the third product feature map with the position embedding information corresponding to the target scene to obtain the feature map finally output by the spectral multi-head self-attention module HS-MSA.
[0024] Further, the loss function L used when training the reconstruction model total is as follows:
[0025]
[0026] where, I s is the ground truth of the hyperspectral image of 31 channels corresponding to the set scene selected from the open-source scene dataset, is the predicted value of the hyperspectral image of 31 channels corresponding to the set scene selected from the open-source scene dataset output by the reconstruction model, α is the open-source scene reconstruction loss weight, K is the number of pixels in the set scene, J is the number of pixels in the target scene, is the predicted value of the hyperspectral image of 31 channels corresponding to the target scene output by the reconstruction model, I tis the true value of the hyperspectral image of 31 channels corresponding to the target scene, ω is the regularization parameter, and β is the regularization weight.
[0027] Furthermore, the method for obtaining the true value of the hyperspectral image of 31 channels corresponding to the target scene or any set scene is as follows:
[0028] S1: Rotate the filter wheel to the narrowband filter of any channel;
[0029] S2: Let the visible light emitted by the current scene pass through the light collection module and then enter the narrowband filter of the current channel, and the light emitted from the narrowband filter of the current channel forms the current filtered image I;
[0030] S3: Control the loading phase modulation of the spatial modulation module to 0, so that the current filtered image directly passes through the spatial modulation module to obtain the current filtered image II without phase modulation;
[0031] S4: Let the current filtered image II enter the color CMOS camera for encoding to obtain the RGB image of the current channel. At the same time, use the RGB image of the current channel as the true value of the hyperspectral image of the current scene in the current channel;
[0032] S5: Rotate the filter wheel to the narrowband filter of the next channel, and re-execute steps S2 to S4 until all the narrowband filters of all channels are traversed to obtain the true value of the hyperspectral image of 31 channels corresponding to the current scene.
[0033] Furthermore, when the application environment changes, it is necessary to correct the tunable computational hyperspectral imaging system, and the correction method is as follows:
[0034] Use the scanning imaging mode to capture the true value image of 31 channels corresponding to the changed application environment as the new input of the reconstruction model, and on the basis of the pre-trained reconstruction model, continue to train the reconstruction model by changing the phase modulation pattern loaded by the spatial modulation module and the network parameters of the reconstruction model until the loss function of the reconstruction model is less than the set threshold, so that the tunable computational hyperspectral imaging system adaptively matches the changed application environment.
[0035] Furthermore, the light collection module is a single lens or an objective lens group.
[0036] Furthermore, the spatial modulation module is implemented by a fully transmissive spatial light modulator;
[0037] The initial image is directly transmitted from the fully transmissive spatial light modulator to the color CMOS camera, and phase modulation is performed during the transmission process to obtain the modulated image.
[0038] Furthermore, the spatial modulation module is implemented by a combination of a total reflection spatial light modulator and a beam splitter;
[0039] The initial image is first reflected by the beam splitter to the total reflection spatial light modulator, then reflected back to the beam splitter by the total reflection spatial light modulator, and finally transmitted from the beam splitter to the color CMOS camera. Among them, the initial image is phase-modulated when reflected by the total reflection spatial light modulator to obtain a modulated image.
[0040] Furthermore, the spatial modulation module is implemented by a combination of a total reflection digital micromirror and a beam splitter;
[0041] The initial image is first reflected by the beam splitter to the total reflection digital micromirror, then reflected back to the beam splitter by the total reflection digital micromirror, and finally transmitted from the beam splitter to the color CMOS camera. Among them, the initial image is phase-modulated when reflected by the total reflection digital micromirror to obtain a modulated image.
[0042] Beneficial effects:
[0043] 1. The present invention provides a scene-adaptive tunable computational hyperspectral imaging system, and proposes a hardware-guided spectral multi-head attention structure module. This module can introduce some parameters of the hardware system as prior knowledge into the spectral reconstruction model, and make full use of the optical characteristics of the hardware system to guide the network feature extraction process; based on the adjustability of the LCoS-SLM, it realizes the dynamic switching between scanning data acquisition and snapshot imaging, and combines the scene-adaptive mechanism to modify the hardware encoding and software decoding parameters of the system, improving the adaptability of the system in complex scenes; that is to say, the present invention adopts an adjustable end-to-end diffraction snapshot hyperspectral imaging system, which can achieve phase encoding without manufacturing diffraction optical elements, and at the same time allows dynamic adjustment of the optical encoder and network decoder; at the same time, the present invention deeply integrates the hardware flexibility with the algorithm, effectively improving the accuracy and dynamic adaptability of spectral imaging, especially in complex dynamic scenes, providing a new solution for the practical application of hyperspectral imaging technology.
[0044] 2. The present invention provides a scene-adaptive tunable computational hyperspectral imaging system. Utilizing the spatial multiplexing technology and the tunability of the LCoS-SLM, TOSHI performs hardware integration on diffraction snapshot hyperspectral imaging and scanning hyperspectral imaging, realizing the switchable acquisition of phase-encoded images and ground truth. At the same time, the ground truth assists the phase-encoded image to improve the spectral reconstruction performance.
[0045] 3. The present invention provides a scene - adaptive tunable computational hyperspectral imaging system, which introduces a scene adaptation mechanism based on CNN - Transformer, adopts spatial - spectral multi - head self - adaptation and hardware - guided spectral multi - head self - adaptation, and makes full use of the spatial and spectral features of the real - world scene, thereby achieving high - quality adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is the original diagram of a scene - adaptive tunable computational hyperspectral imaging system provided by the present invention;
[0047] Figure 2 It is the imaging flow chart of TOSHI provided by the present invention;
[0048] Figure 3 It is the proposed CNN - Transformer spectral reconstruction network training framework provided by the present invention;
[0049] Figure 4 It is the hardware - guided spectral multi - head attention mechanism provided by the present invention;
[0050] Figure 5 It is the spatial attention mechanism provided by the present invention;
[0051] Figure 6 It is the feed - forward neural network provided by the present invention;
[0052] Figure 7 It is the network training strategy of TOSHI provided by the present invention;
[0053] Figure 8 It is the first implementation manner of a scene - adaptive tunable computational hyperspectral imaging system provided by the present invention;
[0054] Figure 9 It is the second implementation manner of a scene - adaptive tunable computational hyperspectral imaging system provided by the present invention;
[0055] Figure 10 It is the third implementation manner of a scene - adaptive tunable computational hyperspectral imaging system provided by the present invention;
[0056] Figure 11 It is the fourth implementation manner of a scene - adaptive tunable computational hyperspectral imaging system provided by the present invention;
[0057] Figure 12 It is the fifth implementation manner of a scene - adaptive tunable computational hyperspectral imaging system provided by the present invention;
[0058] Figure 13The sixth implementation of a scene - adaptive tunable computational hyperspectral imaging system provided by the present invention. Detailed implementation manners
[0059] To enable those skilled in the art to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application.
[0060] To solve the problems that existing snapshot hyperspectral imaging systems are difficult to adapt to complex dynamic scenes, have insufficient spectral reconstruction accuracy, and poor generalization performance under a fixed phase modulation mode, the present invention proposes an adjustable and optimized hyperspectral imaging system based on a liquid crystal on silicon spatial light modulator (LCoS - SLM). This system realizes high - precision spectral information acquisition and reconstruction in dynamic scenes through a cooperative optimization mechanism of programmable optical hardware and deep learning algorithms. The hardware structure of the system is as Figure 1 shown, including a light - receiving module, a filter wheel, a spatial modulation module, a color CMOS camera, and a reconstruction model; among them, the filter wheel includes 1 polarizer and 31 narrow - band filters with different central wavelengths, the wavelength coverage is 400 - 700 nm, and the bandwidth is 10 nm;
[0061] As Figure 2 shown, let the visible light emitted by the target scene pass through the light - receiving module and then enter the polarizer of the filter wheel, then the light emitted from the polarizer forms an initial image; the initial image enters the spatial modulation module for phase modulation to obtain a modulated image; the modulated image enters the color CMOS camera for encoding to obtain an RGB image;
[0062] Let the visible light emitted by the target scene pass through the light - receiving module and then sequentially enter the narrow - band filters of each channel of the filter wheel, then the light emitted from each narrow - band filter forms a ground - truth image respectively;
[0063] Input the RGB image and the ground - truth images of 31 channels into the reconstruction model, and the reconstruction model outputs the hyperspectral images of 31 channels corresponding to the target scene.
[0064] Furthermore, as Figure 3 shown, the reconstruction model includes a first spatial multi - head self - attention module SAMⅠ, a second spatial multi - head self - attention module SAMⅡ, a first hardware - guided spectral multi - head self - attention module HS - MSA Ⅰ, a second hardware - guided spectral multi - head self - attention module HS - MSA Ⅱ, a first feed - forward network module FNNⅠ to a fourth feed - forward network module FNNⅣ, and 1 convolutional module Conv;
[0065] The first feedforward network module FNNⅠ extracts the local features of the RGB image to obtain the first local feature map; the second feedforward network module FNNⅡ extracts the local features of the first local fusion feature map obtained by superimposing the RGB image and the first local feature map to obtain the second local feature map; the third feedforward network module FNNⅢ extracts the local features of the second local fusion feature map obtained by superimposing the first local fusion feature map and the second local feature map to obtain the third local feature map; the fourth feedforward network module FNNⅣ extracts the local features of the third local fusion feature map obtained by superimposing the second local fusion feature map and the third local feature map to obtain the fourth local feature map;
[0066] The first spatial multi-head self-attention module SAMⅠ extracts the spatial features of the true-value image of each channel to obtain the first spatial feature map; the first hardware-guided spectral multi-head self-attention module HS-MSA Ⅰ extracts the spectral channel features of the first spatial feature map to obtain the first spectral channel feature map; the second spatial multi-head self-attention module SAMⅡ extracts the spatial features of the first spectral fusion feature map obtained by superimposing the true-value image of each channel and the first spectral channel feature map to obtain the second spatial feature map; the second hardware-guided spectral multi-head self-attention module HS-MSAⅡ extracts the spectral channel features of the second spatial feature map to obtain the second spectral channel feature map;
[0067] The convolution module Conv performs feature aggregation on the global feature map obtained by superimposing the first spectral fusion feature map, the second spectral channel feature map, and the fourth local feature map, and finally outputs the hyperspectral image of each channel.
[0068] That is to say, the present invention combines the CNN and Transformer architectures, and proposes an end-to-end spectral reconstruction network model CNN-Transformer spectral reconstruction framework, which realizes efficient feature extraction and reconstruction by introducing a spatial attention mechanism (Spatial Attention Mechanism, SAM) and a hardware-guided spectral multi-head attention module (HS-MSA). The CNN part is composed of a feedforward neural network (Feedforward Neural Network, FNN), which is mainly responsible for local feature extraction. Combining with Transformer can enhance the model's attention to key regions, thereby improving the reconstruction effect. SAM can more effectively focus on spatial features, which helps to extract local and global information in hyperspectral images and improve the feature expression ability. HS-MSA performs feature modeling in the spectral dimension and can capture long-range dependencies between spectral channels. This module is designed by combining hardware information (loaded on the LCoS-SLM phase modulation map), making the spectral attention more in line with the characteristics of the actual imaging system, thereby improving the information interaction ability between spectral channels and optimizing the reconstruction quality of spectral features.
[0069] Furthermore, as Figure 4 shown, any hardware-guided spectral multi-head self-attention module HS-MSA includes a feature embedding unit, a three-way feature mapping unit, a first product unit, a second product unit, a third product unit, and a superimposing unit;
[0070] The feature embedding unit is used to convert the feature map received by itself from a high-dimensional sparse state to a low-dimensional dense state, obtaining a low-dimensional dense feature map; the three-way feature mapping unit is used to project the low-dimensional dense feature map through a K transformation matrix, a Q transformation matrix, and a V transformation matrix to respectively obtain a K-transformed feature map, a Q-transformed feature map, and a V-transformed feature map; the first product unit is used to multiply the K-transformed feature map and the Q-transformed feature map to obtain a first product feature map; the second product unit is used to multiply the V-transformed feature map element-wise with a point spread function to obtain a second product feature map; wherein, different phase modulation patterns loaded on the spatial modulation module correspond to different point spread functions; the third product unit is used to multiply the first product feature map and the second product feature map to obtain a third product feature map; the superimposing unit is used to fuse the third product feature map with the position embedding information corresponding to the target scene to obtain the feature map finally output by the spectral multi-head self-attention module HS-MSA.
[0071] It can be seen that the hardware-guided spectral multi-head attention module proposed by the present invention effectively solves the problem that it is difficult for traditional methods to accurately model the complex relationships between spectral channels by introducing the point spread function generated by the hardware system as prior knowledge into the attention calculation process. Specifically, HS-MSA adopts a three-way feature mapping mechanism to project the input features through three transformation matrices of query (Q), key (K), and value (V), and at the same time fuses the PSF information with the value features to generate a hardware-guided attention map. This design enables the module to make full use of the optical characteristics of the hardware system to guide the feature extraction process, so as to more accurately identify and maintain the correlation between different bands during spectral reconstruction.
[0072] The structure of the spatial multi-head self-attention module SAM is as Figure 5 shown, and the structure of the feed-forward neural network FNN is as Figure 6 shown, which will not be elaborated in the present invention.
[0073] Furthermore, Figure 7Shows the network hierarchical training strategy of the TOSHI system, which realizes an efficient scene adaptation mechanism. Specifically, the entire training process is divided into two stages: source scene training and target scene optimization. In the source scene training stage, the system first performs pre-training using a large open-source dataset (such as the ICVL dataset). In this stage, the system will optimize the parameters of the CTHR-net reconstruction network (including two parts: the imaging hardware encoder and the feature extractor). Among them, the hardware encoder is used to optimize the point spread function of the system, and the feature extractor is responsible for extracting the spatio-temporal features and inter-channel correlations of the spectral data. Through the training of this stage, the system will obtain a basic model suitable for general scenes. In the target scene optimization stage, the system will introduce the real spectral data of a specific scene to optimize the basic model. This stage adopts an innovative parameter selective freezing / unfreezing strategy: first, freeze the hardware parameter structure and keep the phase modulation template unchanged; then, through the feature fusion module, fuse the target scene features and the source scene features; finally, adopt the selective parameter freezing strategy to only update specific layers in the network, greatly improving the optimization efficiency. This training strategy not only speeds up the model convergence speed but also effectively prevents the occurrence of overfitting phenomena.
[0074] To achieve efficient scene adaptive optimization, the present invention designs a loss function mechanism with multi-component fusion. This loss function consists of four core components: source scene reconstruction loss, target scene reconstruction loss, gradient alignment loss, and network parameter regularization term. Its mathematical expression is:
[0075]
[0076] Among them, I s is the ground truth of the hyperspectral image of 31 channels corresponding to the set scene selected from the open-source scene dataset, is the predicted value of the hyperspectral image of 31 channels corresponding to the set scene selected from the open-source scene dataset output by the reconstruction model. α is the weight of the open-source scene reconstruction loss, K is the number of pixels in the set scene, and J is the number of pixels in the target scene. is the predicted value of the hyperspectral image of 31 channels corresponding to the target scene output by the reconstruction model, I t is the ground truth of the hyperspectral image of 31 channels corresponding to the target scene, ω is the regularization parameter, and β is the regularization weight.
[0077] It should be noted that in the above expression of the loss function, the source scene reconstruction loss (the first term) evaluates the pre-training performance of the network on the open-source dataset by minimizing the difference between the predicted value of the hyperspectral image and the ground truth I of the hyperspectral image sThe L1 norm difference between them ensures that the model maintains the basic reconstruction ability. The parameter α regulates its weight contribution for normalizing the loss value; the target scene reconstruction loss (the second term) aims at the reconstruction accuracy of a specific application scenario. By minimizing the predicted value of the hyperspectral image in the actual scene and the true value I of the hyperspectral image t is optimized by the L1 norm difference. The weight coefficient (1 - α) dynamically adjusts the adaptability of the model to a specific scene and is also used for loss value normalization; the gradient alignment loss (the third term) enhances the model's ability to maintain scene texture details by calculating the gradient difference of the reconstructed image. The third term is an L2 regularization term, which is used to control the model complexity. Its strength is controlled by the parameter β. By restricting the norm of the network parameter ω, it is used to prevent the model from overfitting and enhance the generalization ability of the network model. This multi-component loss function maintains the basic performance through the source scene loss, guides scene adaptation through the target scene loss, and controls the model complexity through the regularization term. The dynamic balance of the three realizes the unity of generality and specificity.
[0078] The optimization process of the loss function also adopts a phased strategy. In the pre-training stage, it focuses on optimizing the source scene loss and the gradient alignment loss to establish a model with strong basic reconstruction ability; in the transfer adaptation stage, by dynamically adjusting the values of α and β, it balances the retention of source scene features and the learning of target scene features to achieve efficient knowledge transfer; in the fine-tuning stage, it focuses on optimizing the target scene loss to achieve precise adaptation to a specific scene. During the entire optimization process, the regularization strength β is also dynamically adjusted according to the training progress to ensure the stability and generalization ability of the model.
[0079] In summary, the present invention proposes an adjustable snapshot hyperspectral imaging system TOSHI, which can switch to collect phase-encoded images and ground truth information, utilize the current target scene data, and adopt spatial-spectral multi-head adaptive and hardware-guided spectral multi-head adaptive modules to optimize the parameters of the encoding optical hardware and the decoding algorithm in real time, effectively adapting to the actual scene. The present invention innovatively proposes a multi-component fusion loss function mechanism in the spectral reconstruction network. In addition, a new hyperspectral dataset is collected using this system.
[0080] Furthermore, the method for obtaining the true value of the hyperspectral image of 31 channels corresponding to the target scene or any set scene is as follows:
[0081] S1: Rotate the filter wheel to the narrowband filter of any channel;
[0082] S2: Let the visible light emitted by the current scene pass through the light collection module and then enter the narrowband filter of the current channel. The light emitted from the narrowband filter of the current channel forms the current filtered image I;
[0083] S3: Control the loading phase modulation of the spatial modulation module to 0, so that the current filtered image directly passes through the spatial modulation module to obtain the current filtered image II without phase modulation;
[0084] S4: Let the current filtered image II enter the color CMOS camera for encoding to obtain the RGB image of the current channel. At the same time, use the RGB image of the current channel as the true value of the hyperspectral image of the current scene in the current channel;
[0085] S5: Rotate the filter wheel to the narrowband filter of the next channel, and re - execute steps S2 - S4 until the narrowband filters of all channels are traversed to obtain the true values of the hyperspectral images of 31 channels corresponding to the current scene.
[0086] It should be noted that when the application environment changes, the tunable computational hyperspectral imaging system needs to be calibrated, and the calibration method is as follows:
[0087] Capture the true - value images of 31 channels corresponding to the changed application environment in the scanning imaging mode as the new input of the reconstruction model. Based on the pre - trained reconstruction model, continue to train the reconstruction model by changing the phase modulation pattern loaded by the spatial modulation module and the network parameters of the reconstruction model until the loss function of the reconstruction model is less than the set threshold, so that the tunable computational hyperspectral imaging system adaptively matches the changed application environment.
[0088] It should be noted that the present invention proposes a tunable snapshot hyperspectral imaging system (Tunable Optimally - coded Snapshot Hyperspectral Imaging system, TOSHI), and its light - receiving module and spatial modulation module have different implementation forms.
[0089] As Figure 8 shown, the light - receiving module of the present invention is implemented by an objective lens system, and the spatial modulation module is implemented by a lens, a spatial light modulator, and a beam - splitting prism. The initial image is first reflected by the beam splitter to the total - reflection spatial light modulator, then reflected back to the beam splitter by the total - reflection spatial light modulator, and finally transmitted from the beam splitter to the color CMOS camera. Among them, the initial image is phase - modulated when reflected by the total - reflection spatial light modulator to obtain a modulated image.
[0090] It should be noted that the tunable computational hyperspectral imaging system of the present invention has two working modes, namely the scanning hyperspectral imaging mode and the snapshot hyperspectral imaging mode;
[0091] (1) Scanning Hyperspectral Imaging Mode: During the ground truth acquisition process, the narrowband filter is used in conjunction with a rotating filter wheel to collect spatial spectral ground truth across the entire visible spectrum. The visible light emitted by the object first passes through the objective lens and is filtered by a 10-nm narrowband filter, and then an initial image is formed. This initial image is then transmitted to the camera sensor through a lens, a beam splitter prism, and a liquid crystal on silicon - spatial light modulator (LCoS - SLM). Given the high - resolution phase - modulation ability of the LCoS - SLM, it is configured as a planar mirror in this mode. Using a 31 - channel narrowband filter, the TOSHI system scans the target scene band - by - band and captures the light intensity information of each band in sequence. After the scanning is completed, 31 single - channel images are generated, covering the entire spectral range. Then these images are integrated into a hyperspectral data cube, accurately reflecting the spectral characteristics of the scene. This data can be used as ground truth, providing a high - precision reference for subsequent spectral reconstruction and scene adaptation calibration.
[0092] (2) Snapshot Hyperspectral Imaging Mode: During the phase - encoded image capture process, the reflected light first forms an initial image through the objective lens and a polarizer. This initial image is guided to the LCoS - SLM through a beam splitter and undergoes phase modulation on the LCoS - SLM. Different refractive indices of the liquid crystal layer on the LCoS - SLM generate different phase delays, thus achieving spectral phase modulation and encoding. The light emerging from the LCoS - SLM passes through the beam splitter again and reaches a color CMOS camera, generating an encoded RGB image. Subsequently, these images are reconstructed into 31 - channel hyperspectral images using a spectral reconstruction network.
[0093] Furthermore, as Figure 9 shown, the objective lens can be replaced with a single lens, which can simplify the hardware design while reducing the system volume and cost.
[0094] As Figure 10 shown, a transmissive spatial light modulator can be used to replace the reflective spatial light modulator, removing the beam splitter, simplifying the system volume, and making the system more compact.
[0095] As Figure 11 shown, a digital micromirror can be used to replace the spatial light modulator.
[0096] As Figure 12 shown, replacing the objective lens with a single lens and using a transmissive spatial light modulator to replace the reflective spatial light modulator can further simplify the hardware design and make the system more compact.
[0097] As Figure 13 shown, replacing the objective lens with a single lens and using a digital micromirror to replace the spatial light modulator can simplify the hardware design while reducing the system volume and cost.
[0098] In summary, the key technologies of the tunable computational hyperspectral imaging system provided by the present invention are as follows:
[0099] 1. By using switchable scanning hyperspectral imaging and phase-encoded snapshot hyperspectral imaging modes, the adaptability of computational hyperspectral imaging in various application scenarios is solved;
[0100] 2. A scene adaptive optimization mechanism, which includes real-scene spectral data acquisition and correction, online optimization and dynamic adjustment capabilities, and feature distribution adjustment based on batch normalization;
[0101] 3. The CTHR-net spectral reconstruction network, which is jointly composed of a CNN-Transformer hybrid architecture (CTHR-net), a spatio-spectral multi-head self-attention module (SSAM), a hardware-guided spectral multi-head self-attention module (HS-MSA), and a multi-loss fusion optimization mechanism;
[0102] 4. A parameter dynamic adjustment strategy, which includes a phased training mechanism, a selective parameter freezing / unfreezing strategy, and a feature fusion method for source and target scenes;
[0103] 5. The CAESAR hyperspectral dataset, which contains a batch of high-quality spectral data that has been standardized, corrected, annotated, covers diverse scenes, and has a high spatial resolution (5120×5120) and a high spectral resolution (10nm).
[0104] It can be seen that, compared with the prior art, the present invention has the following advantages:
[0105] First of all, the present invention proposes a hardware-guided spectral multi-head attention structure module, which can introduce some parameters of the hardware system as prior knowledge into the spectral reconstruction network, and make full use of the optical characteristics of the hardware system to guide the network feature extraction process; based on the adjustability of the LCoS-SLM, the dynamic switching between scanning data acquisition and snapshot imaging is realized, and the hardware encoding and software decoding parameters of the system are modified in combination with the scene adaptive mechanism, improving the adaptability of the system in complex scenes; the present invention achieves a resolution of up to 5120×5120 pixels, an angular resolution of 0.06 degrees, a spectral resolution of 10nm, and a temporal resolution of up to 61.6fps (spatial resolution of 1024×1024), meeting the requirements of online applications.
[0106] In addition, the present invention provides a brand-new hyperspectral dataset (spectral resolution of 10 nm, image resolution of 5120×5120 pixels, 31 spectral channels, wavelength range of 400 - 700 nm), covering 120 sets of high-resolution spectral data from 15 different indoor and outdoor scenes, which can be used as the simulated input for various spectral imaging systems, providing complete benchmark support for the research and development and evaluation of hyperspectral imaging systems.
[0107] Finally, the present invention also proposes a loss function for multi-component fusion, which realizes the joint optimization of the hardware encoder and software decoder for the actual scene by using the open-source dataset and the captured ground truth information.
[0108] Certainly, the present invention can also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can certainly make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A scene-adaptive tunable computational hyperspectral imaging system, characterized in that: It includes a light receiving module, a filter wheel, a spatial modulation module, a color CMOS camera and a reconstruction model; wherein the filter wheel includes a polarizer and 31 narrow-band filters with different central wavelengths; The visible light emitted by the target scene passes through the light receiving module and enters the polarizer of the filter wheel, and the light emitted from the polarizer forms an initial image; the initial image enters the spatial modulation module for phase modulation to obtain a modulated image; the modulated image enters the color CMOS camera for encoding to obtain an RGB image; The visible light emitted by the target scene passes through the light receiving module and then enters the narrow-band filters of each channel of the filter wheel in sequence, and the light emitted from each narrow-band filter forms a true value image respectively; The RGB image and the true value image of 31 channels are input into the reconstruction model, and the reconstruction model outputs a hyperspectral image of 31 channels corresponding to the target scene.
2. A scene-adaptive tunable computational hyperspectral imaging system as claimed in claim 1, characterized in that: The reconstructed model includes the first spatial multi-head self-attention module SAMⅠ, the second spatial multi-head self-attention module SAMⅡ, the first hardware-guided spectral multi-head self-attention module HS-MSAⅠ, the second hardware-guided spectral multi-head self-attention module HS-MSAⅡ, the first feedforward network module FNNⅠ to the fourth feedforward network module FNNⅣ and a convolution module Conv; The first feedforward network module FNNⅠ extracts local features of the RGB image to obtain a first local feature map; The second feedforward network module FNNⅡ extracts the local features of the first local fusion feature map obtained by superimposing the RGB image and the first local feature map to obtain the second local feature map; the third feedforward network module FNNⅢ extracts the local features of the second local fusion feature map obtained by superimposing the first local fusion feature map and the second local feature map to obtain the third local feature map; the fourth feedforward network module FNNⅣ extracts the local features of the third local fusion feature map obtained by superimposing the second local fusion feature map and the third local feature map to obtain the fourth local feature map; The first spatial multi-head self-attention module SAMⅠ extracts the spatial features of the true value image of each channel to obtain a first spatial feature map; the first hardware-guided spectral multi-head self-attention module HS-MSA Ⅰ extracts the spectral channel features of the first spatial feature map to obtain a first spectral channel feature map; The second spatial multi-head self-attention module SAMⅡ extracts the spatial features of the first spectral fusion feature map obtained by superimposing the true value image of each channel and the first spectral channel feature map to obtain the second spatial feature map; The second hardware-guided spectral multi-head self-attention module HS-MSA II extracts the spectral channel features of the second spatial feature map to obtain the second spectral channel feature map; The convolution module Conv performs feature aggregation on a global feature map obtained by superimposing the first spectral fusion feature map, the second spectral channel feature map and the fourth local feature map, and finally outputs a hyperspectral image of each channel.
3. A scene-adaptive tunable computational hyperspectral imaging system as claimed in claim 2, characterized in that: Any hardware-guided spectral multi-head self-attention module HS-MSA includes a feature embedding unit, a three-way feature mapping unit, a first product unit, a second product unit, a third product unit, and a superposition unit; The feature embedding unit is used to convert the feature map received by itself from a high-dimensional sparse state to a low-dimensional dense state to obtain a low-dimensional dense feature map; The three-way feature mapping unit is used to project the low-dimensional dense feature map through the K transformation matrix, the Q transformation matrix, and the V transformation matrix to obtain the K transformation feature map, the Q transformation feature map, and the V transformation feature map respectively; The first product unit is used to multiply the K transformation feature map and the Q transformation feature map to obtain a first product feature map; The second product unit is used to multiply the V transformation feature map and the point spread function element by element to obtain a second product feature map; wherein the phase modulation patterns loaded on the spatial modulation module are different, and the corresponding point spread functions are different; The third product unit is used to multiply the first product feature map and the second product feature map to obtain a third product feature map; The superposition unit is used to fuse the third product feature map with the position embedding information corresponding to the target scene to obtain the feature map finally output by the spectral multi-head self-attention module HS-MSA.
4. The scene-adaptive tunable computational hyperspectral imaging system according to claim 1, characterized in that: The loss function L used when training the reconstruction model total as follows: Among them, I s is the true value of the hyperspectral image of 31 channels corresponding to the set scene selected from the open source scene dataset, is the predicted value of the hyperspectral image of 31 channels corresponding to the set scene selected from the open source scene dataset output by the reconstruction model, α is the reconstruction loss weight of the open source scene, K is the number of pixels of the set scene, J is the number of pixels of the target scene, is the predicted value of the hyperspectral image of 31 channels corresponding to the target scene output by the reconstruction model, I t is the true value of the hyperspectral image of 31 channels corresponding to the target scene, ω is the regularization parameter, and β is the regularization weight.
5. A scene-adaptive tunable computational hyperspectral imaging system as claimed in claim 4, characterized in that: The method for obtaining the true value of the hyperspectral image of 31 channels corresponding to the target scene or any set scene is as follows: S1: Rotate the filter wheel to select the narrowband filter of any channel; S2: The visible light emitted by the current scene passes through the light receiving module and then enters the narrow-band filter of the current channel. The light emitted from the narrow-band filter of the current channel forms the current filtered image I; S3: Control the loading phase modulation of the spatial modulation module to be 0, so that the current filtered image directly passes through the spatial modulation module, and obtains the current filtered image II without phase modulation; S4: Let the current filtered image II enter the color CMOS camera for encoding to obtain the RGB image of the current channel. At the same time, the RGB image of the current channel is used as the true value of the hyperspectral image of the current scene in the current channel; S5: Rotate the filter wheel to the narrowband filter of the next channel, and re-execute steps S2 to S4 until the narrowband filters of all channels are traversed to obtain the true values of the hyperspectral images of 31 channels corresponding to the current scene.
6. The scene-adaptive tunable computational hyperspectral imaging system according to claim 1, characterized in that: When the application environment changes, the tunable computational hyperspectral imaging system needs to be calibrated, and the calibration method is as follows: The scanning imaging mode is used to capture the true value images of 31 channels corresponding to the changed application environment as new inputs of the reconstruction model, and on the basis of the pre-trained reconstruction model, the reconstruction model is continuously trained by changing the phase modulation pattern loaded by the spatial modulation module and the network parameters of the reconstruction model until the reconstruction model loss function is less than the set threshold, so that the tunable computational hyperspectral imaging system can adaptively match the changed application environment.
7. The scene-adaptive tunable computational hyperspectral imaging system according to claim 1, characterized in that: The light receiving module is a single lens or an objective lens group.
8. The scene-adaptive tunable computational hyperspectral imaging system according to claim 1, characterized in that: The spatial modulation module is implemented by a fully transmissive spatial light modulator; The initial image is directly transmitted from the fully-transmitting spatial light modulator to the color CMOS camera, and is phase modulated during the transmission process to obtain a modulated image.
9. The scene-adaptive tunable computational hyperspectral imaging system according to claim 1, characterized in that: The spatial modulation module is implemented by a combination of a total reflection spatial light modulator and a beam splitter; The initial image is first reflected by the beam splitter to the total reflection spatial light modulator, then reflected back to the beam splitter by the total reflection spatial light modulator, and finally transmitted from the beam splitter to the color CMOS camera, wherein the initial image is phase modulated when reflected by the total reflection spatial light modulator to obtain a modulated image.
10. The scene-adaptive tunable computational hyperspectral imaging system according to claim 1, characterized in that: The spatial modulation module is implemented by a combination of a total reflection digital micro-mirror and a beam splitter; The initial image is first reflected by the beam splitter to the total reflection digital micro-mirror, then reflected back to the beam splitter by the total reflection digital micro-mirror, and finally transmitted from the beam splitter to the color CMOS camera, wherein the initial image is phase modulated when reflected by the total reflection digital micro-mirror to obtain a modulated image.
Citation Information
Cited By
Surrounding rock quality grading method and system suitable for being carried by unmanned aerial vehicle in tunnel
CN120846999A
Monolithic integration calculation reflection spectrometer and spectrum acquisition and reconstruction method
CN120890552A