Video-level lensless spectral imaging system based on color mask and method thereof

The lensless spectral imaging system, which combines color masks with dynamic grayscale masks, solves the problem of low reconstruction accuracy caused by fixed encoding, and achieves high spectral resolution and video-level imaging, making it suitable for lightweight, embedded hyperspectral video applications.

CN121982123APending Publication Date: 2026-05-05NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV
Filing Date
2025-12-16
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

The encoding method of existing lensless spectral imaging systems is fixed and cannot be adjusted, resulting in low reconstruction accuracy, inability to adapt to dynamic scenes, and inability to achieve high-quality video-level spectral imaging.

Method used

A video-grade lensless spectral imaging system based on color masks is adopted, which combines programmable dynamic grayscale masks with fixed random color masks. Through dynamic control coding and multi-frame reconstruction algorithms, high spectral resolution and video-grade imaging are achieved.

Benefits of technology

It achieves high-precision spectral reconstruction, supports real-time monitoring and analysis of dynamic scenes, and is compact and low-cost, making it suitable for portable devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121982123A_ABST
    Figure CN121982123A_ABST
Patent Text Reader

Abstract

The invention discloses a video-level lensless spectral imaging system based on a color mask and a method thereof. The system comprises a filtering window module, a dynamic adjustable spatial spectrum modulation module, a color acquisition module, a sliding window module and a fusion reconstruction module, and the filtering window module, the dynamic adjustable spatial spectrum modulation module and the color acquisition module are sequentially arranged. The fusion reconstruction module is respectively connected with the dynamic adjustable spatial spectrum modulation module, the color acquisition module and the sliding window module, and the sliding window module is connected with the color acquisition module. According to the invention, through dynamic coding and multi-frame information fusion, the space and spectrum quality of a reconstructed image is effectively improved, continuous capture of a hyperspectral video of a dynamic scene is realized, and the system has the remarkable advantages of compact structure, flexible regulation and control and high imaging quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a video-level lensless spectral imaging system and method based on a color mask, belonging to the field of spectral imaging technology. Background Technology

[0002] Snapshot spectral imaging technology can capture the three-dimensional spatial-spectral information of a scene in a single exposure, and has wide application value in fields such as industrial production inspection, smart agriculture, combustion dynamics analysis, and environmental resource exploration. With the rapid development of technology, emerging applications such as wearable devices, autonomous driving, robotics, and virtual / augmented reality are placing higher demands on spectral imaging systems. These applications require spectral imaging systems to be miniaturized, low-cost, high-speed, and have high spectral resolution.

[0003] Currently, snapshot spectral imaging systems exhibit a significant trend towards miniaturization. From early reliance on elongated relay lens assemblies and large-volume beam splitters, to recent advancements in utilizing novel materials such as DOEs and metasurfaces, or pixel-level filtering, snapshot spectral imaging systems have been reduced to the size of traditional lens imaging systems. However, these systems still haven't escaped the limitations of lens-focused imaging; the overall system size cannot be further reduced due to the influence of lens focal length.

[0004] Therefore, the opportunity for further miniaturization lies in its combination with lensless imaging technology. The "point-to-many" imaging configuration and integrated coding paradigm of lensless imaging can further reduce the scale of imaging systems to near-sensor levels, becoming an important direction for the miniaturization of spectral imaging. Existing representative technologies, such as Spectral DiffuserCam, combine frosted glass as a random phase encoder with a color filter array to achieve lensless snapshot spectral imaging. Similarly, other studies have also used linearly variable filters (LVFs) combined with phase masks or fixed amplitude masks as coding devices.

[0005] While these systems have made significant progress in miniaturization, they still suffer from several drawbacks: First, the encoding method is fixed and unadjustable: the physical structure and optical properties of the scattering medium or amplitude mask used in the system are fixed. This fixed encoding mode means that the point spread function of the system cannot be changed once it is manufactured, limiting the dimensionality and diversity of information obtained from a single exposure. Second, the reconstruction problem is highly undefined, resulting in low imaging accuracy: the fixed single-exposure encoding mode provides insufficient constraints when facing the inverse problem of reconstructing a high-dimensional spectral data cube from two-dimensional measurements. This directly leads to problems such as blurred details, low spectral accuracy, and significant noise and artifacts in the reconstructed images. Furthermore, high-quality video-level spectral imaging cannot be achieved: the fixed encoding mode and single-frame reconstruction framework cannot provide sufficient information redundancy and temporal consistency for continuous capture of dynamic scenes; although video can be recorded by increasing the camera frame rate, each frame is an independent reconstruction based on a single exposure with insufficient information, making it difficult to achieve high-quality, high-frame-rate spectral video output while ensuring spectral accuracy. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a lensless spectral imaging system and method based on a highly integrated color mask that can be dynamically adjusted and supports multi-frame reconstruction, so as to solve the problems of low reconstruction accuracy and inability to adapt to dynamic scenes caused by fixed encoding.

[0007] The technical solution adopted by the system of this invention is as follows:

[0008] A video-grade lensless spectral imaging system based on a color mask includes a filtering window module, a dynamically adjustable spatial spectral modulation module, a color acquisition module, a sliding window module, and a fusion reconstruction module. The filtering window module, dynamically adjustable spatial spectral modulation module, and color acquisition module are arranged sequentially. The fusion reconstruction module is connected to the dynamically adjustable spatial spectral modulation module, the color acquisition module, and the sliding window module, respectively. The sliding window module is connected to the color acquisition module. The filtering window module is used to filter out interfering light outside the imaging window band. The dynamically adjustable spatial spectral modulation module includes a fixed random color mask and a programmable dynamic grayscale mask, used for joint encoding and modulation of the incident light field. The color acquisition module is used to detect the encoded and modulated light field to obtain a series of time-related two-dimensional measurements. The sliding window module is used to perform window translation operations on the time series, grouping the continuously captured measurement value sequences and sending them to the fusion reconstruction module. The fusion reconstruction module receives the multi-frame two-dimensional measurement values ​​output by the color acquisition module and reconstructs a hyperspectral data cube sequence using a multi-frame reconstruction algorithm.

[0009] Furthermore, the programmable dynamic grayscale mask in the dynamically adjustable spatial spectrum modulation module is an LCD screen, a digital micromirror device, or an electrically tunable metasurface, which can switch different spatial coding patterns according to a preset time sequence.

[0010] Furthermore, in the dynamically adjustable spatial spectrum modulation module, the programmable dynamic grayscale mask and the random color mask are arranged in close contact with each other in a physical manner, forming a compact joint coding unit.

[0011] Furthermore, the color acquisition module is an RGB color image sensor with a Bayer filter array or a customized multispectral image sensor.

[0012] Furthermore, the filtering window module is placed in close contact with the random color mask, and the size of the filtering window module is larger than the size of the random color mask.

[0013] The present invention also provides an imaging method for the above-mentioned video-level lensless spectral imaging system based on a color mask, the method comprising the following steps:

[0014] S1, the scene light signal first passes through the filtering window module for band pre-filtering;

[0015] S2 utilizes a dynamically adjustable spatial spectrum modulation module to enable a programmable dynamic grayscale mask to switch different encoding patterns according to a time sequence, dynamically encode and diffract the filtered light signal, and project the point light sources in the scene onto the color acquisition module in a multiplexed manner.

[0016] S3, using the color acquisition module, acquires a frame of corresponding two-dimensional measurement value under each encoding pattern, thereby obtaining a sequence of multiple frames of measurement value within a time window;

[0017] S4. Based on the pre-calibrated point spread function corresponding to each coded pattern and the spectral response function of the color acquisition module, the fusion reconstruction module uses a multi-frame reconstruction algorithm to jointly solve the multi-frame measurement value sequence to reconstruct the hyperspectral data cube within the time window.

[0018] S5. By sliding the time window, repeat steps S2 to S4 to continuously output hyperspectral data cube sequences, thereby achieving video-level spectral imaging.

[0019] Furthermore, in step S2, the encoded pattern for the programmable dynamic grayscale mask switching includes a random pattern, a checkerboard pattern, or a striped pattern.

[0020] Furthermore, in step S4, the multi-frame reconstruction algorithm is implemented by combining the alternating direction multiplier method and the convolutional neural network.

[0021] Furthermore, in step S4, the multi-frame reconstruction algorithm first uses the alternating direction multiplier method to perform preliminary reconstruction of the single-frame measurement values, and then inputs the preliminary reconstruction results of multiple frames into a convolutional neural network for fusion and quality enhancement.

[0022] Further, in step S5, the specific steps of the sliding time window include: first, defining and initializing the window, setting a fixed-length time window T, controlling the dynamically adjustable spatial spectral modulation module to switch between T different encoding patterns in sequence, while the color acquisition module synchronously captures the corresponding T frames of two-dimensional measurement values ​​to form the first complete data window; when it is necessary to process the data of the next moment, the window slides forward one step, and the encoding patterns continue to switch in sequence, that is, the oldest frame measurement value in the window is removed, and a newly acquired frame measurement value is added, thereby forming a new data window; each complete time window will trigger a fusion reconstruction process, and the multi-frame reconstruction algorithm regards the multi-frame data in the current time window as a whole and reconstructs the hyperspectral data cube corresponding to the center moment of the current time window.

[0023] The technology of this invention mainly includes the following three points:

[0024] (1) Dynamically adjustable joint coding hardware

[0025] Common spatial spectral modulation coding devices (such as gratings, prisms, and DOEs) typically have fixed structures or high manufacturing costs. This invention creatively employs a joint coding paradigm combining a programmable dynamic grayscale mask (such as an LCD screen) with a fixed random color mask. The fixed random color mask handles the basic spectral coding, while the adjustable grayscale mask enables pixel-level dynamic spatial amplitude coding. This design not only greatly enriches the diversity of coding but also effectively reduces the system's manufacturing cost and complexity, while providing extremely high flexibility for the optimized design of imaging systems.

[0026] (2) Reconstruction algorithm framework based on multi-frame information fusion

[0027] To address the inherent underdeterminacy problem in high-dimensional data reconstruction during lensless spectral imaging, this invention abandons the traditional single-frame reconstruction approach and proposes a reconstruction algorithm framework based on joint solution of multi-frame measurements. This framework acquires a series of two-dimensional measurements under dynamic encoding and effectively fuses multi-frame information at the algorithm level using optimization methods such as the Alternating Direction Multiplier Method (ADMM) and deep learning methods such as Convolutional Neural Networks (CNN). This core algorithmic innovation significantly increases the constraints on the reconstruction problem, thereby greatly improving the spatial resolution and spectral fidelity of the reconstructed spectral image.

[0028] (3) Supports sliding window processing mechanism for video stream output

[0029] To achieve the leap from static imaging to video-level imaging, this invention introduces a sliding time window processing mechanism. This mechanism processes continuously captured measurement sequences through a sliding window and then feeds them sequentially into a reconstruction algorithm, enabling the continuous output of hyperspectral data cube sequences. This technology empowers the system of this invention with the ability to capture real-time, high-quality spectral video of dynamic scenes, greatly expanding its application scenarios.

[0030] Therefore, the solution of this invention can overcome the limitations of fixed encoding, enhance reconstruction quality by dynamically adjusting the encoding mode and utilizing multi-frame information, and ultimately achieve high-precision video-level spectral imaging. Compared with existing technologies, it has the following advantages:

[0031] (1) Dynamic control and performance breakthrough: By changing the spatial coding pattern in real time through a programmable grayscale mask (such as an LCD), the limitations of a fixed mask are broken. This time-varying coding provides multiple sets of differentiated modulation information for the same scene, which significantly reduces the underdeterminacy of the spectral reconstruction problem from the root and lays the foundation for achieving high-precision reconstruction.

[0032] (2) Multi-frame enhancement and quality improvement: Joint reconstruction using different coding information from multiple consecutive frames is equivalent to adding a large number of effective constraint equations to solve the high-dimensional data cube, which can significantly improve the spatial resolution and spectral fidelity of the reconstructed spectral image and effectively suppress noise and artifacts.

[0033] (3) Video-level imaging capability: Through a “sliding window” processing flow, multi-frame reconstruction technology is applied to continuously captured measurement sequences, and hyperspectral video output is achieved under a lensless spectral imaging architecture, enabling it to be applied to real-time monitoring and analysis of dynamic scenes.

[0034] (4) Highly compact and cost-controllable: It inherits the advantages of compact structure and low cost of lensless imaging systems. It uses commercial devices such as LCDs as dynamic masks, avoiding the high processing costs of custom optical components such as DOE and metasurfaces, making the system easier to popularize and integrate into portable devices. Attached Figure Description

[0035] Figure 1 This is a schematic diagram of the overall structure of the video-level lensless spectral imaging system based on a color mask according to the present invention;

[0036] Figure 2 The three figures are schematic diagrams of the video-level lensless spectral imaging system based on color masks according to the present invention: (a) schematic diagram of color mask, (b) schematic diagram of adjustable grayscale mask, and different generated random grayscale mask images.

[0037] Figure 3 This is a schematic diagram of the assembly of the video-level lensless spectral imaging system based on a color mask according to the present invention;

[0038] Figure 4 The actual PSF of the adjustable grayscale mask for different spectral channels in this embodiment of the invention is shown (taking 455nm, 545nm, and 645nm as examples).

[0039] Figure 5 This is a flowchart illustrating the image reconstruction using CNN in this invention;

[0040] Figure 6 This is a schematic diagram of the processing flow for obtaining video-level imaging using a sliding window in this invention;

[0041] Figure 7 Figure (a) shows three measured values ​​collected in embodiment (a) of the present invention; Figure (b) shows a single-frame synthesized RGB image of the reconstructed spectral image corresponding to the measured values ​​in Figure (a); and Figure (c) shows a synthesized RGB image of the three reconstructed spectral images in Figure (b). Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0043] like Figure 1 As shown, the video-level lensless spectral imaging system based on a color mask provided in this embodiment includes a filtering window module, a dynamically adjustable spatial spectral modulation module, a color acquisition module, a sliding window module, and a fusion reconstruction module. The filtering window module, dynamically adjustable spatial spectral modulation module, and color acquisition module are arranged sequentially. The fusion reconstruction module is connected to the dynamically adjustable spatial spectral modulation module, the color acquisition module, and the sliding window module, respectively. The sliding window module is connected to the color acquisition module. The functions of each module are as follows: The filtering window module is used to filter out interfering light outside the imaging window band; the dynamically adjustable spatial spectral modulation module includes a fixed random color mask and a programmable dynamic grayscale mask, used for joint encoding and modulation of the incident light field; the color acquisition module is used to detect the encoded and modulated light field to obtain a series of time-related two-dimensional measurement values; the sliding window module is used to perform window translation operations on the time series, grouping the continuously captured measurement value sequence and sending it to the fusion reconstruction module to support continuous processing and output of the video stream; the fusion reconstruction module is used to receive multi-frame two-dimensional measurement values ​​output by the color acquisition module and reconstruct a hyperspectral data cube sequence through a multi-frame reconstruction algorithm.

[0044] The filtering window module is used to filter out stray light outside the working wavelength band. In this embodiment, a bandpass filter (F1) with a center wavelength of 550nm and a bandwidth of 200nm (450nm-650nm) is used to filter out stray light outside this band, ensuring the quality of the incident light signal. It is important to note that the filter needs to be placed in close contact with the color mask, and the filter size needs to be slightly larger than the color mask to prevent some light from passing through the color mask without being blocked, which would affect the experimental results.

[0045] The dynamically adjustable spatial spectral modulation module is the core of the system. In this embodiment, it employs a combination of an adjustable grayscale mask and a fixed color mask. The adjustable grayscale mask is used to achieve pixel-level dynamic spatial amplitude modulation, precisely controlling the luminous flux at each spatial location. It can be an LCD screen, a digital micromirror device, or an electrically tunable metasurface, etc., and can switch different spatial coding patterns according to a preset time sequence. The coding patterns include random patterns, checkerboard patterns, or striped patterns. This embodiment uses a high-resolution industrial-grade monochrome LCD screen with a pixel size of 0.05mm. This LCD screen is connected to a computer via a driving circuit and can be controlled by a computer program to generate and switch different coding patterns over time. The specific switching method is as follows: the host computer simultaneously controls the camera and the LCD. The host computer sends a trigger signal to control the camera to start exposure. After exposure, the host computer sends a signal to control the LCD screen to sequentially switch the displayed patterns, and then sends another signal to control the camera to expose the next frame. Figure 2 As shown in (b), three different random patterns were generated in the embodiment to spatially modulate light intensity, greatly enriching the diversity of the encoding. A fixed color mask (CM), used to encode different spectral channels and project scene point light sources into the color acquisition module in a multiplexed manner, can be an RGB color image sensor with a Bayer filter array or a customized multispectral image sensor. Figure 2 As shown in (a), this embodiment employs a random color mask fabricated on a fully transparent film using an inkjet printing process. This mask contains multiple spectral filtering units, including R (red), G (green), B (blue), C (cyan), M (magenta), and Y (yellow), arranged in a random distribution. During assembly, the LCD screen and the color mask are physically contacted and tightly fitted together, forming a compact joint encoding unit.

[0046] The color acquisition module is the data sensing end of this system, used to record the encoded two-dimensional measurement values ​​of the scene. It needs to be placed within 1mm of the color mask. Figure 3As shown, its core is a multispectral image sensor. This embodiment preferably uses a consumer-grade or industrial-grade RGB color camera, whose sensor surface integrates a standard Bayer filter array. This array is arranged in a periodic pattern, with the most basic unit being a 2×2 pixel block containing one R (red), one B (blue), and two G (green) filter units. Each unit allows only light of a specific wavelength to pass through, thereby discretizing continuous broadband information into a limited number of spectral channels (R, G, and B channels in this embodiment). The filter array spatially downsamples the diffraction projection into multiple spectral encoded channels, each with a different filter response form. Before imaging, the module undergoes precise spectral response calibration. By using a monochromator, the relative response sensitivity of each color channel of the sensor to incident light of different wavelengths is measured, obtaining a precise spectral response function R. C (λ), where c∈{R,G,B}, spectral response functions for a total of 20 bands (455nm-645nm, selected in 10nm increments) were used in the experiment. The diffraction projection was then recorded by the sensor, and the scene was further spectrally encoded to obtain the two-dimensional measurement value M(x,y). Here, (x,y) corresponds to the planar coordinates of the sensor pixels. The measurement value obtained in this embodiment is as follows: Figure 7 As shown in (a), due to the use of, Figure 2 As shown in (b), there are three different grayscale masks, so three different measurement values ​​will appear.

[0047] like Figure 6 As shown, the sliding window module is a timing control and data management unit. It simulates image motion by implementing window translation over time to support video imaging. It is typically implemented in software and serves as a bridge between data acquisition and continuous reconstruction. The sliding time window processing mechanism updates the data window by removing old frames and adding new frames, outputting a hyperspectral video stream at a stable and continuous rate. Specifically: First, the window is defined and initialized, setting a fixed-length time window T (T = 3 frames in this embodiment). During system initialization, the dynamically adjustable spatial spectral modulation module sequentially switches between T different encoding patterns, while the color acquisition module synchronously captures the corresponding T frame two-dimensional measurement values ​​{M1(x,y),M2(x,y),...,M...}. T The first complete data window is formed by taking a frame (x, y) and then moving forward one step to the next frame. When processing data from the next moment, the window moves forward one step, and the encoded pattern continues to switch sequentially. This means removing the oldest frame measurement M1(x, y) from the window and adding a newly acquired frame measurement M. (T+1) (x,y), thus forming a new data window {M2(x,y),M3(x,y),...,M (T+1)(x,y)}. This process continues, simulating a "window" that translates along the time axis. Each complete window triggers a fusion and reconstruction process. The reconstruction algorithm treats the multiple frames of data within this window as a whole, reconstructing the hyperspectral data cube L corresponding to the center moment of that time window. In this way, the system can output hyperspectral images stably and continuously at a rate lower than the camera's native frame rate, thus forming a hyperspectral video stream.

[0048] The fusion reconstruction module is used to process multiple frames of two-dimensional measurement values ​​based on system calibration parameters and employing a multi-frame reconstruction algorithm to ultimately reconstruct a hyperspectral data cube sequence. Specifically, it receives a multi-frame measurement value sequence {M1(x,y), M2(x,y), ..., M...} from the sliding window module. T (x,y)}, and the point spread function {PSF1,PSF2,...,PSF} retrieved from the system calibration database that corresponds one-to-one with these measurements. T}(like Figure 4 As shown, the three grayscale masks correspond to... Figure 2 (b) shows the mask pattern and the spectral response function R of the sensor. C (λ), input into the reconstruction model, is used to reconstruct the hyperspectral data cube I(x,y,λ) corresponding to the current time window through a multi-frame reconstruction algorithm. For the multi-frame reconstruction algorithm, this embodiment uses a combination of the Alternating Direction Multiplier Method (ADMM) and a CNN network. First, for each single frame of data, a total variational regularization term is introduced into the reconstruction model to constrain the solution in the spatial and spectral three-dimensional domains, thus addressing the underdeterminacy of the reconstruction problem. Then, the Alternating Direction Multiplier Method (ADMM) is used to efficiently iteratively solve the problem, thereby stably reconstructing the single-frame hyperspectral image data. After collecting T frames of images, they are fed together into a pre-trained CNN network for further improvement of image resolution and quality. A schematic diagram of the CNN network performing multi-frame reconstruction is shown below. Figure 5As shown, taking T=3 as an example, the three frames of single-frame hyperspectral images reconstructed using the ADMM algorithm (each frame containing 20 spectral bands) are first stitched together along the channel dimension to form a 60-channel input tensor with a spatial resolution of 512×512. Essentially, this provides the network with multiple observation versions of the same scene exhibiting different error patterns. The network then performs shallow feature extraction using two 3×3 convolutional layers, increasing the number of channels to 64, and learns the local correlation between spatial details and spectral information in different reconstruction results. Subsequently, a channel attention mechanism is introduced, adaptively evaluating and weighting the importance of each feature channel through global average pooling and fully connected layers, thereby enhancing information-rich features and suppressing noise or redundant components. Afterward, the network downsamples the feature map to 256×256 resolution using max pooling, expanding the receptive field while extracting deep abstract features containing contextual information to understand the overall structure of the image. Next, the network uses transposed convolutions to upsample the deep features back to the original resolution of 512×512, and then concatenates them with the previously attention-weighted shallow features along the channel dimension through skip connections, achieving a deep fusion of high-level semantic information and low-level spatial details. The concatenated features are then refined through two convolutional layers and mapped to 20 target spectral channels by a 1×1 convolutional layer. At the end of the entire process, the network innovatively introduces a residual connection, which directly projects the input to an initial estimate of 20 channels through an independent 1×1 convolution and adds it to the output of the main network path. This residual learning strategy allows the network to focus on learning the "correction" between complex high-quality reconstruction results and simple linear projection results, significantly reducing the optimization difficulty and ensuring the direct transmission of valuable information from the input. The final output is a hyperspectral data cube with higher spatial resolution, more accurate spectral fidelity, and lower noise.

[0049] Because each of the three ADMM reconstruction results with different PSFs exhibits a different error pattern, the blurring, noise, and artifacts in the image manifest differently under different PSFs. Therefore, this embodiment constructs a CNN network to learn to select the clearest version at different locations, thereby suppressing various reconstruction artifacts and performing intelligent weighted fusion. Through the fusion of multi-frame information, the network effectively improves spatial resolution by utilizing sub-pixel-level information complementarity and optimizes the reconstruction accuracy of the spectral curve, thus enhancing the quality of image reconstruction.

[0050] This embodiment also provides a method for performing video-level lensless spectral imaging using the above-described imaging system, including the following steps:

[0051] S1, the scene light signal first passes through the filtering window module for band pre-filtering;

[0052] S2 utilizes a dynamically adjustable spatial spectrum modulation module to switch different encoding patterns on the grayscale mask according to the time sequence, performs dynamic amplitude encoding and diffraction projection on the filtered light signal, and projects the point light sources in the scene onto the color acquisition module in a multiplexed manner.

[0053] S3, using the color acquisition module, acquires a frame of corresponding two-dimensional measurement value under each encoding pattern, thereby obtaining a sequence of multiple frames of measurement value within a time window;

[0054] S4. Based on the pre-calibrated point spread function corresponding to each coded pattern and the spectral response function of the color acquisition module, the fusion reconstruction module uses a multi-frame reconstruction algorithm to jointly solve the multi-frame measurement value sequence to reconstruct the hyperspectral data cube within the time window.

[0055] S5. By sliding the time window, repeat steps S2 to S4 to continuously output hyperspectral data cube sequences, thereby achieving video-level spectral imaging.

[0056] In this embodiment, the measured value image obtained by the color acquisition module is as follows: Figure 7 As shown in (a), it can be seen that all the original graphic's morphological features have been completely lost. The single-frame RGB image reconstructed using the fusion algorithm, however, is as follows: Figure 7 As shown in (b), the reconstruction result has restored the features and structure of the original scene. The image contains two lemon slices and a color palette, and the PSNR is around 25, indicating that the reconstruction effect is quite good. The final result of multi-frame reconstruction using a CNN network is shown below. Figure 7 As shown in (c), the image resolution is further improved, with almost no blurry areas or artifacts, and the PSNR is also improved compared to the single-frame result.

[0057] This fusion reconstruction module typically runs on high-performance computing platforms, such as workstations or servers equipped with multi-core CPUs and GPUs. Leveraging the parallel computing capabilities of GPUs can significantly accelerate the aforementioned iterative optimization or neural network inference processes, meeting the real-time or near-real-time requirements of video-level reconstruction.

[0058] This invention innovatively employs a technical approach combining dynamically adjustable coding and multi-frame reconstruction to construct a lensless spectral imaging system that combines high integration, excellent imaging performance, and structural stability, achieving a leap from static snapshots to video-level spectral imaging capabilities. This integrated paradigm provides key technical support for solving the miniaturization and dynamic scene applications of spectral imaging systems, and is particularly suitable for lightweight, embedded hyperspectral video applications.

[0059] It should be noted that the specific embodiments described in this specification are only a part of the embodiments of the present invention, and not all of them. They are only used to assist in understanding the technical content of the present invention. The descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of protection of the present invention. Any technical modifications and equivalent substitutions made by those skilled in the art under the guidance of the technical concept of the present invention should be included within the maximum possible scope of protection claimed by the present invention.

Claims

1. A video-level lensless spectral imaging system based on a color mask, characterized in that, The system includes a filtering window module, a dynamically adjustable spatial spectrum modulation module, a color acquisition module, a sliding window module, and a fusion reconstruction module. The filtering window module, the dynamically adjustable spatial spectrum modulation module, and the color acquisition module are arranged in sequence. The fusion reconstruction module is connected to the dynamically adjustable spatial spectrum modulation module, the color acquisition module, and the sliding window module, respectively. The sliding window module is connected to the color acquisition module. The filtering window module is used to filter out interfering light rays outside the imaging window band; The dynamically adjustable spatial spectrum modulation module includes a fixed random color mask and a programmable dynamic grayscale mask, which are used to jointly encode and modulate the incident light field. The color acquisition module is used to detect the encoded and modulated light field and obtain a series of time-related two-dimensional measurement values. The sliding window module is used to perform window translation operations on the time series, grouping and sending the continuously captured measurement value series into the fusion reconstruction module; The fusion reconstruction module is used to receive multi-frame two-dimensional measurement values ​​output by the color acquisition module and reconstruct a hyperspectral data cube sequence through a multi-frame reconstruction algorithm.

2. The video-level lensless spectral imaging system based on a color mask according to claim 1, characterized in that, The programmable dynamic grayscale mask in the dynamically adjustable spatial spectrum modulation module is an LCD screen, a digital micromirror device, or an electrically tunable metasurface, which can switch different spatial coding patterns according to a preset time sequence.

3. The video-level lensless spectral imaging system based on a color mask according to claim 1, characterized in that, In the dynamically adjustable spatial spectrum modulation module, the programmable dynamic grayscale mask and the random color mask are arranged in close contact with each other in a physical manner, forming a compact joint coding unit.

4. The video-level lensless spectral imaging system based on a color mask according to claim 1, characterized in that, The color acquisition module is an RGB color image sensor with a Bayer filter array or a customized multispectral image sensor.

5. The video-level lensless spectral imaging system based on a color mask according to claim 1, characterized in that, The filtering window module is placed in close contact with the random color mask, and the size of the filtering window module is larger than the size of the random color mask.

6. An imaging method using a video-level lensless spectral imaging system based on a color mask as described in any one of claims 1-5, characterized in that, The method includes the following steps: S1, the scene light signal first passes through the filtering window module for band pre-filtering; S2 utilizes a dynamically adjustable spatial spectrum modulation module to enable a programmable dynamic grayscale mask to switch different encoding patterns according to a time sequence, dynamically encode and diffract the filtered light signal, and project the point light sources in the scene onto the color acquisition module in a multiplexed manner. S3, using the color acquisition module, acquires a frame of corresponding two-dimensional measurement value under each encoding pattern, thereby obtaining a sequence of multiple frames of measurement value within a time window; S4. Based on the pre-calibrated point spread function corresponding to each coded pattern and the spectral response function of the color acquisition module, the fusion reconstruction module uses a multi-frame reconstruction algorithm to jointly solve the multi-frame measurement value sequence to reconstruct the hyperspectral data cube within the time window. S5. By sliding the time window, repeat steps S2 to S4 to continuously output hyperspectral data cube sequences, thereby achieving video-level spectral imaging.

7. The method according to claim 6, characterized in that, In step S2, the encoded pattern for the programmable dynamic grayscale mask switching includes a random pattern, a checkerboard pattern, or a striped pattern.

8. The method according to claim 6, characterized in that, In step S4, the multi-frame reconstruction algorithm is implemented by combining the alternating direction multiplier method and the convolutional neural network.

9. The method according to claim 8, characterized in that, In step S4, the multi-frame reconstruction algorithm first uses the alternating direction multiplier method to perform preliminary reconstruction of the single-frame measurement values, and then inputs the preliminary reconstruction results of multiple frames into a convolutional neural network for fusion and quality enhancement.

10. The method according to claim 6, characterized in that, In step S5, the specific steps of sliding the time window include: First, the window is defined and initialized. A fixed-length time window T is set, and the dynamically adjustable spatial spectrum modulation module is controlled to switch T different coding patterns in sequence. At the same time, the color acquisition module synchronously captures the corresponding T frames of two-dimensional measurement values ​​to form the first complete data window. When it is necessary to process the data of the next moment, the window slides forward one step, and the encoding pattern continues to switch in sequence. That is, the oldest frame of measurement value in the window is removed, and a newly acquired frame of measurement value is added, thus forming a new data window. Each complete time window triggers a fusion reconstruction process. The multi-frame reconstruction algorithm treats the multiple frames of data within the current time window as a whole and reconstructs the hyperspectral data cube corresponding to the center moment of the current time window.