Video snapshot compression imaging method, device and system based on motion consistency verification

By using a dual-optical-path conjugate mask design and motion consistency verification method, the problem of reconstruction results deviating from the physical motion trajectory in snapshot compression imaging technology was solved, achieving efficient light energy utilization and high-quality reconstruction results.

CN122093579APending Publication Date: 2026-05-26JIANGSU YITONG HIGH TECH
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU YITONG HIGH TECH
Filing Date
2026-04-22
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing snapshot compression imaging technology lacks explicit utilization of physical imaging models during the reconstruction process, which makes the reconstruction results prone to deviating from the real physical motion trajectory in the time dimension. Furthermore, it cannot be verified and corrected in real time through the physical characteristics of the imaging system itself, thus limiting the spatiotemporal fidelity and physical reliability of the reconstructed video.

Method used

A dual-optical-path conjugate mask design is adopted. By synchronously observing the dynamic scene frame sequence to be imaged, the original observation image and the complementary observation image are obtained. Combined with differential preprocessing and reconstruction network modules, the joint objective function is constructed using the motion consistency loss function to optimize the reconstruction network, ensuring that the reconstructed video is consistent with the physical measurement domain.

Benefits of technology

It significantly improves light utilization efficiency and spatial resolution, enhances reconstruction robustness, effectively suppresses motion artifacts, and improves the temporal consistency and physical reliability of reconstructed videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122093579A_ABST
    Figure CN122093579A_ABST
Patent Text Reader

Abstract

The invention provides a video snapshot compression imaging method, device and system based on motion consistency verification, and the method comprises the steps: carrying out the synchronous observation of a to-be-imaged frame sequence through a conjugate mask of a dual-optical-path imaging system, and obtaining an original measurement value and a complementary measurement value; carrying out differential preprocessing on the two paths of measurement values, and extracting an actual differential measurement value reflecting the motion characteristics of an object in the scene; inputting the original measurement value, the complementary measurement value and the actual differential measurement value into a reconstruction network to obtain a reconstructed video sequence; a predicted difference measurement value is obtained through back projection of the reconstructed video sequence; and calculating the motion consistency loss between the actual difference measurement value and the predicted difference measurement value, and constructing a joint objective function to optimize the reconstruction network. According to the method, the light utilization efficiency is remarkably improved, the motion artifact of the reconstructed video is effectively eliminated by taking the motion consistency of a physical layer as a constraint, and the temporal-spatial resolution and the reconstruction fidelity of the system in a complex noise environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computational imaging and image processing technology, specifically relating to a video snapshot compression imaging method, apparatus, and system based on motion consistency verification. Background Technology

[0002] Multidimensional imaging technology is an image acquisition method capable of capturing and analyzing multidimensional data beyond two-dimensional spatial images. The information it acquires encompasses three-dimensional space, one-dimensional time, and physical parameters such as spectrum and phase. Among these data, Compressed Time Snapshot Imaging (CACTI) technology, which utilizes low-frame-rate detectors to acquire high-frame-rate video streams within a single exposure, has significant application value in fields such as high-speed machine vision, biomedicine, and the observation of ultrafast physical phenomena.

[0003] Rapid time-compression imaging technology is a combination of acquisition strategies and reconstruction algorithms that achieves higher temporal resolution by trading speed for algorithms, resulting in a more comfortable patient experience in the medical field.

[0004] Traditional snapshot compression imaging technology has significant limitations in hardware modulation and signal acquisition. It often uses a single-path binary mask to modulate dynamic scenes, resulting in low light energy utilization and severe loss of spatial information. However, existing deep learning-based reconstruction algorithms have become the mainstream solution for snapshot compression imaging. These algorithms leverage powerful data modeling capabilities to more effectively decouple spatiotemporal information from highly compressed measurements, overcoming the bottlenecks of traditional algorithms in reconstruction quality and speed.

[0005] However, existing deep learning-based reconstruction methods typically treat the physical encoding process of imaging separately from the data-driven reconstruction algorithm. The reconstruction process relies solely on statistical regularities learned from observational data, lacking explicit utilization of the physical imaging model. This leads to reconstruction results that easily deviate from the true physical motion trajectory in the temporal dimension, and cannot be verified and corrected in real time using the physical characteristics of the imaging system itself. Due to the lack of a closed-loop constraint mechanism to map the reconstruction results back to the physical measurement domain, the system struggles to incorporate physical priors such as motion consistency into the reconstruction process, thus limiting the spatiotemporal fidelity and physical reliability of the reconstructed video. Therefore, how to overcome single-channel signal loss while constructing a closed-loop reconstruction framework that integrates physical models and motion consistency constraints has become a key challenge for improving the performance of snapshot compressed imaging. Summary of the Invention

[0006] This invention provides a video snapshot compression imaging method, apparatus, and system based on motion consistency verification to solve the above-mentioned problems.

[0007] The technical solution adopted in this invention is: A video snapshot compression imaging method based on motion consistency verification is provided, including the following steps: Using conjugate masks, dual-optical-path synchronous observations are performed on the dynamic scene frame sequence to be imaged to acquire the original observation image and complementary observation image. Differential preprocessing is performed on the original observation image and the complementary observation image to extract the actual differential measurement values ​​that reflect the motion characteristics of the scene; The original observation image, complementary observation image, and actual difference measurement value are input into the reconstruction network module to obtain the reconstructed video sequence; The reconstructed video sequence is back-projected onto the physical coding model to obtain the predicted difference measurement value; The motion consistency loss between the actual and predicted differential measurements is calculated, and a joint objective function is constructed to optimize the reconstructed network.

[0008] Furthermore, the specific calculation formula for using conjugate masks to perform dual-optical-path synchronous observation of the dynamic scene frame sequence to be imaged is as follows: , ,in, These are the original observation image and the complementary observation image, respectively. Let t be the scene image of the t-th frame. Let n be a random binary mask for the t-th frame, and n represent the compression ratio. and These are the width and height of the image frame, respectively. This is for the Hadamard product operation.

[0009] Furthermore, the formula for calculating the actual difference measurement values ​​that reflect the motion characteristics of the scene is as follows: ,in, This represents the actual differential measurement value, whose physical meaning is equivalent to that obtained through... The mask modulates the scene to reveal the degree of fluctuation in scene pixel values ​​over time.

[0010] Furthermore, the reconstructed video sequence is back-projected onto the physical coding model, and the formula for calculating the predicted difference measurement is as follows: ,in, Indicates the predicted difference measurement value. This indicates the reconstruction of the video sequence. It is its t-th frame.

[0011] Furthermore, the formula for calculating the motion consistency loss between the actual differential measurement value and the predicted differential measurement value is as follows: ,in, This represents the motion consistency loss, used to constrain the motion trajectory of the reconstructed video to remain consistent with the physical measurement domain.

[0012] Furthermore, constructing a joint objective function to optimize the reconstruction network involves: using the reconstruction fidelity loss function and the motion consistency loss function to construct a joint loss function, which forces the dynamic evolution trajectory of the reconstructed video to remain consistent with the physical measurement domain until the preset convergence criterion is met.

[0013] Furthermore, the reconstruction fidelity loss function The calculation formula is: ,in, To reconstruct the video sequence, This represents the true value of the original scene.

[0014] Furthermore, construct a joint objective function. The reconstructed network is optimized using the following formula: ,in, is the weighting coefficient for motion consistency loss, used to balance the quality of the reconstructed image with the accuracy of the physical motion trajectory.

[0015] Furthermore, the convergence criteria include: the joint objective function value is less than a preset threshold, or the decrease in the loss function over several consecutive iterations is less than a preset minimum value.

[0016] A video snapshot compression imaging device based on motion consistency verification is provided, including: A dual-optical-path conjugate coding module is used to acquire the original observation image and the complementary observation image; The differential feature preprocessing module is used to extract the actual differential measurement values; The reconstruction network module is used to generate reconstructed video sequences; The verification and optimization module is used to calculate the predicted difference measurement value and adjust the reconstruction network module based on the consistency loss between the predicted difference measurement value and the actual difference measurement value.

[0017] Furthermore, the dual-optical-path conjugate coding module includes: A digital micromirror device is used to load a random binary mask and use its deflection state to split the incident light into two conjugate outgoing light rays; The controller is used to control the loading of a time-varying random binary mask onto the digital micromirror device and send a synchronization trigger signal to the two image sensors to ensure that spatial modulation and image capture are highly synchronized in the time domain. Two synchronously triggered image sensors are placed in the two outgoing optical paths respectively, for synchronous integration to capture the original observation image and the complementary observation image.

[0018] Furthermore, the reconstructed network modules include: The feature extraction unit is used to extract spatiotemporal complementary features from the input original observation image, complementary observation image, and actual difference measurement value; The reconstructed network unit is used to generate the corresponding reconstructed video sequence based on the extracted spatiotemporal complementary feature maps.

[0019] Furthermore, the verification optimization module includes: The forward reprojection unit is used to perform Hadamard product and accumulation operations on the reconstructed video sequence and the corresponding mask to obtain the predicted difference measurement value. The loss calculation unit is used to calculate the motion consistency loss function and feed the loss value back to the reconstruction network module to reduce the reconstruction bias through gradient descent or iterative correction.

[0020] Furthermore, reducing bias through gradient descent means: using the backpropagation algorithm to calculate the gradient of the joint objective function with respect to the parameters of each layer of the reconstructed network, using the Adam optimizer to dynamically adjust the learning rate, and updating the network weights through multiple iterations to make the predicted difference measurement value approximate the actual difference measurement value.

[0021] Furthermore, reducing reconstruction bias through iterative correction means that during the reconstruction process, the reconstructed video sequence generated in the current round is fed back to the verification and optimization module. Based on the residual map generated by motion consistency loss, local weight compensation in the spatial domain is performed on the deep feature map of the reconstruction network to eliminate motion artifacts.

[0022] A video snapshot compression imaging system based on motion consistency verification is provided, comprising: a memory and a processor, wherein the processor executes instructions stored in the memory to implement the above-described dual-optical-path video snapshot compression imaging method based on motion consistency verification.

[0023] The beneficial effects of the video snapshot compression imaging method, apparatus, and system based on motion consistency verification of the present invention are: 1. To address the problems of spatial information loss and low light efficiency in existing technologies, this invention employs a dual-optical-path conjugate mask coding design, which ensures that the light from the scene to be imaged can be captured by the detector through one of the optical paths at each sampling moment. By combining the original observation image with the complementary observation image, the spatial information loss caused by the "0" value in the traditional single-optical-path mask is effectively compensated, significantly improving the light utilization efficiency of the system and the spatial resolution of the reconstructed image. 2. To address the difficulty of motion feature extraction in low-light or high-noise environments using existing methods, this invention extracts the actual differential measurement value by calculating the difference between the original observation and the complementary observation. This measurement value not only explicitly reflects the motion features of scene pixels fluctuating over time, but also effectively cancels the readout noise of the two detectors by utilizing the common-mode suppression characteristics of differential operations, thereby significantly enhancing the reconstruction robustness of the system in weak signal and strong noise environments. 3. To address the problem of motion inconsistency caused by existing reconstruction methods deviating from physical model constraints, this invention proposes a reconstruction mechanism based on dual optical paths for motion consistency verification. By reprojecting the reconstructed video sequence onto the differential measurement domain and comparing it with the actual differential measurement values, a closed-loop constraint at the physical level is constructed. This mechanism forces the reconstruction results to be consistent with the imaging physical process, effectively suppressing artifacts such as motion blur, ghosting, and double images, and significantly improving the temporal consistency and physical reliability of the reconstructed video. Attached Figure Description

[0024] Figure 1 This is a diagram illustrating the overall system framework of the video snapshot compression imaging method based on motion consistency verification according to the present invention. Figure 2 This is a schematic diagram illustrating the method of the present invention for obtaining a reconstructed video sequence; Figure 3 This is a schematic diagram of the reconstructed network module in this invention; Figure 4 This is the residual block used in the first embodiment of the present invention. Detailed Implementation

[0025] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby providing a clearer and more explicit definition of the scope of protection of the present invention.

[0026] Additionally, it should be noted that the terms "comprising," "including," or any other variations thereof used in the embodiments of the present invention are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.

[0027] Combination Figure 1 The method of this invention includes dual-optical-path conjugate sampling, differential feature preprocessing, video reconstruction, and motion consistency verification. The specific steps are as follows: 1. Dual-optical-path conjugate sampling First, the conjugate mask of the dual-optical-path imaging system is used to image the dynamic scene frame sequence to be imaged. Simultaneous observation specifically refers to using a digital micromirror device to load a random binary mask. This modulates and divides the incident light into two complementary optical paths in space; First optical path corresponding mask The transmission direction is captured by integration using the first image sensor to obtain the original observation image. The second optical path corresponds to the conjugate mask. The direction is captured by synchronous integration from the second image sensor to obtain complementary observation images. .

[0028] The specific calculation formula for its physical encoding process is as follows: , , in, Let t be the dynamic scene image of the t-th frame, n be the compression ratio, and ⊙ represent the Hadamard product operation.

[0029] 2. Differential Feature Preprocessing The original observation images acquired synchronously and complementary observation images The input differential feature preprocessing module extracts actual differential measurements that reflect the motion characteristics of objects within the scene by performing matrix subtraction on the two observations. The specific calculation formula is as follows: , This step is physically equivalent to utilizing The ternary mask modulates the scene, which can significantly suppress the common-mode noise of the two sensors and reveal the dynamic gradient information of pixels in the scene as they fluctuate over time.

[0030] 3. Video reconstruction Original observation image Complementary observation images and actual differential measurement values As a common input, it is fed into a pre-defined reconstruction network, which learns from the input. and The provided spatial information, combined with The provided enhanced motion features are used for end-to-end nonlinear mapping reconstruction to obtain the reconstructed video sequence. .

[0031] 4. Motion consistency verification To ensure that the reconstructed video conforms to the physical laws of imaging and eliminates motion artifacts, a closed-loop verification mechanism is introduced, the specific process of which is as follows: Physical reprojection: reconstructing video sequences Substitute the system's physical coding model back into the forward projection, calculate the predicted measurements of the reconstructed video under the conjugate mask, and obtain the predicted difference measurements accordingly. : ; Loss calculation: Calculate the actual differential measurement value Differential measurement with prediction Loss of motion consistency between : ; Feedback optimization: A joint loss function is constructed using reconstruction fidelity loss and motion consistency loss, where motion consistency loss... As an auxiliary loss function, it forces the dynamic evolution trajectory of the reconstructed video to remain consistent with the physical measurement domain until the preset convergence criterion is met, thereby obtaining the final high-fidelity reconstruction result. Specifically, reconstruct the fidelity loss function The calculation formula is: ,in, To reconstruct the video sequence, This represents the true value of the original scene.

[0032] Specifically, construct the joint objective function. The calculation formula is: ,in, is the weighting coefficient for motion consistency loss, used to balance the quality of the reconstructed image with the accuracy of the physical motion trajectory.

[0033] Specifically, there are two convergence criteria: one is "the joint objective function value is less than a preset threshold", and the other is "the decrease in the loss function over several consecutive iterations is less than a preset minimum value".

[0034] Combination Figure 2 The video reconstruction network module includes: The preprocessing unit is used to process the input raw measurements. Complementary measurement values and actual differential measurement value Feature extraction and dimension mapping are performed to obtain multi-path corresponding observation feature maps, and then concatenation is performed in the channel dimension to obtain a fused feature map. ; Reconstructing network units for fusing feature maps Cascaded residual processing is performed to restore the spatiotemporal details of high frame rate video through deep nonlinear mapping, ultimately obtaining the reconstructed video sequence. .

[0035] Combination Figure 3 and Figure 4The reconstructed network unit consists of multiple sequentially concatenated residual blocks (RBs) (taking three as an example), which receive the fused feature map output by the preceding channel concatenation (Concat) layer. And perform deep feature extraction; Each residual block in the reconstructed network unit consists of a first convolutional layer (Conv), a ReLU activation function, and a second convolutional layer (Conv) connected sequentially. The kernel size of each convolutional layer (Conv) is preset to 3×3, the stride is set to 1, and the padding mode is set to "same" to maintain the stability of the feature dimension. By using residual connections, the input features of the residual block can be directly passed to the output position at the end of the block and accumulated at the pixel level (that is, by directly passing the network input to the network output to build a direct path), making the gradient easier to propagate through the entire network, which helps to avoid the problems of gradient vanishing or gradient exploding when training deep reconstruction networks. Meanwhile, this structure forces the network to learn the residual mapping between input features and output video, effectively preserving the original full-light energy features and explicit motion gradient features extracted by the preprocessing unit, and finally outputting a high-quality reconstructed video sequence through channel restoration. .

[0036] The beneficial effects of the video snapshot compression imaging method, apparatus, and system based on motion consistency verification of the present invention are: 1. To address the problems of spatial information loss and low light efficiency in existing technologies, this invention employs a dual-optical-path conjugate mask coding design, which ensures that the light from the scene to be imaged can be captured by the detector through one of the optical paths at each sampling moment. By combining the original observation image with the complementary observation image, the spatial information loss caused by the "0" value in the traditional single-optical-path mask is effectively compensated, significantly improving the light utilization efficiency of the system and the spatial resolution of the reconstructed image. 2. To address the difficulty of motion feature extraction in low-light or high-noise environments using existing methods, this invention extracts the actual differential measurement value by calculating the difference between the original observation and the complementary observation. This measurement value not only explicitly reflects the motion features of scene pixels fluctuating over time, but also effectively cancels the readout noise of the two detectors by utilizing the common-mode suppression characteristics of differential operations, thereby significantly enhancing the reconstruction robustness of the system in weak signal and strong noise environments. 3. To address the problem of motion inconsistency caused by existing reconstruction methods deviating from physical model constraints, this invention proposes a reconstruction mechanism based on dual optical paths for motion consistency verification. By reprojecting the reconstructed video sequence onto the differential measurement domain and comparing it with the actual differential measurement values, a closed-loop constraint at the physical level is constructed. This mechanism forces the reconstruction results to be consistent with the imaging physical process, effectively suppressing artifacts such as motion blur, ghosting, and double images, and significantly improving the temporal consistency and physical reliability of the reconstructed video.

[0037] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A video snapshot compression imaging method based on motion consistency verification, characterized in that, Includes the following steps: Using conjugate masks, dual-optical-path synchronous observations are performed on the dynamic scene frame sequence to be imaged to acquire the original observation image and complementary observation image. The original observation image and the complementary observation image are subjected to differential preprocessing to extract the actual differential measurement values ​​that reflect the motion characteristics of the scene; The original observation image, complementary observation image, and actual difference measurement values ​​are input into the reconstruction network to obtain the reconstructed video sequence; The reconstructed video sequence is back-projected onto the physical coding model to obtain the predicted difference measurement value; The motion consistency loss between the actual differential measurement value and the predicted differential measurement value is calculated, and a joint objective function is constructed to optimize the reconstructed network.

2. The method according to claim 1, characterized in that, The specific calculation formula for the dual-optical-path synchronous observation is as follows: , , in, These are the original observation image and the complementary observation image, respectively. Let t be the scene image of the t-th frame. Let n be a random binary mask for the t-th frame, and n represent the compression ratio. and These are the width and height of the image frame, respectively. This is for the Hadamard product operation.

3. The method according to claim 1, characterized in that, The formula for calculating the actual difference measurement value is as follows: , in, This represents the actual differential measurement value, whose physical meaning is equivalent to that obtained through... The mask modulates the scene to reveal the degree of fluctuation in scene pixel values ​​over time.

4. The method according to claim 1, characterized in that, The formula for calculating the predicted difference measurement value is as follows: , in, Indicates the predicted difference measurement value. This indicates the reconstruction of the video sequence. It is its t-th frame.

5. The method according to claim 1, characterized in that, The formula for calculating the motion consistency loss is as follows: , in, This represents the motion consistency loss, used to constrain the motion trajectory of the reconstructed video to remain consistent with the physical measurement domain.

6. The method according to claim 1, characterized in that, The optimization of the reconstruction network by constructing a joint objective function refers to: constructing a joint loss function using the reconstruction fidelity loss function and the motion consistency loss function, and forcibly constraining the dynamic evolution trajectory of the reconstructed video to remain consistent with the physical measurement domain until the preset convergence criterion is met.

7. The method according to claim 6, characterized in that, The reconstruction fidelity loss function The calculation formula is: ,in, To reconstruct the video sequence, This represents the true value of the original scene.

8. The method according to claim 6, characterized in that, The joint objective function The calculation formula is: ,in, is the weighting coefficient for motion consistency loss, used to balance the quality of the reconstructed image with the accuracy of the physical motion trajectory.

9. The method according to claim 6, characterized in that, The convergence criteria include: the joint objective function value is less than a preset threshold, or the decrease in the loss function over several consecutive iterations is less than a preset minimum value.

10. A video snapshot compression imaging device based on motion consistency verification, characterized in that, include: A dual-optical-path conjugate coding module is used to acquire the original observation image and the complementary observation image; The differential feature preprocessing module is used to extract the actual differential measurement values; The reconstruction network module is used to generate reconstructed video sequences; The verification and optimization module is used to calculate the predicted difference measurement value and adjust the reconstruction network module based on the consistency loss between the predicted difference measurement value and the actual difference measurement value.

11. The apparatus according to claim 10, characterized in that, The dual-optical-path conjugate encoding module includes a digital micromirror device, a controller, and two synchronously triggered image sensors.

12. The apparatus according to claim 10, characterized in that, The reconstructed network module includes: The feature extraction unit is used to extract spatiotemporal complementary features from the input original observation image, complementary observation image, and actual difference measurement value; The reconstruction network unit, consisting of multiple sequentially cascaded residual blocks, is used to generate the corresponding reconstructed video sequence based on the extracted spatiotemporal complementary feature maps.

13. The apparatus according to claim 10, characterized in that, The verification optimization module includes: The forward reprojection unit is used to perform Hadamard product and accumulation operations on the reconstructed video sequence and the corresponding mask to obtain the predicted difference measurement value. The loss calculation unit is used to calculate the motion consistency loss function and feed the loss value back to the reconstruction network module to reduce the reconstruction bias through gradient descent or iterative correction.

14. The apparatus according to claim 13, characterized in that, The method of reducing bias through gradient descent refers to: using the backpropagation algorithm to calculate the gradient of the joint objective function with respect to the parameters of each layer of the reconstructed network, using the Adam optimizer to dynamically adjust the learning rate, and updating the network weights through multiple iterations to make the predicted difference measurement value approximate the actual difference measurement value.

15. The apparatus according to claim 13, characterized in that, The method of reducing reconstruction deviation through iterative correction refers to the following: during the reconstruction process, the reconstructed video sequence generated in the current round is fed back to the verification and optimization module. Based on the residual map generated by motion consistency loss, local weight compensation in the spatial domain is performed on the deep feature map of the reconstruction network to eliminate motion artifacts.

16. A video snapshot compression imaging system based on motion consistency verification, characterized in that, It includes a memory and a processor, the processor executing instructions stored in the memory to implement the method as claimed in any one of claims 1 to 9.