A High Frame Rate Reflection-Free Video Reconstruction Method and System Based on a High-Speed Rotating Polarizer Modulated Pulse Camera
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-08
- Publication Date
- 2026-08-14
AI Technical Summary
[0006]本发明的目的在于提供一种基于高速旋转偏振片调制脉冲相机的高帧率无反射视频重建方法及系统,以解决上述背景技术中存在的至少一项技术问题
[0027]本发明有益效果:通过“高速旋转偏振片+脉冲相机”的成像方式,在时间域上连续获取偏振调制信息,再结合混合脉冲帧对齐和局部时窗建模,使系统能够在动态场景下构建更充分的物理分离约束,因此能够实现高帧率无反射视频重建,而不仅仅是单幅图像的反射去除;能够解决视频级、高速动态场景下的反射分离问题。在动态场景、强反射场景和高噪声场景下不仅能更稳定地完成透射层与反射层分离,而且能够获得更好的纹理恢复效果和更强的复杂场景适应能力。
Smart Images

Figure CN122573709A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video reconstruction technology, specifically to a high frame rate, reflection-free video reconstruction method and system based on a high-speed rotating polarizer modulated pulse camera. Background Technology
[0002] When a camera captures a scene through glass, the light entering the imaging system typically includes both transmitted and reflected light, resulting in a mixture of transmitted and reflected information on the image plane. Without additional prior information or hardware modulation, reconstructing the true transmitted layer from a single frame is a typical ill-posed problem. Under these conditions, it's difficult to obtain high-fidelity, high-frame-rate video without reflection interference when shooting high-speed dynamic scenes using traditional equipment. Existing methods can be broadly categorized into image-prior-based methods, deep learning-based methods, and physical cue-based methods. The first two types typically rely on gradient sparsity, smoothness, ghosting features, or semantic generation capabilities. However, in scenes with strong reflections, sharp reflections, or complex dynamics, they can easily mistake reflections for real textures, making it difficult to guarantee the physical consistency between the reconstructed result and the actual scene. To improve solvability, existing research has introduced physical cues such as flash, polarization, depth of field, and dual pixels. Among these, the polarization method, due to its lack of active emission and greater flexibility in imaging distance, is more valuable in practical applications. However, existing polarization dereflection methods still mainly rely on traditional frame cameras to acquire several discrete polarization angles, or to acquire polarized and non-polarized image pairs simultaneously. In dynamic scenes, the observation no longer meets the strict registration conditions due to the movement of objects at different times, thus limiting their applicability in high-speed scenes.
[0003] like Figure 1 As shown, the existing implementation scheme 1, "Reflection Separation using a Pair of Unpolarized and Polarized Images," first acquires a pair of polarized and unpolarized images from the same viewpoint as input. Assuming that the semi-reflective medium, such as glass, is approximately planar, the orientation parameters of the glass plane are estimated first. Then, based on the polarization physical imaging model, the incident angle and polarization direction corresponding to each pixel are calculated, thereby obtaining the initial separation result of the reflective and transmissive layers. A refinement network is then introduced to refine and restore the initial separation result, thereby improving the reflection removal effect and edge quality. The implementation process is as follows: A pair of polarized and unpolarized images are input into the "semi-reflective medium orientation estimation module" to obtain the predicted orientation parameters of the glass plane. Based on the parameters predicted in the previous step, camera intrinsic parameters, and the physical imaging model, the initial solutions for the reflective and transmissive layers are calculated. The initial solutions, the original input images, and the corresponding physical parameters are then fed into an encoder-decoder refinement network to obtain the final result.
[0004] like Figure 2As shown, the existing technical solution 2, "Polarized Reflection Removal with PerfectAlignment in the Wild," employs a two-stage deep learning framework: the first stage estimates the reflection layer, and the second stage combines the mixed image with the estimated reflection result to recover the transmission layer. Simultaneously, PNCC loss and perceptual loss are introduced to reduce the similarity between the reflection and transmission layers and improve visual quality. Implementation process: RAW data is acquired using a polarizing camera. This camera can obtain images at four polarization angles (0°, 45°, 90°, 135°) in a single shot and calculate the total intensity map based on these images. polarization degree and polarization angle Preprocessing of the RAW polarization image: Four polarization images are extracted from the mixed image, and the total intensity map, degree of polarization, polarization angle, and overexposure mask are calculated to form the multi-channel features of the network input. The multi-channel features from the previous step are then input into the reflection estimation network. The reflection layer estimate R is obtained, and then a refined network is used in the second stage. The mixed image M and R are used together as The input is further used to obtain the final recovered transmission layer T.
[0005] Both of the aforementioned existing technologies demonstrate the effectiveness of polarization information for reflection removal, but their technical approaches are primarily aimed at static or low-dynamic image-level reflection separation, and still have the following shortcomings: they are still based on discrete frame-based acquisition, making them difficult to apply to high-speed dynamic scenes. Essentially, both belong to frame-based image reflection removal methods. These methods can only obtain a limited number of polarization observations within a finite amount of time, and cannot continuously record the dense polarization state during the polarizer's rotation. Therefore, it is difficult to construct stable and sufficiently redundant separation constraints in high-speed motion scenes, and even more difficult to directly output high-frame-rate reflection-free video sequences. In the target scene, the high-speed rotation of the polarizer, the extremely short time window, and the light attenuation caused by the polarizer itself will significantly reduce the signal-to-noise ratio of the observed signal. Although existing technologies 1 and 2 utilize polarization information, their input is still in the form of conventional images. They do not design specific reliability modeling and robust solution mechanisms for the high noise problem under the conditions of asynchronous binary output from pulse cameras and extremely short integration, making it difficult to directly transfer to high-speed, low-signal-to-noise-ratio polarization modulation scenes. Summary of the Invention
[0006] The purpose of this invention is to provide a high frame rate, high-fidelity, reflection-free video reconstruction method and system based on a high-speed rotating polarizer-modulated pulse camera, to solve at least one of the technical problems existing in the background art. By combining a high-speed rotating polarizer with a pulse camera, polarization modulation information in dynamic scenes is continuously acquired, and a physical modeling, robust solution, and fine-tuning mechanism suitable for this type of observation is established to achieve high frame rate, high-fidelity, reflection-free video reconstruction in dynamic transparent glass scenes.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] In a first aspect, the present invention provides a high frame rate, reflection-free video reconstruction method based on a high-speed rotating polarizer-modulated pulse camera, comprising:
[0009] The system receives the modulated hybrid optical signal and performs asynchronous integration on it. When the accumulated charge reaches a threshold, it outputs a pulse signal and converts the asynchronous pulse stream into a hybrid pulse frame.
[0010] Hybrid pulse frame alignment: Establish a harmonic fitting model for each pulse frame within the local time window, and calculate the degree of polarization of each pixel accordingly; construct an effective low-degree polarization mask, perform masked phase correlation operation only in the low-degree polarization and reliable intensity region, estimate the translation amount of adjacent pulse frames relative to the center frame, and register each frame to the center frame coordinate system;
[0011] After aligning the hybrid pulse frames, a linear observation model of the reflective and transmissive layers is constructed within a local time window. A lightweight parameter estimation branch is introduced, and the hybrid coefficient and reliability weight of each frame are output based on the aligned observation sequence. A diagonal weighted matrix is then constructed, and the initial solution is obtained by minimizing the reliability weighted ridge regression objective function.
[0012] Based on the physical-guided refinement network, the initial results are further reconstructed, and the initial separation results are back-projected back to the observation domain to obtain the physical reconstruction observations. Then, the back-projection residual map is calculated to explicitly reflect the observation error areas that the physical model cannot explain.
[0013] As a further limitation of the first aspect of the present invention, in order to suppress outliers in high-noise impulse observations, an iterative reweighted least squares strategy is introduced: first, the initial weights of the branch outputs are estimated based on the parameters to complete the first solution, and then the effective weights are updated based on the residuals in the observation domain of the current estimation results, so that the contribution of observations with large residuals in the next round of solution is reduced, and then the process of "residual evaluation - weight update - re-solution" is repeated several times.
[0014] As a further limitation of the first aspect of the present invention, the physical-guided refinement network adopts a shared encoder and a dual-decoding branch structure, wherein the encoder fuses the aligned local time window observation sequence and the corresponding polarization angle information, and enhances the feature representation related to the polarization modulation state through the angle attention module; the two decoding branches predict the transmission layer residual and the reflection layer residual respectively, thereby obtaining the final output.
[0015] As a further limitation of the first aspect of the present invention, a linear observation model of the reflective and transmissive layers is constructed within a local time window, based on the observation of the first time frame near the center frame. Given the conditions satisfied by each observation, all observations can be stacked and written in matrix form. A lightweight parameter estimation branch is introduced, and the mixture coefficient and reliability weight of each frame are output based on the aligned observation sequence. A diagonal weighted matrix is constructed, and the initial solution is obtained by minimizing the reliability weighted ridge regression objective function.
[0016] As a further limitation of the first aspect of the present invention, the center frame transmission layer and reflection layer respectively calculate reconstruction loss and gradient loss to constrain the pixel error and edge error between the output result and the true value; using angular domain reconstruction loss and frame-by-frame rendering loss, the output result at each time moment is constrained to satisfy the polarization physical imaging relationship; explicit supervision is applied to the mixing coefficients predicted by the network to avoid parameter branch degradation; in areas with strong reflection, weighted high-frequency texture loss, phase loss and interlayer repulsion loss are further applied to improve the restoration quality of strong reflection areas and detail areas.
[0017] As a further limitation of the first aspect of the present invention, the network training adopts a phased progressive training strategy. The first phase is the warm-up phase, in which a strong teacher forcing is used in the first 5 epochs, and all the true mixing coefficients are injected into the physics solver, so that the network can learn a stable physics solution process first. The second phase is the transition phase, in which the proportion of teacher forcing is gradually reduced in the 6th to 15th epochs, so that the network partially relies on its own predicted mixing coefficients for solving. The third phase is the autonomous optimization phase, in which the network prediction parameters are used completely in the 16th to 50th epochs to complete the physics decoupling and refined reconstruction, and the physical consistency constraints and detail recovery constraints are enhanced simultaneously.
[0018] In a second aspect, the present invention provides a high frame rate, reflection-free video reconstruction system based on a high-speed rotating polarizer modulated pulse camera, comprising:
[0019] The acquisition module receives the modulated mixed optical signal and performs asynchronous integration on the modulated mixed optical signal. When the accumulated charge reaches a threshold, it outputs a pulse signal and converts the asynchronous pulse stream into a mixed pulse frame.
[0020] The alignment module is used for hybrid pulse frame alignment: it establishes a harmonic fitting model for each pulse frame within a local time window and calculates the degree of polarization of each pixel accordingly; it constructs an effective low-degree polarization mask, performs masked phase correlation operations only in regions with low-degree polarization and reliable intensity, estimates the translation amount of adjacent pulse frames relative to the center frame, and registers each frame to the coordinate system of the center frame.
[0021] The solution module is used to construct a linear observation model of the reflective and transmissive layers within a local time window after the alignment of the hybrid pulse frames is completed. A lightweight parameter estimation branch is introduced to output the hybrid system and reliability weights of each frame based on the aligned observation sequence. A diagonal weighted matrix is further constructed, and the initial solution results are obtained by minimizing the reliability weighted ridge regression objective function.
[0022] The reconstruction module is used to refine the network based on physical guidance, further reconstruct the initial results, backproject the initial separation results back to the observation domain to obtain the physical reconstruction observations; then calculate the backprojection residual map to explicitly reflect the observation error areas that the physical model could not explain.
[0023] Thirdly, the present invention provides a non-transitory computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the high frame rate reflection-free video reconstruction method based on a high-speed rotating polarizer modulation pulse camera as described in the first aspect.
[0024] Fourthly, the present invention provides a computer device including a memory and a processor, wherein the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor invokes the program instructions to execute the high frame rate reflection-free video reconstruction method based on a high-speed rotating polarizer modulation pulse camera as described in the first aspect.
[0025] Fifthly, the present invention provides an electronic device, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions for implementing the high frame rate reflection-free video reconstruction method based on a high-speed rotating polarizer modulation pulse camera as described in the first aspect.
[0026] Terminology Explanation: FPS: Frames Per Second; DoP: Degree of Polarization, used to characterize the strength of a pixel's influence by polarization modulation; RAW: Raw data from the camera sensor, referring to the raw output data before image signal processing; CNN: Convolutional Neural Network; ReLU: Rectified Linear Unit; AdamW: A parameter optimization algorithm for neural network training; Teacher Forcing: A strategy that injects some ground truth parameters into the network or physics solver module in the early stages of training to improve training stability; PNCC: A constraint loss used to reduce the similarity between reflective and transmissive layers; Fresnel: The Fresnel reflection / transmission model, used to describe the reflection and transmission of light at the interface of a medium; Epoch: A training cycle, referring to the process of all training samples completing one full iteration; U-Net: A convolutional neural network with an encoder-decoder structure; Transformer: A neural network structure based on an attention mechanism; Huber: Huber loss or Huber regression, a robust optimization method that takes into account both L1 and L2 characteristics.
[0027] The beneficial effects of this invention are as follows: By employing an imaging method combining a high-speed rotating polarizer and a pulse camera, polarization modulation information is continuously acquired in the time domain. Combined with hybrid pulse frame alignment and local time window modeling, the system can construct more comprehensive physical separation constraints in dynamic scenes. Therefore, it can achieve high frame rate reflection-free video reconstruction, rather than simply removing reflections from a single image; it can solve the reflection separation problem in video-level, high-speed dynamic scenes. In dynamic, highly reflective, and high-noise scenes, it can not only more stably separate the transmission and reflection layers but also achieve better texture restoration and stronger adaptability to complex scenes.
[0028] The advantages of additional aspects of the invention will be set forth more clearly in the following description or will be learned by practice of the invention. Attached Figure Description
[0029] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 The diagram shows the framework of the existing technical solution 1.
[0031] Figure 2 This is a schematic diagram of the existing technical solution 2.
[0032] Figure 3 This is a flowchart of the high frame rate reflection-free video reconstruction method based on a high-speed rotating polarizer modulated pulse camera according to an embodiment of the present invention.
[0033] Figure 4 This is a block diagram illustrating the principle of a high frame rate, reflection-free video reconstruction system based on a high-speed rotating polarizer-modulated pulse camera, as described in an embodiment of the present invention. Detailed Implementation
[0034] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0035] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0036] It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as here.
[0037] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or groups thereof.
[0038] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0039] To facilitate understanding of the present invention, the present invention will be further explained and described below with reference to the accompanying drawings and specific embodiments. However, the specific embodiments do not constitute a limitation on the embodiments of the present invention.
[0040] Those skilled in the art should understand that the accompanying drawings are merely schematic diagrams of embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.
[0041] Existing methods for video dereflection through glass or transparent media are mostly based on traditional frame-by-frame image acquisition, which can only capture a limited number of discrete polarization states. This makes it difficult to maintain physical consistency between observations in dynamic scenes, thus hindering the stable reconstruction of high-frame-rate, realistic transmission layer videos. Especially in fast-moving scenes, the discrete sampling of traditional cameras leads to temporal misalignment between the reflective and transmission layers, making it difficult to establish linear equations based on a few discrete polarization angles. Although high-speed pulse cameras can continuously record the polarization modulation process, the extremely short integration time and light attenuation caused by the polarizer result in low signal-to-noise ratios for pulse observations. Directly performing layer separation based on a continuous polarization physical model easily leads to numerical instability, noise amplification, and local solution failures. Furthermore, even if an initial separation result is obtained using a physical model, high-frequency textures, fine edges, and local structures in dynamic scenes are prone to oversmoothing, residual reflections, and detail loss. Therefore, a refined reconstruction mechanism that both adheres to physical constraints and can perform data-driven compensation for unreliable areas is needed.
[0042] Therefore, the technical problems to be solved by this invention include how to continuously acquire sufficiently dense polarization states in dynamic scenes to support high frame rate reflection-free video reconstruction; and how to robustly recover the transmission layer and its high-frequency texture and structural details under low signal-to-noise ratio pulse observation conditions.
[0043] Example 1
[0044] like Figure 3 , Figure 4As shown in Embodiment 1, a high frame rate reflection-free video reconstruction system based on a high-speed rotating polarizer modulated pulse camera is first provided. This system includes: an acquisition module, which receives the modulated mixed optical signal and performs asynchronous integration on the signal, outputting a pulse signal when the accumulated charge reaches a threshold; and converts the asynchronous pulse stream into mixed pulse frames. An alignment module is used for aligning the mixed pulse frames: establishing a harmonic fitting model for each pulse frame within a local time window and calculating the degree of polarization of each pixel accordingly; constructing a low-degree polarization effective mask, performing masked phase correlation operations only in low-degree polarization and reliable intensity regions, estimating the translation amount of adjacent pulse frames relative to the center frame, and registering each frame to the center frame coordinate system. A solution module is used to construct a linear observation model about the reflective and transmissive layers within a local time window after aligning the mixed pulse frames; introducing a lightweight parameter estimation branch, outputting the mixing system and reliability weights of each frame based on the aligned observation sequence, further constructing a diagonal weighted matrix, and obtaining the initial solution result by minimizing the reliability weighted ridge regression objective function. The reconstruction module is used to refine the network based on physical guidance, further reconstruct the initial results, backproject the initial separation results back to the observation domain to obtain the physical reconstruction observations; then calculate the backprojection residual map to explicitly reflect the observation error areas that the physical model could not explain.
[0045] In this embodiment, the above-described system is used to implement a high frame rate, reflection-free video reconstruction method based on a high-speed rotating polarizer-modulated pulse camera. The method includes: receiving a modulated mixed optical signal and asynchronously integrating the modulated mixed optical signal; outputting a pulse signal when the accumulated charge reaches a threshold; converting the asynchronous pulse stream into mixed pulse frames; aligning the mixed pulse frames: establishing a harmonic fitting model for each pulse frame within a local time window and calculating the degree of polarization of each pixel accordingly; constructing a low-degree polarization effective mask; performing masked phase correlation operations only in low-degree polarization regions with reliable intensity; and estimating the flatness of adjacent pulse frames relative to the center frame. The model is shifted and each frame is registered to the coordinate system of the center frame. After the hybrid pulse frames are aligned, a linear observation model about the reflective and transmissive layers is constructed within a local time window. A lightweight parameter estimation branch is introduced, and the hybrid system and reliability weights of each frame are output according to the aligned observation sequence. A diagonal weighted matrix is further constructed, and the initial solution is obtained by minimizing the reliability weighted ridge regression objective function. Based on the physically guided refinement network, the initial results are further reconstructed, and the initial separation results are back-projected back to the observation domain to obtain the physically reconstructed observations. The back-projection residual map is then calculated to explicitly reflect the observation error areas that the physical model cannot explain.
[0046] The overall technical solution of the method described in this embodiment mainly consists of four steps: preprocessing of the hybrid pulse observation signal, alignment of the hybrid pulse frame, solving with an iterative weighted physical solver, and physical-guided refined network reconstruction and recovery. Each step is specifically implemented through physical modeling or a module in the neural network that designs the response.
[0047] Step 1: Preprocessing of the mixed pulse observation signal:
[0048] A pulse camera receives a mixed light signal modulated by a high-speed rotating polarizer. Let the number of pixels be... At any moment The light intensity of the transmission layer is The light intensity of the reflective layer is Then the intensity of the modulated mixed light can be expressed as:
[0049]
[0050] in, The mixing coefficient is the time-varying factor related to the polarizer angle, incident angle, and reflected polarization component; 1 / 2 in the equation represents the inherent attenuation of unpolarized light caused by the linear polarizer. (Pulse camera...) Asynchronous integration; outputs a pulse signal when the accumulated charge reaches a threshold. .
[0051] In order for the subsequent network to process the pulse observation data, this embodiment segments the pulse stream within a fixed time window. Integral reconstruction is performed within the system to convert the asynchronous pulse stream into a mixed pulse frame:
[0052]
[0053] Pick This satisfies the "relatively zero motion observation" condition within a local time window, meaning that the object displacement between adjacent pulse reconstruction frames is negligible, while the intensity change caused by polarization modulation is still sufficiently significant. (Based on the center time...) Based on this, a local observation sequence can be constructed. Simultaneously record the corresponding polarizer angle sequence. This serves as the input for the subsequent solution module.
[0054] Step 2 Hybrid Pulse Frame Alignment: Since there may still be slight object movement or flickering changes caused by polarization modulation within the local time window, directly stacking multiple pulse frames pixel by pixel would destroy the pixel correspondence required by the physical solution. Therefore, this embodiment performs hybrid pulse frame alignment before solving.
[0055] Specifically, a harmonic fitting model is established for each pulse frame within the local time window:
[0056]
[0057] in, It is the angle-independent baseline term at pixel p, which can also be understood as the average level, DC component, and constant term of this pixel across all polarization angles. It is the polarization angle. It is at pixel p The modulation coefficient of the direction, which is the amplitude of the cosine double-angle component, is similarly... It is at pixel p The modulation coefficient of the direction is the amplitude of the double-angle component of the sine wave. Therefore... and Together, they determine the polarization angle response intensity and direction of the pixel.
[0058] And based on this, the degree of polarization of each pixel is calculated:
[0059]
[0060] in, This represents the regularization constant, used to prevent the denominator from being zero. Regions with higher polarization tend to correspond to regions more strongly affected by reflection modulation, while relatively stable transmission background regions have lower polarization values.
[0061] Based on this, this embodiment constructs a low-polarization effective mask, performing masked phase correlation operations only in regions with low polarization and reliable intensity to estimate the translation of adjacent pulse frames relative to the center frame, and registering each frame to the center frame coordinate system. To ensure consistency between subsequent monitoring and physical solution, the same geometric transformation is also applied synchronously to the mixing coefficient map at the corresponding time. This step significantly reduces the impact of polarization flicker and slight motion on subsequent physical solution.
[0062] Step 3: Solve using an iterative weighted physics solver: After alignment, construct a linear observation model of the reflective and transmissive layers within a local time window.
[0063] Let the center frame be near the first One observation satisfies:
[0064]
[0065] All observations can be stacked and written in matrix form:
[0066]
[0067] in, For the reflective and transmissive layers to be determined, ,matrix Each line is composed of the corresponding time. and Composition. To characterize the different contributions of various observations to the solution results, a lightweight parameter estimation branch is introduced, which outputs the mixture coefficients of each frame based on the aligned observation sequence. and reliability weight And further construct a diagonal weighted matrix:
[0068]
[0069] Based on this, the initial solution is obtained by minimizing the reliability-weighted ridge regression objective function:
[0070]
[0071] Its closed-form solution is:
[0072]
[0073] Since the unknowns only include two components, the reflective layer and the transmissive layer, the inversion of this matrix is very small, numerically stable, and the entire process is differentiable, making it easy to train end-to-end with neural networks.
[0074] Furthermore, to suppress outliers in high-noise pulse observations, this embodiment introduces an iterative reweighted least squares strategy into the solver. Specifically, the first solution is completed based on the initial weights output by the parameter estimation branch. Then, the effective weights are updated based on the residuals of the current estimation results in the observation domain, reducing the contribution of observations with large residuals to the next round of solution. This process of "residual evaluation—weight update—resolution" is repeated several times. In a preferred embodiment, this process can be performed three times to obtain a more robust initial value for the reflector layer. and initial value of the transmission layer Furthermore, this invention also constructs a confidence graph based on the determinant of the weighted second-order matrix to characterize the separability and reliability of the current local polarization decomposition problem, and then passes this confidence graph to the subsequent refinement network.
[0075] Step 4: Physics-guided refinement network: obtained from the above physics solver and Basic layer separation has been achieved, but due to factors such as the local time window assumption, impulse noise, and physical model approximation, the initial results may still contain missing details, local oversmoothing, and residual reflections.
[0076] Therefore, this embodiment designs a physically guided refinement network to further reconstruct the initial results. First, the initial separation results are back-projected back into the observation domain to obtain the physically reconstructed observations:
[0077]
[0078] Then calculate the back-projection residual plot:
[0079]
[0080] This residual plot can explicitly reflect the regions of observational error that the physical model fails to explain. (Initial transmission layer) Initial reflective layer Central Hybrid Pulse Frame Multi-angle minimum value map Confidence graph of physical solver and back projection residual A common input refinement network is used. The network employs a shared encoder and dual decoding branch structure. The encoder fuses the aligned local time window observation sequence and corresponding polarization angle information, and enhances the feature representation related to polarization modulation state through an angle attention module. The two decoding branches predict the transmission layer residuals respectively. and reflective layer residual Thus, the final output is obtained:
[0081] ,
[0082] In this embodiment, through this "physical solution + refinement" structure, the refinement network retains more analytical physical results in high-confidence regions and focuses on restoring texture and structural information in low-confidence or high-error regions, thereby ensuring both physical consistency and improving the reconstruction quality of high-frequency details. Finally, by repeating the above process for all center moments, a high frame rate, reflection-free video sequence can be continuously output.
[0083] In this embodiment, synthetic data is used to train the neural network. The specific training process is as follows:
[0084] (1) Synthetic Training Data: Due to the extreme difficulty in obtaining pixel-by-pixel aligned, strictly aligned, real-labeled data for the transmission layer, reflection layer, and pulse observations in real high-speed dynamic scenes, a physically consistent forward rendering method is used to construct synthetic training data. Specifically, firstly, non-reflection video sequences are collected and used as transmission layer videos. and reflective layer video The material; then continuously sampled the virtual polarizer angle. The system covers a 360° rotation range and simulates changes in the incident plane and polarization state at different spatial locations. Then, based on the Fresnel polarization mixing model, it synthesizes a mixed video frame by frame to obtain continuous-time mixed light observations. After obtaining the mixed video, it is input into a pulse simulator, which generates a corresponding pulse stream according to the pulse camera's integration-trigger mechanism, and then executes the pulse stream within a fixed time window. These are aggregated into a hybrid pulse frame. This allows us to simultaneously obtain: the pulse observation sequence, the hybrid pulse frame sequence, the corresponding polarization angle sequence, the ground truth values for the reflective and transmissive layers, and the ground truth values for the mixing coefficients needed to train the physics branch. In one embodiment, 50 sets of synthetic data, each with a continuous duration of 1 second, are constructed, with one set used for training and the other for testing.
[0085] (2) Training of neural networks
[0086] The entire network employs a training approach of "parameter estimation—differentiable physics decomposition—refinement." Specifically, the network mainly consists of a spatiotemporal coding module, a mixing coefficient and weight prediction module, an iterative weighted physics solution module, and a refinement network. The spatiotemporal coding module is used to extract parameters of length [missing information]. The characteristics of the mixed pulse frame sequence; the mixing coefficient and weight prediction module is used to output the mixing coefficient at each time step. The system includes frame reliability weights; an iterative weighted physical solver module calculates the baseline transmission and reflection layers based on the polarization physics model; and a refinement network combines physical residuals, confidence levels, and local context to further recover high-frequency textures and edge details from the baseline results.
[0087] During training, a length of Hybrid pulse frame stack The corresponding polarization angle sequence True value of the center frame transmission layer The true value of the center frame reflection layer True value of mixing coefficient The frame reliability weights are input into the network. The network first extracts joint features from the pulse sequence using the encoder, and then predicts the frame reliability weights for each frame using the angle attention parameter head. The prediction results are then fed into a differentiable physics solver to calculate the baseline transmission layer and the reflection layer. Finally, the baseline results, minimum map, mean map, backprojection residual map, and confidence map are input into the refinement network to output the final reconstruction results of the transmission layer and the reflection layer.
[0088] To ensure training stability, a phased, progressive training strategy is adopted. The first phase is the warm-up phase, in which strong teacher forcing is used in the first 5 epochs, injecting all ground truth mixing coefficients into the physics solver, allowing the network to learn a stable physics solution process. The second phase is the transition phase, in which the proportion of teacher forcing is gradually reduced from the 6th to the 15th epoch, allowing the network to partially rely on its own predicted mixing coefficients for solving. The third phase is the autonomous optimization phase, in which the network's predicted parameters are used entirely from the 16th to the 50th epoch to complete the physics decoupling and refined reconstruction, while simultaneously enhancing the physics consistency constraints and detail recovery constraints.
[0089] The training loss employs multiple joint constraints. Specifically, reconstruction loss and gradient loss are calculated for the transmission and reflection layers of the center frame, respectively, to constrain pixel and edge errors between the output and the ground truth. Angle domain reconstruction loss and frame-by-frame rendering loss are used to constrain the output at each time step to satisfy the polarization physical imaging relationship. Explicit supervision is applied to the mixing coefficients predicted by the network to avoid parameter branch degradation. In areas with strong reflection, weighted high-frequency texture loss, phase loss, and interlayer repulsion loss are further applied to improve the restoration quality of areas with strong reflection and details.
[0090] For training, the AdamW optimizer was used with a batch size of 8 and a total training duration of 50 epochs. The learning rate, number of physics iterations, ridge regularization parameters, and gradient pruning threshold were dynamically adjusted at different stages. Specifically, the learning rate was higher and the regularization was stronger during the warm-up stage to prioritize solution stability; the regularization was gradually reduced and more detailed constraints were introduced during the transition stage; and the learning rate was further reduced and the number of physics iterations was increased in the later stage to improve the final reconstruction quality.
[0091] Example 2
[0092] This embodiment 2 provides a non-transitory computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, they implement the high frame rate, reflection-free video reconstruction method based on a high-speed rotating polarizer modulation pulse camera as described above. The method includes:
[0093] The system receives the modulated hybrid optical signal and performs asynchronous integration on it. When the accumulated charge reaches a threshold, it outputs a pulse signal and converts the asynchronous pulse stream into a hybrid pulse frame.
[0094] Hybrid pulse frame alignment: Establish a harmonic fitting model for each pulse frame within the local time window, and calculate the degree of polarization of each pixel accordingly; construct an effective low-degree polarization mask, perform masked phase correlation operation only in the low-degree polarization and reliable intensity region, estimate the translation amount of adjacent pulse frames relative to the center frame, and register each frame to the center frame coordinate system;
[0095] After aligning the hybrid pulse frames, a linear observation model of the reflective and transmissive layers is constructed within a local time window. A lightweight parameter estimation branch is introduced, and the hybrid coefficient and reliability weight of each frame are output based on the aligned observation sequence. A diagonal weighted matrix is then constructed, and the initial solution is obtained by minimizing the reliability weighted ridge regression objective function.
[0096] Based on the physical-guided refinement network, the initial results are further reconstructed, and the initial separation results are back-projected back to the observation domain to obtain the physical reconstruction observations. Then, the back-projection residual map is calculated to explicitly reflect the observation error areas that the physical model cannot explain.
[0097] Example 3
[0098] This embodiment 3 provides a computer device, including a memory and a processor, wherein the processor and the memory communicate with each other, and the memory stores program instructions that can be executed by the processor. The processor calls the program instructions to execute the high frame rate reflection-free video reconstruction method based on a high-speed rotating polarizer modulation pulse camera as described above, the method including:
[0099] The system receives the modulated hybrid optical signal and performs asynchronous integration on it. When the accumulated charge reaches a threshold, it outputs a pulse signal and converts the asynchronous pulse stream into a hybrid pulse frame.
[0100] Hybrid pulse frame alignment: Establish a harmonic fitting model for each pulse frame within the local time window, and calculate the degree of polarization of each pixel accordingly; construct an effective low-degree polarization mask, perform masked phase correlation operation only in the low-degree polarization and reliable intensity region, estimate the translation amount of adjacent pulse frames relative to the center frame, and register each frame to the center frame coordinate system;
[0101] After aligning the hybrid pulse frames, a linear observation model of the reflective and transmissive layers is constructed within a local time window. A lightweight parameter estimation branch is introduced, and the hybrid coefficient and reliability weight of each frame are output based on the aligned observation sequence. A diagonal weighted matrix is then constructed, and the initial solution is obtained by minimizing the reliability weighted ridge regression objective function.
[0102] Based on the physical-guided refinement network, the initial results are further reconstructed, and the initial separation results are back-projected back to the observation domain to obtain the physical reconstruction observations. Then, the back-projection residual map is calculated to explicitly reflect the observation error areas that the physical model cannot explain.
[0103] Example 4
[0104] This embodiment 4 provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes instructions to implement the high frame rate reflection-free video reconstruction method based on a high-speed rotating polarizer modulation pulse camera as described above. The method includes:
[0105] The system receives the modulated hybrid optical signal and performs asynchronous integration on it. When the accumulated charge reaches a threshold, it outputs a pulse signal and converts the asynchronous pulse stream into a hybrid pulse frame.
[0106] Hybrid pulse frame alignment: Establish a harmonic fitting model for each pulse frame within the local time window, and calculate the degree of polarization of each pixel accordingly; construct an effective low-degree polarization mask, perform masked phase correlation operation only in the low-degree polarization and reliable intensity region, estimate the translation amount of adjacent pulse frames relative to the center frame, and register each frame to the center frame coordinate system;
[0107] After aligning the hybrid pulse frames, a linear observation model of the reflective and transmissive layers is constructed within a local time window. A lightweight parameter estimation branch is introduced, and the hybrid coefficient and reliability weight of each frame are output based on the aligned observation sequence. A diagonal weighted matrix is then constructed, and the initial solution is obtained by minimizing the reliability weighted ridge regression objective function.
[0108] Based on the physical-guided refinement network, the initial results are further reconstructed, and the initial separation results are back-projected back to the observation domain to obtain the physical reconstruction observations. Then, the back-projection residual map is calculated to explicitly reflect the observation error areas that the physical model cannot explain.
[0109] In summary, the high frame rate, reflection-free video reconstruction method based on a high-speed rotating polarizer-modulated pulse camera described in this invention employs a high-speed rotating polarizer to continuously polarize the mixed light and utilizes a pulse camera to sample the modulation process at high temporal resolution, obtaining an imaging scheme with a dense polarization state sequence that is difficult to obtain using traditional frame-based polarization methods. First, a hybrid pulse frame alignment method based on degree polarization masks is used to reduce local motion and reflection flicker interference. Then, an iterative weighted physical solver is constructed using the mixing coefficients predicted by the network and reliability weights to achieve robust separation of the transmission and reflection layers. The physical solution results are reprojected onto the observation domain, and combined with residual maps, confidence maps, and local prior information, the transmission and reflection layers are refined and restored, thereby improving the reconstruction scheme's ability to suppress high-frequency textures, edge structures, and residual reflections.
[0110] This invention utilizes a high-speed rotating polarizer + pulse camera imaging method to continuously acquire polarization modulation information in the time domain. Combined with hybrid pulse frame alignment and local time window modeling, the system can construct more comprehensive physical separation constraints in dynamic scenes, thus achieving high frame rate reflection-free video reconstruction, rather than simply removing reflections from single images. In other words, existing technologies, due to discrete observations and sparse temporal sequences, can only solve image-level problems; this invention, with continuous observations and dense temporal sequences, can solve video-level reflection separation problems in high-speed dynamic scenes. Existing polarization dereflection methods typically use conventional images as input and lack robust solution mechanisms specifically designed for extremely short integration times, polarization light attenuation, and low signal-to-noise ratio conditions under pulse observation. Therefore, they are prone to unstable separation, loss of detail, or residual reflections under high noise or weak modulation conditions. This invention, however, not only achieves more stable separation of the transmission and reflection layers in dynamic, highly reflective, and high-noise scenes but also achieves better texture restoration and stronger adaptability to complex scenes.
[0111] In practical applications, the sensor section can be replaced with an event camera, micro-polarization camera, high-speed polarization camera, or other imaging devices capable of providing high temporal resolution / polarization information. The alignment module can also be replaced with optical flow alignment, feature point registration, ECC registration, or deep learning registration methods. The iterative weighted physical solver and refinement network can be replaced with weighted least squares, Huber regression, L1 robust regression, Bayesian estimation, or other differentiable optimization methods, and U-Net, Transformer, diffusion models, or other residual reconstruction networks.
[0112] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0113] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0114] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0115] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment, whereby a series of operational steps are performed to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0116] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that, based on the technical solutions disclosed in the present invention, various modifications or variations that can be made by those skilled in the art without creative effort should be included within the scope of protection of the present invention.
Claims
1. A high frame rate, reflection-free video reconstruction method based on a high-speed rotating polarizer-modulated pulse camera, characterized in that, include: It receives the modulated mixed optical signal and performs asynchronous integration on the modulated mixed optical signal. When the accumulated charge reaches the threshold, it outputs a pulse signal. Convert asynchronous pulse streams into mixed pulse frames; Hybrid pulse frame alignment: Establish a harmonic fitting model for each pulse frame within the local time window, and calculate the degree of polarization of each pixel accordingly; construct an effective low-degree polarization mask, perform masked phase correlation operation only in the low-degree polarization and reliable intensity region, estimate the translation amount of adjacent pulse frames relative to the center frame, and register each frame to the center frame coordinate system; After aligning the hybrid pulse frames, a linear observation model of the reflective and transmissive layers is constructed within a local time window. A lightweight parameter estimation branch is introduced, and the hybrid coefficient and reliability weight of each frame are output based on the aligned observation sequence. A diagonal weighted matrix is then constructed, and the initial solution is obtained by minimizing the reliability weighted ridge regression objective function. Based on the physical-guided refinement network, the initial results are further reconstructed, and the initial separation results are back-projected back to the observation domain to obtain the physical reconstruction observations. Then, the back-projection residual map is calculated to explicitly reflect the observation error areas that the physical model cannot explain.
2. The high frame rate, reflection-free video reconstruction method based on a high-speed rotating polarizer-modulated pulse camera according to claim 1, characterized in that, To suppress outliers in high-noise impulse observations, an iterative reweighted least squares strategy is introduced: first, the initial weights of the branch outputs are estimated based on the parameters to complete the first solution; then, the effective weights are updated based on the residuals in the observation domain of the current estimation results, so that the contribution of observations with large residuals in the next round of solution is reduced; then, the process of "residual evaluation - weight update - re-solution" is repeated several times.
3. The high frame rate, reflection-free video reconstruction method based on a high-speed rotating polarizer-modulated pulse camera according to claim 1, characterized in that, The physics-guided refinement network employs a shared encoder and a dual-decoding branch structure. The encoder fuses the aligned local time window observation sequence and the corresponding polarization angle information, and enhances the feature representation related to the polarization modulation state through an angle attention module. The two decoding branches predict the transmission layer residual and the reflection layer residual, respectively, to obtain the final output.
4. The high frame rate, reflection-free video reconstruction method based on a high-speed rotating polarizer-modulated pulse camera according to claim 1, characterized in that, A linear observation model of the reflective and transmissive layers is constructed within a local time window, based on the observation near the center frame. Given the conditions satisfied by each observation, all observations can be stacked and written in matrix form. A lightweight parameter estimation branch is introduced, and the mixture coefficient and reliability weight of each frame are output based on the aligned observation sequence. A diagonal weighted matrix is constructed, and the initial solution is obtained by minimizing the reliability weighted ridge regression objective function.
5. The high frame rate reflection-free video reconstruction method based on a high-speed rotating polarizer-modulated pulse camera according to claim 1, characterized in that, The reconstruction loss and gradient loss are calculated for the central frame transmission layer and reflection layer, respectively, to constrain the pixel error and edge error between the output result and the ground truth. The angular domain reconstruction loss and frame-by-frame rendering loss are used to constrain the output result at each time step to satisfy the polarization physical imaging relationship. Explicit supervision is applied to the mixing coefficients predicted by the network to avoid parameter branch degradation. In areas with strong reflection, weighted high-frequency texture loss, phase loss and interlayer repulsion loss are further applied to improve the restoration quality of strong reflection areas and detail areas.
6. The high frame rate, reflection-free video reconstruction method based on a high-speed rotating polarizer-modulated pulse camera according to claim 1, characterized in that, The network training adopts a phased and progressive training strategy. The first phase is the warm-up phase, in which strong teacher forcing is used in the first 5 epochs, and all the ground truth mixing coefficients are injected into the physics solver, so that the network can learn a stable physics solution process. The second phase is the transition phase, in which the proportion of teacher forcing is gradually reduced in the 6th to 15th epochs, so that the network partially relies on its own predicted mixing coefficients for solving. The third phase is the autonomous optimization phase, in which the network prediction parameters are used to complete the physics decoupling and refined reconstruction in the 16th to 50th epochs, and the physical consistency constraints and detail recovery constraints are enhanced at the same time.
7. A high frame rate, reflection-free video reconstruction system based on a high-speed rotating polarizer-modulated pulse camera, characterized in that, include: The acquisition module receives the modulated mixed optical signal and performs asynchronous integration on the modulated mixed optical signal. When the accumulated charge reaches the threshold, it outputs a pulse signal. Convert asynchronous pulse streams into mixed pulse frames; The alignment module is used for hybrid pulse frame alignment: it establishes a harmonic fitting model for each pulse frame within a local time window and calculates the degree of polarization of each pixel accordingly; it constructs an effective low-degree polarization mask, performs masked phase correlation operations only in regions with low-degree polarization and reliable intensity, estimates the translation amount of adjacent pulse frames relative to the center frame, and registers each frame to the coordinate system of the center frame. The solution module is used to construct a linear observation model of the reflective and transmissive layers within a local time window after the alignment of the hybrid pulse frames is completed. A lightweight parameter estimation branch is introduced to output the hybrid system and reliability weights of each frame based on the aligned observation sequence. A diagonal weighted matrix is further constructed, and the initial solution results are obtained by minimizing the reliability weighted ridge regression objective function. The reconstruction module is used to refine the network based on physical guidance, further reconstruct the initial results, backproject the initial separation results back to the observation domain to obtain the physical reconstruction observations; then calculate the backprojection residual map to explicitly reflect the observation error areas that the physical model could not explain.
8. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement the high frame rate reflection-free video reconstruction method based on a high-speed rotating polarizer modulation pulse camera as described in any one of claims 1-6.
9. A computer device, characterized in that, The method includes a memory and a processor, the processor and the memory communicating with each other, the memory storing program instructions that can be executed by the processor, and the processor calling the program instructions to execute the high frame rate reflection-free video reconstruction method based on a high-speed rotating polarizer modulation pulse camera as described in any one of claims 1-6.
10. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions for implementing the high frame rate reflection-free video reconstruction method based on a high-speed rotating polarizer modulation pulse camera as described in any one of claims 1-6.