Fourier single-pixel imaging method and system based on space-time three-dimensional joint prior

By employing a spatiotemporal three-dimensional joint prior Fourier single-pixel imaging method, which combines sinusoidal structured light and phase-shifting techniques with 3D Hessian regularization and local low-rank temporal priors, the problems of reconstruction blur and poor inter-frame consistency in dynamic scenes are solved, achieving efficient and robust dynamic scene imaging.

CN121010702APending Publication Date: 2025-11-25CHONGQING UNIV OF EDUCATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511109771.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing Fourier single-pixel imaging technology suffers from problems such as blurred reconstructed images, poor inter-frame consistency, and discontinuous motion trajectories in dynamic scenes, especially at extremely low sampling rates. Furthermore, deep learning-based methods have poor generalization ability for non-trained distributed scenes.

Method used

A Fourier single-pixel imaging method based on spatiotemporal three-dimensional joint prior is adopted. The Fourier spectrum is obtained by combining sinusoidal structured light patterns with phase shift technology. A spatiotemporal three-dimensional joint prior objective function is constructed. Iterative optimization and reconstruction are performed using 3D Hessian regularization and local low-rank time prior terms. High-fidelity reconstruction of dynamic scenes is achieved by combining the ADMM algorithm.

Benefits of technology

Achieving high-fidelity reconstruction of dynamic scenes at extremely low sampling rates improves spatial detail preservation and inter-frame consistency in the temporal dimension. The temporal resolution of the reconstructed video can reach 20-60 frames per second, making it suitable for industrial dynamic inspection and real-time robot vision, while reducing hardware modification costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010702A_ABST
    Figure CN121010702A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computational imaging, and particularly discloses a Fourier single-pixel imaging method based on space-time three-dimensional joint priori, which comprises the following steps: S1, Fourier spectrum measurement: projecting a sine structured light pattern to a dynamic target scene through a spatial light modulator, collecting a corresponding light intensity signal by using a single-pixel detector, and measuring the Fourier spectrum; fourier spectrum sampling data of the scene is obtained; wherein the sine structured light pattern generates a measurement value of a complex Fourier coefficient through phase shift technology combination; s2, constructing a reconstruction objective function: taking Fourier spectrum consistency of measurement data and a reconstruction result as a data fidelity constraint, and introducing space-time three-dimensional joint priori as a regularization item; the space-time three-dimensional joint priori comprises a 3D Hessian regularization item and a local low-rank time priori item; and S3, an optimization solution step. By adopting the technical scheme of the invention, high-fidelity reconstruction of a dynamic scene can be realized by simultaneously utilizing spatial structure priori and time redundancy information at an extremely low sampling rate.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computational imaging, in particular to a Fourier ptychographic imaging method and system based on a spatio-temporal three-dimensional joint prior. BACKGROUND

[0002] As a new type of computational imaging technology, single-pixel imaging (SPI) combines a single high-sensitivity detector with a spatial light modulator, and has shown significant advantages in scenarios where traditional pixelated sensors are difficult to apply (such as mid-infrared, terahertz, X-ray, etc. wavebands) and extreme conditions (such as ultra-weak light, non-line-of-sight imaging). The core principle is to use a digital micromirror device (DMD) or other modulator to project a coded pattern, collect light intensity information through a single-pixel detector, and then recover the image through a reconstruction algorithm, thereby significantly reducing the imaging cost in a specific spectral range and improving the signal-to-noise ratio in a weak light environment.

[0003] However, the spatio-temporal resolution of SPI is limited by the trade-off between sampling efficiency and reconstruction quality: point-by-point scanning can achieve high resolution, but is extremely slow; compressed sampling can reduce the number of measurements through coded patterns and computational reconstruction, but can easily lead to image blurring and loss of details at low sampling rates. To break this limitation, researchers have proposed various sampling schemes and reconstruction algorithms, among which Fourier ptychographic imaging (FSI) has become one of the mainstream technologies due to its direct acquisition of the Fourier spectrum of the scene and its use of the natural image energy concentrated in the low frequency to significantly improve the sampling efficiency.

[0004] The development of existing FSI technology can be divided into two directions: sampling strategy optimization and reconstruction algorithm improvement. In terms of sampling strategy, a high-sampling-efficiency Fourier ptychographic imaging method is disclosed in Chinese patent publication No. CN113114882A, the core of which is non-uniform density sampling of the Fourier spectrum: by analyzing the Fourier spectrum of a large number of natural images, the Fourier coefficients are arranged in descending order of amplitude (importance), and high-probability sampling is performed on the high-importance coefficients based on a Gaussian function, and the L1-Magic compressed sensing algorithm is used to recover the unsampled high-importance coefficients. This method optimizes the sampling distribution, reduces the number of measurements while preserving key spectral information, and effectively improves the imaging efficiency of static scenes. However, its limitations are: only the sampling strategy of a single frame of image Fourier spectrum is optimized, without considering the temporal correlation between frames in dynamic scenes, and in extremely low sampling rates (such as below 10%) and fast motion scenes, there are still problems such as blurred reconstructed images and poor inter-frame consistency.

[0005] In terms of reconstruction algorithm, traditional FSI mostly uses total variation (TV) regularization to deal with underdetermined linear system, and enhances sharpness by constraining first-order derivative of image, but it is easy to produce "staircase artifact", that is, smooth gradient is quantized as block region, resulting in loss of details. To improve this problem, researchers proposed high-order prior (such as Hessian norm regularization), which preserves curvature information of image by constraining second-order derivative to avoid artifact. However, existing high-order prior is only applied to single-frame spatial domain, ignoring time redundancy of dynamic scene: when there is moving object in the scene, independent reconstruction of each frame will cause inter-frame flicker and incoherent motion trajectory due to lack of time constraint, especially under low sampling rate, the above problems are more prominent.

[0006] In addition, although the SPI method based on deep learning can directly map the measurement value and the image through the neural network, it relies on large-scale training data and has poor generalization ability for scenes with non-training distribution; and other dynamic scene SPI researches mostly reduce the sampling rate of single frame to exchange speed, sacrificing image quality.

[0007] Therefore, there is an urgent need for a Fourier ptychographic imaging method and system based on spatiotemporal three-dimensional joint prior, which can realize high-fidelity reconstruction of dynamic scene under extremely low sampling rate (extremely low sampling rate refers to 10% sampling rate) while utilizing spatial structure prior and temporal redundancy information. SUMMARY

[0008] The present application provides a Fourier ptychographic imaging method based on spatiotemporal three-dimensional joint prior, which can realize high-fidelity reconstruction of dynamic scene under extremely low sampling rate while utilizing spatial structure prior and temporal redundancy information.

[0009] To solve the above technical problems, the present application provides the following technical solutions:

[0010] The Fourier ptychographic imaging method based on spatiotemporal three-dimensional joint prior comprises the following contents:

[0011] S1 Fourier spectrum measurement step: a sinusoidal structured light pattern is projected onto a dynamic target scene through a spatial light modulator, and a single-pixel detector is used to collect corresponding light intensity signals to obtain Fourier spectrum sampling data of the scene; wherein the sinusoidal structured light pattern is combined to generate measurement values of complex Fourier coefficients through phase shift technology;

[0012] S2 construction of reconstruction target function step: the consistency of measurement data and Fourier spectrum of reconstruction result is taken as data fidelity constraint, and a spatiotemporal three-dimensional joint prior is introduced as a regularization term; the spatiotemporal three-dimensional joint prior comprises a 3D Hessian regularization term and a local low-rank temporal prior term, wherein the 3D Hessian regularization term constrains the second-order derivative of the scene in space and time dimensions to maintain smoothness, and the local low-rank temporal prior term utilizes the redundancy information between consecutive frames of the dynamic scene;

[0013] S3 optimization solving step: using plug and play alternating direction multiplier method to iteratively optimize the objective function, and solving to obtain the video reconstruction result of the dynamic scene.

[0014] The principle and beneficial effects of the basic scheme are as follows: in step S1 of the present application, the Fourier spectrum of the scene is directly collected by using a sinusoidal structured light pattern combined with a phase shift technique. Compared with random sampling or Hadamard basis sampling, the Fourier basis pattern naturally matches the characteristic that the spectral energy of a natural image is concentrated in the low frequency, and the key frequency components can be preserved in a small amount of measurement; the phase shift technique (such as four-step phase shift) accurately extracts the real part and the imaginary part of the complex Fourier coefficient by combining sinusoidal patterns of different phases, and provides a high-quality frequency domain measurement basis for subsequent reconstruction.

[0015] The objective function constructed in step S2 realizes information complementation through double regularization terms, uses a 3D Hessian regularization term to constrain the second-order derivatives (curvatures) in the spatial (x, y) and time (t) dimensions, avoids the ladder artifacts caused by traditional total variation (TV) regularization, and ensures the smooth transition of spatial details and the continuity of inter-frame changes in the time dimension (such as the coherence of the trajectory of a slowly moving object).

[0016] A local low-rank temporal prior term is used to utilize the redundancy between consecutive frames of a dynamic scene, and it is assumed that a local spatio-temporal block (such as multiple frames of pixels at the same spatial position) has a low-rank structure, and the inter-frame correlation is mined through a kernel norm constraint to supplement the missing high-frequency information under a low sampling rate. The combination of the two forms a synergistic constraint of spatial structure preservation and temporal redundancy utilization, and solves the problem of spatio-temporal information fragmentation in traditional frame-by-frame reconstruction.

[0017] The plug and play ADMM algorithm used in step S3 decomposes the objective function to transform the complex optimization problem into a sub-problem (Fourier domain data fidelity, Hessian domain smoothing, low-rank matrix approximation) that can be solved in a closed form, improves the iteration efficiency while ensuring the reconstruction accuracy, and effectively realizes the constraint effect of the spatio-temporal joint prior.

[0018] Compared with the traditional Fourier single-pixel imaging method, the present method can still realize high-fidelity reconstruction of a dynamic scene under a sampling rate of 10% or below, in the spatial dimension, the 3D Hessian regularization retains the image edges and details (such as a texture-rich background), and the average PSNR (peak signal-to-noise ratio) is improved by about 5dB, avoiding ladder artifacts; in the time dimension, the local low-rank prior ensures the inter-frame consistency, and reduces the blur and flicker of fast motion scenes.

[0019] By mining temporal redundancy information, this invention overcomes the limitations of static scene sampling strategies (such as Gaussian sampling) in dynamic scenes, and can handle challenging content such as fast motion and complex textures. The temporal resolution of the reconstructed video can reach 20-60 frames / second, meeting the needs of real-time dynamic imaging.

[0020] Model-based spatiotemporal priors do not rely on large-scale training data (unlike deep learning methods) and have stronger generalization ability for unseen scenes (such as novel dynamic textures); the plug-and-play Alternating Direction Multiplier (ADMM) algorithm is robust to measurement noise (such as Gaussian noise in low-light environments), ensuring imaging reliability under extreme conditions.

[0021] The low sampling rate reduces the number of measurements, and combined with efficient algorithms, it can shorten the single-frame imaging time to the millisecond level, making it suitable for fields such as industrial dynamic detection and real-time robot vision. The high-fidelity reconstruction results provide higher quality input for downstream tasks such as target recognition and tracking, which can theoretically improve the accuracy of dynamic target recognition. Furthermore, this invention is compatible with existing digital micromirror devices (DMD) and single-pixel detectors, and can be deployed without additional hardware modifications, reducing industrialization costs.

[0022] In summary, this invention achieves high-fidelity reconstruction of dynamic scenes at extremely low sampling rates by simultaneously utilizing spatial structure priors and temporal redundancy information.

[0023] Furthermore, the phase shifting technique described in S1 is a four-step phase shifting method, which involves projecting phase φ. i Four sinusoidal patterns, representing 0°, 90°, 180°, and 270° respectively, are combined to obtain the complex Fourier coefficients corresponding to the spatial frequencies (u, v), satisfying:

[0024]

[0025] Where, α i p represents the weighting coefficient. i It is a sine pattern, φ i Let j be the phase and j be the imaginary unit.

[0026] Furthermore, the imaging model of the Fourier spectrum sampling data described in S1 is an underdetermined linear system:

[0027] y = F p f+n

[0028] Where y is the measurement vector acquired by the single-pixel detector, F p is a partial Fourier matrix, f is the vectorized scene video volume, and n is the measurement noise vector.

[0029] Furthermore, the 3D Hessian regularization term mentioned in S2 is a combination of the l1 norms of the spatial and temporal second derivatives of the scene video volume f(x,y,t), expressed as:

[0030]

[0031] Among them, D xx D yy D tt The second-order partial derivative operators in the x, y, and t directions, respectively, D xy D xt D yt For cross-second-order partial derivative operators, σ is the weight coefficient for balancing spatial and temporal curvature constraints, used to adjust the relative weights of spatial smoothing and temporal smoothing regularization.

[0032] Furthermore, the value of the weighting coefficient σ is determined based on the spatial and temporal sampling rates. When the spatial and temporal sampling rates are comparable, σ = 1.

[0033] Furthermore, the local low-rank temporal prior mentioned in S2 is the sum of the kernel norms of the matrix formed by the pixel groups corresponding to the spatial locations in consecutive frames, and its expression is:

[0034]

[0035] Among them, P i (f) is the matrix of the i-th group, which consists of pixels at the same spatial location or small spatial neighborhood blocks in consecutive frames, ||·|| * For nuclear norm.

[0036] Furthermore, the expression for the objective function described in S2 is:

[0037]

[0038] The first item is the data fidelity item, λ H and λ L These are the weight coefficients of the 3D Hessian regularization term and the local low-rank time prior term, respectively.

[0039] Furthermore, the iterative process for optimizing the objective function using the plug-and-play alternating direction multiplier method described in S3 includes:

[0040] S31 introduces auxiliary variables to decompose the objective function and constructs an augmented Lagrange function;

[0041] S32 Alternating Variable Update: Solve the data fidelity subproblem in the Fourier domain, solve the 3D Hessian subproblem in the spatial domain by constrained second derivatives, and solve the local low-rank subproblem by singular value decomposition;

[0042] S33 iterates until convergence, then outputs the reconstruction result.

[0043] Furthermore, the "continuous frames" in the local low-rank temporal prior term are 3-10 frames, and the size of the "small spatial neighborhood block" is 8×8 to 16×16 pixels; the weighting coefficient λ H and λ L The value ranges from 0.01 to 0.1. The optimal value is determined by testing the test set to balance reconstruction accuracy and smoothness. Attached Figure Description

[0044] Figure 1 A flowchart illustrating an embodiment of a Fourier single-pixel imaging method based on spatiotemporal three-dimensional joint priors;

[0045] Figure 2 A comparison chart of reconstruction results from simulation experiments on the DAVIS-2017-test-dev-480p dataset;

[0046] Figure 3 A comparison chart of reconstruction results from simulation experiments on the DAVIS-2017-test-dev-480p dataset;

[0047] Figure 4 This is a schematic diagram of a Fourier single-pixel imaging system. Detailed Implementation

[0048] The following detailed description illustrates the specific implementation method:

[0049] Fourier single-pixel imaging methods based on spatiotemporal 3D joint priors (such as...) Figure 1 As shown), it includes the following:

[0050] S1 Fourier spectrum measurement steps: A sinusoidal structured light pattern is projected onto the dynamic target scene through a spatial light modulator, and the corresponding light intensity signal is collected using a single-pixel detector to obtain the Fourier spectrum sampling data of the scene; wherein, the sinusoidal structured light pattern is combined using phase-shifting technology to generate the measured values ​​of complex Fourier coefficients.

[0051] S2 steps for constructing the reconstruction objective function: Using the consistency of the Fourier spectrum between the measurement data and the reconstruction results as the data fidelity constraint, a spatiotemporal three-dimensional joint prior is introduced as a regularization term; the spatiotemporal three-dimensional joint prior includes a 3D Hessian regularization term and a local low-rank temporal prior term, wherein the 3D Hessian regularization term constrains the second derivative of the scene in the spatial and temporal dimensions to maintain smoothness, and the local low-rank temporal prior term utilizes the redundant information between consecutive frames of the dynamic scene;

[0052] S3 optimization solution steps: The objective function is iteratively optimized using the plug-and-play alternating direction multiplier method to obtain the video reconstruction results of the dynamic scene.

[0053] Specifically, Fourier single-pixel imaging (FSI) directly obtains the Fourier spectrum samples of the target scene by using structured sinusoidal illumination. In a typical FSI system, the scene f(x,y) is sequentially illuminated by sinusoidal fringe patterns with different spatial frequencies (u,v) and phases. A single-pixel detector integrates the reflected light intensity of each pattern, generating a Fourier coefficient measurement at that spatial frequency (u,v). Mathematically described, the measurement for the spatial frequency (u,v) can be expressed as the inner product (plus noise) of the scene and the corresponding illumination pattern:

[0054]

[0055] In actual operation, in order to obtain the complex Fourier coefficients at the frequencies (v,u), a set of phase-shift patterns with i = 1,..., 4 (e.g., φ i = 0°, 90°, 180°, 270°) need to be projected sequentially. By appropriately choosing the weights α i , the corresponding four intensity readings can be combined into a complex Fourier sample at that frequency. For example, in the four-step phase-shift scheme, the weighted sum of the four patterns satisfies:

[0056]

[0057] where the right side of Equation (2) represents the complex sinusoidal basis function at the frequency (u,v), which indicates that the combination of the four phase-shift patterns effectively constitutes the complex Fourier basis function corresponding to that frequency. After obtaining a set of measurements for the spatial frequencies , the results are assembled into a system of linear equations. Let f represent the vectorized scene image (with length N), and y = (y1, y2,..., y M ) T represent the vector containing M single-pixel measurements. If the complete set of Fourier basis patterns is used (covering the required bandwidth, so M = N), then there is:

[0058] y = Ff + n, (3)

[0059] where F is an M×N sensing matrix, and n is the measurement noise vector. In the ideal case of no noise and full sampling (M = N), F is invertible, and applying the Fourier inverse transform to the measurements can accurately recover the original image f. In practical situations, the sampling rate is usually less than 1, which means that only a partial set of Fourier basis patterns are projected (i.e., M < N). Let F p represent the sensing matrix formed by selecting M rows from the complete Fourier matrix F. Then the imaging model under partial Fourier sampling can be expressed as:

[0060] y = F p f + n, (4)

[0061] This is an underdetermined linear system. Equivalently, the sampling process can be described in the Fourier domain as a pointwise mask of the complete Fourier spectrum.

[0062] Let S(u,v) be a binary sampling function, taking the value 1 for the sampled frequency (u,v) and 0 otherwise. Treating F as a complete N×N discrete Fourier transform matrix, then... Let f represent the Fourier transform of image f. Therefore, the measurement process can be modeled as:

[0063]

[0064] Where ⊙ represents the Hadamard (element-wise) product. Here, S and Ff can be regarded as the sampling mask and the complete Fourier spectrum of the tiled image, respectively (if S(u,v)=1, it means that the frequency (u,v) is sampled; otherwise, it is 0). Substituting into equation (5), we get:

[0065]

[0066] This indicates that the measurement vector y essentially represents the sampled (and noisy) Fourier spectrum of the scene. Once the Fourier domain sampled values ​​y are obtained, a direct reconstruction method is to zero-padded the spectrum and then perform an inverse Fourier transform (i.e., set all unmeasured frequency components to zero and apply F...). -1 This allows for a quick acquisition of the initial reconstruction result f. ifft However, when the sampling rate is very low (missing a large number of high-frequency components), due to insufficient spectral information, direct reconstruction will result in severe blurring and ringing artifacts, and important image details will also be lost.

[0067] In other words, recovering f from y at low sampling rates is an ill-conditioned problem. To address this difficulty, regularized reconstruction methods need to be introduced, incorporating prior knowledge into the recovery process.

[0068] To stabilize the reconstruction of highly undersampled FSI data, this embodiment employs a 3D Hessian norm regularization method, extending the previous two-dimensional Hessian image prior to the spatiotemporal three-dimensional domain. Hessian regularization is a second-order image prior that penalizes the second derivative of the image, encouraging solutions to be sufficiently smooth in curvature while still allowing for sharp edges. Compared to first-order TV regularization, which easily leads to piecewise constant block artifacts, Hessian regularization, by penalizing the second derivative, can more effectively preserve smooth intensity transitions and details.

[0069] 2D Hessian Prior: For a two-dimensional image f(x,y), let D xx f and D yyf represents the second-order partial derivatives in the horizontal and vertical directions, respectively, and D... xy f = D yx f denotes the mixed second-order partial derivative. The Hessian regularization term can be defined as the l1 norm obtained by applying the Hessian operation to f. The two-dimensional Hessian norm of f is defined as:

[0070] ||f|| Hessian =||D xx f||1+||D yy f||1+2||D xy f||1,(7)

[0071] Where ||·||1 represents the summation of absolute values ​​over all pixel positions. The mixed second derivative D in this definition... xy f is multiplied by 2 to simultaneously consider D in the Hessian matrix. xy f and D yx f has two components. Minimize ||f|| Hessian It tends to obtain images where the second derivative is almost zero everywhere (i.e., the image intensity changes linearly), thus avoiding blocky artifacts caused by TV regularization while preserving smooth gradients.

[0072] Extending to 3D: The above Hessian prior is extended to a video volume f(x,y,t) containing two spatial and two temporal dimensions t. The 3D Hessian of f consists of the second derivatives of each dimension: D xx f、D yy f、D tt f (the pure second derivative in both spatial and temporal directions), and the cross derivative D xy f、D xt f、D yt f. By directly extending equation (7), the second derivative in the time direction can be incorporated in a similar form ||·||1. However, since the units of measurement or noise levels in the spatial and temporal dimensions may differ, a weighting factor σ is introduced to balance the spatial curvature regularization and the temporal curvature regularization. The 3DHessian norm of the video f(x,y,t) is defined as:

[0073]

[0074] Where σ is a balance coefficient (for example, σ = 1 when the spatial and temporal sampling rates are comparable), used to adjust the relative weights of spatial smoothing and temporal smoothing regularization. 3D Hessian regularization term. Penalizing high curvature both spatially and temporally encourages the reconstruction to approach piecewise linearity in the xy-plane and evolve smoothly over time. This leverages the similarity (temporal redundancy) between adjacent frames as a spatiotemporal prior. When an object moves or changes slowly, its second-order time derivative D... tt f will approach zero.

[0075] Low-rank temporal prior: In addition to the Hessian prior, an explicit temporal redundancy prior is incorporated. Cross-frame local low-rank constraints significantly improve the reconstruction quality of high-speed SPI video. It is assumed that the local 3D image patches unfolding along the time axis in the video are low-rank. Pixels (or small spatial neighborhoods) at corresponding positions in consecutive frames are grouped, and each group's data matrix is ​​encouraged to have a low-rank structure. This is typically achieved by imposing a nuclear norm penalty on each group's matrix (this idea is often called tensor nuclear norm when applied to spatiotemporal data).

[0076] Let P i (f) represents the matrix of the i-th group (formed by blocks or small objects in the same spatial location in consecutive t frames). Introducing Σ i ||P i (f)|| * This term promotes a low-rank structure in the time dimension. This local low-rank prior, combined with the Hessian smoothness prior, utilizes redundant information in dynamic scenes to recover finer details with fewer measurements.

[0077] Optimization problem: Combining the data fidelity term in equation (6) with the two regularization terms mentioned above, we obtain the overall reconstruction objective function:

[0078]

[0079] Where λ H and λ L These are the regularization weights of the 3D Hessian prior and the low-rank prior, respectively.

[0080] First item It is the least squares data fidelity term that ensures the reconstruction result is consistent with the measured Fourier coefficients y.

[0081] The second term is a penalty for the 3D Hessian norm defined in equation (8), and the third term is a penalty for the nuclear norm of the cross-frame pixel group matrix. The optimization problem is to find a video f that both conforms to the measurement data and is as smooth as possible in curvature, while having the maximum temporal redundancy. Directly solving (9) is difficult because the e1 term and the nuclear norm term are non-smooth.

[0082] The Alternating Direction Multiplier Method (ADMM) is used to solve this optimization problem. It introduces Hessian components and auxiliary variables of the block to iteratively minimize the augmented Lagrangian function.

[0083] Specifically, a splitting strategy similar to plug-and-play regularization is used: Hessian prior constraints and low-rank prior constraints are applied alternately while keeping the data fidelity terms unchanged, so that the two priors complement and reinforce each other.

[0084] This embodiment experimentally validates and evaluates the proposed 3D Hessian-based reconstruction method on a publicly available dynamic scene dataset, and compares it with the basic Fourier single-pixel imaging (FSI) method without advanced regularization. All experiments used a 10% sampling rate (i.e., only 10% of the full Fourier coefficients were captured) to simulate extreme undersampling conditions. The DAVIS-2017 video dataset (480p resolution) was used as the test platform, which provides diverse real-world scenes and their corresponding ground truth video frames, facilitating reconstruction quality evaluation.

[0085]

[0086] Table 1. Reconstruction results of the DAVIS-2017 dataset (PSNR in dB)

[0087] Quantitative and qualitative results: Table 1 lists the peak signal-to-noise ratio (PSNR) reconstructed at a 10% sampling rate for four representative video sequences selected from the DAVIS dataset, comparing the performance of the baseline FSI method and this embodiment.

[0088] These four sequences—"aerobatics," "car-race," "carousel," and "cats-car"—include challenging content such as fast-paced movement and rich detail.

[0089] The results show that this embodiment achieved significantly higher PSNR in all cases.

[0090] For example, in the "car-race" sequence, the PSNR reconstructed by traditional FSI is approximately 26.8 dB, while this embodiment achieves 34.9 dB, an improvement of over 8 dB. Across these video segments, this embodiment achieves an average PSNR of approximately 32.1 dB, significantly higher than FSI's 26.8 dB, representing an improvement of approximately 5.3 dB. This substantial improvement demonstrates the effectiveness of the 3D Hessian prior in ensuring image fidelity at extremely low sampling rates.

[0091] This improvement is particularly noticeable for sequences with complex motion or texture (such as the detailed background in "cats-car"): in these cases, the reconstruction quality of FSI drops below 26 dB, while the method in this embodiment can still maintain a PSNR of about 30 dB.

[0092] The above results indicate that introducing spatiotemporal regularization is key to achieving high-fidelity reconstruction of video SPI.

[0093] Figure 2 and Figure 3 Examples of reconstruction results for partial frames in the test sequence are shown, comparing the image quality of traditional FSI at a 10% sampling rate with that of the method in this embodiment. Each image contains three rows: the first row is the original frame, the second row is the reconstruction result of conventional FSI (obtained by directly performing an inverse Fourier transform on the spectrum with only 10% sampling), and the third row is the reconstruction result of the method in this embodiment. It can be seen that the image reconstructed by the method in this embodiment is clearer, sharper, and more closely matches the original scene.

[0094] Another significant advantage of this embodiment is that it maintains temporal consistency. As the video progresses, the FSI-reconstructed image exhibits flickering and inconsistencies in detail between frames (this is because it reconstructs each frame independently without utilizing temporal priors); in contrast, this embodiment generates a temporally stable and consistent video by utilizing inter-frame redundancy. Moving objects in this embodiment exhibit coherent and smooth motion trajectories in the reconstruction, without the abrupt changes seen in the baseline method.

[0095] This embodiment combines spatial Hessian priors with temporal low-rank constraints to preserve fine spatial details and ensure inter-frame consistency, even at a sampling rate of only 10%, resulting in higher PSNR and sharper video reconstruction. Importantly, these improvements are entirely derived from model-based priors, requiring no deep learning training. This embodiment provides a robust and data-efficient solution for obtaining high-quality single-pixel video under limited measurements. This approach will help open up new possibilities in computational imaging, enabling single-pixel sensors to capture dynamic scenes with unprecedented clarity even under extremely constrained perception conditions.

[0096] To achieve the above method, this embodiment also discloses a Fourier single-pixel imaging system (such as...). Figure 4 As shown, in a Fourier single-pixel imaging system, a digital projector or digital micromirror device projects a sinusoidal illumination pattern onto the scene (test target). A single-pixel detector collects the reflected light corresponding to each pattern (either directly or through a scattering medium). A computer controls the projected pattern and reconstructs the image based on the Fourier coefficients measured by a data acquisition board.

[0097] The above are merely embodiments of the present invention. The invention is not limited to the fields covered by these embodiments. Commonly known structures and characteristics in the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are able to access all existing technologies in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. A Fourier single-pixel imaging method based on spatiotemporal three-dimensional joint priors, characterized by the following: S1 Fourier spectrum measurement steps: A sinusoidal structured light pattern is projected onto the dynamic target scene using a spatial light modulator, and the corresponding light intensity signal is collected using a single-pixel detector to obtain the Fourier spectrum sampling data of the scene; among which, The sinusoidal structured light pattern is combined using phase-shifting technology to generate the measured values ​​of the complex Fourier coefficients; S2 steps for constructing the reconstruction objective function: Using the consistency of the Fourier spectrum between the measurement data and the reconstruction results as the data fidelity constraint, a spatiotemporal three-dimensional joint prior is introduced as a regularization term; the spatiotemporal three-dimensional joint prior includes a 3D Hessian regularization term and a local low-rank temporal prior term, wherein the 3D Hessian regularization term constrains the second derivative of the scene in the spatial and temporal dimensions to maintain smoothness, and the local low-rank temporal prior term utilizes the redundant information between consecutive frames of the dynamic scene; S3 optimization solution steps: The objective function is iteratively optimized using the plug-and-play alternating direction multiplier method to obtain the video reconstruction results of the dynamic scene.

2. The Fourier single-pixel imaging method based on spatiotemporal three-dimensional joint prior as described in claim 1, characterized in that, The phase shifting technique described in S1 is a four-step phase shifting method, which involves projecting phase φ. i Four sinusoidal patterns, representing 0°, 90°, 180°, and 270° respectively, are combined to obtain the complex Fourier coefficients corresponding to the spatial frequencies (u, v), satisfying: Where, α i p represents the weighting coefficient. i It is a sine pattern, φ i Let j be the phase and j be the imaginary unit.

3. The Fourier single-pixel imaging method based on spatiotemporal three-dimensional joint prior as described in claim 2, characterized in that, The imaging model of the Fourier spectrum sampling data described in S1 is an underdetermined linear system: y=F p f+n Where y is the measurement vector acquired by the single-pixel detector, F p is a partial Fourier matrix, f is the vectorized scene video volume, and n is the measurement noise vector.

4. The Fourier single-pixel imaging method based on spatiotemporal three-dimensional joint prior as described in claim 3, characterized in that, The 3D Hessian regularization term mentioned in S2 is the l1 norm combination of the spatial and temporal second derivatives of the scene video volume f(x,y,t), expressed as: Among them, D xx D yy D tt The second-order partial derivative operators in the x, y, and t directions, respectively, D xy D xt D yt For cross-second-order partial derivative operators, σ is the weight coefficient for balancing spatial and temporal curvature constraints, used to adjust the relative weights of spatial smoothing and temporal smoothing regularization.

5. The Fourier single-pixel imaging method based on spatiotemporal three-dimensional joint prior as described in claim 4, characterized in that, The value of the weighting coefficient σ is determined based on the spatial and temporal sampling rates. When the spatial and temporal sampling rates are comparable, σ = 1.

6. The Fourier single-pixel imaging method based on spatiotemporal three-dimensional joint prior as described in claim 5, characterized in that, The local low-rank temporal prior mentioned in S2 is the sum of the kernel norms of the matrix formed by the pixel groups corresponding to the spatial locations in consecutive frames, and its expression is: Among them, P i (f) is the matrix of the i-th group, which consists of pixels at the same spatial location or small spatial neighborhood blocks in consecutive frames, ||·|| * For nuclear norm.

7. The Fourier single-pixel imaging method based on spatiotemporal three-dimensional joint prior as described in claim 6, characterized in that, The expression for the objective function described in S2 is: The first item is the data fidelity item, λ H and λ L These are the weight coefficients of the 3D Hessian regularization term and the local low-rank time prior term, respectively.

8. The Fourier single-pixel imaging method based on spatiotemporal three-dimensional joint prior as described in claim 7, characterized in that, The iterative process for optimizing the objective function using the plug-and-play alternating direction multiplier method described in S3 includes: S31 introduces auxiliary variables to decompose the objective function and constructs an augmented Lagrange function; S32 Alternating Variable Update: Solve the data fidelity subproblem in the Fourier domain, solve the 3D Hessian subproblem in the spatial domain by constrained second derivatives, and solve the local low-rank subproblem by singular value decomposition; S33 iterates until convergence, then outputs the reconstruction result.

9. The Fourier single-pixel imaging method based on spatiotemporal three-dimensional joint prior as described in claim 8, characterized in that, The "continuous frames" in the local low-rank temporal prior term are 3-10 frames, and the size of the "small spatial neighborhood block" is 8×8 to 16×16 pixels; the weighting coefficient λ H and λ L The value ranges from 0.01 to 0.

1. The optimal value is determined by testing the test set to balance reconstruction accuracy and smoothness.

10. A Fourier single-pixel imaging system based on spatiotemporal three-dimensional joint prior, characterized in that, The method described in any one of claims 1-9 was adopted.

Citation Information

Patent Citations

  • Fourier single-pixel imaging method with high sampling efficiency

    CN113114882A