A light field microscope three-dimensional image reconstruction system and method with integrated event camera
By combining event cameras and deep learning methods, event stream data is generated and SSL-E2VID and VCD-Net networks are used to solve the image authenticity and resolution problems of light field microscopes in high dynamic range and high speed scenes, and efficient three-dimensional image reconstruction is achieved.
Patent Information
- Application Number
- CN202411358588.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2044-09-27
AI Technical Summary
The existing light field microscopy technology has poor image authenticity in high dynamic range and high speed scenarios, and the lack of large-scale event data and corresponding ground truth image data sets, resulting in insufficient quality and resolution of 3D reconstruction.
Combining event cameras and deep learning methods, event stream data is generated through data simulation modules, image reconstruction is performed using SSL-E2VID and VCD-Net networks, including event image reconstruction network module and three-dimensional image reconstruction network module, and optical flow estimation and image reconstruction are performed using self-supervised learning and optical flow prediction deep network units.
The image reconstruction resolution and signal-to-noise ratio in high dynamic range and high speed scenes are improved, artifacts and axial deformation are reduced, and three-dimensional image reconstruction with high video speed is realized.
Smart Images

Figure CN119313813B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of optical technology, relates to the intersection of microscopy and computational imaging technology, and specifically relates to a light field microscope three-dimensional image reconstruction system and method integrated with an event camera. Background Art
[0002] Optical imaging methods using fluorescent indicators can convert biophysical changes into fluorescence changes, thereby directly monitoring cellular activities in living organisms with high spatial resolution and non-invasively, enabling visualization of complex dynamic biological processes. Therefore, in recent years, three-dimensional rapid microscopy technology has become an essential means for biological dynamic observation and medical research.
[0003] Light-field microscopy (LFM) is a novel optical microscopy technique. LFM uses a pinhole array or microlens array to record the distribution of light in free space, encoding three-dimensional information into two-dimensional, multi-path measurements. This allows for capturing scene information about the incident light from a single snapshot and computationally reconstructing the three-dimensional structure of the object space, enabling scanless imaging of the target three-dimensional region at video frame rates. However, LFM is a single-lens, three-dimensional (3D) wide-field-of-view technique limited by the synchronous readout of conventional CMOS cameras, resulting in an inherent trade-off between maintaining high resolution and high data throughput at limited frame rates. Furthermore, due to the high-dimensional data volume of light-field signals, computational post-processing of the signals faces a trade-off between efficient data perception and computational efficiency. Existing LFM techniques typically operate below 100 Hz at full frame resolution, a limitation that hinders their application in capturing ultrafast dynamic biological processes that can exceed kilohertz, such as voltage signals in the mammalian brain, hemodynamics, and muscle contraction. Ultrafast imaging strategies can address these technical limitations. As a novel biomimetic visual sensor, event cameras offer ultra-high temporal resolution, can provide measurements with latency as low as 1 microsecond, record ultrafast signal changes at speeds exceeding 10 kHz, and can be flexibly integrated with various platforms. They can capture high-quality data under challenging visual conditions such as extreme lighting or high-speed motion. They have great potential for the reconstruction and restoration of high dynamic range images and high frame rate videos, providing a new paradigm for high dynamic range, low-latency image sensors and showing promise in various applications such as autonomous driving, gesture recognition, and single-molecule localization microscopy. Currently, a large amount of research is focused on optimizing the performance of light-field microscopy. Combining light-field microscopy with event cameras has the potential to overcome the limitations of frame-based cameras in light-field microscopy imaging.
[0004] In 2006, Levoy et al. placed a microlens array at the natural image plane of a widefield microscope to create the first light-field microscope. Initial light-field microscopes sacrificed lateral resolution to capture sufficient angular information and suffered from inherent artifacts in the focal plane. In 2013, Broxton et al. proposed a wave optics model for light-field microscopy and the Richardson-Lucy (RL) 3D deconvolution algorithm for reconstructing 3D volumes from 2D light-field images. In 2019, Guo et al. introduced a Fourier lens into the optical path and proposed the Fourier light-field microscope (Fourier LFM). This method transfers the various physical transformations of the optical system from the time domain to the frequency domain, improving reconstruction speed and effectively mitigating artifacts caused by angular crosstalk due to sample redundancy or scattering. In 2021, Wang et al. combined a view-channel-depth (VCD) deep neural network with a light-field microscope to design a high-speed, real-time 3D reconstruction algorithm that does not require an R–L deconvolution step. In 2023, Guo et al. creatively combined an event camera with a Fourier light-field microscope to develop EventLFM, a simple and economical system. Operating on a novel asynchronous readout architecture, EventLFM bypasses the inherent frame rate limitations of conventional CMOS systems and significantly alleviates the low SBR challenges typically encountered in scattering environments. It has the potential to become a tool for visualizing complex, dynamic 3D biological phenomena in various biomedical applications. However, Guo et al. used a simple light-field refocusing algorithm for 3D reconstruction, which resulted in artifacts and axial elongation. In recent years, data-driven algorithms have been widely applied in various signal and image processing tasks, demonstrating excellent image reconstruction capabilities in the field of neuroimaging. Deep neural networks are capable of modeling ambiguous or even unknown physical systems. Their universal approximation capabilities are suitable for solving highly complex functions. Using deep neural networks for reconstruction can significantly improve the reconstructed imaging resolution and signal-to-noise ratio. Parallel computing also greatly increases the reconstruction speed and provides stronger fitting capabilities.
[0005] To address the 3D reconstruction quality and resolution issues of EventLFM, we proposed a new method for 3D reconstruction of target areas using LFM combined with an event camera based on EventLFM. We also proposed a new deep learning method combining the neural networks SSL-E2VID and VCD-Net. This method trains the network with 2D light field image events generated by the original 3D events, and then recovers intensity information from the event data for reconstruction. Recent advances in deep learning rely on training with simulated data, requiring the network to match the training data in the form of event sequences with corresponding real image sequences. However, large-scale event data and corresponding ground-truth image datasets are currently unavailable. Furthermore, in high dynamic range and high-speed scenarios, the image authenticity of traditional cameras is poor. Summary of the Invention
[0006] In order to address the problem that there is currently no large-scale event data and corresponding ground-truth image datasets, and at the same time to solve the problem that the images acquired by traditional cameras in high dynamic range and high-speed scenes have poor authenticity, the present invention provides a light field microscope three-dimensional image reconstruction system and method integrated with an event camera.
[0007] The technical solution of the light field microscope 3D image reconstruction system and method integrated with an event camera of the present invention is as follows:
[0008] A light field microscope 3D image reconstruction system with an integrated event camera includes a data simulation module that simulates event data for training a network using raw light field data, an event image reconstruction network module that can convert a continuous event stream into an image, and a 3D image reconstruction network module that reconstructs the image into a 3D scene.
[0009] A further improvement to the technical solution of the present invention is that the data simulation module generates event stream data of image sequences with different movement paths based on a random process, providing training samples for the event image reconstruction network module. The data simulation module uses an event simulator based on a random process, namely a dynamic visual sensor voltmeter, to obtain the event sequence. The specific steps are as follows:
[0010] S1. Observe the three-dimensional area of the target using a light field microscope, intercept the target's movement trajectory within a certain time period, obtain two-dimensional light field images of the same three-dimensional area at different times, and determine the timestamp of each image;
[0011] S2, inputs the two-dimensional light field image with time stamp obtained in S1 into the dynamic vision sensor voltmeter, detects its spatiotemporal brightness changes in the form of asynchronous "event" stream when the intensity occurs, and outputs the corresponding event sequence;
[0012] S3. Encapsulate the event sequence and the light field image to obtain the training samples required by the optical flow prediction deep network unit and the reconstruction deep neural network training unit;
[0013] S4. Repeat S1, S2, and S3 to capture images in multiple three-dimensional regions with different motion trajectories to obtain a large amount of event training data.
[0014] A further improvement of the above technical solution of the present invention is that the specific process of step S2 is:
[0015] The event camera samples light when each pixel brightness L changes, and each event with index i can be encoded as e i =(x i ,t i ,p i), where x i =(h i ,l i ) T Represents the horizontal and vertical coordinates of the pixel, t i Record timestamp, polarity p i ∈{+,-} records the increase and decrease of brightness; the kth time window Δt k The brightness increase ΔL k Can be coded as:
[0016]
[0017] Where x represents the set of pixel horizontal and vertical coordinates, and C represents the contrast sensitivity threshold.
[0018] A further improvement of the above technical solution of the present invention is that: the event image reconstruction network module includes an optical flow prediction deep network unit for performing optical flow estimation from event data, and an image reconstruction deep network unit for performing three-dimensional reconstruction through learning; the optical flow prediction deep network unit and the image reconstruction deep network unit are jointly trained.
[0019] The further improvement of the above technical solution of the present invention is that the optical flow prediction deep network unit uses contrast maximization proxy loss to learn to estimate optical flow by compensating for motion blur in input events; the image reconstruction deep network unit learns to perform image reconstruction through image registration
[0020] A further improvement of the above technical solution of the present invention is that the optical flow prediction deep network unit uses contrast maximization proxy loss to learn to estimate optical flow by compensating for motion blur in input events. The specific process is:
[0021] On the Lambert surface, constant illumination, Δt k Under smaller assumptions, the event-based luminosity constant is expressed as:
[0022]
[0023] Where x represents the set of horizontal and vertical coordinates of pixels, u k (x) represents the pixel in the kth time window Δt k The optical flow vector inside, The luminance signal L k As the optical flow vector u=(u h ,u l ) T The spatial gradient of movement, δ h L k , δ l L k Respectively represent L k The partial gradient in the horizontal and vertical directions, ≈ means approximately equal;
[0024] The input of the optical flow prediction deep network unit is a voxel grid E with B time windows k , each time window is filled with event streams consecutive non-overlapping partitions of N events, each containing N events; for each partition, each event e i Its polarity p i Assigned to the two closest windows as follows:
[0025]
[0026] κ(a)=max(0,1-|a|) (4)
[0028]
[0029] Where b is the window index, κ(a) is the maximum function operator, a is a scalar value, To normalize event timestamps;
[0030] The optical flow prediction deep network unit retrieves the accurate pixel-by-pixel optical flow u(x i ), the event is propagated to the reference time t in the following way ref :
[0031] x′ i =x i +(t ref -t i )u(x i ) (6)
[0033] x′ i is x i Optical flow propagation at t ref Position after the moment;
[0034] By bilinear interpolation at each propagation t ref The deblurring quality is evaluated by generating an image with the average timestamp at each pixel of polarity p′ after time instant:
[0035]
[0036]
[0037] in, is the normalized t ref Timestamp, j = {i|p i =p′}, p′∈{+,-}, parameter∈≈0;
[0038] Contrast maximization proxy loss Defined as:
[0039]
[0040] The total loss for training FlowNet is:
[0041]
[0042] in, is the Charbonnier smoothing prior, and λ1 is a scalar that balances the effects of the two losses.
[0043] A further improvement of the above technical solution of the present invention is that the image reconstruction deep network unit learns to perform image reconstruction through image registration, and the specific process is:
[0044] The input of the image reconstruction deep network unit is the same as that of the optical flow prediction deep network unit. The reconstruction problem uses the incremental reference image ΔL and the predicted The difference between the two is used to reconstruct the brightness signal of the input event, is the brightness after reconstruction; ΔL and The two incremental images are warped to a common time frame*, yielding ΔL * and
[0045] Defined as:
[0046]
[0047] Among them, x represents the pixel position, u is the optical flow vector, G + and G - The definition is as follows:
[0048]
[0049] Where P is a two-channel image containing the number of pixel positions of the image H that received events during the event warping process;
[0050] The spatial gradient of is warped to the current time instance and is defined as:
[0051]
[0052] in, In the time window Δt k Reconstructed optical flow within The warping function of
[0053] The photometric reconstruction loss is defined as:
[0054]
[0055] The resulting unbounded brightness estimate is first obtained by Transformed to intensity space, exp() represents the exponential function, is in the time window Δt k The reconstruction results within the reconstruction time period; the final reconstruction Expressed as:
[0056]
[0057] Among them, m and M are The 1% and 99% percentiles, Clipped to the range [0,1];
[0058] The temporal loss is defined as the photometric error between two consecutive reconstructed frames:
[0059]
[0060] ‖‖1 represents the L1 norm;
[0061] The total loss for training ReconNet is:
[0062]
[0063] Among them, S represents the number of steps of expanding the recurrent network during training, is a smooth total variation constraint, and λ2 and λ3 are scalars that balance the effects of the three losses.
[0064] A further improvement of the technical solution of the present invention is that the three-dimensional image reconstruction network module is trained based on a three-dimensional reference image and a light field image synthesized by convolving the three-dimensional reference image with a point spread function calculated based on a wave optics model, and the network input of the three-dimensional image reconstruction network module is the output of the event image reconstruction network module.
[0065] A method for reconstructing three-dimensional images of a light field microscope using an integrated event camera, using the above-mentioned reconstruction system, comprises the following steps:
[0066] S1, the original light field image and its corresponding timestamp simulation package into a series of event data for subsequent network training;
[0067] S2, the event image reconstruction network module reconstructs the continuous event stream into an intensity image and inputs it into the 3D image reconstruction network module;
[0068] S3. The three-dimensional image reconstruction network module reconstructs the event light field image reconstructed by the event image reconstruction network module into a three-dimensional scene.
[0069] Due to the adoption of the above technical solution, the technical advancements achieved by the present invention include:
[0070] 1. This paper proposes a neural network-driven light field microscope 3D image reconstruction system integrated with an event camera. By replacing the traditional CMOS camera of a light field microscope with an event camera, this system bypasses the inherent frame rate limitations of traditional CMOS systems. Furthermore, the event camera employs a novel asynchronous readout architecture. Instead of recording intensity frames at fixed intervals, it dynamically samples light by asynchronously measuring the brightness change of each pixel. This allows sparse event stream encoding to perceive the polarity of changes. This paradigm shift enables the event camera to maintain high signal-to-background ratio data throughput within the limited frame rate compared to traditional frame-based cameras.
[0071] 2. The present invention uses a deep neural network to perform event reconstruction, combines the estimated optical flow and event-based photometric constant to train the neural network, and uses the trained network to reconstruct the spatiotemporal measurement data into a light field event image. The self-supervised method adopted improves the reconstruction performance and significantly improves the reconstructed imaging resolution and signal-to-noise ratio.
[0072] 3. The present invention uses a visual field depth (VCD) neural network to reliably reconstruct rapidly changing three-dimensional scenes from two-dimensional light field event images. The VCD network uses high-resolution three-dimensional scenes and two-dimensional light field images as training data, and iteratively reduces the loss of spatial resolution by integrating rich structural information from the training data. This reconstruction method has uniform spatial resolution, and the resulting three-dimensional image sequence effectively avoids artifacts and axial deformation. It also uses parallel computing to achieve high video rate reconstruction throughput, thereby improving the speed of reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1 1. SSL-E2VID network architecture diagram of a light field microscope 3D image reconstruction system and method integrated with an event camera according to the present invention;
[0074] Figure 2 1. It is a training diagram of VCD-Net of a light field microscope three-dimensional image reconstruction system and method integrated with an event camera of the present invention. DETAILED DESCRIPTION
[0075] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. In the following description, the description of the public structure and technology is omitted to avoid unnecessary confusion of the concept of the present invention.
[0076] The light field microscope three-dimensional image reconstruction system with integrated event camera in the present invention includes a data simulation module that simulates event data for training the network through original light field data, an event image reconstruction network module that can convert a continuous event stream into a series of intensity images, and a three-dimensional image reconstruction network module that reconstructs the image into a high-resolution three-dimensional scene. The three-dimensional image reconstruction network module is trained based on a three-dimensional reference image and a light field image synthesized by convolving the three-dimensional reference image with a point spread function calculated based on a wave optics model, and the network input of the three-dimensional image reconstruction network module is the output of the event image reconstruction network module.
[0077] The data simulation module generates event stream data of image sequences with different movement paths based on a random process, providing training samples for the event image reconstruction network module. The data simulation module uses a dynamic vision sensor voltmeter (DVS-Voltmeter), an event simulator based on a random process, to obtain the event sequence. The specific steps are as follows:
[0078] S1. Use a light field microscope to observe the three-dimensional area of the target, intercept the movement trajectory of the target within a certain time period, obtain two-dimensional light field images of the same three-dimensional area at different times, and determine the timestamp of each image.
[0079] S2, input the two-dimensional light field image with timestamp obtained in S1 into DVS-Voltmeter, detect its spatiotemporal brightness change in the form of asynchronous "event" stream when the intensity occurs, and output the corresponding event sequence; the specific process is: the event camera samples the light when each pixel brightness L changes, and each event with index i can be encoded as e i =(x i ,t i ,p i ), where x i =(h i ,l i ) T Represents the horizontal and vertical coordinates of the pixel, t i Record timestamp, polarity p i ∈{+,-} records the increase or decrease of brightness; each pixel x i In response to the change of the brightness signal L(t), when the brightness change reaches the contrast sensitivity threshold C(C>0), at the pixel point x i and timestamp t i Trigger event e i Therefore, the k-th time window Δt k The brightness increase ΔL k It can be encoded in the event data by pixel-by-pixel accumulation, ΔL kCan be coded as:
[0080]
[0081] Where x represents the set of pixel horizontal and vertical coordinates, and C represents the contrast sensitivity threshold.
[0082] S3. Encapsulate the event sequence and light field image to obtain the training samples required for the optical flow prediction deep network unit and the reconstruction deep neural network training unit.
[0083] S4. Repeat S1, S2, and S3 to capture images in multiple three-dimensional regions with different motion trajectories to obtain a large amount of event training data.
[0084] The above-mentioned event image reconstruction network module includes an optical flow prediction deep network unit for performing optical flow estimation from event data, and an image reconstruction deep network unit for performing three-dimensional reconstruction through learning; at the same time, the optical flow prediction deep network unit and the image reconstruction deep network unit are jointly trained.
[0085] The above optical flow prediction deep network unit uses contrast maximization proxy loss to learn to estimate optical flow by compensating for motion blur in input events. The specific process is:
[0086] On the Lambert surface, constant illumination, Δt k Under the smaller assumption, Equation (1) can be linearized to obtain the event-based luminosity constant, which is expressed as:
[0087]
[0088] Where x represents the set of horizontal and vertical coordinates of pixels, u k (x) represents the pixel in the kth time window Δt k The optical flow vector inside; The luminance signal L k As the optical flow vector u=(u h ,u l ) T The spatial gradient of movement, which triggers the encoding event, δ h L k , δ l L k Respectively represent L k In the horizontal and vertical directions, the deviation gradient is approximately equal to. When , it indicates a vertical relationship and no events will be generated; When , it indicates a parallel relationship, and events will be generated at the highest rate.
[0089] These events encode the timing, location, and polarity of brightness changes. In theory, the event stream contains the entire visual signal in a highly compressed form, and can therefore be decompressed to recover video at any high frame rate and resolution. We select the original 3D dataset and perform data augmentation operations such as cropping, rotation, and mirroring. We then simulate and generate light field images using Broxton's point spread function. We then determine the event duration and use the DVS-Voltmeter to generate a series of 2D light field events, which serve as training data for the network.
[0090] The neural network SSL-E2VID converts a continuous event stream into a series of intensity images in a self-supervised learning (SSL) manner, where two neural networks are jointly trained to perform optical flow estimation and image reconstruction from event data using a contrast maximization proxy loss and an event-based photometric constant, respectively. Figure 1 As shown in Figure 2, the optical flow network (FlowNet) learns to estimate optical flow by compensating for motion blur in input events, and the reconstruction network (ReconNet) learns to perform image reconstruction by learning a photometric constant based on the events.
[0091] The input of the optical flow prediction deep network unit is a voxel grid E with B time windows k , each time window is filled with event streams consecutive non-overlapping partitions of N events, each containing N events; for each partition, each event e i Its polarity p i Assigned to the two closest windows as follows:
[0092]
[0093] κ(a)=max(0,1-|a|) (4)
[0095]
[0096] Where b is the window index, κ(a) is the maximum function operator, a is a scalar value, Normalized event timestamps. This representation adaptively normalizes the time dimension of the input based on the timestamp of each event partition.
[0097] L is reconstructed by learning the photometric constant in equation (2), which depends on the optical flow u in addition to the spatiotemporal derivatives of the brightness itself. This ill-posed problem can be solved using the true optical flow value. FlowNet training performs flow estimation in a self-supervised manner, and using contrast maximization proxy loss for motion compensation can solve the problem of the lack of event camera datasets with accurate real data. When there is spatiotemporal misalignment between events, events generated by the same part of the motion edge will be captured by different timestamps and pixel locations. The optical flow prediction deep network unit retrieves the accurate pixel-by-pixel optical flow u(x i ), the event is propagated to the reference time t in the following way ref :
[0098] x′ i =x i +(t ref -t i )u(x i ) (6)
[0100] x′ i is x i Optical flow propagation at t ref Position after time.
[0101] The deblurring quality metric is judged by the pixel-by-pixel and polarity-by-polarity average timestamps of the image H of the distorted event. The lower the metric, the better the deblurring effect. ref The deblurring quality is evaluated by generating an image with the average (normalized) timestamp at each pixel of polarity p′ after time instant:
[0102]
[0103] in, is the normalized t ref Timestamp, j = {i|p i =p′}, p′∈{+,-}, parameter∈≈0.
[0104] Minimize the sum of the squared images produced by the forward and backward warping events To address the scaling issue during backpropagation, the contrast maximization proxy loss Defined as:
[0105]
[0106] The total loss for training FlowNet is:
[0107]
[0108] in, is the Charbonnier smoothing prior, and λ1 is a scalar that balances the effects of the two losses.
[0109] At the same time, the above-mentioned image reconstruction deep network unit learns to perform image reconstruction through image registration. The specific process is as follows:
[0110] The input of the image reconstruction deep network unit is the same as that of the optical flow prediction deep network unit. The reconstruction problem uses the incremental reference image ΔL and the predicted The difference between the two is used to reconstruct the brightness signal of the input event, is the brightness after reconstruction; ΔL and The two incremental images are warped to a common time frame (indicated by superscript *), yielding ΔL * and
[0111] Defined as:
[0112]
[0113] in, It is defined as x represents the pixel position, u is the optical flow vector, G + and G - The definition is as follows:
[0114]
[0115] Where P is a two-channel image containing the number of pixel positions that receive events during the event warping process. is a deblurred representation of the contrast changes encoded in the input events.
[0116] On the other hand, the event-based luminosity constant is used to reconstruct the image through spatial transformation. The spatial gradient of is warped to the current time instance, that is:
[0117]
[0118] in, In the time window Δt k Reconstructed optical flow within The warping function, Means defined as.
[0119] According to the maximum likelihood method, the photometric reconstruction loss is defined as the square of the L2 norm of the difference between the warped brightness increments:
[0120]
[0121] The contrast threshold C is unknown. To relax the dependence on this parameter, ReconNet uses linear activation in its last layer. The obtained unbounded brightness estimate is first obtained by Transformed to intensity space, exp() represents the exponential function, is in the time window Δt k The reconstruction results within the reconstruction time period; the final reconstruction Expressed as:
[0122]
[0123] Among them, m and M are The 1% and 99% percentiles, is clipped to the range [0,1]. This min / max normalization allows training with any value of C, as long as the ratio of the positive and negative contrast thresholds is similar to the ratio of the evaluation sequences. We assume that most event camera datasets are trained with C + / C - ≈1 record, C + 、C - are the positive and negative signal thresholds, both of which are set to 1, and ≈ indicates approximately equal. The absence of input events can be vaguely understood as the lack of obvious motion or the lack of spatial image gradient. This problem can be solved by introducing an explicit temporal consistency loss based on the frame-based photometric constant formula. The temporal loss is defined as the photometric error between two consecutive reconstructed frames:
[0124]
[0125] ‖‖1 represents the L1 norm.
[0126] The total loss for training ReconNet is:
[0127]
[0128] Among them, S represents the number of steps of expanding the recurrent network during training, is a smooth total variation constraint, and λ2 and λ3 are scalars that balance the effects of the three losses.
[0129] FlowNet uses a self-supervised deep learning network EV-FlowNet for event-based camera optical flow estimation. Input voxel grid E kAfter four cross-row convolutional layers, the output channels start at 64 and double at each layer. The resulting activations pass through two residual blocks and four decoding layers that perform bilinear upsampling and convolution. After each decoding layer, there is a serial skip connection from the corresponding encoder, and another depthwise convolution to produce a lower-scale flow estimate, which is then concatenated with the activation of the previous decoder. EV-FlowNet has no convolution kernels and ReLU (rectified linear unit) activations, except for the flow prediction layer, which uses tanh (hyperbolic tangent function) activation.
[0130] ReconNet uses a recurrent network E2VID to reconstruct intensity images from an event stream. Input voxel grid E k A convolutional head layer and three recurrent encoding layers perform cross-row convolutions, followed by bilinear upsampling through a ConvLSTM (convolutional long short-term memory network), two residual modules, three decoding layers, and finally a convolutional and depthwise convolutional prediction layer. There are tandem skip connections between the symmetric encoding and decoding layers. The number of output channels in the head layer is 32, and the number of output channels doubles after each encoder. The head layer, encoder, and decoder layers use 5*5 kernels, while the remaining layers use 3*3 kernels. All layers use ReLU activations except the final prediction layer, which uses linear activations.
[0131] The 2D images reconstructed by SSL-E2VID can be reconstructed into 3D images by VCD-Net. Compared with traditional reconstruction methods, VCD-Net can image transient biological dynamics with higher spatial resolution, minimal reconstruction artifacts and higher reconstruction throughput. It is robust, versatile and widely applicable. The training process of this method is as follows Figure 2 As shown. It requires using a confocal microscope to obtain high-resolution three-dimensional reference images from static samples or synthetic data, and then convolving them with a point spread function calculated based on a wave optics model to synthesize a light field image. After the data is prepared, the view channel deep neural network is trained on the training dataset to reconstruct the input original light field into a three-dimensional image. By setting an appropriate loss function, the VCD neural network is trained by iteratively minimizing the loss between its intermediate output and the reference image. When the network is well trained, the light field image can be converted into an image stack with a millisecond time scale through the network forward process. Since high-resolution three-dimensional images have strong priors, this method can remove the artifacts caused by traditional deconvolution algorithms and obtain uniform spatial resolution along the depth.
[0132] In the above-mentioned embodiments, the present invention provides a light-field microscopy 3D image reconstruction system and method integrated with an event camera. By replacing the CMOS camera of a conventional light-field microscopy system with an event camera, the inherent frame rate limitations of conventional CMOS systems are circumvented. Furthermore, the event camera of the present invention employs a novel asynchronous readout architecture. Instead of recording intensity frames at fixed time intervals, it dynamically samples light by asynchronously measuring the brightness change of each pixel. This allows sparse event stream encoding to perceive the polarity of changes. This paradigm shift enables the event camera to maintain high signal-to-background ratio data throughput at a limited frame rate compared to conventional frame-based cameras. Furthermore, the present invention employs a visual field depth (VCD) neural network to reliably reconstruct rapidly changing 3D scenes from 2D light-field event images. The VCD network uses high-resolution 3D scenes and 2D light-field images as training data, and iteratively reduces spatial resolution loss by integrating the rich structural information from the training data. This reconstruction method achieves uniform spatial resolution, effectively avoiding artifacts and axial distortion in the resulting 3D image sequence. Parallel computing is used to achieve high video-rate reconstruction throughput, thereby improving reconstruction speed.
[0133] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the concept and scope of the present invention. Any modifications and improvements made to the technical solution of the present invention by a person of ordinary skill in the art without departing from the design concept of the present invention shall fall within the scope of protection of the present invention. The technical content for which protection is sought in the present invention is fully set forth in the claims.
Claims
1. A light field microscope 3D image reconstruction system with an integrated event camera, characterized by: It includes a data simulation module that simulates event data for training the network through original light field data, an event image reconstruction network module that can convert a continuous event stream into an image, and a three-dimensional image reconstruction network module that reconstructs the image into a three-dimensional scene; The data simulation module generates event stream data of image sequences with different movement paths based on a random process, providing training samples for the event image reconstruction network module. The data simulation module uses a random process-based event simulator, namely a dynamic visual sensor voltmeter, to obtain the event sequence. The specific steps are as follows: S1. Observe the three-dimensional area of the target using a light field microscope, intercept the target's movement trajectory within a certain time period, obtain two-dimensional light field images of the same three-dimensional area at different times, and determine the timestamp of each image; S2, inputs the two-dimensional light field image with time stamp obtained in S1 into the dynamic vision sensor voltmeter, detects its spatiotemporal brightness changes in the form of asynchronous "event" stream when the intensity occurs, and outputs the corresponding event sequence; S3. Encapsulate the event sequence and the light field image to obtain the training samples required by the optical flow prediction deep network unit and the reconstruction deep neural network training unit; S4, repeat S1, S2 and S3, capture images in multiple three-dimensional areas with different motion trajectories, and obtain a large amount of event training data; The event image reconstruction network module includes an optical flow prediction deep network unit for performing optical flow estimation from event data, and an image reconstruction deep network unit for performing three-dimensional reconstruction by learning; the optical flow prediction deep network unit and the image reconstruction deep network unit are jointly trained; The three-dimensional image reconstruction network module is trained based on a three-dimensional reference image and a light field image synthesized by convolving the three-dimensional reference image with a point spread function calculated based on a wave optics model, and the network input of the three-dimensional image reconstruction network module is the output of the event image reconstruction network module.
2. The light field microscope 3D image reconstruction system integrated with an event camera according to claim 1, characterized in that: The specific process of step S2 is: The event camera samples light when each pixel brightness L changes, and each event with index i can be encoded as e i =(x i ,t i ,p i ), where x i =(h i ,l i ) T Represents the horizontal and vertical coordinates of the pixel, t i Record timestamp, polarity p i ∈{+,-} records the increase and decrease of brightness; the kth time window Δt k The brightness increase ΔL k Can be coded as: Where x represents the set of pixel horizontal and vertical coordinates, and C represents the contrast sensitivity threshold.
3. The light field microscope 3D image reconstruction system integrated with an event camera according to claim 1, characterized in that: The optical flow prediction deep network unit learns to estimate optical flow by compensating for motion blur in input events using a contrast maximization proxy loss; the image reconstruction deep network unit learns to perform image reconstruction through image registration.
4. The light field microscope 3D image reconstruction system integrated with an event camera according to claim 3, characterized in that: The optical flow prediction deep network unit uses contrast maximization proxy loss to learn to estimate optical flow by compensating for motion blur in input events. The specific process is: On the Lambert surface, constant illumination, Δt k Under smaller assumptions, the event-based luminosity constant is expressed as: Where x represents the set of horizontal and vertical coordinates of pixels, u k (x) represents the pixel in the kth time window Δt k The optical flow vector inside, The luminance signal L k As the optical flow vector u=(u h ,u l ) T The spatial gradient of movement, δ h L k , δ l L k Respectively represent L k The partial gradient in the horizontal and vertical directions, ≈ means approximately equal; The input of the optical flow prediction deep network unit is a voxel grid E with B time windows k , each time window is filled with event streams consecutive non-overlapping partitions of N events, each containing N events; for each partition, each event e i Its polarity p i Assigned to the two closest windows as follows: κ(a)=max(0,1-|a|) (4) Where b is the window index, κ(a) is the maximum function operator, a is a scalar value, To normalize event timestamps; The optical flow prediction deep network unit retrieves the accurate pixel-by-pixel optical flow u(x i ), the event is propagated to the reference time t in the following way ref : x′ i =x i +(t ref -t i )u(x i ) (6) x' i is x i Optical flow propagation at t ref Position after the moment; By bilinear interpolation at each propagation t ref The deblurring quality is evaluated by generating an image with the average timestamp at each pixel of polarity p' after time instant: in, is the normalized t ref Timestamp, j = {i|p i =p'}, p'∈{+,-}, parameter∈≈0; Contrast maximization proxy loss Defined as: The total loss for training FlowNet is: in, is the Charbonnier smoothing prior, and λ1 is a scalar that balances the effects of the two losses.
5. The light field microscope 3D image reconstruction system integrated with an event camera according to claim 3, characterized in that: The image reconstruction deep network unit learns to perform image reconstruction through image registration. The specific process is as follows: The input of the image reconstruction deep network unit is the same as that of the optical flow prediction deep network unit. The reconstruction problem uses the incremental reference image ΔL and the predicted The difference between the two is used to reconstruct the brightness signal of the input event, is the brightness after reconstruction; ΔL and The two incremental images are warped to a common time frame*, yielding ΔL * and Defined as: Among them, x represents the pixel position, u is the optical flow vector, G + and G - The definition is as follows: Where P is a two-channel image containing the number of pixel positions of the image H that received events during the event warping process; The spatial gradient of is warped to the current time instance and is defined as: in, In the time window Δt k Reconstructed optical flow within The warping function of The photometric reconstruction loss is defined as: The resulting unbounded brightness estimate is first obtained by Transformed to intensity space, exp() represents the exponential function, is in the time window Δt k The reconstruction results within the reconstruction time period; the final reconstruction Expressed as: Among them, m and M are The 1% and 99% percentiles, Clipped to the range [0,1]; The temporal loss is defined as the photometric error between two consecutive reconstructed frames: || ||1 represents L1 norm; The total loss for training ReconNet is: Among them, S represents the number of steps of expanding the recurrent network during training, is a smooth total variation constraint, and λ2 and λ3 are scalars that balance the effects of the three losses.
6. A method for 3D image reconstruction using a light field microscope with an integrated event camera, characterized in that: Use of the reconstruction system according to any one of claims 1 to 5 comprises the following steps: S1, the original light field image and its corresponding timestamp simulation package into a series of event data for subsequent network training; S2, the event image reconstruction network module reconstructs the continuous event stream into an intensity image and inputs it into the 3D image reconstruction network module; S3. The three-dimensional image reconstruction network module reconstructs the event light field image reconstructed by the event image reconstruction network module into a three-dimensional scene.
Citation Information
Patent Citations
Simulation system and method of high-dimensional light field event camera
CN115730423A
Event camera image reconstruction method based on brightness constancy
CN118485735A