A Deep Learning-Based Multimodal High-Resolution Light Field Reconstruction Method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-02
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]当前,光场成像主要面临的挑战有:(1)对于时间分辨率在毫秒级别的样本,记录这类信号对于成像系统的帧率有极高要求,一般要求在kHz以上,也即对硬件要求极高;(2)传统的光场成像后处理方法在重建三维图像的过程中存在空间分辨率低、信噪比低和重建伪影的限制,并且需要较高计算迭代重建三维图像,也进一步对相机的量子效率和读出噪声提出了要求,在硬件方面限制了记录帧率的提升
Smart Images

Figure CN117078850B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of 3D reconstruction technology, and in particular to a multimodal high-resolution light field reconstruction method based on deep learning. Background Technology
[0002] Light field imaging, by installing microlens arrays in an optical system, can reconstruct three-dimensional image information with subcellular spatial resolution. It holds great application potential in biological imaging, medical image processing, and computer vision-based 3D reconstruction, and represents a new direction for development. Compared to other 3D imaging technologies, light field imaging can obtain three-dimensional spatial information from a single image, and its imaging rate is limited only by the camera's frame rate. It is more sensitive to sample activity and has lower latency, thus holding significant importance for research in biological imaging, medical image processing, and computer vision-based 3D reconstruction.
[0003] Currently, the main challenges facing light field imaging are: (1) For samples with a temporal resolution at the millisecond level, recording such signals places extremely high demands on the frame rate of the imaging system, generally requiring a frame rate above kHz, which means extremely high hardware requirements; (2) Traditional light field imaging post-processing methods suffer from limitations such as low spatial resolution, low signal-to-noise ratio, and reconstruction artifacts in the process of reconstructing three-dimensional images. Furthermore, they require high computational iterations to reconstruct three-dimensional images, which further places demands on the quantum efficiency and readout noise of the camera, thus limiting the improvement of the recording frame rate in terms of hardware. The three-dimensional volume reconstructed based on light field imaging serves as the main basis for research and analysis in biological imaging, medical image processing, and computer vision three-dimensional reconstruction. How to reconstruct the three-dimensional image of light field imaging with higher temporal resolution, and thus more realistically and accurately reflect the dynamic change process of the sample, is an important problem that urgently needs to be solved. This is of great significance to research in biological imaging, medical image processing, and computer vision three-dimensional reconstruction. Summary of the Invention
[0004] Purpose of the invention: To address the shortcomings of low spatial resolution and reconstruction artifacts in light field imaging systems, this invention provides a multimodal high-resolution light field reconstruction method based on deep learning.
[0005] Technical Solution: To solve the above problems, this invention employs a deep learning-based multimodal high-resolution light field reconstruction method, comprising the following steps:
[0006] Step 1: Acquire the light field image and the 3D image;
[0007] Step 2: Construct a training set of light field and 3D image pairs: Calculate the point spread function based on optical parameters, and project the randomly cropped 3D image into a light field image based on the wave optics model. Use the projected light field image as the dataset for network training.
[0008] Step 3: Construct and train a deep learning network; construct a deep learning network based on conditional generative adversarial networks and perform mapping training;
[0009] Step 4: Use the trained deep learning network to perform 3D volume reconstruction on the captured light field image.
[0010] Furthermore, step 1 specifically involves: capturing images of the sample using a light field camera or light field microscope to obtain light field images; and obtaining three-dimensional images of the sample through optical slicing / simulation.
[0011] Furthermore, step 2, which involves calculating the point spread function based on optical parameters, specifically involves:
[0012] First, calculate the optical impulse response function:
[0013]
[0014] Represents the Fourier transform, ω x and ω y It is the spatial frequency along the x and y directions in the sensor plane; U i(x, p) is the wavefront propagation equation for the natural plane, and its expression is as follows:
[0015]
[0016] Where x is the position on the sensor plane, p is the position of the imaging volume, i is the integral of the voxel over the cube centered at point p, and f obj λ is the focal length of the objective lens, M is the magnification of the objective lens, P(θ) represents the apodization function of the device, J0 represents the zeroth-order Bessel function of the first kind, the variables u and v represent the normalized radial and axial optical coordinates, and α is the numerical aperture half angle.
[0017] Φ(x) is a planar microlens array, and its expression is as follows:
[0018]
[0019] k is the sample wavenumber, f μlens d and d represent the focal length and spacing of the microlenses, respectively, and * denotes convolution;
[0020] Calculate the element h in the point spread function matrix H using the optical impulse response function. ij :
[0021]
[0022] α j Let β be the area of pixel j. j Let i be the volume of voxel i.
[0023] Furthermore, the projection of the randomly cropped 3D image into a light field image based on the wave optics model, as described in step 2, is obtained using the following formula:
[0024]
[0025] Where f is the light field image, N p N is the number of pixels, g is the volume of the image, and N is the image size. v H is the number of voxels, and H is the point spread function matrix calculated in step 2.
[0026] Furthermore, the deep learning network in step 3 includes a generator and a discriminator.
[0027] Furthermore, the generator includes an encoder module, a decoder module, and jump connections.
[0028] Furthermore, the discriminator includes convolutional layers and linear layers.
[0029] Furthermore, the training described in step 3 specifically involves: using the dataset from step 2, taking the light field image obtained in step 1 as the generator input, the 3D image as the generator output, and the discriminator evaluating the generator output and feeding it back to the generator.
[0030] Furthermore, the three-dimensional volume reconstruction in step 4 specifically involves: during the inference phase, loading the network weights obtained in step 3, and reconstructing the captured light field image through a trained deep learning network to obtain a three-dimensional image.
[0031] Furthermore, the network weights are the network weights after training, when the loss function is reduced to a preset value.
[0032] Beneficial effects: Compared with the prior art, the significant advantages of this invention are (1) by constructing and training a deep learning network, the spatial resolution of the light field imaging system is improved, and the reconstruction artifact problem is solved; (2) the real-time dynamic change details of the captured sample are processed with higher temporal resolution. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of the multimodal high-resolution light field reconstruction method based on degree learning of the present invention;
[0034] Figure 2 This is a schematic diagram of the neural network training of the present invention;
[0035] Figure 3 This is a schematic diagram of the inference stage network of the present invention reconstructing the light field image. Detailed Implementation
[0036] like Figure 1 As shown in the figure, the specific steps of the deep learning-based multimodal high-resolution light field reconstruction method in this embodiment are as follows:
[0037] Step 1: Acquire light field images and 3D images. Light field images are obtained by photographing the sample using a light field camera or light field microscope. These light field images have low angular-spatial resolution. 3D images of the sample are obtained through optical slicing / simulation. These 3D images have high spatial resolution but low temporal resolution.
[0038] Step 2: Construct a training set of light field and 3D image pairs. First, calculate the point spread function based on the optical parameters. The specific steps are as follows:
[0039] Calculate the optical impulse response function:
[0040]
[0041] Represents the Fourier transform, ω x and w y It is the spatial frequency along the x and y directions in the sensor plane; U i (x, p) is the wavefront propagation equation for the natural plane, and its expression is as follows:
[0042]
[0043] Where x is the position on the sensor plane, p is the position of the imaging volume, i is the integral of the voxel over the cube centered at point p, and f obj λ is the focal length of the objective lens, M is the magnification of the objective lens, P(θ) represents the apodization function of the device, J0 represents the zero-order Bessel function of the first kind, the variables u and υ represent the normalized radial and axial optical coordinates, and α is the numerical aperture half angle.
[0044] Φ(x) is a planar microlens array, and its expression is as follows:
[0045]
[0046] k is the sample wavenumber, f μlens d and d represent the focal length and spacing of the microlenses, respectively, and * denotes convolution.
[0047] Calculate the element h in the point spread function matrix H using the optical impulse response function. ij :
[0048]
[0049] α j Let β be the area of pixel j. j Let i be the volume of voxel i.
[0050] After calculating the point spread function, the randomly cropped 3D image is projected into a light field image according to the wave optics model, specifically obtained through the following formula:
[0051]
[0052] Where f is the light field image, N p N is the number of pixels, g is the volume of the image, and N is the image size. v H is the number of voxels, and H is the point spread function matrix calculated in step 2.
[0053] The light field image calculated above is used as the dataset for network training.
[0054] Step 3: Construct and train the deep learning network. For example... Figure 2 As shown, the deep learning network in this embodiment includes a generator and a discriminator; the generator includes an encoder module, a decoder module, and skip connections; the discriminator includes convolutional layers and linear layers. The specific training steps are as follows: using the dataset from step 2, the low-angle-spatial resolution light field image obtained in step 1 is used as the generator input, and the high spatial resolution 3D image is used as the generator output. The discriminator evaluates the generator output and feeds it back to the generator, further optimizing the generator's results. The trained deep learning network achieves millisecond-level inference speed, thereby improving the temporal resolution of the 3D image. At this point, the output 3D image has both high spatial and temporal resolution.
[0055] Step 4: Use the trained network to perform 3D volume reconstruction on the captured light field image. For example... Figure 3 As shown, in the inference phase, the network weights obtained in step 3 are loaded. These network weights are network weights after training and the loss function is reduced to a preset value. The captured light field image is reconstructed through the trained deep learning network to obtain a three-dimensional image with high spatial and temporal resolution.
Claims
1. A multimodal high-resolution light field reconstruction method based on deep learning, characterized in that, Includes the following steps: Step 1: Acquire the light field image and the 3D image; Step 2: Construct a training set of light field and 3D image pairs: Calculate the point spread function based on optical parameters, and project the randomly cropped 3D image into a light field image based on the wave optics model. Use the projected light field image as the dataset for network training. The calculation of the point spread function based on optical parameters specifically involves: First, calculate the optical impulse response function: Indicates Fourier transform, and It is the spatial frequency along the x and y directions in the sensor plane; The wavefront propagation equation for the natural plane is expressed as follows: in, It is the position on the sensor plane. It is the location of the imaging volume. For voxels at points Integrate over the product of the cube centered at the center. It is the focal length of the objective lens. It is the emission wavelength of the sample. It is the magnification of the objective lens. The apodization function represents the device. Represents the zeroth-order Bessel function of the first kind, variable and Represents the normalized radial and axial optical coordinates. The numerical aperture half-angle; For a planar microlens array, its expression is as follows: For the sample wavenumber, and For the focal length and spacing of the microlenses, Represents convolution; Calculate the point spread function matrix using the optical impulse response function. elements in : For pixels area, voxels Volume; The projection of the randomly cropped 3D image into a light field image based on the wave optics model is obtained by the following formula: in For light field images, where g is the number of pixels and g is the image volume. The number of voxels The point spread function matrix calculated in step 2; Step 3: Construct and train a deep learning network; construct a deep learning network based on conditional generative adversarial networks and perform mapping training; Step 4: Use the trained deep learning network to perform 3D volume reconstruction on the captured light field image.
2. The deep learning-based multimodal high-resolution light field reconstruction method as described in claim 1, characterized in that, Step 1 specifically involves: capturing a light field image of the sample using a light field camera or light field microscope; and obtaining a three-dimensional image of the sample through optical slicing / simulation.
3. The deep learning-based multimodal high-resolution light field reconstruction method as described in claim 1, characterized in that, The deep learning network in step 3 includes a generator and a discriminator.
4. The deep learning-based multimodal high-resolution light field reconstruction method as described in claim 3, characterized in that, The generator includes an encoder module, a decoder module, and jump connections.
5. The deep learning-based multimodal high-resolution light field reconstruction method as described in claim 4, characterized in that, The discriminator includes convolutional layers and linear layers.
6. The deep learning-based multimodal high-resolution light field reconstruction method as described in claim 5, characterized in that, The training described in step 3 is as follows: using the dataset from step 2, the light field image obtained in step 1 is used as the generator input, the 3D image is used as the generator output, and the discriminator evaluates the generator output and feeds it back to the generator.
7. The deep learning-based multimodal high-resolution light field reconstruction method as described in claim 6, characterized in that, The three-dimensional volume reconstruction in step 4 specifically involves: during the inference phase, loading the network weights obtained in step 3, and reconstructing the captured light field image through a trained deep learning network to obtain a three-dimensional image.
8. The deep learning-based multimodal high-resolution light field reconstruction method as described in claim 7, characterized in that, The network weights are the network weights after training, when the loss function is reduced to a preset value.