A holographic camera based on a physics model driven network and a liquid lens
By using a holographic camera based on a physical model-driven network and a liquid lens, the problems of rapid acquisition of real 3D scenes and reconstruction of real depth in existing technologies have been solved, achieving high-quality holographic 3D reconstruction effects.
Patent Information
- Application Number
- CN202311728574.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2043-12-15
AI Technical Summary
Existing holographic 3D display technologies struggle to rapidly acquire and reconstruct realistic 3D scenes, and existing deep neural network methods cannot effectively reconstruct the true depth of 3D scenes.
A holographic camera based on a physical model-driven network and a liquid lens is used to quickly acquire multi-focal depth information through the liquid camera, and then the physical model-driven network is used to calculate pure phase holograms to reconstruct a holographic image with real depth.
It enables rapid acquisition of real 3D scenes and high-quality holographic 3D reconstruction, and can reconstruct holographic images with real depth.
Smart Images

Figure CN117930614B_ABST
Abstract
Description
I. TECHNICAL FIELD
[0001] The present application relates to holographic display technology, more particularly, the present application relates to a holographic camera based on a physical model driven network and a liquid lens. II. BACKGROUND
[0002] Holographic 3D display technology can reconstruct the complete wavefront information of a 3D scene, and is therefore considered as one of the most ideal naked-eye 3D display technologies. However, for the holographic reconstruction of a real 3D scene, the existing holographic 3D display technology has the problems of slow hologram acquisition speed and low quality of reconstructed image, which limits the further development of holographic 3D display technology. In order to acquire the hologram of a real 3D scene, researchers use a camera to directly acquire the interference fringes of a real 3D scene based on digital holographic technology. However, this method requires the use of complex optical scanning devices and coherent light active illumination, making it difficult to realize real-time calculation of the hologram. Some researchers use the fast zooming capability of an electrowetting liquid lens to realize fast acquisition of a real 3D scene, however, due to the small aperture of the electrowetting liquid lens, the quality of the acquired real 3D scene is limited. In terms of high-fidelity hologram generation, in recent years, with the rapid development of deep learning technology, domestic and foreign researchers have used deep neural networks to realize fast calculation of high-fidelity holograms through data-driven or model-driven methods, effectively improving the quality of the reconstructed image of holographic 3D display. However, the existing hologram calculation methods based on deep neural networks usually estimate the depth value of a 3D scene, and cannot reconstruct the real depth of a 3D scene. At present, the fast acquisition and real depth reconstruction of a hologram of a real 3D scene are difficult problems that have not been solved in the field of holographic display. III. SUMMARY
[0003] The present application proposes a holographic camera based on a physical model driven network and a liquid lens. As shown in FIG. 1, the holographic camera is composed of a hardware part and a software part, wherein the hardware part is a liquid camera based on a liquid lens, which is used to quickly acquire the multi-focal depth information of a real 3D scene, and the software part is a physical model driven network, which is used to quickly calculate the pure phase hologram of a real 3D scene, and the pure phase hologram can reconstruct a holographic reconstructed image with real depth. Figure 1 As shown in FIG. 2, the liquid camera is composed of a camera shell, an electrode wire, a liquid lens, a solid lens group and a CMOS target surface. The liquid lens is used to adjust the focal length, the solid lens group is used to provide the main optical power, and the electrode wire is connected to the electrodes of the liquid lens. The camera shell is used to fix the liquid lens, the solid lens group and the electrode wire, and the CMOS target surface is used to receive the information of a real 3D scene. The total optical power Φ of the liquid camera is expressed as:
[0004] Figure 2 As shown in FIG. 3, the liquid camera is composed of a camera shell, an electrode wire, a liquid lens, a solid lens group and a CMOS target surface. The liquid lens is used to adjust the focal length, the solid lens group is used to provide the main optical power, and the electrode wire is connected to the electrodes of the liquid lens. The camera shell is used to fix the liquid lens, the solid lens group and the electrode wire, and the CMOS target surface is used to receive the information of a real 3D scene. The total optical power Φ of the liquid camera is expressed as:
[0005] Φ = Φs +Φ l -d sl Φ s Φ l (1)
[0006] Where, Φ s and Φ l The optical power of the solid lens group and the liquid lens group are respectively; d sl This is the distance between the optical principal plane of the solid lens group and the optical principal plane of the liquid lens.
[0007] As attached Figure 2 As shown in (b), in the liquid camera, the liquid lens consists of a lens housing, two glass substrates, a hydrophobic layer, a dielectric layer, an upper electrode, a lower electrode, a conductive liquid, and an insulating liquid. A sealed space is formed between the upper electrode, the lower electrode, and the two glass substrates, filled with both insulating and conductive liquids, forming a naturally curved liquid-liquid interface. The inner surface of the upper electrode is sequentially coated with a dielectric layer and a hydrophobic layer. The lower electrode is directly connected to the conductive liquid and is insulated from the upper electrode. By changing the voltage of the upper and lower electrodes of the liquid lens, the liquid camera can quickly acquire multi-focal depth information of a real 3D scene without requiring additional mechanical movement.
[0008] As attached Figure 3 As shown, in a holographic camera, the process of using a physical model to drive the network to calculate and train a pure phase hologram of a real 3D scene consists of four steps: First, using depth calculation and scene fusion methods, the multi-focal depth information of the real 3D scene captured by the liquid camera is processed into a depth map and a full-focus image of the real 3D scene. Then, the depth map and full-focus image of the real 3D scene are input into encoder-decoder network 1. Second, using the selected physical model, the amplitude and phase information of the real 3D scene generated by encoder-decoder network 1 is propagated forward by a distance z0, thus obtaining the amplitude and phase information on the spatial light modulator. Then, the amplitude and phase information on the spatial light modulator are input into encoder-decoder network 2, thereby quickly generating a pure phase hologram of the real 3D scene. Third, based on the grayscale values of the depth map of the real 3D scene and the selected number of layers i, the depth map is divided into grayscale intervals to obtain i binary masks. Then, using the selected physical model, the pure phase hologram of the real 3D scene is propagated backward by z0. idistance to obtain the i-th layer holographic reconstruction image, then, calculate the dot product between the i-th binary mask and the i-th layer holographic reconstruction image to obtain the reconstruction image of the i-th layer band mask; the fourth step is that the all-in-focus image of the real 3D scene is input into the unsharpened mask filter to enhance the high-frequency information, so as to obtain an enhanced 3D scene, then, the reconstruction images of the i-th layer band mask are added pixel by pixel to obtain an all-in-focus reconstruction image, finally, a loss function between the enhanced 3D scene and the all-in-focus reconstruction image is calculated, and the encoder-decoder network 1 and the encoder-decoder network 2 in the physical model driven network are optimized based on the loss function. The pure phase hologram of the new real 3D scene is recalculated by using the optimized physical model driven network, and the above steps are repeated until the value of the loss function is less than a set threshold, at this time, the pure phase hologram of the real 3D scene output by the encoder-decoder network 2 is the final hologram. The final hologram is loaded onto the spatial light modulator, and when laser irradiates the spatial light modulator, the holographic reconstruction image of the real 3D scene with real depth information is reconstructed.
[0009] In step one, the flow chart of the depth calculation and scene fusion of the physical model driven network of the holographic camera is as shown in the accompanying Figure 4 For the multi-focal plane depth information of the real 3D scene collected by the liquid camera, first, the sharpness evaluation value of each focal plane is calculated by using the Laplacian operator, so as to determine the position of the focal plane where the in-focus image is located. Then, the defocus estimation is performed on the in-focus image, and the binaryzation and morphological dilation processing are performed, so as to obtain the mask information of the in-focus image. Finally, the point product calculation is performed on the multi-focal plane depth information of the real 3D scene and the mask information of the in-focus image, so as to obtain the in-focus image of the multi-focal plane. All the in-focus images of the multi-focal plane are added pixel by pixel to obtain the all-in-focus image of the real 3D scene. At the same time, according to the actual depth when the image is focused by the liquid camera, the depth map of the real 3D scene is obtained by the method of gray value assignment to the mask of the in-focus image.
[0010] The structural schematic diagram of the encoder-decoder network 1 and the encoder-decoder network 2 in the physical model driven network is as shown in the accompanying Figure 5 (a) and the accompanying Figure 5(b) shown. The encoder-decoder network 1 comprises four down-sampling blocks, four perception field blocks, four up-sampling blocks, three skip connection layers, a hyperbolic tangent activation function layer and a normalization layer. The encoder-decoder network 2 comprises four down-sampling blocks, four perception field blocks, four up-sampling blocks, three skip connection layers and a hyperbolic tangent activation function layer. The down-sampling block is used to increase the channel number of the input image and reduce the height and width of the image, so as to fully extract the feature information of the input image. The up-sampling block is used to reduce the channel number of the input image and increase the height and width of the image, so as to restore the input image to the matrix size at the time of input. The perception field block is used to fully perceive the features of the input image, so that the network learns more information of the input image. The skip connection layer is used between the down-sampling block and the up-sampling block, so as to avoid gradient disappearance and network degradation, and ensure that the physical model driven network can be effectively trained. The normalization layer is used to limit the pixel value of the output image to the range of [-1, 1], and the hyperbolic tangent activation function layer is used to limit the pixel value of the output hologram to the range of [-π, π].
[0011] As shown in FIG. 3, the height and width of the image are halved after the input image passes through each down-sampling block. Figure 5 As shown in FIG. 4, each down-sampling block comprises three sequential networks, a skip connection layer and two 1x1 convolution layers. After the image enters the down-sampling block, the image information is merged after passing through the sequential network 1, the sequential network 2 and a 1x1 convolution layer, respectively. The merged image information is merged again after passing through the sequential network 3 and the skip connection layer, and the image information directly passing through another 1x1 convolution layer is merged, and finally outputted. As shown in FIG. 5, the height and width of the image are doubled after the input image passes through each up-sampling block. Figure 5 As shown in FIG. 6, each up-sampling block comprises three sequential networks, a skip connection layer and two 2x2 transpose convolution layers. After the image enters the up-sampling block, the image information is merged after passing through the sequential network 4, the sequential network 5 and the 2x2 transpose convolution layer, respectively. The merged image information is merged again after passing through the sequential network 6 and the skip connection layer, and the image information directly passing through the 2x2 transpose convolution layer is merged, and finally outputted. As shown in FIG. 7, the height and width of the image are unchanged after the input image passes through each perception field block. Figure 5 As shown in FIG. 8, each perception field block comprises four branch networks, two 1x1 convolution layers and a channel connection. After the image enters the perception field block, the image information is merged after passing through the branch network 1, the branch network 2, the branch network 3 and the branch network 4, respectively, and passing through the channel connection. The merged image information is merged again after passing through a 1x1 convolution layer, and the image information directly passing through another 1x1 convolution layer is merged, and finally outputted.
[0012] The specific structures of the sequential network 1, the sequential network 2 and the sequential network 3 are shown in FIG. 9, FIG. 10 and FIG. 11, respectively. Figure 6(a)-(c) are shown. Sequence network 1 is composed of a 3x3 convolutional layer with a step size of 1, a 3x3 convolutional layer with a step size of 2, and two parameter linear rectifier function layers; sequence network 2 is composed of two 3x3 convolutional layers with a step size of 1, a 3x3 convolutional layer with a step size of 2, and three parameter linear rectifier function layers; and sequence network 3 is composed of two 3x3 convolutional layers with a step size of 1 and two parameter linear rectifier function layers. The specific structures of sequence network 4, sequence network 5, and sequence network 6 are shown in Figs. 4(a)-(c), respectively. Figure 6 (d)-(f) are shown. Sequence network 4 is composed of a 3x3 convolutional layer with a step size of 1, a 3x3 convolutional layer with a step size of 2, a pixel recombination layer, a bicubic interpolation layer, and two parameter linear rectifier function layers; sequence network 5 is composed of two 3x3 transposed convolutional layers with a step size of 1, a 3x3 transposed convolutional layer with a step size of 2, and two parameter linear rectifier function layers; and sequence network 6 is composed of two 3x3 transposed convolutional layers with a step size of 1 and two parameter linear rectifier function layers. The specific structures of branch network 1, branch network 2, branch network 3, and branch network 4 are shown in Figs. 4(d)-(f), respectively. Figure 6 (g)-(j) are shown. Branch network 1 is composed of a 1x1 convolutional layer with a step size of 1, a 3x3 convolutional layer with a dilation rate of 1 and a step size of 1, and a parameter linear rectifier function layer; branch network 2 is composed of two 1x1 convolutional layers with a step size of 1, a 3x3 convolutional layer with a dilation rate of 3 and a step size of 1, and two parameter linear rectifier function layers; branch network 3 is composed of a 1x1 convolutional layer with a step size of 1, a 3x3 convolutional layer with a step size of 1, a 3x3 convolutional layer with a dilation rate of 3 and a step size of 1, and two parameter linear rectifier function layers. Branch network 4 is composed of two 1x1 convolutional layers with a step size of 1, a 3x3 convolutional layer with a step size of 1, a 3x3 convolutional layer with a dilation rate of 5 and a step size of 1, and three parameter linear rectifier function layers. The parameter linear rectifier function layer is used to solve the overfitting problem and make the training of the model more stable.
[0013] Preferably, in the liquid camera, the conductive liquid of the liquid lens is selected from a 1,3-propanediol solvent doped with tetrabutylammonium chloride, the insulating liquid is selected from an isomeric alkane solvent doped with 1-bromo-4-ethylbenzene, and the electrode material of the liquid lens is selected from aluminum.
[0014] Preferably, the calculation method of the sharpness evaluation value C in the physical model driving network of the holographic camera uses the Laplace operator, and the expression of the sharpness evaluation value C is:
[0015]
[0016] wherein F(p, q) is an image filtered using a Gaussian filter with a standard deviation of 1 on the input image; l is the number of pixels in the horizontal and vertical directions of the focused image; L 3×3is a Laplacian operator with size 3x3.
[0017] Preferably, when performing depth calculation and defocus estimation in scene fusion, it is necessary to first calculate the gradient amplitude ratio R of the edge position of the image, and the expression is:
[0018]
[0019] wherein, and are the gradients along the x and y directions after Gaussian filtering of the image with a standard deviation of σ1. and are the gradients along the x and y directions after Gaussian filtering of the image with a standard deviation of σ2. The size of R reflects the degree of defocus of the object, and the expression of the defocus estimation value σ is:
[0020]
[0021] After the defocus estimation value of the image is calculated, a sparse defocus estimation map d' is obtained. Then, the sparse defocus estimation map d' is interpolated by using the Laplacian matrix interpolation method to obtain a complete defocus estimation map d, and the expression is:
[0022] d = λDd' (L + λD) -1 (5)
[0023] wherein, L is a Laplacian matrix, D is a diagonal matrix, and λ is a constraint constant.
[0024] Preferably, in the physical model driven network, the distance z i backpropagated is expressed as:
[0025] z i = z0+ Δd × i (6)
[0026] wherein, i is the selected number of layers; z0is the reference recording distance; and Δd is the depth interval of the holographic reconstruction image of the real 3D scene. The physical model in the physical model driven network of the holographic camera is a band-limited angular spectrum diffraction model, and in the training process of the physical model driven network, the loss function is a combination of perceptual loss, multi-scale structural similarity loss, total variation loss and mean square error loss. The expression of the loss function Loss is:
[0027]
[0028] wherein, ɑ, β and η are the coefficients of perceptual loss, multi-scale structural similarity loss and total variation loss, respectively, and Loss PE , Loss MS and Loss TVrespectively are perceptual loss, multi-scale structural similarity loss and total variation loss, denotes the arithmetic mean of the amplitude information on the spatial light modulator, O SLM denotes the amplitude information on the spatial light modulator. The generated image is made visually closer to the image recognized by the human eye using the loss function composed of perceptual loss, multi-scale structural similarity loss, total variation loss and mean square error loss, thereby improving the quality of the holographic reconstruction image. IV. BRIEF DESCRIPTION OF DRAWINGS
[0029] Figure 1 is a structural schematic diagram of a holographic camera of the present application. Figure 1 Figure 1 is a structural schematic diagram of a holographic camera of the present application.
[0030] Figure 2 is a structural schematic diagram of a liquid camera in the holographic camera of the present application. Figure 2 Figure 2 is a structural schematic diagram of a liquid camera in the holographic camera of the present application. Figure 2 (a) is a general structural schematic diagram of the liquid camera; and Figure 2 (b) is a structural schematic diagram of a liquid lens in the liquid camera.
[0031] Figure 3 is a flow chart of the present application using a physical model to drive network calculation and training of a pure phase hologram of a real 3D scene. Figure 3
[0032] Figure 4 is a flow chart of the present application using a physical model to drive network for depth calculation and scene fusion. Figure 4
[0033] Figure 5 is a structural schematic diagram of an encoder-decoder network in the physical model driven network of the present application. Figure 6 is a structural schematic diagram of an encoder-decoder network 1. Figure 5 (a) is a structural schematic diagram of an encoder-decoder network 1; and Figure 5 (b) is a structural schematic diagram of an encoder-decoder network 2; and Figure 5 (c) is a structural schematic diagram of a down-sampling block; and Figure 5 (d) is a structural schematic diagram of an up-sampling block; and Figure 5 (e) is a structural schematic diagram of a perceptual field block. Figure 5
[0034] Figure 7 is a structural schematic diagram of a sequence network and a branch network of the present application. Figure 6 (a) is a structural schematic diagram of a sequence network 1; and Figure 6 (b) is a structural schematic diagram of a sequence network 2; and Figure 6 (c) is a structural schematic diagram of a sequence network 3; and Figure 6 (d) is a structural schematic diagram of a sequence network 4; and Figure 6 (e) is a structural schematic diagram of a sequence network 5; and Figure 6 (f) is a structural schematic diagram of a sequence network 6; and Figure 6 (f) is a structural schematic diagram of a sequence network 6; and Figure 6 (g) is a schematic diagram of the structure of branched network 1; Fig. Figure 6 (h) is a schematic diagram of the structure of branched network 2; Fig. Figure 6 (i) is a schematic diagram of the structure of branched network 3; Fig. Figure 6 (j) is a schematic diagram of the structure of branched network 4.
[0035] Fig. 1 is a schematic diagram of the structure of a holographic camera according to the present application; Fig. Figure 7 Fig. 2 is a holographic 3D reconstruction effect drawing of a real 3D scene collected by the holographic camera according to the present application under real depth; Fig. Figure 7 (a) is a full-focus image of a real 3D scene; Fig. Figure 7 (b) is a depth map of a real 3D scene; Fig. Figure 7 (c) - (f) are holographic reconstruction images of four signboards in a real 3D scene respectively focused.
[0036] In the above figures, the figure numbers are as follows:
[0037] (1) camera shell, (2) electrode wire, (3) liquid lens, (4) solid lens group, (5) CMOS target surface, (6) lens shell, (7) glass substrate, (8) hydrophobic layer, (9) dielectric layer, (10) upper electrode, (11) lower electrode, (12) conductive liquid, (13) insulating liquid.
[0038] It should be understood that the above figures are only schematic and not drawn to scale. V. DETAILED DESCRIPTION
[0039] The following will describe in detail an embodiment of a holographic camera based on a physical model driving network and a liquid lens according to the present application, and further describe the present application. It is necessary to point out here that the following embodiment is only used to further illustrate the present application, and cannot be understood as limiting the protection scope of the present application. Those skilled in the art can make some non-essential improvements and adjustments to the present application according to the above description of the present application, which still belongs to the protection scope of the present application.
[0040] In the experiment, the solid lens group of the solid lens in the liquid camera has a main focal power of 2.8, a focal length of 12 mm, a CMOS target surface area of 1 / 1.8", a pixel size of 2.4 μm, a dielectric layer thickness of 3 μm on the inner surface of the upper electrode, and a hydrophobic layer thickness of 100 nm. In the physical model driven network, the laser wavelengths used to train the blue, green and red holograms are 473 nm, 532 nm and 671 nm respectively, the recorded real 3D scene is "no parking", "sidewalk", "intersection" and "traffic light" four signs, the actual depth interval of the four signs is 50 mm, the resolution is 1849x2773, the reference recording distance is 300 mm, and the ɑ, β and η in the loss function are set to 0.05, 1 and 1x10 -5 In the optical reconstruction experiment, the spatial light modulator used has a resolution of 1920x1080 and a pixel pitch of 6.4 μm, and the refresh rate is 60 Hz. After training the physical model driven network, the training loss value of the red channel is 0.6157, and the verification loss value is 0.5447. The training loss value of the green channel is 0.5755, and the verification loss value is 0.5037. The training loss value of the blue channel is 0.5119, and the verification loss value is 0.4244. The resolution of the output pure phase hologram is 1920x1072.
[0041] The all-focus image and the depth map of the real 3D scene processed by the physical model driven network of the holographic camera are shown in Figs. Figure 7 (a) and 7(b). In order to verify that the holographic camera proposed in the application can realize the holographic 3D reconstruction of the real 3D scene at the real depth, the output pure phase hologram is loaded onto the spatial light modulator, and high-quality holographic reconstruction effects are shot after time sequence illumination by color lasers. The holographic reconstruction images of the "no parking", "sidewalk", "intersection" and "traffic light" four signs are shown in Figs. Figure 7 (c)-(f). The holographic camera proposed in the application realizes high-quality holographic 3D reconstruction of the real 3D scene at the real depth.
Claims
1. A holographic camera based on a physics model driven network and a liquid lens, characterized in that, The holographic camera is composed of a hardware part and a software part, wherein the hardware part is a liquid camera based on a liquid lens, which is used for quickly acquiring multi-focal depth information of a real 3D scene, and the software part is a physical model driven network, which is used for quickly calculating a pure phase hologram of the real 3D scene, and the pure phase hologram is used for reconstructing a holographic reconstruction image with real depth; The liquid camera is composed of a camera shell, electrode wires, a liquid lens, a solid lens group and a CMOS target surface, wherein the liquid lens is used for adjusting focal length, the solid lens group is used for providing main optical power, the electrode wires are connected to electrodes of the liquid lens, the camera shell is used for fixing the liquid lens, the solid lens group and the electrode wires, and the CMOS target surface is used for receiving information of a real 3D scene, and total optical power Φ of the liquid camera is expressed as: Φ = Φ s + Φ l - d sl Φ s Φ l wherein Φ s and Φ l are the optical powers of the solid lens group and the liquid lens, respectively, and d sl is the distance between the optical principal plane of the solid lens group and the optical principal plane of the liquid lens; In the liquid camera, the liquid lens is composed of a lens shell, two glass substrates, a hydrophobic layer, a dielectric layer, an upper electrode, a lower electrode, a conductive liquid and an insulating liquid, the upper electrode, the lower electrode and the two glass substrates form a sealed space, the space is filled with the insulating liquid and the conductive liquid, a natural curved liquid-liquid interface is formed between the two liquids, an inner surface of the upper electrode is coated with the dielectric layer and the hydrophobic layer in sequence, the lower electrode is directly connected with the conductive liquid, and the lower electrode is insulated from the upper electrode, when the voltage of the upper electrode and the lower electrode of the liquid lens is changed, the liquid camera can quickly acquire multi-focal depth information of a real 3D scene without additional mechanical movement; In a holographic camera, the process of using a physical model to drive network computation and training of pure phase holograms of a real 3D scene consists of four steps: First, using depth calculation and scene fusion methods, the multi-focal depth information of the real 3D scene captured by the liquid camera is processed into a depth map and a full-focus image of the real 3D scene. Then, the depth map and full-focus image of the real 3D scene are input into encoder-decoder network 1. Second, using the selected physical model, the amplitude and phase information of the real 3D scene generated by encoder-decoder network 1 is propagated forward by a distance z0, thus obtaining the amplitude and phase information on the spatial light modulator. Then, the amplitude and phase information on the spatial light modulator are input into encoder-decoder network 2, thereby quickly generating a pure phase hologram of the real 3D scene. Third, the depth map is divided into grayscale intervals according to the grayscale values of the depth map of the real 3D scene and the selected number of layers i, thus obtaining i binary masks. Then, using the selected physical model, the pure phase hologram of the real 3D scene is propagated backward by z0. i The distance is calculated to obtain an i-layer holographic reconstruction image. Then, the dot product between i binary masks and the i-layer holographic reconstruction image is calculated to obtain an i-layer masked reconstruction image. In the fourth step, the fully focused image of the real 3D scene is input into an unsharpened mask filter to enhance high-frequency information, thereby obtaining an enhanced 3D scene. Then, the i-layer masked reconstruction image is added pixel by pixel to obtain a fully focused reconstruction image. Finally, the loss function between the enhanced 3D scene and the fully focused reconstruction image is calculated, and the encoder and decoder network 1 and encoder and decoder network 2 in the physical model driving network are optimized based on the loss function. The optimized physical model driving network is used to recalculate the pure phase hologram of the new real 3D scene, and the above steps are repeated until the value of the loss function is less than the set threshold. At this time, the pure phase hologram of the real 3D scene output by encoder and decoder network 2 is the final hologram. The final hologram is loaded onto the spatial light modulator. When the laser illuminates the spatial light modulator, a holographic reconstruction image of the real 3D scene with real depth information is reconstructed.
2. The holographic camera based on physical model driven network and liquid lens according to claim 1, characterized in that, When the physical model driven network of the holographic camera performs depth calculation and scene fusion, for the multi-focal depth information of a real 3D scene collected by the liquid camera, first, a Laplacian operator is used to calculate a sharpness evaluation value of each focal plane, so as to determine the position of the focal plane of the in-focus image, then, defocus estimation is performed on the in-focus image, and binarization and morphological dilation processing are performed, so as to obtain mask information of the in-focus image, finally, point product calculation is performed on the multi-focal depth information of the real 3D scene and the mask information of the in-focus image, so as to obtain in-focus images of the multi-focal planes, all the in-focus images of the multi-focal planes are added pixel by pixel, so as to obtain a full-focus image of the real 3D scene, and simultaneously, according to the actual depth when the image is focused by the liquid camera, a depth map of the real 3D scene is obtained by a method of gray value assignment to the mask of the in-focus image. The encoder-decoder network 1 comprises four down-sampling blocks, four perception field blocks, four up-sampling blocks, three skip connection layers, a hyperbolic tangent activation function layer and a normalization layer, the encoder-decoder network 2 comprises four down-sampling blocks, four perception field blocks, four up-sampling blocks, three skip connection layers and a hyperbolic tangent activation function layer, the down-sampling block is used to increase the channel number of the input image and reduce the height and width of the image, so as to fully extract the feature information of the input image, the up-sampling block is used to reduce the channel number of the input image and increase the height and width of the image, so as to restore the input image to the matrix size at the time of input, the perception field block is used to fully perceive the features of the input image, so that the network learns more information of the input image, the skip connection layer is used between the down-sampling block and the up-sampling block, so as to avoid gradient disappearance and network degradation, ensure that the physical model driven network can be effectively trained, and the normalization layer is used to limit the pixel value of the output image in the range of [-1, 1], and the hyperbolic tangent activation function layer is used to limit the pixel value of the output hologram in the range of [-π, π]; The height and width of the input image are halved after each down-sampling block, each down-sampling block comprises three sequence networks, a skip connection layer and two 1*1 convolution layers, after the image enters the down-sampling block, the image information is combined after passing through sequence network 1, sequence network 2 and a 1*1 convolution layer, the combined image information is combined after passing through sequence network 3 and the skip connection layer, and then combined with the image information after the input image directly passing through another 1*1 convolution layer, and finally output; the height and width of the image are doubled after each up-sampling block, each up-sampling block comprises three sequence networks, a skip connection layer and two 2*2 transpose convolution layers, after the image enters the up-sampling block, the image information is combined after passing through sequence network 4, sequence network 5 and a 2*2 transpose convolution layer, the combined image information is combined after passing through sequence network 6 and the skip connection layer, and then combined with the image information after directly passing through a 2*2 transpose convolution layer, and finally output, the height and width of the image are unchanged after each perception field block; Each perception field block comprises four branch networks, two 1*1 convolution layers and a channel connection, after the image enters the perception field block, the image information is combined after passing through branch network 1, branch network 2, branch network 3 and branch network 4 through the channel connection, the combined image information is combined after passing through a 1*1 convolution layer, and then combined with the image information directly passing through another 1*1 convolution layer, and finally output. The sequence network 1 is composed of a 3*3 convolution layer with a step of 1, a 3*3 convolution layer with a step of 2 and two parameter linear rectifier function layers; the sequence network 2 is composed of two 3*3 convolution layers with a step of 1, a 3*3 convolution layer with a step of 2 and three parameter linear rectifier function layers; the sequence network 3 is composed of two 3*3 convolution layers with a step of 1 and two parameter linear rectifier function layers; the sequence network 4 is composed of a 3*3 convolution layer with a step of 1, a 3*3 convolution layer with a step of 2, a pixel recombination layer, a bicubic interpolation layer and two parameter linear rectifier function layers; the sequence network 5 is composed of two 3*3 transpose convolution layers with a step of 1, a 3*3 transpose convolution layer with a step of 2 and two parameter linear rectifier function layers; the sequence network 6 is composed of two 3*3 transpose convolution layers with a step of 1 and two parameter linear rectifier function layers; the branch network 1 is composed of a 1*1 convolution layer with a step of 1, a 3*3 convolution layer with a step of 1 and an expansion rate of 1 and a parameter linear rectifier function layer; the branch network 2 is composed of two 1*1 convolution layers with a step of 1, a 3*3 convolution layer with a step of 1 and an expansion rate of 3 and two parameter linear rectifier function layers; the branch network 3 is composed of a 1*1 convolution layer with a step of 1, a 3*3 convolution layer with a step of 1, a 3*3 convolution layer with a step of 1 and an expansion rate of 3 and two parameter linear rectifier function layers; the branch network 4 is composed of two 1*1 convolution layers with a step of 1, a 3*3 convolution layer with a step of 1, a 3*3 convolution layer with a step of 1 and an expansion rate of 5 and three parameter linear rectifier function layers, and the parameter linear rectifier function layers are used to solve the problem of overfitting and make the training of the model more stable.
3. The holographic camera based on physical model driven network and liquid lens according to claim 1, characterized in that, The calculation method of the definition evaluation value C in the physical model driven network of the holographic camera uses a Laplace operator, and the expression of the definition evaluation value C is: where F(p, q) is an image filtered using a Gaussian filter with a standard deviation of 1 for the input image; l is the number of pixels in the horizontal and vertical directions of the focused image; L 3×3 is a Laplacian operator of size 3 x 3.
4. The holographic camera based on physical model driven network and liquid lens according to claim 1, characterized in that, When performing depth calculation and defocus estimation in scene fusion, the gradient amplitude ratio R of the edge position of the image needs to be calculated first, and the expression is: where, and are the gradients along x and y directions after Gaussian filtering with a standard deviation of σ1to the image, respectively; and are the gradients along x and y directions after Gaussian filtering with a standard deviation of σ2to the image, respectively, and the size of R reflects the degree of defocus of the object, and the expression of the defocus estimation value σ is: After the defocus estimation value of the image is calculated, a sparse defocus estimation graph d' is obtained, then the sparse defocus estimation graph d' is interpolated by using the Laplace matrix interpolation method to obtain a complete defocus estimation graph d, and the expression is: d = λDd'(L + λD) -1 Wherein, L is a Laplace matrix, D is a diagonal matrix, and lambda is a constraint constant.
5. The holographic camera based on physical model driven network and liquid lens according to claim 1, characterized in that, In a physical model driven network, the distance z of back propagation i is expressed by the equation: z i = z0+ Ad x i Wherein, i is the selected number of layers; z0 is the reference recording distance; Δd is the depth interval of the holographic reconstruction image of the real 3D scene; the physical model in the physical model driven network of the holographic camera is a band-limited angular spectrum diffraction model, and in the training process of the physical model driven network, the loss function is a combination of perceptual loss, multi-scale structural similarity loss, total variation loss and mean square error loss, and the expression of the loss function Loss is: wherein a, β and η are coefficients of the perceptual loss, the multi-scale structural similarity loss and the total variation loss, respectively, Loss PE , Loss MS and Loss TV are the perceptual loss, the multi-scale structural similarity loss and the total variation loss, respectively, denotes the arithmetic mean of the amplitude information on the spatial light modulator, O SLM denotes the amplitude information on the spatial light modulator, the generated image is made visually closer to the image recognized by the human eye using a loss function composed of the perceptual loss, the multi-scale structural similarity loss, the total variation loss and the mean square error loss, thereby improving the quality of the holographic reconstruction image.