Learning device, depth map generation device, and program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-04-14
- Publication Date
- 2026-08-13
AI Technical Summary
【0022】 本発明におけるデプスマップ生成装置及びプログラムによれば、デジタルホログラフィの技術で撮影したホログラムから、高精度に被写体の撮影距離情報を抽出することができ、且つ、被写体の撮影距離情報の2次元分布を推定することができる。また、本発明における学習装置及びプログラムによれば、高精度に被写体の2次元距離情報を抽出するための学習済みモデルを生成することができる。
Smart Images

Figure 0007904715000005 
Figure 0007904715000006 
Figure 0007904715000007
Abstract
Description
[Technical Field]
[0001] The present invention relates to a learning device, a depth map generation device, and a program, and more particularly to a learning device, a depth map generation device, and a program for acquiring three-dimensional information based on digital holography technology. [Background technology]
[0002] Digital holography technology allows for the capture and display of three-dimensional images of subjects. However, because distance information in holography is represented by the wavefront (phase) of light waves, it is difficult to directly extract quantitative distance information of the subject from the captured data.
[0003] Conventional techniques involve numerically reconstructing a hologram (the captured image data) by applying propagation calculations that propagate over an arbitrary distance, and then obtaining the subject's shooting distance information from the propagation distance at which the image's sharpness is maximized. However, this method requires multiple 2D Fourier transform operations, resulting in a high computational load and long computation times for obtaining distance information.
[0004] In recent years, research has been conducted on methods for extracting shooting distance information from digital holographic image data using machine learning, with the aim of overcoming the shortcomings of this conventional technology, namely reducing the computational load and increasing speed. Non-Patent Document 1 discloses a method for extracting distance information of a subject using machine learning of a regression model with the input being a captured hologram, in the field of digital holography. Non-Patent Document 2 also discloses a method for extracting distance information of a subject using machine learning of a classification model with the input being a captured hologram.
[0005] Figure 11 illustrates a conventional method for extracting distance information from a hologram using machine learning. In Figure 11, the device on the left is a learning device 30 that trains a neural network model for extracting depth (shooting distance) from a hologram, and the device on the right is a depth extraction device 40 that uses the trained neural network.
[0006] The learning device 30 comprises a depth extraction unit 31 composed of a neural network and an error determination unit 32. A learning hologram is input to the depth extraction unit 31. The depth extraction unit 31 extracts depth from the input learning hologram based on a depth extraction model (neural network model) and outputs it to the error determination unit 32.
[0007] The error determination unit 32 receives the depth output from the depth extraction unit 31 and the correct depth data corresponding to the training hologram as input, and calculates the error of the depth extracted by the depth extraction unit 31. The calculated error is fed back to the depth extraction unit 31.
[0008] Then, in the depth extraction unit 31, the parameters of the depth extraction model (neural network) are adjusted based on the feedback error to minimize the error. The learning device 30 repeatedly adjusts the parameters based on the training hologram and the set of ground truth depth data, and terminates the learning process when the optimal parameters, i.e., the trained depth extraction model, are obtained.
[0009] The depth extraction device 40 includes a depth extraction unit 41. The depth extraction unit 41 is composed of a neural network and has a trained depth extraction model (optimal parameters) output from the learning device 30. The depth extraction unit 41 receives the captured hologram as input and outputs depth data (shooting distance) based on the trained depth extraction model.
[0010] Thus, in conventional technology, the captured hologram was used directly as input, and the shooting distance was used as output to train the model of the depth extraction unit 31. Furthermore, the depth extraction device 40 generated a single piece of distance information (depth data) within the field of view in response to the hologram input. [Prior art documents] [Non-patent literature]
[0011] [Non-Patent Document 1] T. Shimobaba, et al., “Convolutional neural network-based regression for depth prediction in digital holography”, 2018 IEEE 27th International Symposium on Industrial Electronics (ISIE), pp.1323-1326, (2018) [Non-Patent Document 2] Z. Ren, et al., “Learning-based nonparametric autofocusing for digital holography”, Optica, Vol.5, No.4, pp.337-344, (2018) [Non-Patent Document 3] S. Jegou,et al.,“The one hundred layers tiramisu: fully convolutional DenseNets for semantic segmentation”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pp.11-19, (2017) [Non-Patent Document 4] S. Li, et al., “Imaging through glass diffusers using densely connected convolutional networks”, Optica, Vol.5, No.7, pp.803-813, (2018)
Summary of the Invention
Problems to be Solved by the Invention
[0012] [[ID=I0]]However, the conventional method using machine learning directly uses the hologram data acquired by the imaging device as input data. Therefore, in addition to the contrast of the hologram decreasing according to the brightness of the light from the subject, the noise of the imaging element is included, resulting in the problem of low quality of the training data. Furthermore, the conventional method only extracts one distance information for the captured image and cannot be applied when there are multiple subjects at different distances. [[ID=I1]] [[ID=I2]]
[0013] [[ID=I3]] [[ID=I4]]Therefore, in view of the above problems, an object of the present invention is to provide a depth map generation device and program that can accurately extract the shooting distance information of a subject from a hologram taken by digital holography technology and can estimate the two-dimensional distribution of the shooting distance information of the subject. Another object is to provide a learning device and program that generate a trained model for accurately extracting two-dimensional distance information of a subject. [[ID=I5]] [[ID=I6]]
Means for Solving the Problems
[0014] In order to solve the above problems, the learning device according to the present invention is (1) A learning device that generates a machine learning model for estimating the two-dimensional distribution of the shooting distance from a hologram, which takes as input the complex amplitude distribution extracted from the acquired hologram and the virtual hologram created from the virtual reference light, and uses the depth image of the reproduction distance of the object light corresponding to the hologram as the correct answer data for the output to perform the learning of the machine learning model.
[0015] (2) The learning device described in (1) above is further preferably configured such that the machine learning model is composed of a convolutional neural network.
[0016] (3) The learning device described in (1) or (2) above preferably further includes a reference light whose amplitude is the maximum value of the amplitude distribution of the acquired hologram object light and whose phase is the average value of the phase of the object light.
[0017] To solve the above problems, the depth map generation device according to the present invention (4) A depth map generation device for estimating a two-dimensional distribution of shooting distance from a hologram, comprising a trained depth image generation model that takes a complex amplitude distribution extracted from the acquired hologram and a virtual hologram created from a virtual reference light as input and generates a depth image of the object light's reconstructed distance.
[0018] (5) The depth map generation device described in (4) above preferably further comprises a distance conversion unit that converts the reproduction distance of the object light to the shooting distance of the subject.
[0019] (6) The depth map generating apparatus described in (4) or (5) above preferably further includes a reference light whose amplitude is the maximum value of the amplitude distribution of the acquired hologram object light and whose phase is the average value of the phase of the object light.
[0020] To solve the above problems, the program according to the present invention (7) A program that causes a computer to function as one of the learning devices described in (1) through (3) above.
[0021] Furthermore, in order to solve the above problems, the program according to the present invention is (8) A program that causes a computer to function as a depth map generating device as described in (4) through (6) above. [Effects of the Invention]
[0022] The depth map generation device and program of the present invention can extract subject shooting distance information with high accuracy from holograms captured using digital holography technology, and can also estimate the two-dimensional distribution of the subject shooting distance information. Furthermore, the learning device and program of the present invention can generate a trained model for extracting two-dimensional distance information of a subject with high accuracy. [Brief explanation of the drawing]
[0023] [Figure 1] This is a block diagram showing an example of the configuration of the learning device according to Embodiment 1. [Figure 2] This diagram illustrates the extraction of complex amplitude distributions. [Figure 3] This is a diagram illustrating the generation of virtual holograms. [Figure 4] This figure shows an example of a neural network configuration. [Figure 5] This is a flowchart showing an example of the learning process by the learning device of Embodiment 1. [Figure 6] This block diagram shows an example of the configuration of the depth map generation device according to Embodiment 2. [Figure 7] This is an example of a digital holography imaging device that uses coherent light as illumination. [Figure 8] This is an example of a digital holography imaging device that uses incoherent light as illumination. [Figure 9] This figure illustrates the effectiveness of the depth map generation device of the present invention. [Figure 10] This figure shows the evaluation results of the generated depth map. [Figure 11] This is an example of a block diagram of a conventional depth extraction device and its learning device. [Modes for carrying out the invention]
[0024] Hereinafter, embodiments of the present invention will be described with reference to the drawings.
[0025] (Embodiment 1) Figure 1 is a block diagram showing an example of the configuration of a learning device according to Embodiment 1 of the present invention. The learning device according to Embodiment 1 of the present invention comprises a learning data creation device 100 and a learning device (machine learning unit) 110. The learning data creation device 100 generates learning data to be used for machine learning, and the learning device (machine learning unit) 110 performs learning (machine learning) of a depth image generation model. As described later, the learning data creation device 100 and the learning device (machine learning unit) 110 may be integrated and configured as a learning device 120.
[0026] [Building training data] The training data creation device 100 takes a training hologram as input and outputs a training virtual hologram. This training virtual hologram becomes the input to the training device 110 and serves as training data (teaching data) for the machine learning of the depth image generation model. The training data creation device 100 includes a complex amplitude extraction unit 11 and a virtual hologram generation unit 12.
[0027] This section outlines the construction of the training data. First, multiple holograms with different phase shift amounts are acquired. These can be collected experimentally or through simulation. Here, "experiment" refers to, for example, actually photographing holograms with an imaging device. "Simulation" refers to generating holograms through a simulation of optical wave propagation calculations incorporating the imaging conditions. In addition to the holograms, a depth image of the reproduction distance of the object light corresponding to the hologram is acquired. In this invention, the depth image refers to an image showing the two-dimensional distribution of the reproduction distance of the object light, and there may be a distance discrepancy with the depth map, which is the two-dimensional distribution of the actual subject's shooting distance. The conversion between reproduction distance and shooting distance will be described later. For multiple holograms, for example, the phase shift method is applied to obtain the amplitude and phase information of the object light, and a virtual hologram is generated from this information. Then, the "virtual hologram image" and the "depth image" of the reproduction distance of the object light (which will be the ground truth data during training) are paired to form the "virtual hologram dataset," which is the training data.
[0028] Next, we will explain the processing performed by the learning data creation device 100.
[0029] In constructing the training data, multiple (N) holograms, either captured using a digital holography system or generated through simulation, are prepared as training holograms and input to the complex amplitude extraction unit 11. The complex amplitude extraction unit 11 extracts the complex amplitude distribution, which is the amplitude and phase information of object light, from the input training holograms.
[0030] Figure 2 illustrates the extraction of the complex amplitude distribution (sometimes called "complex amplitude information"). The intensity distribution of N holograms (N=4 in Figure 2) is extracted using I j Let (x,y) (j=1, 2, … N), and the phase shift amount applied to each hologram is φ. j Let (x,y) (j=1, 2, … N). The complex amplitude obtained by combining the amplitude A(x,y) and phase φ(x,y) of the object light is obtained by calculation based on the phase shift method shown in equation (1). The imaginary number is denoted as i. This process removes unwanted light components, such as direct light and conjugate image, contained in the captured hologram. Furthermore, it can reduce noise components from the image sensor. If the reference light contains spherical phase, linear phase, or optical system aberration components, these may be removed. Moreover, the method for obtaining complex amplitude information is not limited to the phase shift method; other methods such as Fourier fringe analysis, iterative phase retrieval, and intensity transport equations may also be used.
[0031]
number
[0032] Furthermore, the amplitude A(x,y) can be extracted separately by calculating the absolute value of the complex amplitude obtained by equation (1), and the phase φ(x,y) can be extracted separately by calculating the argument. Figure 2 (bottom panel) shows the complex amplitude distribution with the amplitude A(x,y) and phase φ(x,y) separated.
[0033] The complex amplitude extraction unit 11 outputs the extracted complex amplitude distribution (amplitude and phase information of object light) to the virtual hologram generation unit 12.
[0034] Next, the virtual hologram generation unit 12 generates a virtual hologram from the complex amplitude distribution obtained by the complex amplitude extraction unit 11. Figure 3 is a diagram illustrating the generation of the virtual hologram. Amplitude R(x,y), Phase θ k Prepare M virtual reference beams (k=1, 2, … M). The amplitude and phase of the reference beams are arbitrary, but R(x,y) is the maximum value of the distribution of the object beam's amplitude A(x,y), phase θ1 is the average value of the object beam's phase φ(x,y), and θ2 ~ θ M If we set it to 2π(k-1) / M+θ1(k=2, … M), we can create a virtual hologram with the best contrast.
[0035] This virtual reference light and the object light (complex amplitude distribution) obtained in equation (1) are interfered with based on equation (2) to form virtual interference fringes (virtual hologram) H k M (M=4 in Figure 3) (x,y) (k=1, 2, … M) are generated. Note that the calculation in equation (1) reduces the direct light component and the noise component of the image sensor in the captured hologram, so the virtual hologram retains a stronger light signal from the subject with the excess light removed compared to the captured hologram itself. Furthermore, by setting the reference light optimally as described above, the virtual hologram has higher contrast compared to the captured hologram.
[0036]
number
[0037] The generated virtual holograms are used as training virtual holograms for machine learning, and are combined with depth images (sometimes called "ground truth depth images"), which are two-dimensional distributions of the reproduction distance of the corresponding object light obtained separately, to create training data (virtual hologram dataset) for a depth image generation model. Compared to holograms obtained by actual photography, virtual holograms have reduced noise and improved contrast, thus enhancing the learning effect of machine learning.
[0038] Furthermore, for depth images, a depth map representing the two-dimensional distribution of shooting distances is generated, and if necessary (depending on the hologram acquisition method), the shooting distance is converted to the object light reproduction distance to obtain the ground truth depth image. The method for generating the depth map may be to measure the distance from the image sensor to the subject using means such as LiDAR (Light Detection And Ranging) and create a map, or to place the subject at a predetermined distance from the image sensor and create a planned depth map, or to create a depth map using simulation. The constructed training data (virtual hologram dataset) becomes the input to the training device 110.
[0039] [Machine learning processing] The learning device (machine learning unit) 110 uses training data, namely training virtual holograms and ground truth depth images, to perform machine learning on a depth image generation model and generate a trained depth image generation model (optimal parameters for the model). The learning device (machine learning unit) 110 includes a depth image generation unit 13 and an error determination unit 14.
[0040] The training virtual hologram is input to the depth image generation unit 13. The depth image generation unit 13 includes a depth image generation model, for example, composed of a neural network. An example of the configuration of the depth image generation model will be described later. The depth image generation unit 13 converts the training virtual hologram image into a depth image. The generated depth image is output to the error determination unit 14.
[0041] The error determination unit 14 calculates the error between the depth image generated by the depth image generation unit 13 and the depth image of the correct data provided as training data. It is desirable to use a function such as MSE (Mean Squared Error) that yields larger values as the difference increases when calculating the error. The error determination unit 14 then feeds the calculated error back to the depth image generation unit 13.
[0042] The depth image generation unit 13 updates the parameters of the depth image generation model based on the feedback error to reduce the error.
[0043] After this, the learning device 110 repeatedly loads training data and updates (learns) parameters to find the optimal parameters. When the predetermined learning conditions are met, it terminates the learning process and outputs the trained depth image generation model.
[0044] The depth image generation model will be described below. The depth image generation unit (depth image generation model) 13 takes a virtual hologram as input and outputs a depth image with a resolution of 2x2 or higher. The depth image generation model is a machine learning model (a model with learning capabilities), and for example, a regression model or a clustering model neural network can be used. The input virtual hologram is a set of multiple virtual holograms H that have been created. k Any number of (x,y) images can be used. If two or more images are used as input, the number of channels corresponding to the number of images should be provided. The accuracy of the output is improved by inputting a large number of virtual holograms simultaneously. In this embodiment, a neural network is used to estimate the depth image from the input virtual holograms. The neural network consists of an encoder unit that performs downsampling and a decoder unit that performs upsampling.
[0045] Figure 4 shows an example of a neural network configuration. Here, an example is shown where a CNN architecture (DenseNet) similar to the CNN (Convolutional Neural Network) described in Non-Patent Documents 3 and 4 is adopted. The CNN applied in this invention is not limited to DenseNet, as long as it satisfies the above-mentioned input / output conditions. Furthermore, configurations similar to Fully convolutional networks, SegNet, ResNet, UNet, GAN (Generative Adversarial Network), etc., or improvements based on these architectures may also be used. In addition, it may be constructed using any machine learning model that can be trained, not just CNNs.
[0046] Let's briefly explain Figure 4. The upper part of Figure 4 shows an example of the configuration of a neural network that serves as a depth image generation model. The neural network comprises an input layer, hidden layers, and output layers, each with a predetermined number of pixels, along with processing blocks represented by thick arrows between each layer and skip connections indicated by dashed arrows. Each processing step indicated by a thick arrow is shown in the lower left of Figure 4, and the content of the processing differs depending on the arrow pattern. For example, a white arrow indicates a Dilated Convolution layer and a pooling layer, while a cross-hatched arrow indicates Dense Block 1 and a Downsampling transition block. The processing content of each block is shown in the lower right of Figure 4. For example, Dense Block 1 consists of a processing block combining BN (Batch Normalization), ReLU (Rectified Linear Unit), and convolution, and skip connections that directly connect each layer. Dense Block 2, Downsampling transition block, and Upsampling transition block each have the configurations shown in the diagram. In the learning process of this embodiment, the error function used is MSE, the Adam (Adaptive Moment Estimation) optimization method is used for parameter updates, and the number of epochs is set to 50.
[0047] In this way, a depth image generation model can be constructed, and by using the learning device 110 to perform machine learning on it, a trained depth image generation model can be generated.
[0048] Although the learning data creation device 100 and the learning device 110 have been described as separate devices, they may be integrated and configured as a single learning device 120. In this case, the learning device 120 can be configured to include a learning data generation unit 100 and a machine learning unit 110.
[0049] Next, an example of the learning process of the learning device of Embodiment 1 will be explained using the flowchart shown in Figure 5. In the flowchart of Figure 5, steps S11 to S14 are processes of the learning data creation device 100, and steps S15 to S19 are processes of the learning device (machine learning unit) 110.
[0050] Step S11: The training data creation device 100 acquires training holograms. The training holograms may be captured by the imaging device or generated by simulation. Multiple (N) holograms are acquired for each object light.
[0051] Step S12: The complex amplitude extraction unit 11 extracts the complex amplitude distribution from the acquired multiple (N) training holograms. The complex amplitude distribution is extracted using, for example, the phase shift method, but other methods may also be used.
[0052] Step S13: The virtual hologram generation unit 12 generates virtual holograms from the extracted complex amplitude distribution. M virtual reference lights are prepared and interfered with the object light (complex amplitude distribution) obtained by the complex amplitude extraction unit 11 to generate M virtual holograms.
[0053] Step S14: The training data creation device 100 measures the shooting distance to the subject corresponding to the acquired training hologram to obtain a depth map, and uses this as a depth image, or, if necessary, converts the shooting distance to the reproduction distance of the object light to obtain a depth image. This is used as the ground truth depth image and is combined with the generated virtual hologram (training virtual hologram) to create a virtual hologram dataset.
[0054] Steps S11 to S14 create one virtual hologram dataset. This virtual hologram dataset becomes the input data for the learning device (machine learning unit) 110. It is desirable to repeat the process from steps S11 to S14 to create the number of virtual hologram datasets required for machine learning in advance.
[0055] Step S15: The learning device 110 reads a virtual hologram dataset from the learning data creation device 100, obtains a virtual hologram for learning from it, and uses it as input data for the depth image generation unit 13.
[0056] Step S16: The depth image generation unit 13 converts the training virtual hologram using a depth image generation model to generate a depth image. The depth image generation model is composed of, for example, a convolutional neural network.
[0057] Step S17: The error determination unit 14 obtains a ground truth depth image from the virtual hologram dataset and compares it with the depth image input from the depth image generation unit 13. The error determination unit 14 calculates the error of the depth image generated by the depth image generation unit 13 and feeds it back to the depth image generation unit 13. Any error function (loss function) can be used for error determination, but in this embodiment, MSE (Mean Squared Error) is used.
[0058] Step S18: The depth image generation unit 13 updates the parameters of the depth image generation model (neural network) to minimize the error in the generated depth image. General neural network optimization methods such as Adam or SGD (Stochastic Gradient Descent) can be used to update the parameters.
[0059] Step S19: The learning device 110 determines whether the learning termination conditions are met. The learning termination conditions may include, for example, whether the parameters have been updated a predetermined number of times, or whether the error amount of the generated depth image or the amount of parameter updates has fallen below a predetermined threshold. If the termination conditions are met, the learning process is terminated. If the termination conditions are not met, the process returns to step S15, loads a new virtual hologram dataset again, and continues machine learning (parameter updates).
[0060] The above is an example of the learning process in the learning device of Embodiment 1. After the learning process is completed, the learning device 110 may output the learned depth image generation model (optimal parameters of the model).
[0061] Furthermore, a computer can be suitably used to function as the learning device 110 (or learning device 120) described above. Such a computer can be realized by storing a program describing the processing content that realizes each function of the learning device in the computer's memory, and having the computer's central processing unit (CPU) read and execute this program. Also, a computer can be suitably used to perform the learning process shown in Figure 5. The learning process can be realized by storing a program describing the processing content of each step of the learning process in the computer's memory, and having the computer's central processing unit (CPU) read and execute this program. This program can be recorded on a computer-readable recording medium.
[0062] (Embodiment 2) Figure 6 is a block diagram showing an example of the configuration of a depth map generation device according to Embodiment 2 of the present invention. The depth map generation device according to Embodiment 2 of the present invention comprises a preprocessing device 200 and a depth map generation device (main unit) 210. The preprocessing device 200 performs preprocessing (virtual hologram generation) for depth map generation, and the depth map generation device (main unit) 210 generates the depth map. As described later, the preprocessing device 200 and the depth map generation device (main unit) 210 may be integrated to form a depth map generation device 220 compatible with image holograms.
[0063] [Generate virtual holograms] The preprocessor 200 takes the captured hologram as input and outputs a virtual hologram. This virtual hologram becomes the input to the depth map generation device (main unit) 210. The preprocessor 200 includes a complex amplitude extraction unit 21 and a virtual hologram generation unit 22.
[0064] First, a digital hologram imaging device captures a plurality of holograms with different phase shift amounts and uses them as inputs to the complex amplitude extraction unit 21 of the preprocessing device 200. The complex amplitude extraction unit 21 applies, for example, the phase shift method to these holograms to obtain the amplitude and phase information of the object light. The specific processing content is the same as that of the complex amplitude extraction unit 11 of the learning data creation device 100 in FIG. 1. For example, the intensity distributions of N holograms are represented as I j (x,y) (j = 1, 2, … N), and the phase shift amounts applied to each hologram are represented as φ j (x,y) (j = 1, 2, … N). Based on the operation according to the phase shift method shown in the above formula (1), the complex amplitude distribution combining the amplitude A(x,y) and phase φ(x,y) of the object light is obtained (see FIG. 2). Note that, as a method for obtaining the complex amplitude information, a method other than the phase shift method may be used. The complex amplitude extraction unit 21 outputs the extracted complex amplitude distribution (amplitude and phase information of the object light) to the virtual hologram generation unit 22.
[0065] The virtual hologram generation unit 22 generates a virtual hologram from the complex amplitude distribution obtained by the complex amplitude extraction unit 21. The specific processing content is the same as that of the virtual hologram generation unit 12 of the learning data creation device 100 in FIG. 1. For example, M virtual reference lights with amplitudes R(x,y) and phases θ k (k = 1, 2, … M) are prepared. This virtual reference light and the object light (complex amplitude distribution) obtained by the formula (1) are interfered based on the above formula (2) to generate virtual interference fringes (virtual holograms) H k (x,y) (k = 1, 2, … M) (see FIG. 3). Note that the magnitudes of the amplitude and phase of the reference light are arbitrary. If R(x,y) is set as the maximum value of the distribution of the amplitude A(x,y) of the object light, the phase θ1 is set as the average value of the phase φ(x,y) of the object light, and θ2 ~ θ M is set as 2π(k - 1) / M + θ1 (k = 2, … M), the virtual hologram with the best contrast can be created. By using the virtual hologram with good contrast, the output accuracy can be improved. The generated virtual hologram is used as an input to the depth map generation device 210.
[0066] [Generate depth map] The depth map generation device (main unit) 210 takes a virtual hologram as input and outputs a depth map. The depth map generation device 210 includes a depth image generation unit 23 and, if necessary, a distance conversion unit 24.
[0067] The depth image generation unit 23 is equipped with a pre-trained depth image generation model, for example, composed of a neural network. This pre-trained depth image generation model either directly utilizes the depth image generation model after training is completed in the learning device 110 shown in Figure 1, or it is obtained by extracting the depth image generation model (optimal parameters) that has been trained by the learning device 110 shown in Figure 1 and porting it to the depth image generation unit 23. The pre-trained depth image generation model performs the conversion from a virtual hologram image to a depth image. That is, the depth image generation unit 23 inputs a virtual hologram into the pre-trained depth image generation model and outputs a depth image, which is a two-dimensional distribution of the reproduction distance of object light. The depth image is output to the distance conversion unit 24 as needed.
[0068] The distance conversion unit 24 converts the depth image of the object light's reproduction distance into a depth map of the subject's shooting distance. Therefore, when the object light's reproduction distance and the subject's shooting distance are equal (when the depth image and depth map are the same), the distance conversion unit 24 is unnecessary. The relationship between the object light's reproduction distance and the subject's shooting distance will be explained below.
[0069] In digital holography, the reproduction distance z of object light depends on the configuration of the optical elements and optical system of the imaging device. r and the shooting distance z of the subject s There are cases where they are equal and cases where they are not necessarily equal.
[0070] Figure 7 shows an example of a digital holography system that uses coherent light as illumination. The digital holography system 10 includes a laser light source 2, a spatial filter 3, a beam splitter 4, a plane mirror 5, and an image sensor 6. sThis is the distance of the optical path from the subject 1 to the imaging surface of the image sensor 6. Coherent light generated from the laser light source 2 has its distorted wavefronts and noise removed by the spatial filter 3, and is split into two beams by the beam splitter 4. One split beam is directed toward the subject and becomes the object light reflected by the subject. The other split beam is directed toward the plane mirror 5 and becomes the reference light reflected by the plane mirror 5. The object light and the reference light interfere with each other at the imaging surface of the image sensor 6, and the resulting hologram is acquired by the image sensor 6.
[0071] From multiple holograms acquired by the image sensor 6, a complex amplitude distribution on the imaging plane is obtained using a phase shift method, etc., and by applying diffraction propagation calculations to this complex amplitude distribution, an arbitrary distance z in the reconstructed system is obtained. r The image of the object can be reconstructed. The reconstruction distance z of the object light is determined from the propagation distance at which the sharpness of the subject's image is highest. r This can be obtained. When using a digital holography imaging device 10 that uses coherent light as illumination light, z r =z s In other words, since the distance at which the object light is reproduced is equal to the distance at which the subject is photographed, the depth image and the depth map are equal.
[0072] In contrast, Figure 8 shows an example of a digital holography apparatus that uses incoherent light as illumination. The digital holography apparatus 20 includes a lens 7, a beam splitter 4, a concave mirror 8, a plane mirror 9, and an image sensor 6. s is the distance from subject 1 to lens 7, f0 is the focal length of lens 7, z l f is the distance between lens 7 and concave mirror 8 or plane mirror 9. d1 ,f d2 The focal lengths of the concave mirror 8 and the plane mirror 9, z h is the distance from the concave mirror 8 or the plane mirror 9 to the image sensor 6. The incoherent illumination light passes from the subject 1 through the lens 7 and is split into two beams by the beam splitter 4. One beam is reflected by the concave mirror 8 and heads towards the image sensor 6, while the other beam is reflected by the plane mirror 9 and heads towards the image sensor 6. The two beams interfere with each other on the imaging surface of the image sensor 6, and the resulting hologram is acquired by the image sensor 6.
[0073] The object light regeneration distance z is similar to the method described above. r However, when using a digital holography imaging device 20 that uses incoherent light as illumination light, z r ≠z s In other words, the reproduction distance of object light and the shooting distance of the subject do not match. Therefore, in order to obtain a depth map, the reproduction distance of object light obtained by the depth image generation unit 23 is calculated using equations (3) and (4) below. r The subject's shooting distance z s Convert to.
[0074]
number
[0075]
number
[0076] If the type or element of the optical elements constituting the optical system changes, equations (3) and (4) above also need to be changed. The distance conversion unit 24 changes z depending on the optical system. r and z s The relationship is correctly converted (corrected). When using a hologram with incoherent light as illumination, the depth image reproduction distance z r Shooting distance z s By converting to z, a depth map is generated. When coherent light is used as illumination light, etc. r =z s In this case, the processing of the distance conversion unit 24 may be skipped.
[0077] Furthermore, when generating depth maps using a learning model with a hologram illuminated by incoherent light, directly using the depth map of the shooting distance as ground truth data for the input of the virtual hologram did not yield sufficient learning results. In contrast, using the depth image of the object light's regeneration distance as ground truth data for the input of the virtual hologram yielded good learning results. Therefore, by using a trained depth image generation model and performing distance transformations as needed, high-quality depth maps can be generated.
[0078] Although the preprocessing device 200 and the depth map generation device 210 have been described as separate devices, they may be integrated to form a single depth map generation device 220 that supports image hologram input. In this case, the device can be configured as an image hologram-compatible depth map generation device 220 comprising a preprocessing unit 200 and a depth map generation unit 210.
[0079] Furthermore, a computer can be suitably used to function as the depth map generation device 210 (or depth map generation device 220) described above. Such a computer can be realized by storing a program in the computer's memory that describes the processing content for realizing each function of the depth map generation device, and by having the computer's central processing unit (CPU) read and execute this program. This program can be recorded on a computer-readable recording medium.
[0080] (Verification of effectiveness) Figure 9 shows images used to verify the effectiveness of the depth map generation device of the present invention. The subjects were objects shaped like the MNIST (Modified National Institute of Standards and Technology (database): image data of handwritten digits) dataset, and samples were created by randomly placing them at positions at 10 mm intervals within a range of 10 to 100 mm from the imaging device. Multiple holograms with different phases were acquired for each sample. Figure 9(a) shows an example of a hologram captured at each imaging distance. The numbers above each hologram represent the distance from the imaging device (unit: mm). The captured holograms have a large difference in intensity between the object light and the reference light, resulting in low contrast.
[0081] Next, a virtual hologram was generated. Figure 9(b) shows the virtual hologram generated using the virtual hologram generation procedure described above, based on the captured hologram (a), using equations (1) and (2). It can be seen that the virtual hologram has improved contrast compared to the original hologram.
[0082] Furthermore, a correct depth image was generated from each sample. A predetermined number of virtual hologram datasets, which combine virtual holograms and depth images, were used as training data to perform machine learning on the aforementioned neural network depth map generation model. A depth map generation device 210 was constructed using the trained depth map generation model.
[0083] The depth map generation device 210 was compared with the results obtained by inputting both a captured hologram and a virtual hologram. Figure 9(c) shows the depth map generated when the captured hologram itself (Figure 9(a)) is input. Figure 9(d) shows the depth map generated when the virtual hologram of the present invention (Figure 9(b)) is input.
[0084] Figure 9(e) shows the true depth map values created from the placement positions of the objects in each sample. In the depth map of Figure 9, the background area other than the objects is set to a distance of 0 mm (values other than 10-100 mm) and is displayed in black.
[0085] Figure 10 shows the evaluation results of the generated depth maps. The vertical axis represents the RMSE (Root Mean Squared Error) of the generated depth map, and the horizontal axis represents the shooting distance. The graph in Figure 10 shows the error between the true depth map (e) and the depth map generated using the captured hologram itself as input (c) and the depth map generated using a virtual hologram as input (d) for each sample at each shooting distance. Note that this RMSE is the average value of the RMSE when generating depth maps by changing the shape of the subject multiple times at each shooting distance. It can be seen that by using the virtual hologram of the present invention, the RMSE is reduced for subjects at all shooting distances, and the accuracy of distance measurement is improved. Therefore, it has been confirmed that the present invention enables the realization of a depth map generation device that can extract shooting distance information of a subject with high accuracy and estimate the two-dimensional distribution of the shooting distance information of a subject with higher accuracy compared to the conventional technology.
[0086] In the above embodiment, the configuration and operation of the learning device 110 and the depth map generation device 210 have been described. However, the present invention is not limited to this, and may be configured as a learning method for creating a trained depth image generation model, or as a method for generating a depth map using a trained depth image generation model.
[0087] Although the embodiments described above are representative examples, it will be apparent to those skilled in the art that many modifications and substitutions are possible within the spirit and scope of the present invention. Therefore, the present invention should not be interpreted as being limited by the embodiments described above, and various modifications or changes are possible without departing from the scope of the claims. For example, the functions, etc., included in each block, step, etc., described in the embodiments can be rearranged in a logically consistent manner, and multiple constituent blocks, steps, etc., can be combined into one or divided. [Explanation of Symbols]
[0088] 1. Subject 2. Laser light source 3 Spatial Filters 4-beam splitter 5 plane mirror 6 Image sensor 7 lenses 8 concave mirror 9 plane mirror 10 Digital holography equipment 11 Complex Amplitude Extraction Unit 12 Virtual Hologram Generation Unit 13 Depth Image Generation Unit 14 Error judgment section 20 Digital Holography Equipment 21 Complex Amplitude Extraction Unit 22 Virtual Hologram Generation Unit 23 Depth Image Generation Unit 24 Distance conversion unit 30 Learning device 31 Depth extraction section 32 Error judgment section 40 Depth Extraction Device 41 Depth extraction section 100 Training data creation device 110 Learning Device (Machine Learning Section) 200 Pre-treatment device 210 Depth map generation device (main unit)
Claims
1. A learning device that generates a machine learning model for estimating the two-dimensional distribution of shooting distance from a hologram, The machine learning model is trained using a virtual hologram created from a complex amplitude distribution extracted from an acquired hologram and a virtual reference light as input, and a depth image of the reproduction distance of the object light corresponding to the hologram as output ground truth data. A learning device comprising a reference light whose amplitude is the maximum value of the amplitude distribution of the acquired object light of the hologram, and whose phase is the average value of the phase of the object light.
2. In the learning device according to claim 1, The aforementioned machine learning model is a learning device composed of a convolutional neural network.
3. A depth map generation device that estimates the two-dimensional distribution of shooting distance from a hologram, The system includes a trained depth image generation model that takes a complex amplitude distribution extracted from an acquired hologram and a virtual hologram created from a virtual reference light as input, and generates a depth image of the object light's reproduction distance. A depth map generating device that includes a reference light whose amplitude is the maximum value of the amplitude distribution of the acquired hologram object light and whose phase is the average value of the phase of the object light.
4. In the depth map generation apparatus according to claim 3, A depth map generation device further comprising a distance conversion unit that converts the reproduction distance of the object light into the shooting distance of the subject.
5. A program that causes a computer to function as a learning device according to claim 1 or 2.
6. A program that causes a computer to function as a depth map generating device according to claim 3 or 4.
Citation Information
Patent Citations
Digital holography device
JP2013246424A
Hologram recording device and hologram reproduction device
JP2020060709A
Apparatus extracting three-dimensional location of object using optical scanning holography and convolutional neural network and method thereof
KR1020200072158A
Incoherent digital holography based depth camera
US20210166409A1