Image generation processing system, image generation program, image generation method and image generation processing device

The combination of NeRF and active stereo techniques in the image generation processing system addresses the challenge of generating accurate 3D images in extreme environments by using pattern projection and neural network processing, achieving high accuracy without requiring precise sensor calibration or position information.

JP2025141318APending Publication Date: 2025-09-29KYUSHU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024041198
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-09-29

AI Technical Summary

Technical Problem

Existing 3D image reconstruction methods, such as Neural Radiance Fields (NeRF), struggle to generate accurate images in extreme environments with low illumination, low contrast, or low texture, such as the ocean floor or inside a living organism, due to the inability to reconstruct 3D images or poor accuracy.

Method used

An image generation processing system that combines NeRF with active stereo techniques, using a projection device to project patterns and an imaging device to capture images, processing them with a neural network to generate highly accurate 3D images by reducing photometric loss through volume rendering and pattern information.

Benefits of technology

Enables the generation of highly accurate 3D images even in extreme environments with low illumination, low contrast, or low texture, reducing the number of required captured images and eliminating the need for precise sensor calibration and position information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025141318000001_ABST
    Figure 2025141318000001_ABST
Patent Text Reader

Abstract

To provide an image generation processing system, an image generation program, an image generation method and an image generation processing device which allow generation of highly accurate three-dimensional image even under extreme environment.SOLUTION: An image generation processing system comprises: a projection means device for projecting a pattern to an object having a three-dimensional shape: imaging device means for acquiring an image obtained by capturing the object; means for processing the image obtained by the imaging device means by a neural network to obtain information regarding a three-dimensional space; and three-dimensional image generation means for generating a three-dimensional image of the object by using pattern information about the pattern projected in projection device means and the information regarding the three-dimensional space.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image generation processing system, an image generation program, an image generation method, and an image generation processing device. [Background technology]

[0002] Acquiring precise (high-density) and accurate 3D images of various environments is important for applications in a variety of environments that are difficult for humans to access, such as scanning the inside of the human body with an endoscope, creating 3D maps of the ocean floor, and acquiring 3D shapes from images of planets such as Mars and satellite images. In recent years, the demand for accurate 3D measurements has also increased in fields such as autonomous vehicle control and industrial inspection.

[0003] In recent years, Neural Radiance Fields (NeRF) has attracted attention. NeRF is a neural network that can reconstruct complex 3D scenes from a series of 2D images. By optimizing deep neural networks (DNNs) to minimize photometric losses in an end-to-end manner, NeRF enables 3D image reconstruction and super-resolution.

[0004] For example, Patent Document 1 discloses an information processing device including one or more memories and one or more processors, where the one or more processors use an autoregressive model to generate information about a plurality of discrete codes that represent three-dimensional information, and a neural network model is used to improve the accuracy of the representation of three-dimensional information. Patent Document 2 also discloses an image processing device including: an acquisition unit that acquires an image of an index captured by an imaging device; a detection unit that detects the index from the captured image; an estimation unit that estimates the resolution of the imaging device based on the image of the index in the captured image; and a generation unit that generates information indicating the relationship between the distance from the imaging device and the resolution based on the resolution estimated by the estimation unit from a plurality of captured images of an index at a plurality of positions that are different distances from the imaging device. A technology for estimating the resolution of the captured image is also discussed. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Publication No. 2023-79022 [Patent Document 2] Japanese Patent Publication No. 2023-26245 Summary of the Invention [Problem to be solved by the invention]

[0006] For example, environments that are difficult for humans to access, such as the ocean floor or inside a living organism, are often extreme environments with low illumination, low contrast, low texture, etc. Because the principle of 3D image reconstruction using NeRF as described above is the same as that of conventional passive stereo methods, when 3D image reconstruction is performed in extreme environments such as low illumination, low contrast, or low texture, there are issues such as the inability to reconstruct 3D images or the accuracy of the 3D image reconstruction being poor.

[0007] Therefore, in order to solve these problems of the conventional technology, the inventors conducted research with the aim of building an image generation processing system that can generate high-precision three-dimensional images even in extreme environments such as low illumination, low contrast, and low texture. [Means for solving the problem]

[0008] Examples of specific embodiments of the present invention are given below.

[0009] [1] A projection device that projects a pattern onto an object having a three-dimensional shape; an imaging device for capturing an image of the target; a means for processing the image obtained by the imaging device using a neural network to obtain information about a three-dimensional space; and a three-dimensional image generating means for generating a three-dimensional image of the object using pattern information about the pattern projected by the projection device and information about the three-dimensional space. [2] The image generation processing system according to [1], wherein the neural network is trained to reduce photometric loss between the three-dimensional image and the image. [3] The image generation processing system according to [1] or [2], wherein the three-dimensional image generation means performs volume rendering processing. [4] The image generation processing system described in [3], wherein the three-dimensional image generation means acquires projection color information from information of the imaging device, information of the projection device, and color information of the projected pattern. [5] The image generation processing system according to any one of [1] to [4], wherein the information about the three-dimensional space includes color information and SDF information. [6] The image generation processing system according to any one of [1] to [5], wherein the three-dimensional image generation means further generates a three-dimensional image of the object using information about reflectance and brightness. [7] The image generation processing system described in [6], wherein the information on reflectance and brightness is obtained by training to reduce photometric loss between the three-dimensional image and the image. [8] The image generation processing system according to any one of [1] to [7], wherein the projection device projects laser light. [9] The image generation processing system according to any one of [1] to [8], wherein the projection device projects a plane-crossing laser beam.

[10] The image generation processing system according to any one of [1] to [9], further comprising an estimation unit for estimating information about the imaging device and information about the projection device.

[11] The image generation processing system described in

[10] , wherein the estimation means estimates the information of the imaging device and the information of the projection device using an encoding means that combines encoding using a hash table and encoding using a Fourier series.

[12] Pattern information about a pattern projected in a projection function that projects a pattern onto an object having a three-dimensional shape; An image generation program that causes a computer to execute a process of generating a three-dimensional image of an object using information about three-dimensional space obtained by processing images obtained by an imaging function that acquires images of the object using a neural network.

[13] generating a three-dimensional image of the object using information about the three-dimensional space obtained by weighting the image with the neural network and the pattern information;

[12] The image generation program described in

[12] , which trains the computer to perform weighting processing in the neural network so as to reduce photometric loss between the generated three-dimensional image of the object and the image.

[14] A step of processing an image of an object having a three-dimensional shape using a neural network to obtain information about three-dimensional space; generating a three-dimensional image of the object using pattern information about the pattern projected onto the object and information about the three-dimensional space.

[15] weighting the image with the neural network to obtain information about the three-dimensional space; generating a three-dimensional image of the object using the pattern information and information about the three-dimensional space;

[14] The image generation method according to

[14] , comprising the step of generating a three-dimensional image of the object and training a weighting process in the neural network so as to reduce photometric loss between the generated three-dimensional image and the image.

[16] A means for capturing an image of an object having a three-dimensional shape onto which a pattern is projected by an imaging means, and processing the image obtained by the imaging means using a neural network to obtain information about the three-dimensional space; and a three-dimensional image generating means for generating a three-dimensional image of the object using pattern information about the projected pattern and information about the three-dimensional space. [Effects of the Invention]

[0010] According to the present invention, it is possible to provide an image generation processing system, an image generation program, an image generation method, and an image generation processing device that enable the generation of highly accurate three-dimensional images even under extreme environments. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a schematic diagram illustrating the configuration of an image generation processing system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram illustrating a pipeline of the image generation processing system according to the present embodiment. [Figure 3] The cross laser pattern is projected onto the synthetic data under normal light and no light source. [Figure 4] We show the results of generating 3D images of the NeRF-Synthetic synthetic dataset under normal light or no light. [Figure 5] The results of generating 3D images of a synthetic BlendedMVS dataset under normal or no light source are shown. [Figure 6] The results of generating three-dimensional images of synthetic data with and without illuminance estimation are shown. [Figure 7] The composite data is shown when various laser patterns are applied. [Figure 8] 10 shows the results of generating three-dimensional images of the synthetic data when various laser patterns are irradiated onto the synthetic data. [Figure 9] This shows how a laser is irradiated and imaged on an underwater object. [Figure 10] The results of generating a three-dimensional image of an underwater object are shown. DETAILED DESCRIPTION OF THE INVENTION

[0012] The present invention will be described in detail below. The following description may be based on representative embodiments and specific examples, but the present invention is not limited to such embodiments.

[0013] (Image generation processing system) A first embodiment of the present invention relates to an image generation processing system having a projection device that projects a pattern onto an object (subject) having a three-dimensional shape, an imaging device that acquires an image of the object, a means for processing the image acquired by the imaging device using a neural network to obtain information about three-dimensional space, and a three-dimensional image generation means that generates a three-dimensional image of the object using pattern information about the pattern projected by the projection device and information about the three-dimensional space.

[0014] The image generation processing system of this embodiment combines NeRF and active stereo techniques. It generates (reconstructs) a 3D image of an object using pattern information and information about the 3D space obtained by processing images captured by an imaging device using a neural network. This enables the generation of highly accurate 3D images. Furthermore, in this embodiment, the combination of NeRF and active stereo techniques enables the generation of highly accurate 3D images even in extreme environments, such as those with low illumination, low contrast, and low texture. Furthermore, in this embodiment, the combination of NeRF and active stereo techniques enables the generation of highly accurate 3D images. As a result, it is possible to reduce the number of captured images required to obtain information about the 3D space. In other words, accurate information about the 3D space can be obtained even when the number of captured images is small. In this embodiment, it is also possible to generate (reconstruct) a 3D image of an object using information obtained from a small number of captured images, such as 10 or fewer, 5 or fewer, 3 or fewer, or even 1, and pattern information.

[0015] As shown in FIG. 1, the image generation processing system of this embodiment includes a plurality of projection devices (e.g., projectors) and imaging devices (e.g., cameras). The number of projection devices included in the image generation processing system is preferably 1 to 30, more preferably 2 to 20, and even more preferably 2 to 10. The image generation processing system preferably includes one imaging device, but may include two or more. The plurality of projection devices and imaging devices may be fixed so as to be movable as a single unit, or may be arranged so as to be movable individually. In this embodiment, by combining NeRF and active stereo, it is possible to generate (reconstruct) a three-dimensional image of an object even if the positions of the projection devices and imaging devices and their relative positions are unclear.

[0016] Although not shown, the image generation processing system of this embodiment may also include an image generation processing device. For example, the image generation processing device is a personal computer, and includes a processing circuit (such as a processor) and a storage device (such as a memory or storage). The image generation processing device executes an image generation program as described below. Specifically, the image generation processing device preferably includes: a means for processing an image obtained by capturing an image of an object having a three-dimensional shape onto which a pattern is projected using an imaging means, using a neural network to obtain information about the three-dimensional space; and a three-dimensional image generation means for generating a three-dimensional image of the object using pattern information about the projected pattern and information about the three-dimensional space. The present invention may also relate to such an image generation processing device. The image generation processing device, the projection device, and the imaging device may be electrically connected to each other via a path, or may be connected to each other via a network such as a client-server system or a cloud system.

[0017] The projection device is preferably a device that projects laser light, and the projected laser light may be laser light of a single wavelength or laser light of multiple wavelengths. The projection device may project the laser light in parallel at a predetermined interval, or may project the laser light so that the laser light forms intersections, and the projection pattern is not particularly limited. For example, the projection device can project any pattern onto an object (subject), such as a gray code, a grid pattern, or a cross laser pattern. In particular, the projection device is preferably a device that projects planar crossing laser light. The planar crossing laser light is composed of two line lasers, and the two laser planes are fixed approximately perpendicular. The two line lasers are preferably laser light of different wavelengths. By projecting planar crossing laser light from the projection device, an image with good contrast can be obtained, thereby obtaining more accurate pattern information.

[0018] The imaging device captures an image of an object having a three-dimensional shape. In this embodiment, the image obtained by the imaging device is processed using a neural network to obtain information about the three-dimensional space. Here, the information about the three-dimensional space is, for example, color information and SDF information. In this embodiment, color information and SDF information are obtained by weighting the image obtained by the imaging device using a neural network, and these are combined with pattern information to generate a three-dimensional image of the object.

[0019] A neural network is given "training data" and "learns" it, thereby outputting information about three-dimensional space from an image of a target (ground-truth image). The neural network used in this embodiment is NeRF, which outputs information about three-dimensional space (color information and SDF information) by inputting training data (the coordinates of the target image and parameters of the imaging device). NeRF learns based on the training data so that it can output accurate information about three-dimensional space (color information and SDF information).

[0020] FIG. 2 is a diagram illustrating the pipeline of the image generation processing system of this embodiment. Note that FIG. 2 omits some of the hierarchical sampling and distributed networks used in NeuS (NeurIPS, 2021), but these networks can also be used as appropriate. As shown in FIG. 2, the neural network used in this embodiment preferably has a trained SDF network and a color network.

[0021] The neural network is preferably trained to reduce the photometric loss between the generated 3D image based on the information about the 3D space (color information and SDF information) and the image of the object, where the photometric loss is the difference in brightness values ​​between the generated 3D image and the image of the object.

[0022] Specific training methods include, for example, the following methods. (1) Using the camera parameters, randomly sample rays cast from the optical center of each camera. (2) Three-dimensional points on the ray from the near clipping plane to the far clipping plane are sampled at regular or weighted intervals. (3) Pass the 3D points through the SDF network / color network to obtain the density and diffuse color of the points. (4) Reproject the 3D points onto the projector pattern and use the projector parameters to obtain 2D projected points. (5) Using the two-dimensional projected points, calculate the projected color from the projection pattern onto the three-dimensional points. (6) Render the image using volume rendering with pattern projection. (7) Calculate the photometric loss between the rendered image and the target image. (8) Use Adam, RMSProp, or AdamW to update the network parameters to minimize photometric loss. By updating the network parameters using the above training method and optimizing the SDF, it becomes possible to generate higher-resolution 3D images.

[0023] 2, the three-dimensional image generating means performs volume rendering processing by combining information about the three-dimensional space (color information and SDF information) obtained by neural network processing with pattern information about the pattern projected by the projection device. By performing volume rendering processing, a three-dimensional image can be generated from the information about the three-dimensional space (color information and SDF information).

[0024] The three-dimensional image generating means preferably further includes means for acquiring projection color information (calculating the color of the projected pattern observed on the object) from information about the imaging device, information about the projection device, and color information about the projected pattern. The three-dimensional image generating means preferably performs volume rendering processing by combining the projection color information in addition to information about the three-dimensional space (color information and SDF information) and pattern information. In this embodiment, performing volume rendering processing by combining the pattern information and projection color information in addition to information about the three-dimensional space (color information and SDF information) enables the generation of a higher-resolution three-dimensional image.

[0025] The following describes functions for volume rendering processing, but the functions used in the image generation processing system of this embodiment are not limited to the following.

[0026] In volume rendering, a two-dimensional point p is projected onto a three-dimensional point P on a ray cast from p. i (i=1~n) and is rendered as in equation (1).

[0027]

number

[0028] In equation (1), α(Pi ) is the three-dimensional point P i opacity, c(P i ,v) is P seen from the direction v i 's color, C is the final rendering color. Then, according to NeuS, α(P i ) is calculated using equation (2).

[0029]

number

[0030] In equation (2), f(P) is the SDF value of the three-dimensional point P, Φ s is a sigmoid function. Usually, c(P i ,v) is the color network itself, but we modify it to calculate the color by pattern projection. i Diffuse color (D(P i , v) and the color projected by the kth projector (Q k (P i ) are blended as shown in equation (3).

[0031]

number

[0032] In equation (3), i r is the reflectance coefficient, i b is the bias coefficient. i r HA P i Controls the reflection level of i b is Q(P i ) is controlled by the light emission level. As shown in equation (3), if the image can be correctly rendered by pattern projection, it means that the correct depth and texture can be obtained. And, i r and i b By appropriately setting the parameters and combining them with pattern information, it is possible to render highly accurate 3D images, especially in texture-less areas where depth information is lacking or in low-light environments.

[0033] In this embodiment, D(P i ,v) is the color network itself, and Q k (P i ) is calculated as equation (4).

[0034]

number

[0035] In equation (4), I k (x) returns the color of the pattern of the kth projector at a particular point x by bilinear sampling. K k are the internal parameters of the projector, and R k , t k are the k-th projector's relative transformations. They transform 3D points in the world coordinate system into the projector screen coordinate system.

[0036] As mentioned above, the three-dimensional image generating means preferably uses information about reflectance and brightness to generate a three-dimensional image of the object, and illumination parameters (i r and i b It is important to set the illumination parameter (i r and i b ) may be trained and optimized, i.e., the reflectance and brightness information may be trained to reduce photometric loss between the 3D images.

[0037] Furthermore, the image generation processing system of this embodiment performs the following steps when generating a three-dimensional image after learning is completed: r , i b By substituting 0, an image without a pattern projected can be generated, which can be used to create 3D mesh data, etc.

[0038] The image generation processing system of this embodiment may further include an estimation unit for estimating information on the image capture device and information on the projection device. For example, as shown in FIG. 2, the image capture device may have K c , R, t, and the projection device has projector parameters K k , R k , t k These parameters are used in the volume rendering process. For this reason, the image generation processing system of this embodiment preferably includes means for estimating projector parameters and camera parameters. The estimation means estimates information about the imaging device (camera parameters) and information about the projection device (projector parameters) using encoding using a hash table and / or encoding using a Fourier series. The estimation means may use a combination of encoding using a hash table and encoding using a Fourier series. When the image generation processing system further includes estimation means for estimating information about the imaging device and information about the projection device, it becomes possible to generate a three-dimensional image with higher accuracy. When the image generation processing system includes such estimation means, there is no need for highly accurate calibration of the imaging device in advance, and there is also no need to obtain position information of each device.

[0039] As described above, in this embodiment, we have successfully constructed a system that combines NeRF and active stereo. In this specification, this type of image generation system is sometimes referred to as ActiveNeuS.

[0040] The objective function of the pipeline as a whole is expressed by equation (5).

[0041]

number

[0042] In equation (5), L color is the photometric loss (L1), L reg is the eikonal term, L mask is the mask loss term, and λ and β are balance coefficients.

[0043] (Image generation method) A second embodiment of the present invention relates to an image generation method including the steps of: processing an image of an object having a three-dimensional shape using a neural network to obtain information about the three-dimensional space; and generating a three-dimensional image of the object using pattern information about a pattern projected onto the object and the information about the three-dimensional space. In this embodiment, the information about the three-dimensional space is, for example, color information or SDF information.

[0044] As described above, the image generation method of this embodiment is a system that combines NeRF and active stereo, making it possible to generate highly accurate 3D images. Furthermore, by combining NeRF and active stereo in this embodiment, it is possible to generate highly accurate 3D images even in extreme environments such as low illumination, low contrast, and low texture. Furthermore, by combining NeRF and active stereo in this embodiment, it is possible to reduce the number of captured images required to obtain information about 3D space.

[0045] The image generation method of this embodiment preferably further includes a step of training a weighting process in a neural network so as to reduce photometric loss between the generated three-dimensional image of the object and the image. The step of obtaining information about the three-dimensional space preferably includes a step of weighting the captured image with a trained neural network to obtain information about the three-dimensional space (color information and SDF information). Specific training techniques include the methods described above.

[0046] The step of generating a three-dimensional image preferably includes a volume rendering process. In the volume rendering process, volume rendering is performed by combining information about the three-dimensional space (color information and SDF information) obtained by the neural network process with pattern information about the pattern projected by the projection device. This allows a three-dimensional image to be generated from the information about the three-dimensional space (color information and SDF information).

[0047] When obtaining pattern information, a laser beam is projected onto the target. Examples of the projected laser beam include the laser beams described above. Among them, the laser beam is preferably a plane-crossing laser beam.

[0048] The step of generating a three-dimensional image preferably includes acquiring projection color information (calculating the color of the projected pattern observed on the object) from information about an imaging device that captures an image of the object, information about a projection device that projects a pattern onto the object, and color information about the projected pattern. In the step of generating a three-dimensional image, it is preferable to perform volume rendering processing by combining the projection color information in addition to information about the three-dimensional space (color information and SDF information) and pattern information. In this embodiment, performing volume rendering processing by combining the pattern information and projection color information in addition to information about the three-dimensional space (color information and SDF information) enables the generation of a higher-resolution three-dimensional image. The specific volume rendering processing method is as described above.

[0049] The step of generating a three-dimensional image preferably uses information about reflectance and brightness to generate a three-dimensional image of the object, and illumination parameters (i r and i b It is important to set the illumination parameter (i r and i b ) may be trained and optimized, i.e., the reflectance and brightness information may be trained to reduce photometric loss between the 3D images.

[0050] The image generating method of this embodiment may further include a step of estimating information about an imaging device that captures an image of the target and information about a projection device that projects a pattern onto the target. For example, as shown in FIG. 2, the imaging device may have K c , R, t, and the projection device has projector parameters K k, R k , t k These parameters are used in the volume rendering process. For this reason, the image generation method of this embodiment preferably includes a step of estimating projector parameters and camera parameters. In the estimation step, information on the image capture device (camera parameters) and information on the projection device (projector parameters) are estimated using a combination of encoding using a hash table and encoding using a Fourier series. When the image generation processing system further includes estimation means for estimating information on the image capture device and information on the projection device, it becomes possible to generate a three-dimensional image with higher accuracy.

[0051] (Image generation program) A third embodiment of the present invention relates to an image generation program that causes a computer to execute a process of generating a three-dimensional image of an object using pattern information about a pattern projected by a projection function that projects a pattern onto an object having a three-dimensional shape and information about the three-dimensional space obtained by processing an image obtained by an imaging function that captures an image of the object using a neural network. In this embodiment, the information about the three-dimensional space is, for example, color information or SDF information. Furthermore, the computer that executes the image generation program is an image generation processing device (e.g., a personal computer) included in the above-mentioned image generation processing system.

[0052] In the image generation program of this embodiment, it is preferable to train a computer to perform weighting processing in a neural network so as to reduce the loss of photometric measurement between the generated 3D image of the object and the image. In the image generation program of this embodiment, a 3D image of the object is generated using information about the 3D space obtained by weighting the captured image with a trained neural network and pattern information. Specific training methods include the methods described above.

[0053] The image generation program of this embodiment includes executing a volume rendering process. In the volume rendering process, volume rendering is performed by combining information about the three-dimensional space (color information and SDF information) obtained by neural network processing with pattern information about the pattern projected by the projection device. This makes it possible to generate a three-dimensional image from the information about the three-dimensional space (color information and SDF information). The specific volume rendering process method is as described above.

[0054] (Application) By using the image generation processing system, image generation method, and image generation program of this embodiment, it is possible to generate high-resolution three-dimensional images even underwater, in low-light, or texture-less environments. Furthermore, by using the image generation processing system, image generation method, and image generation program of this embodiment, highly accurate calibration of sensors and the like in advance is not required, and highly accurate position information acquisition is also not required. Therefore, the image generation processing system, image generation method, and image generation program of this embodiment can be used for the inspection and maintenance of underwater structures, the automatic generation of ocean floor maps and dynamic terrestrial maps, and the measurement of three-dimensional shapes in various environments that are difficult for humans to access, such as outer space, the deep sea, and inside living organisms.

[0055] In particular, the image generation processing system, image generation method, and image generation program of this embodiment are suitable for application to underwater scenes, and are preferably used in environments where photography is performed from underwater using an underwater drone or the like (such as underwater, the seabed, harbors, riverbanks, lakeshores, and other underwater structure inspections and maintenance, and photographing underwater objects and organisms). Furthermore, the image generation processing system, image generation method, and image generation program of this embodiment can generate detailed and accurate three-dimensional images even in target areas where external calibration is difficult, such as murky water. When applied to underwater scenes, high accuracy can be achieved by correctly reproducing refraction. [Example]

[0056] The features of the present invention will be explained in more detail below with reference to examples and comparative examples. The materials, amounts used, ratios, treatment contents, treatment procedures, etc. shown in the following examples can be changed as appropriate without departing from the spirit of the present invention. Therefore, the scope of the present invention should not be construed as being limited by the specific examples shown below.

[0057] <3D image generated> The three-dimensional image was generated using the following synthetic data: (1) NeRF-Synthetic: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. 1, 2, 4). (2) BlendedMVS: Synthetic data of "stone," "dog," "bear," and "sculpture" (A large-scale dataset for generalized multi-view stereo networks. Computer Vision and Pattern Recognition (CVPR), 2020. 4)

[0058] Laser light was projected from a projection device (virtual projector) onto the above synthetic data set. Specifically, cross laser patterns (red and green) were projected from four projection devices (virtual projectors) arranged to the left and right of the imaging device (camera). The imaging device (camera) and projection device (virtual projector) were fixed to have a predetermined positional relationship, and the projection device (virtual projector) moved together with the imaging device (camera). Specifically, the baseline lengths of the projection devices (virtual projectors) were set to 20 cm and 60 cm from the imaging device (camera), and the field of view was set to 60°. Figure 3 shows the cross laser pattern projected onto the synthetic data. Figure 3(a) shows the appearance of "Lego" in NeRF-Synthetic under normal light (Normal Illumination Scene), Figure 3(b) shows the appearance of "Lego" in NeRF-Synthetic under no light (Dark Illumination Scene), Figure 3(c) shows the appearance of "Stone" in BlendedMVS under normal light, and Figure 3(b) shows the appearance of "Stone" in BlendedMVS under no light.

[0059] In the example (ActiveNeuS), a 3D image of a synthetic dataset was generated using the scheme shown in Figure 2, which uses pattern information about a pattern projected by a projection device (virtual projector) and information about the 3D space obtained by processing images obtained by an imaging device (camera) with a neural network. In the comparative example (NeuS), a laser pattern was not projected onto the synthetic dataset, and only information about the 3D space obtained by processing images obtained by an imaging device (camera) with a neural network was used to generate a 3D image of the synthetic dataset. The results of the NeRF-Synthetic synthetic dataset are shown in Figure 4, and the results of the BlendedMVS synthetic dataset are shown in Figure 5.

[0060] In the NeRF-Synthetic synthetic dataset shown in Figure 4, 20 images were sampled under normal illumination (Normal Illumination Scene) and no illumination (Dark Illumination Scene). In the BlendedMVS synthetic dataset shown in Figure 5, 56 images of "stone," 31 images of "dog," 123 images of "bear," and 79 images of "sculpture" were sampled under normal illumination (Normal Illumination Scene) and no illumination (Dark Illumination Scene). As shown in Figures 4 and 5, the image generation system of the present invention accurately generated 3D images under both normal illumination (Normal Illumination Scene) and no illumination (Dark Illumination Scene). Using the image generation system of the present invention, 3D images were accurately generated even under no illumination (Dark Illumination Scene), including the background. Note that "Failed" in Figures 4 and 5 indicates that meaningful results were not obtained. In particular, as shown in Figure 4, when using the NeRF-Synthetic synthetic dataset, many "Failed" results were obtained in the dark illumination scene, and the accuracy of the generated 3D images was poor. Similarly, in Figure 5, although 3D images were generated in the dark illumination scene, the accuracy was poor.

[0061] <Quantitative analysis> Next, a quantitative analysis of the generated 3D image was performed using the Chamfer distance measurement method, in which the sum of the distances between the nearest neighboring points was used to calculate the distance between point clouds.

[0062] [Table 1]

[0063] [Table 2]

[0064] As shown in Tables 1 and 2, the Chamfer distance values ​​are small and the accuracy is high. In particular, in this embodiment, images were generated with high accuracy even when there was no light source.

[0065] <Illuminance estimation> To confirm the effectiveness of illuminance estimation, we compared the reconstructed shapes with and without illuminance estimation. We set the initial illuminance parameters as shown in Table 3 and trained the neural network shown in Figure 2. As shown in Figure 6, when accurate initial illuminance parameters were set (Figure 6(b)), a more accurate 3D image was generated compared to when illuminance estimation was not performed (Figure 6(a)).

[0066] <Pattern variations> The cross laser pattern was changed to a gray code (20 patterns), a grid pattern (one shot), and a color full line pattern (one shot), as shown in Figure 7, and evaluations were performed. All of these patterns were projected by a single virtual projector with a baseline length of 60 cm. The results of the 3D image generation are shown in Figure 8. As shown in Figure 8, with the gray code, multiple captured images were required to generate a 3D image, but with the grid pattern and color full line pattern, a 3D image could be generated with just one captured image.

[0067] <Verification using real scenes> To verify the feasibility of the proposed method in a real environment, we conducted experiments using real data captured underwater. Using a Remotely Operated Vehicle (ROV) in a pool, we captured images of an underwater scene containing a table and a mannequin wearing a swimsuit (Figure 9). The ROV was equipped with four pre-calibrated green cross-linked laser projectors. The above verification was also performed in a no-light environment (right image of Figure 9).

[0068] The captured images contained distortion due to refraction on the waterproof housing, which was removed by approximating it to radial lens distortion. We then used COLMAP to perform time-series transformation and a conventional algorithm (In Conference on Computer Vision and Pattern Recognition (CVPR), July 2016, In European Conference on Computer Vision (ECCV), July 2016) to reconstruct the 3D shape. For projector calibration, we obtained the laser plane parameters using the method proposed in (IEEE International Conference on Image Processing (ICIP), July 2022). We then manually set the field of view and adjusted its position along the depth axis. Since the laser line is always on the refracting surface, the refraction of the projector can be ignored, except for changes in the field of view.

[0069] The results are shown in Figure 10. As shown in Figures 10(a) and 10(b), the conventional method (COLMAP) only produced sparse reconstruction results, and the conventional method (NeuS) produced low reconstruction accuracy in areas with little texture. On the other hand, as shown in Figures 10(c) to 10(e), the image generation system of the present invention generated denser and more accurate 3D images, including the background, compared to COLMAP and NeuS. When illuminance estimation was performed, as shown in Figure 10(d), a larger area was reconstructed compared to Figure 10(c), in which illuminance estimation was not performed. This is because accurate estimation of illuminance parameters enabled accurate rendering of the floor, even at the edge of the screen. Furthermore, as shown in Figure 10(e), even in dark environments, a plausible shape could be reconstructed, although with reduced accuracy. Regarding quantitative error, the chamfer distance between the ground truth and the NeuS was 3.645 mm, while that of ActiveNeuS was 2.258 mm.

Claims

1. a projection device that projects a pattern onto an object having a three-dimensional shape; an imaging device for capturing an image of the target; a means for processing the image obtained by the imaging device using a neural network to obtain information about a three-dimensional space; and a three-dimensional image generating means for generating a three-dimensional image of the object using pattern information about the pattern projected by the projection device and information about the three-dimensional space.

2. The image generation and processing system of claim 1 , wherein the neural network is trained to reduce photometric loss between the three-dimensional image and the image.

3. 2. The image generation processing system according to claim 1, wherein said three-dimensional image generating means performs volume rendering processing.

4. 4. The image generation processing system according to claim 3, wherein said three-dimensional image generating means acquires projection color information from information on said imaging device, information on said projection device, and color information on said projected pattern.

5. The image generation processing system of claim 1 , wherein the information about the three-dimensional space includes color information and SDF information.

6. 2. The image generation processing system according to claim 1, wherein said three-dimensional image generating means further uses information about reflectance and brightness to generate a three-dimensional image of said object.

7. 7. The image generation and processing system of claim 6, wherein the reflectance and brightness information is trained to reduce photometric loss between the three-dimensional image and the image.

8. The image generation and processing system according to claim 1 , wherein the projection device projects laser light.

9. The image generating and processing system of claim 1 , wherein the projection device projects a plane-crossing laser beam.

10. The image generation processing system according to claim 1 , further comprising an estimation unit for estimating information about the image capture device and information about the projection device.

11. 11. The image generation processing system according to claim 10, wherein the estimation means estimates the information about the imaging device and the information about the projection device using encoding means that combines encoding using a hash table and encoding using a Fourier series.

12. Pattern information about a pattern projected by a projection function that projects a pattern onto an object having a three-dimensional shape; An image generation program that causes a computer to execute a process of generating a three-dimensional image of an object using information about three-dimensional space obtained by processing images obtained by an imaging function that acquires images of the object using a neural network.

13. generating a three-dimensional image of the object using information about the three-dimensional space obtained by weighting the image with the neural network and the pattern information; 13. The image generation program of claim 12, further comprising training the computer to weight the neural network so as to reduce photometric loss between the generated three-dimensional image of the object and the image.

14. A step of processing an image of an object having a three-dimensional shape using a neural network to obtain information about the three-dimensional space; generating a three-dimensional image of the object using pattern information about the pattern projected onto the object and information about the three-dimensional space.

15. weighting the image with the neural network to obtain information about the three-dimensional space; generating a three-dimensional image of the object using the pattern information and information about the three-dimensional space; 15. The method of claim 14, further comprising: generating a three-dimensional image of the object; and training a weighting process in the neural network to reduce photometric loss between the generated three-dimensional image and the image.

16. a means for capturing an image of an object having a three-dimensional shape onto which a pattern is projected by an imaging means, and processing the image obtained by the imaging means using a neural network to obtain information about the three-dimensional space; and a three-dimensional image generating means for generating a three-dimensional image of the object using pattern information about the projected pattern and information about the three-dimensional space.

Citation Information

Patent Citations

  • Image processing apparatus, image processing method, and program

    JP2023026245A

  • Information processing device and information generation method

    JP2023079022A