Three-dimensional reconstruction method and three-dimensional reconstruction system

The method enhances three-dimensional reconstruction accuracy for monochromatic and textureless objects by projecting structured light and updating neural network parameters based on calculated coordinates, addressing the limitations of conventional methods.

JP2026042099APending Publication Date: 2026-03-11PREFERRED NETWORKS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-11-18
Publication Date
2026-03-11

AI Technical Summary

Technical Problem

Conventional three-dimensional reconstruction methods suffer from low accuracy, particularly for monochromatic and textureless objects.

Method used

A three-dimensional reconstruction method that projects structured light onto an object from multiple viewpoints, identifies corresponding pixels using a neural network, and updates the neural network parameters based on calculated coordinates to enhance reconstruction accuracy.

Benefits of technology

Improves the reconstruction accuracy of objects by aligning corresponding pixels and updating neural network parameters, resulting in a more precise three-dimensional model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026042099000001_ABST
    Figure 2026042099000001_ABST
Patent Text Reader

Abstract

To provide a three-dimensional reconstruction method capable of increasing the reconstruction accuracy of an object. [Solution] A three-dimensional reconstruction method acquires images of an actual object onto which structured light is projected, from at least two viewpoints including a first viewpoint and a second viewpoint; identifies a first pixel in the image from the first viewpoint and a second pixel in the image from the second viewpoint that correspond to the same point on the surface of the actual object onto which the structured light is projected, based on the projection pattern of the structured light; calculates a first coordinate corresponding to the first pixel and a second coordinate corresponding to the second pixel on the surface of the object reconstructed by the object shape neural network, based on current parameters in an object shape neural network that reconstructs the three-dimensional shape of the object; and updates the parameters of the object shape neural network using at least the first coordinate and the second coordinate.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a three-dimensional reconstruction method and a three-dimensional reconstruction system. [Background technology]

[0002] Conventionally, a 3D reconstruction method is known that generates a 3D model of an object based on multiple images obtained by capturing the object from different directions using a camera. Such a 3D reconstruction method is used in 3DCG (3 Dimensional Computer Graphics) and the like.

[0003] There is also a technology for performing three-dimensional reconstruction by volume rendering using a neural network. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Peng Wang et al., "35th Conference on Neural Information Processing Systems", 16,December,2021, Internet<URL:https: / / arxiv.org / pdf / 2106.10689.pdf> ,internet<URL:https: / / lingjie0206.github.io / papers / NeuS / index.htm> Summary of the Invention [Problem to be solved by the invention]

[0005] However, in conventional three-dimensional reconstruction methods, the accuracy of reconstruction of an object may be low, for example, when the object is monochromatic and textureless.

[0006] An object of the present disclosure is to provide a three-dimensional reconstruction method and a three-dimensional reconstruction system that can improve the reconstruction accuracy of an object. [Means for solving the problem]

[0007] A three-dimensional reconstruction method according to one aspect of an embodiment of the present disclosure acquires images of an actual object onto which structured light is projected, from at least two viewpoints including a first viewpoint and a second viewpoint; identifies a first pixel in the image from the first viewpoint and a second pixel in the image from the second viewpoint, which correspond to the same point on the surface of the actual object onto which the structured light is projected, based on a projection pattern of the structured light; calculates a first coordinate corresponding to the first pixel and a second coordinate corresponding to the second pixel on the surface of the object reconstructed by the object shape neural network, based on current parameters in an object shape neural network that reconstructs the three-dimensional shape of the object; and updates the parameters of the object shape neural network using at least the first coordinate and the second coordinate. [Effects of the Invention]

[0008] According to the present disclosure, it is possible to provide a three-dimensional reconstruction method and a three-dimensional reconstruction system that can improve the reconstruction accuracy of an object. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a diagram illustrating an example of the overall configuration of a three-dimensional reconstruction system according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a functional configuration of a control device according to the first embodiment. [Figure 3] 3 is a diagram illustrating an example of the functional configuration of a model generation unit in the control device of FIG. 2. [Figure 4] 10A and 10B are diagrams illustrating examples of structured light projection patterns. [Figure 5] 3A and 3B are diagrams showing examples of images captured by the first and second cameras. [Figure 6]1 is a diagram illustrating an example of an image of an object without structured light projected thereon. [Figure 7] FIG. 1 is a diagram illustrating an example of a three-dimensional scene. [Figure 8] 10A and 10B are diagrams showing examples of results of identifying corresponding pixels in the first and second images. [Figure 9] FIG. 10 is a diagram illustrating an example of the relationship between corresponding pixels and corresponding coordinates. [Figure 10] 3 is a flowchart of an example of processing by the control device of FIG. 2. [Figure 11] 3 is a flowchart of an example of processing by a rendering unit of the control device of FIG. 2. [Figure 12] 3 is a flowchart of an example of processing by a model output unit of the control device of FIG. 2. [Figure 13] FIG. 10 is a diagram illustrating an example of the functional configuration of a control device according to a second embodiment. [Figure 14] 14 is a diagram illustrating an example of the functional configuration of a model generation unit in the control device of FIG. 13. [Figure 15] FIG. 10 is a diagram illustrating an example of a first loss according to the second embodiment. [Figure 16] FIG. 10 is a diagram illustrating an example of a second loss according to the second embodiment. [Figure 17] 14 is a flowchart of an example of processing by the control device of FIG. 13. [Figure 18] FIG. 2 is a block diagram of an example of a hardware configuration of a control device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. To facilitate understanding of the description, the same components in the drawings will be denoted by the same reference numerals as much as possible, and duplicate descriptions will be omitted where appropriate.

[0011] [Embodiment] <Overall configuration example of the 3D reconstruction system 100> The configuration of a three-dimensional reconstruction system 100 according to this embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the overall configuration of the three-dimensional reconstruction system 100. In Fig. 1, the X direction, Y direction, and Z direction are perpendicular to each other. The Z direction is the normal direction of the mounting surface and is typically the vertical direction. The positive Z direction side is referred to as the upper side, and the negative Z direction side is referred to as the lower side. The X direction and Y direction are extension directions of the mounting surface and are typically horizontal directions.

[0012] 1, the three-dimensional reconstruction system 100 includes a support table 2, a first camera 3, a second camera 4, a rotation unit 5, a projection unit 6, and a control device 70. Note that the support table 2 and the rotation unit 5 are not essential components, and the three-dimensional reconstruction system 100 does not necessarily have to include them.

[0013] The three-dimensional reconstruction system 100 generates a three-dimensional model of the object 10 based on multiple images of the object 10 taken from different directions by the first camera 3 and the second camera 4. The object 10 may be, for example, a textureless cylindrical object as exemplified in Fig. 1, or an object of any shape that is sized to be placed inside the outer edge of the support table 2 when viewed from above.

[0014] In this specification and claims, the term "object" includes both an object that exists in reality and is the subject of three-dimensional reconstruction, i.e., a subject, and a reconstructed object based on multiple images of the subject. Whether "object" refers to a subject that exists in reality or a reconstructed object can be distinguished appropriately depending on the context. For example, when describing the handling of an object in real space, such as "placing an object on a support base," "object" refers to a subject that exists in reality. On the other hand, when describing the handling of an object in virtual space, such as "storing an object" or "reconstructing an object," "object" refers to a reconstructed object that exists in reality. An "actual object" corresponds to a subject that exists in reality and is the subject of three-dimensional reconstruction.

[0015] The support table 2 is a base on which the object 10 is placed and rotated on its upper surface 2A. The support table 2 is a transparent plate such as an acrylic plate. Because the support table 2 is a transparent plate, it is possible to photograph the object 10 placed on the upper surface 2A from the lower surface 2B, which is the back side of the upper surface 2A, as shown in FIG. 2. By using a transparent plate for the support table 2, the three-dimensional reconstruction system 100 can photograph the object 10 from all directions, thereby obtaining three-dimensional reconstruction information of the object 10 without any information loss.

[0016] The first camera 3 and the second camera 4 are each an example of a camera capable of photographing an object 10 placed on the upper surface 2A from different directions. The first camera 3 is installed toward the upper surface 2A of the support base 2, and can be directed toward the object 10 from a diagonally upward direction relative to the upper surface 2A to photograph the object 10. The first cameras 3 are installed at different tilt angles. The number of cameras may be one or more.

[0017] The second camera 4 is installed below the first camera 3 and is directed toward the object 10 from a position lower than the first camera 3, so that it can photograph the object 10. The second cameras 4 are installed at different tilt angles.

[0018] Each of the first camera 3 and the second camera 4 may be, but is not limited to, a so-called RGB camera that can capture an RGB image including the colors R (Red), G (Green), and B (Blue) of the object 10. For example, the first camera 3 and the second camera 4 may be an infrared camera.

[0019] In FIG. 1, the dashed arrow extending from the first camera 3 represents a first gaze direction vector v1, which is the direction of a line of sight from the first camera 3. The first gaze direction vector v1 corresponds to the first viewpoint. The dashed arrow extending from the second camera 4 represents a second gaze direction vector v2, which is the direction of a line of sight from the second camera 4. The second gaze direction vector v2 corresponds to the second viewpoint. In this specification, when there is no need to distinguish between the first gaze direction vector v1 and the second gaze direction vector v2, they may be collectively referred to as the gaze direction vector v.

[0020] In this embodiment, the term "camera" refers to an element capable of capturing an RGB image of the object 10. This "camera" encompasses the entire camera device, sensors such as a CMOS sensor or depth sensor built into the camera device, and standalone sensors. The "camera" preferably has a spatial resolution corresponding to the spatial frequency of the projection pattern contained in the structured light 60 so as to be able to capture the structured light 60 projected onto the object 10. For example, in this embodiment, it is preferable to use a camera with high spatial resolution, such as a 4K camera, so as to be able to capture the structured light 60 with a high spatial frequency.

[0021] Each of the first camera 3 and the second camera 4 has internal parameters and external parameters. The internal parameters include information related to the distortion of the lens included in each camera. In this embodiment, the internal parameters are known based on simulation results or the like. The external parameters include information related to the attitude (pose) of each camera. The attitude of a camera includes the relative attitude of the camera with respect to a predetermined reference and the absolute attitude in the world coordinate system. The attitude includes position and orientation. The attitude of a camera also corresponds to the tilt of the optical axis of the optical system, such as a lens, included in each camera. The external parameters are known based on simulation results or the like, or can be calculated by any method for each reconstruction operation by the 3D reconstruction system 100. In this embodiment, a case where the external parameters are known based on simulation results or the like is illustrated.

[0022] Rotation unit 5 rotates object 10 by rotating support base 2. For example, rotation unit 5 rotates support base 2 in the direction of arrow 20. First camera 3 and second camera 4 each capture images of object 10 at each of a plurality of rotation angles set by rotation unit 5, thereby being able to capture images of object 10 from different directions.

[0023] A known power transmission system can be used for the mechanism of the rotating unit 5. For example, the rotating unit 5 has a motor and a gear mechanism. The rotating unit 5 may be configured so that the driving force of the motor is transmitted to the rotation shaft of the support base 2 via the gear mechanism. Alternatively, the rotating unit 5 may be configured so that a driving force is applied to the outer edge of the support base 2 to rotate the support base 2.

[0024] The projection unit 6 projects structured light 60 onto the real object 10. The structured light 60 is light having a predetermined pattern. In the example shown in FIG. 1 , the structured light 60 includes a plurality of linear patterns, each of which is a linear pattern extending in the Z direction, i.e., the vertical direction, and which are arranged in a direction perpendicular to the Z direction, i.e., the horizontal direction. The shape of the cross-sectional light intensity distribution in the direction perpendicular to the extension direction of the linear patterns included in the structured light 60 is substantially rectangular here, but may be another shape, such as a substantially sinusoidal wave. The vertical direction is an example of a first direction, and the horizontal direction is an example of a second direction. The projection pattern of the structured light 60 here has a predetermined spatial frequency in the horizontal direction. When the projection unit 6 projects the structured light 60, the predetermined pattern of the structured light 60 is projected onto the surface of the real object 10.

[0025] In this embodiment, the structured light 60 may have a plurality of predetermined projection patterns. The plurality of projection patterns may be composed of, for example, a plurality of images. The structured light 60 may have a plurality of projection patterns in which the spatial frequencies of the linear patterns are different from one another. Alternatively, the structured light 60 may have a plurality of projection patterns in which the linear patterns extend in different directions from one another. Each pixel (the position of each pixel) of the projection pattern image can be encoded into a plurality of projection patterns. By decoding the projection patterns included in each image captured with each projection pattern of the structured light 60 projected, it is possible to calculate which pixel of the projection pattern image each pixel of the captured image corresponds to. In this embodiment, based on the plurality of predetermined projection patterns of the structured light 60, it is possible to calculate which pixel of the projection pattern image each pixel of the images captured by the first camera 3 and the second camera 4 corresponds to. Therefore, using these plurality of projection patterns as clues, it is possible to associate a first pixel in the first image captured by the first camera 3 with a second pixel in the second image captured by the second camera 4, which correspond to the same point on the object surface. However, as long as the structured light 60 can correspond to the first pixel and the second pixel, it does not necessarily have to have multiple projection patterns made up of multiple images, and may have multiple projection patterns made up of a single image.

[0026] In this embodiment, the structured light 60 may include a plurality of linear patterns that are each extending in the horizontal direction and arranged side by side in the vertical direction. In this case, the projection pattern of the structured light 60 has a predetermined spatial frequency in the vertical direction.

[0027] In this embodiment, the structured light 60 may include a plurality of linear patterns, each extending vertically and arranged horizontally, and a plurality of linear patterns, each extending horizontally and arranged vertically. In this manner, if the camera orientation is accurate, corresponding points can be determined by determining the former plurality of linear patterns and the epipolar line. However, even if the camera orientation is not so accurate, the accuracy of the correspondence between the first pixel and the second pixel can be improved by using both the former plurality of linear patterns and the latter plurality of linear patterns. Note that the structured light 60 may include linear patterns extending diagonally instead of or in addition to the linear patterns extending vertically and horizontally, or may include patterns other than linear patterns, such as circular patterns and curved patterns. Furthermore, the pattern projected by the structured light 60 is not limited to a pattern in which a predetermined geometric pattern is repeated, and may be any pattern that can be encoded as described above.

[0028] The structured light 60 will be described in detail later with reference to FIG.

[0029] The projection unit 6 projects structured light 60 onto the real object 10 so that each of the first camera 3 and the second camera 4 can capture the structured light 60 projected onto the real object 10. For example, as shown in FIG. 1 , the projection unit 6 is disposed between the first camera 3 and the second camera 4, and projects the structured light 60 onto the real object 10 from this position. The projection unit 6 is preferably disposed in accordance with the installation positions of the first camera 3 and the second camera 4, particularly so that shadows of the structured light 60 projected onto the surface of the object 10 disappear in the images captured by the first camera 3 and the second camera 4.

[0030] The projection unit 6 projects structured light 60 onto the real object 10 in synchronization with the photographing by each of the first camera 3 and the second camera 4. For example, the projection unit 6 projects the structured light 60 onto the real object 10 in synchronization with the exposure in the photographing by each of the first camera 3 and the second camera 4. By synchronizing the projection by the projection unit 6 with the photographing by each of the first camera 3 and the second camera 4, the 3D reconstruction system 100 can acquire a photographed image of the real object 10 onto which a desired pattern is projected. Note that in order to eliminate the need to strictly synchronize the projection of the structured light 60 with the photographing by each camera, the photographing by each camera may be performed as a video rather than a still image.

[0031] The projection unit 6 is, for example, a projector capable of projecting an image including a predetermined pattern. Alternatively, the projection unit 6 may be configured by combining a display unit, such as a liquid crystal panel or an organic EL (Electro Luminescence) panel, that displays the predetermined pattern with a light source that irradiates the display unit with light. The projection unit 6 may also be configured by including one or more combinations of one or more projectors, display units, and light sources. From the perspective of projecting structured light 60 having a plurality of patterns, it is preferable that the projection unit 6 be capable of easily changing the projection pattern.

[0032] The control device 70 controls the operation of the three-dimensional reconstruction system 100. Specifically, the control device 70 controls the projection operation by the projection unit 6, the photographing operation of the object 10 by the first camera 3 and the second camera 4, etc. The control device 70 also reconstructs the object 10 by generating a three-dimensional model of the object 10 based on the photographed image of the object 10. In this embodiment, the control device 70 is configured to output reconstruction information of the object 10 based on a plurality of RGB images photographed by the first camera 3 and the second camera 4.

[0033] In addition to the above components, the three-dimensional reconstruction system 100 may further include other components, such as an illumination unit that illuminates the object 10. The illumination unit is preferably positioned according to the installation positions of the first camera 3 and the second camera 4, particularly so that there is no shadow on the surface of the object 10 in the images captured by the first camera 3 and the second camera 4.

[0034] [First embodiment] <Example of functional configuration of control device 70> 2 is a block diagram showing an example of the functional configuration of the control device 70 according to the first embodiment. The control device 70 has a projection control unit 701, an imaging control unit 702, an attitude acquisition unit 703, and a model generation unit 704. Each of these functions may be realized by a processor such as a CPU (Central Processing Unit) or an electric circuit, or may be realized by one or more electric circuits or one or more processors. Furthermore, each of the above functions may be realized by distributed processing between the control device 70 and components other than the control device 70.

[0035] The projection control unit 701 controls the operation of the projection unit 6. For example, the projection control unit 701 controls the operation of the projection unit 6 so as to project structured light having a predetermined pattern onto the object 10 in synchronization with the capture of images by the first camera 3 and the second camera 4, respectively.

[0036] The photographing control unit 702 controls the operations of the rotation unit 5, the first camera 3, and the second camera 4 so that the first camera 3 and the second camera 4 photograph the object 10 and acquire multiple images during rotation by the rotation unit 5. The photographing control unit 702 may also control lighting.

[0037] The attitude acquisition unit 703 acquires attitude information of the first camera 3 and the second camera 4 that has been acquired in advance and stored in memory.

[0038] The model generation unit 704 reconstructs the object 10 by generating a three-dimensional model of the object 10 based on multiple images of the object 10 captured by the first camera 3 and the second camera 4 and information about the orientations of the first camera 3 and the second camera 4 acquired by the orientation acquisition unit 703. The model generation unit 704 outputs the generated three-dimensional model of the object 10 as reconstruction information of the object 10. The three-dimensional model may represent only the object 10 or may represent a scene including the object 10. In this embodiment, the model generation unit 704 reconstructs the object 10 using a machine learning method based on the error (i.e., loss) between an image obtained by rendering the three-dimensional information and the captured image.

[0039] Furthermore, in this embodiment, the model generation unit 704 acquires images of the actual object 10 onto which the structured light 60 is projected, for at least two viewpoints including a first viewpoint (e.g., a viewpoint including the first line-of-sight vector v1 in FIG. 1 ) and a second viewpoint (e.g., a viewpoint including the second line-of-sight vector v2 in FIG. 1 ). Based on the projection pattern of the structured light 60, the model generation unit 704 identifies a first pixel in the image from the first viewpoint and a second pixel in the image from the second viewpoint, which correspond to the same point on the surface of the actual object 10 onto which the structured light 60 is projected. Based on a current parameter θ in an object shape neural network (hereinafter referred to as an object shape NN (Neural Network)) that reconstructs the three-dimensional shape of the object 10, the model generation unit 704 calculates a first coordinate corresponding to the first pixel and a second coordinate corresponding to the second pixel on the surface of the object 10 reconstructed by the object shape NN. The first coordinate and the second coordinate are three-dimensional coordinates. The model generation unit 704 updates the parameters of the object shape NN using at least the first coordinate and the second coordinate. That is, as will be described later, model generation unit 704 updates the parameters of the object shape NN so that, with regard to the three-dimensional shape of object 10 reconstructed by the object shape NN, a point on the surface of object 10 corresponding to a first pixel coincides with a point on the surface of object 10 corresponding to a second pixel. Note that "updating the parameters of the object shape NN using the first coordinates and the second coordinates" conceptually includes directly incorporating information on the three-dimensional coordinates of the first coordinates and the second coordinates into a loss function related to updating the parameters of the object shape NN, and incorporating other information obtained based on the three-dimensional coordinates of the first coordinates and the second coordinates into a loss function related to updating the parameters of the object shape NN. For example, this concept includes incorporating information on two-dimensional coordinates obtained by projecting the three-dimensional coordinates of the first coordinates and the second coordinates onto a two-dimensional plane into a loss function related to updating the parameters of the object shape NN.

[0040] Furthermore, in this embodiment, the model generation unit 704 may further acquire images of the actual object 10 from at least two viewpoints onto which structured light 60 is not projected. The model generation unit 704 further calculates the color of the object 10 represented by the object shape NN and the object color NN, based on the current parameter θ in the object shape NN and the current parameter φ in an object color neural network (hereinafter referred to as object color NN) that represents the color of the object 10. The model generation unit 704 further updates the parameters of the object shape NN and the object color NN using information about the calculated color of the object 10. By using images of the actual object 10 onto which structured light 60 is not projected to update the object shape NN, it is possible to learn the shape of an object 10 made of a transparent or metallic material onto which structured light 60 is difficult to project well.

[0041] <Example of functional configuration of model generation unit 704> Next, the functional configuration of the model generation unit 704 will be described with reference to FIGS. 3 to 9. FIG. 3 is a block diagram showing an example of the functional configuration of the model generation unit 704. FIG. 4 is a diagram showing an example of a projection pattern of structured light 60. FIG. 5 is a diagram showing an example of images captured by the first camera 3 and the second camera 4. FIG. 6 is a diagram showing an example of an image captured of an object onto which structured light 60 is not projected. FIG. 7 is a diagram illustrating an example of a three-dimensional scene 55. FIG. 8 is a diagram illustrating an example of a result of identifying corresponding pixels in the first image 30 and the second image 40. FIG. 9 is a diagram illustrating an example of the relationship between corresponding pixels and corresponding coordinates.

[0042] As shown in FIG. 3, the model generation unit 704 includes an acquisition unit 41, an object data storage unit 42, a rendering unit 43, a corresponding pixel identification unit 44, a corresponding coordinate calculation unit 45, a model update unit 46, and a model output unit 47.

[0043] The acquisition unit 41 acquires images of the actual object 10 onto which the structured light 60 is projected, for at least two viewpoints including a first viewpoint and a second viewpoint. Specifically, the acquisition unit 41 acquires the first image 30 captured by the first camera 3 as the image of the first viewpoint via the imaging control unit 702 shown in FIG. 2. The acquisition unit 41 also acquires the second image 40 captured by the second camera 4 as the image of the second viewpoint via the imaging control unit 702.

[0044] 4, structured light 60 includes N different projection patterns consisting of structured light 60-1, structured light 60-2, ..., and structured light 60-N. Structured light 60-1, structured light 60-2, ..., and structured light 60-N each include a linear pattern extending in the vertical direction, and the spatial frequencies of the linear patterns are different from one another. The spatial frequencies of the linear patterns increase in the order of structured light 60-1, structured light 60-2, ..., and structured light 60-N.

[0045] The projection unit 6 synchronizes with the photographing by the first camera 3 and the second camera 4 and sequentially projects structured light 60-1, structured light 60-2, ..., and structured light 60-N onto the object 10. The first camera 3 photographs first images 30 (30-1, 30-2, ..., 30-N) of the object 10 onto which structured light 60 (60-1, 60-2, ..., 60-N) of different projection patterns is projected. The second camera 4 photographs second images 40 (40-1, 40-2, ..., 40-N) of the object 10 onto which structured light 60 (60-1, 60-2, ..., 60-N) of different projection patterns is projected.

[0046] 5, the first image 30 includes first image 30-1, first image 30-2, ..., and first image 30-N. First image 30-1 is a captured image of object 10 with structured light 60-1 projected thereon. First image 30-2 is a captured image of object 10 with structured light 60-2 projected thereon. First image 30-N is a captured image of object 10 with structured light 60-N projected thereon. The same applies to the other first images 30.

[0047] The second images 40 include second images 40-1, 40-2, ..., and 40-N. The second image 40-1 is a captured image of the object 10 with structured light 60-1 projected thereon. The second image 40-2 is a captured image of the object 10 with structured light 60-2 projected thereon. The second image 40-N is a captured image of the object 10 with structured light 60-N projected thereon. The same applies to the other second images 40.

[0048] In this specification, an example is shown in which each of the first image 30 and the second image 40 is an RGB image, but this is not limiting, and each of the first image 30 and the second image 40 may be a monochrome image or a grayscale image. Each of the first image 30 and the second image 40 may be an image having three or more gradations per pixel.

[0049] In the example shown in this specification, the three-dimensional reconstruction system 100 intermittently rotates the support table 2 shown in FIG. 1 at predetermined angle intervals (for example, by rotating 360° in 30° increments), thereby intermittently rotating the object 10 on the support table 2 at predetermined angle intervals. The three-dimensional reconstruction system 100 stops the rotation of the object 10 at each predetermined angle, and then captures images of the object 10 onto which N patterns of structured light 60 are projected, using the first camera 3 and the second camera 4. The acquisition unit 41 acquires N first images 30 and N second images 40 at each of the predetermined angles.

[0050] The three-dimensional reconstruction system 100 can capture images of the object 10 from multiple capture directions by rotating the object 10 and capturing images at predetermined angles. That is, the three-dimensional reconstruction system 100 can acquire captured images by increasing the number of viewpoints of the object 10 in addition to the two viewpoints of the first camera 3 and the second camera 4. Increasing the number of viewpoints can improve the accuracy of the three-dimensional reconstruction. However, the three-dimensional reconstruction system 100 may increase the number of viewpoints of the object 10 by increasing the number of cameras without rotating the object 10, or by intermittently rotating the cameras around the object 10. Furthermore, the three-dimensional reconstruction system 100 may capture images by continuously rotating the object 10, rather than by intermittent rotation.

[0051] In this embodiment, the acquisition unit 41 may further acquire images of the actual object 10, for at least one viewpoint, onto which structured light is not projected. In FIG. 6, the third image 50 is a captured image of the actual object 10 onto which structured light is not projected. The third image 50 is an "image for at least one viewpoint" acquired by at least one of the first camera 3 and the second camera 4 via the image capture control unit 702. The third image 50 is a color image (e.g., an RGB image). In the example shown in this specification, the three-dimensional reconstruction system 100 continuously rotates the support table 2, thereby continuously rotating the object 10 on the support table 2. The three-dimensional reconstruction system 100 captures images of the object 10 onto which structured light is not projected, at predetermined angles, using at least one of the first camera 3 and the second camera 4 during the continuous rotation. The acquisition unit 41 acquires the third image 50 at the predetermined angles.

[0052] The three-dimensional reconstruction system 100 can reconstruct not only the shape of the object 10 but also the color of the object 10 by acquiring the third image 50 in addition to the first image 30 and the second image 40. Furthermore, by acquiring the third image 50 in addition to the first image 30 and the second image 40, the three-dimensional reconstruction system 100 can improve the accuracy of reconstructing the shape of the object 10 compared to when reconstruction is based only on the image of the object 10 onto which the structured light 60 is projected.

[0053] In addition to the above, acquisition unit 41 acquires the orientation information of first camera 3 and second camera 4 acquired by orientation acquisition unit 703 in FIG.

[0054] The acquisition unit 41 outputs the first image 30, the second image 40, the third image 50, and information relating to the orientations of the first camera 3 and the second camera 4, respectively, to the rendering unit 43 and the corresponding pixel identification unit 44.

[0055] The object data storage unit 42 in Fig. 3 stores data representing the shape and color of the object 10. An example of the data representing the shape of the object 10 will be described with reference to Fig. 7. Fig. 7 shows a three-dimensional scene 55 including a first camera 3, the object 10, a line-of-sight vector v, a camera view 51, etc. Note that although the first camera 3 is shown here as an example, it can be replaced with the second camera 4.

[0056] An example of data representing the shape of the object 10 is a signed distance field (SDF), which represents a field formed corresponding to each location in a three-dimensional scene. Each point included in the SDF contains information about the signed distance from that point to the nearest surface of the object 10. The distance d in FIG. 7 represents this signed distance. The signed distance takes a negative (-) value for points inside the object 10 and a positive (+) value for points outside the object 10. The zero isosurface of the SDF represents the surface of the object 10.

[0057] The line of sight vector v corresponds to the line of sight from the camera and is also called a ray. The camera view 51 is an image generated (rendered) as an image seen from the first camera 3. The color of a pixel at position P in the camera view 51 is determined by the color of light (point color) emitted from each point in the three-dimensional scene on the line of sight vector v from the camera that passes through that pixel. The color of this light can be calculated by a weighted sum, which is an accumulation of color information at each three-dimensional coordinate on the line of sight vector v.

[0058] The sampling points Sm represent sampling points of pixel colors on the gaze direction vector v. A weight is assigned to each sampling point Sm. The weight increases as the sampling point Sm approaches the surface of the object 10. The weight assigned to a sampling point Sm may be expressed as a function of the SDF calculated for that sampling point Sm. The area Ar indicates an area where the weight is relatively large compared to other areas.

[0059] The SDF may be approximately represented as a neural field by a neural network. For example, when a point x (three-dimensional coordinates: x=(x, y, z)) in a three-dimensional scene 55 is input, this neural network outputs a signed distance d from the point x to the surface of the object 10. This neural network corresponds to an object shape NN. The object shape NN can represent the zero isosurface of the SDF, i.e., the object surface, and therefore corresponds to data representing the shape of the object 10.

[0060] An example of data representing the color of the object 10 is a vector field (color field) representing the color (e.g., RGB value) of each point x in the three-dimensional scene 55, conditioned by the line of sight. This vector field may be approximately represented as a neural field by a neural network. For example, when a point x in the three-dimensional scene and a line of sight vector v for that point x are input, this neural network outputs the color c of the pixel corresponding to that point x. This neural network corresponds to an object color NN. The object color NN corresponds to data representing the color of the object 10.

[0061] 7, the rendering unit 43 in Fig. 3 generates the camera view 51 in Fig. 7 using the object shape NN and object color NN stored in the object data storage unit 42 as data representing the shape and color of the object 10. In this embodiment, the camera view 51 may be generated for each pixel according to a volume rendering method, for example, using the method of the NeuS (Neural Surface Reconstruction Method) paper (see, for example, Non-Patent Document 1).

[0062] Based on the multiple projection patterns of the structured light 60, the corresponding pixel identification unit 44 identifies a first pixel in the first image 30 from the first viewpoint and a second pixel in the image from the second viewpoint that correspond to the same point on the surface of the actual object 10 onto which the structured light 60 is projected.

[0063] For example, a case will be described with reference to FIG. 5 where the first pixel Pa0 in the first image 30 and the second pixel Pb0 in the second image 40 are pixels obtained by capturing an image of the same point (same position) on the surface of the actual object 10. The luminance of the first pixel Pa0 in the first image 30 changes depending on the projection pattern of the structured light 60. That is, the first pixel Pa0 has a dark luminance in the first image 30-1, a dark luminance in the first image 30-2, ..., and a bright luminance in the first image 30-N. On the other hand, the second pixel Pb0 in the second image 40, like the first pixel Pa0, has a dark luminance in the second image 40-1, a dark luminance in the second image 40-2, ..., and a bright luminance in the second image 40-N. The corresponding pixel identification unit 44 identifies multiple pixels (e.g., the first pixel Pa0 and the second pixel Pb0) that represent the same point (same position) on the actual object 10 in multiple captured images (e.g., the first image 30 and the second image 40) as corresponding pixels.

[0064] The above-mentioned dark luminance refers to the luminance of a pixel in the captured image that corresponds to a position on the object 10 onto which a dark pattern in the structured light 60 is projected. The bright luminance refers to the luminance of a pixel in the captured image that corresponds to a position on the object 10 onto which a bright pattern in the structured light 60 is projected. Pixels with dark luminance and pixels with bright luminance may be called dark pixels and bright pixels, respectively.

[0065] In the three-dimensional reconstruction system 100, the encoding process described above corresponds to a process of expressing the position of each pixel as a combination of dark luminance and bright luminance using structured light 60-1, structured light 60-2, ..., structured light 60-N. On the other hand, the decoding process described above corresponds to a process of acquiring the position of each pixel included in the captured image based on the combination of dark luminance and bright luminance using multiple projection patterns.

[0066] 8, the corresponding pixel identification unit 44 can identify corresponding pixels such as the first pixel Pa and the second pixel Pb from all pixels in the first image 30 and the second image 40 that reflect the structured light 60 projected onto the actual object 10. The corresponding line P0 in FIG. 8 is a line that connects corresponding pixels such as the first pixel Pa0 and the second pixel Pb0, and visualizes the corresponding pixels for the sake of explanation.

[0067] The corresponding coordinate calculation unit 45 in Fig. 3 will be described with reference to Fig. 9. In Fig. 9, when a point x is given in a three-dimensional scene 55, a function S(x) that returns the signed distance between the point x and the surface of the object 10 that is closest to the point x can be approximated by a function S(x; θ) that is expressed by an object shape NN having a parameter θ. The function S(x) corresponds to the above-mentioned SDF, and the function S(x; θ) corresponds to an approximation of the SDF. The coordinates of the point x when S(x) = 0 correspond to the coordinates on the surface of the object 10.

[0068] The first pixel Pa represents a pixel in the first image 30 captured by the first camera 3. The second pixel Pb represents a pixel in the second image 40 captured by the second camera 4. The first pixel Pa and the second pixel Pb are pixels that have been associated with each other by the corresponding pixel identification unit 44.

[0069] The first pixel Pa and the second pixel Pb are images of the same point on the structured light 60 projected onto the surface of the object 10. Therefore, the first coordinate xa of the intersection of the surface of the object 10 and the line of sight vector passing through the first pixel Pa, expressed by S(x;θ)=0, and the second coordinate xb of the intersection of the surface of the object 10 and the line of sight vector passing through the pixel Pb, should be the same coordinate. Therefore, the distance loss L is an index (term) whose value decreases as the distance |xa-xb|, i.e., the first coordinate xa and the second coordinate xb, become closer. DBy optimizing the parameter θ in the function S(x;θ) of the object shape NN using a loss function including the following, the zero isosurface of S(x;θ) moves in the direction in which the first coordinate xa and the second coordinate xb coincide, and it is possible to obtain a function S(x;θ) that accurately represents the shape of the surface of the object 10. Note that the distance loss L D can be the sum of distances for all corresponding pixels. Here, the loss function refers to a function that may include, in addition to the losses represented by the above-mentioned indices, the error (i.e., loss) between the captured image and the image (camera view) obtained by rendering based on the 3D information (object shape and color) represented by the object shape NN and object color NN. Note that in FIG. 9, the first coordinate xa' represents the first coordinate xa after the zero isosurface of S(x;θ) is shifted to make the first coordinate xa and the second coordinate xb the same coordinate. Similarly, the second coordinate xb' represents the second coordinate xb after the zero isosurface of S(x;θ) is shifted to make the first coordinate xa and the second coordinate xb the same coordinate.

[0070] The set of points x that satisfy the function S(x;θ) = 0 represents the surface of the object, so the intersection point x between the ray from the camera with the pose ο and the surface of the object can be calculated by the following equation (1).<URL:https: / / arxiv.org / pdf / 2003.09852.pdf> (see [1]). The ray is a ray that passes from the center of the camera through pixel P in the camera view, and its line of sight vector is v. This equation (1) is found by using the implicit differential of the equation S(x;θ) = 0, where θ0 is the current parameter and τ0 is the camera parameter τ0 (τ0 is known in this embodiment), which is information about the camera's posture, and where ο0 = ο0(τ0), t0 = t(θ0, τ0), ν0 = ν(τ0), and x0 = ο0 + t0 · ν0t. Note that instead of using equation (1), multiple sample points x may be set on the ray, S(x;θ) may be calculated for each sample point x, and the sample point x that can be considered as S(x;θ) ≒ 0 may be found as the intersection of the ray and the surface of the object.

[0071]

number

[0072] 3 calculates a first coordinate xa corresponding to a first pixel Pa and a second coordinate xb corresponding to a second pixel Pb on the surface of the object 10 reconstructed by the object shape NN, based on the current parameter θ in the object shape NN that reconstructs the three-dimensional shape of the object 10. The corresponding coordinate calculation unit 45 outputs the calculated first coordinate xa and second coordinate xb to the model update unit 46.

[0073] The model update unit 46 updates the parameter θ of the object shape NN using at least the first coordinate xa and the second coordinate xb. The model update unit 46 also updates the parameters of the object shape NN and the object color NN stored in the object data storage unit 42 using information about the calculated color of the object 10. For example, for each of the multiple camera views 51 in FIG. 7 generated by the rendering unit 43, the model update unit 46 calculates the loss L between the camera view 51 and the actually captured third image 50. R Calculate (θ, φ, τ) and calculate the loss L R can be calculated by accumulating the color differences between corresponding pixels for all pixels between the camera view 51 and the third image 50, for example. R The definition of loss L is not limited to the above-mentioned cumulative color difference between pixels for all pixels, but may be other definitions. R The calculation method of is not limited to the above-mentioned method of cumulatively calculating the color difference between pixels for all pixels, and other methods may be used. When reconstructing only the shape of the object 10 without reconstructing the color of the object 10, the model update unit 46 calculates the loss L between the camera view 51 and either the first image 30 or the second image 40 that was actually captured. R may be calculated.

[0074] In this embodiment, the model update unit 46 updates the above loss L R The distance loss L D , the weight λD and calculate the loss function L(θ, φ, τ) using the following equation. L(θ, φ, τ)=L R +λ D L D

[0075] The loss function L(θ, φ, τ) is given by the SDF regularization constraint E(||∇S(x;θ)||-1) 2 with a given coefficient λ E It is also possible to add weights with E. Note that E means taking the expectation value with respect to x, and ∇ means taking the gradient with respect to x.

[0076] The model update unit 46 updates each parameter of the object shape NN and object color NN by backpropagation based on the loss function L(θ, φ, τ). The model update unit 46 repeats updating each NN until a predetermined criterion is met. As a result of updating by the model update unit 46, the object shape NN and object color NN finally obtained represent the three-dimensional shape and color of the actual object 10, i.e., a three-dimensional model. The predetermined criterion may be, for example, until the loss L(θ, φ, τ) becomes less than a predetermined value. Alternatively, the predetermined criterion may be until a predetermined number of updates have been performed.

[0077] For example, the method in the NeuS paper reconstructs the 3D shape of an object by optimizing the parameter θ of the object shape neural network based on images captured from multiple viewpoints. However, with the method in the NeuS paper, if the object contains at least a portion of a monochromatic textureless area, the accuracy of the 3D reconstruction may be reduced due to the inability to uniquely determine a solution (an ill-posed problem).

[0078] In this embodiment, a first image 30 and a second image 40 of the object 10 onto which the structured light 60 is projected are used to calculate a distance loss L Dis included in the loss function L(θ, φ, τ) to optimize the parameter θ. This adds constraints that bring the object closer to a desired solution, even when the object includes at least a portion of a monochromatic textureless region, making it possible to increase the accuracy of 3D reconstruction. As a result, this embodiment can provide a 3D reconstruction method and a 3D reconstruction system 100 that can increase the accuracy of reconstruction of the object 10.

[0079] The model update unit 46 stores the three-dimensional model obtained as a result of the update as reconstruction information in the object data storage unit 42. The model generation unit 704 can also output the reconstruction information to an external device such as a display device or a PC (Personal Computer).

[0080] The model output unit 47 uses the data stored in the object data storage unit 42 to generate three-dimensional mesh data representing the shape of the three-dimensional model of the object 10 and a texture map to be applied to this three-dimensional mesh data. The model output unit 47 converts these into a data format viewable by a three-dimensional viewer and outputs them.

[0081] In this embodiment, the finally obtained object shape NN represents an SDF, and the model output unit 47 generates a mesh that approximates the zero isosurface of this SDF. The model output unit 47 can convert the zero isosurface of the SDF into mesh data using, for example, the marching cubes method or other known methods. Furthermore, for each polygon constituting the mesh, the model output unit 47 inputs the polygon's three-dimensional coordinates and the inverse vector of the normal vector of the SDF's zero isosurface at these three-dimensional coordinates to the object color NN. The model output unit 47 obtains the polygon color as an output from the object color NN. The model output unit 47 can generate a texture map for the mesh using the obtained polygon color.

[0082] The method of acquiring the polygon colors is not limited to acquiring one color per polygon. For example, the method of acquiring the polygon colors may be to acquire one color for each of a plurality of unit areas obtained by further dividing one polygon. In this case, the spatial resolution can be improved compared to acquiring one color per polygon. For example, when the reconstructed object 10 is displayed as an image, the unit area corresponds to one pixel.

[0083] In this embodiment, the model update unit 46 may update only the object shape NN without updating the object color NN. Also, the 3D reconstruction system 100 may reconstruct only the object shape without reconstructing the object color. This also has the effect of providing a 3D reconstruction method and a 3D reconstruction system 100 that can improve the reconstruction accuracy of the object 10.

[0084] <Example of processing by the control device 70> Fig. 10 is a flowchart showing an example of processing by the control device 70. For example, the control device 70 starts the processing of Fig. 10 when it receives an operation input of a reconstruction start instruction from a user via its operation unit. Before the processing of Fig. 10 starts, an object 10 is placed on the upper surface 2A of the support table 2 in Fig. 1. The position where the object 10 is placed is preferably the center of rotation of the support table 2. In addition, in this embodiment, it is assumed that external parameters such as attitude information of each of the first camera 3 and the second camera 4 with respect to the object 10 have been acquired in advance.

[0085] First, in step S91, the control device 70 controls the operation of the projection unit 6 using the projection control unit 701, thereby projecting structured light 60 from the projection unit 6 onto the object 10. At this time, the support base 2 is not rotating, and the object 10 placed on the upper surface 2A of the support base 2 is stationary. The projection unit 6 continues to project the structured light 60 until it is stopped by the projection control unit 701.

[0086] Next, in step S92 (step of capturing the first and second images), the control device 70 controls the operation of each of the first camera 3 and the second camera 4 using the capturing control unit 702, thereby capturing images of the stationary object 10 using each of the first camera 3 and the second camera 4, and obtaining the first image 30 and the second image 40 of the object 10.

[0087] Next, in step S93, the control device 70 determines whether or not the object 10 has been photographed using all of the predetermined projection patterns included in the structured light 60. In other words, the control device 70 determines whether or not the object 10 has been photographed by projecting all of the N projection patterns.

[0088] If it is determined in step S93 that no image has been captured (step S93, NO), in step S94, the control device 70 changes the projection pattern of the structured light 60 from the projection unit 6 by controlling the operation of the projection unit 6 with the projection control unit 701. After changing the projection pattern, the control device 70 performs the processes from step S91 onwards again.

[0089] On the other hand, if it is determined in step S93 that an image has been captured (step S93, YES), in step S95, the control device 70 controls the operation of the rotation unit 5 with the imaging control unit 702 to rotate the support base 2 by a predetermined angle (for example, 30°), thereby rotating the object 10 on the upper surface 2A. After the support base 2 has rotated by the predetermined angle, the control device 70 controls the operation of the rotation unit 5 with the imaging control unit 702 to stop the rotation unit 5 and stop the rotation of the support base 2.

[0090] Subsequently, in step S96, the control device 70 determines whether or not the object 10 has been photographed from all photographing directions (for example, 12 viewpoints obtained by dividing a full 360° circumference into 30° increments).

[0091] If it is determined in step S96 that an image has not been captured (step S96, NO), the control device 70 repeats the processes from step S91 onwards. On the other hand, if it is determined in step S96 that an image has been captured (step S96, YES), the control device 70 controls the operation of the projection unit 6 by the projection control unit 701 in step S97, thereby stopping the projection of structured light 60 from the projection unit 6 onto the object 10.

[0092] Subsequently, in step S98, the control device 70 controls the operation of the rotation unit 5 by the imaging control unit 702 to start continuous rotation of the support base 2 (for example, continuous rotation through one revolution of 360°). The continuous rotation of the support base 2 causes the object 10 placed on the upper surface 2A of the support base 2 to continuously rotate. The rotation unit 5 continues to rotate the support base 2 until it is stopped by the imaging control unit 702.

[0093] Next, in step S99 (step of capturing a third image), the control device 70 controls the operation of at least one of the first camera 3 and the second camera 4 using the image capturing control unit 702 to capture an image of the rotating object 10 and acquire a third image 50. The first camera 3 captures an image of the object 10 every time the rotating support base 2 rotates a predetermined angle (for example, 30°), and acquires a predetermined number of third images 50 (for example, 12 images obtained by dividing a 360° rotation into 30° increments), i.e., images of the object 10 onto which the structured light 60 is not projected.

[0094] When the third image 50 is captured by each of the first camera 3 and the second camera 4, the capture control unit 702 causes the first camera 3 and the second camera 4 to capture images at approximately the same timing when the object 10 is rotating. This allows the capture control unit 702 to acquire RGB images from each camera in parallel when the object 10 is at a position of any rotation angle. However, the captures by the first camera 3 and the second camera 4 do not necessarily have to be captured at approximately the same timing, and may be intentionally staggered.

[0095] In step S99, the photography control unit 702 continues the photography operation to acquire images while the rotation unit 5 rotates the object 10 once. This allows at least one of the first camera 3 and the second camera 4 to photograph the object 10 from directions tilted at different angles relative to the upper surface 2A and from multiple directions along the rotation direction. Each camera can acquire an RGB image in each photography direction. In other words, the photography control unit 702 can photograph the object 10 simultaneously from different directions using multiple cameras. The photography control unit 702 outputs the third image 50 to the model generation unit 704.

[0096] Subsequently, in step S100, the control device 70 controls the operation of the rotation unit 5 by the imaging control unit 702, thereby stopping the rotation unit 5 and stopping the continuous rotation of the support base 2.

[0097] Next, in step S101 (camera view generation step), the control device 70 causes the model generation unit 704 to generate a camera view 51 for each shooting direction.

[0098] Subsequently, in step S102 (corresponding pixel specifying step), the control device 70 causes the corresponding pixel specifying unit 44 of the model generating unit 704 to specify corresponding pixels for each imaging direction.

[0099] Subsequently, in step S103 (corresponding coordinate calculation step), the control device 70 causes the corresponding coordinate calculation unit 45 of the model generation unit 704 to calculate corresponding coordinates for each imaging direction.

[0100] Subsequently, in step S104, the control device 70 calculates the loss function L(θ, φ, τ) by the model update unit 46 of the model generation unit 704. After the calculation is completed, the model update unit 46 determines whether the loss function L(θ, φ, τ) is less than a predetermined value.

[0101] If it is determined in step S104 that the difference is not less than the predetermined value (step S104, NO), in step S105 (update step), the control device 70 calculates the loss function L(θ, φ, τ) for each shooting direction using the model update unit 46 of the model generation unit 704. The model update unit 46 updates each parameter of the object shape NN and the object color NN using the backpropagation algorithm based on this loss function L(θ, φ, τ). Thereafter, the control device 70 performs the processes from step S101 onwards again. The control device 70 repeats updating each NN using the model update unit 46 until a predetermined criterion is met.

[0102] On the other hand, if it is determined in step S104 that the difference is less than the predetermined value (step S104, YES), in step S106 (output step), the control device 70 causes the model output unit 47 of the model generation unit 704 to generate a 3D model (3D mesh and texture map) based on the object shape NN and object color NN finally obtained as a result of the update, and outputs the generated model as reconstruction information. When the processing of step S106 is completed, this processing flow ends. Note that, because the corresponding pixels do not change during the repeated processing of steps S101 to S105, the processing of step S102 to identify the corresponding pixels may be executed only once outside the repeated processing of steps S101 to S105. For example, after executing steps S100 and S102, the processing of steps S101, S103, S104, and S105 may be repeated.

[0103] 10, the process by control device 70 may be such that, after performing the process of projecting structured light 60 and capturing an image multiple times, the process of capturing an image without projecting structured light 60 is repeated for each predetermined rotation angle. In other words, in the process by control device 70, the capturing processes of steps S97 and S99 may be inserted between steps S93 and S95 in FIG. 10, but the processes of steps S97 to S100 may be omitted.

[0104] <Example of Generation Process of Camera View 51 by Rendering Unit 43> Fig. 11 is a flowchart showing an example of the generation process of camera view 51 shown in step S101 of Fig. 10 by rendering unit 43 of model generation unit 704. For example, in the generation process of camera view 51 in model generation unit 704, rendering unit 43 starts the process of Fig. 11 when position P in camera view 51 is selected.

[0105] First, in step S111, the rendering unit 43 determines a line-of-sight vector v that passes through the pixel at the selected position P.

[0106] Subsequently, in step S112, the rendering unit 43 obtains the three-dimensional coordinates of each sample point x_i on the determined line-of-sight direction vector v.

[0107] Subsequently, in step S113, the rendering unit 43 obtains the color c_obj of each sample point x_i based on the line-of-sight vector v, the three-dimensional coordinates of each sample point x_i, and the object color NN.

[0108] Subsequently, in step S114, the rendering unit 43 acquires information on the signed distance d_i from each sample point x_i to the surface of the object 10 based on the three-dimensional coordinates of each sample point x_i and the object shape NN.

[0109] Subsequently, in step S115, the rendering unit 43 calculates the color weight w_i of each sample point x_i based on the signed distance d_i of each sample point x_i.

[0110] Next, in step S116, the rendering unit 43 calculates the weighted sum of the colors of each sample point based on the color c_obj of each sample point x_i and the color weight w_i of each sample point x_i, and obtains the final color c_obj of the pixel. In this case, the weights are normalized so that the sum is 1. Also, it is designed in advance that the weight becomes larger when the signed distance d_i is closer to 0, that is, closer to the surface of the object 10.

[0111] Next, in step S117, rendering unit 43 outputs color c_obj of the pixel at position P in camera view 51. When the processing of step S117 is completed, the processing of acquiring color c_obj of one pixel at position P in camera view 51 ends.

[0112] By performing the above processing for all pixels of the camera view 51, the rendering unit 43 can generate the camera view 51.

[0113] <Example of a three-dimensional model generation process by the model output unit 47> Fig. 12 is a flowchart showing an example of a process for generating a three-dimensional model (three-dimensional mesh and texture map) by the model output unit 47. The model output unit 47 starts the process of Fig. 12 when it is determined in step S104 of Fig. 10 that the difference is less than a predetermined value (step S104, YES). Note that this flowchart describes an example in which one color is calculated for one polygon, but as described above, it is also possible to calculate colors for each of multiple locations within one polygon.

[0114] First, in step S121, the model output unit 47 generates a three-dimensional mesh that approximates the shape of the object 10 based on the object shape NN.

[0115] Next, in step S122, the model output unit 47 calculates, for each polygon constituting the three-dimensional mesh, the inverse vector u of the normal vector of the polygon at the three-dimensional coordinate q that represents the position of the polygon. Note that the model output unit 47 may, for example, calculate the three-dimensional coordinate corresponding to the position of the center of gravity of the polygon and acquire the coordinate as the three-dimensional coordinate q of the polygon, or may acquire another coordinate value relating to the position of the polygon in three-dimensional space as the three-dimensional coordinate of the polygon.

[0116] Subsequently, in step S123, the model output unit 47 acquires the three-dimensional coordinate q of each polygon by, for example, the calculation described above. Note that the order of execution of steps S122 and S123 may be reversed. Steps S122 and S123 may also be executed in parallel.

[0117] Subsequently, in step S124, the model output unit 47 calculates and obtains the color c_p of each polygon based on the three-dimensional coordinates q of the polygon, the inverse vector u, and the object color NN.

[0118] Subsequently, in step S125, the model output unit 47 outputs the polygon color c_p acquired for each polygon in step S124. That is, in this step, the model output unit 47 generates a texture map to be applied to the three-dimensional mesh generated in step S121. When the processing of step S125 is completed, this processing flow ends.

[0119] [Second embodiment] Next, a description will be given of a three-dimensional reconstruction system according to the second embodiment. Note that the same names and symbols as those in the already described embodiments indicate the same or equivalent members or components, and detailed descriptions will be omitted as appropriate.

[0120] This embodiment differs from the first embodiment in that it simultaneously optimizes an object shape NN, an object color NN, and information related to the orientations of the first and second cameras in the 3D reconstruction system. More specifically, this embodiment differs from the first embodiment mainly in that it acquires an image of the actual object 10 at a first viewpoint using the first camera 3, acquires an image of the actual object 10 at a second viewpoint using the second camera 4, calculates a first point obtained by projecting the second coordinates onto the first camera 3, and a second point obtained by projecting the first coordinates onto the second camera 4, and further updates the parameters of the object shape NN and object color NN, and the camera parameters, using at least the first point and the second point.

[0121] <Example of functional configuration of the control device 70a> 13 is a block diagram showing an example of the functional configuration of a control device 70a included in a three-dimensional reconstruction system 100a according to this embodiment. The three-dimensional reconstruction system 100a has substantially the same configuration as the three-dimensional reconstruction system 100 according to the first embodiment, except that it has a control device 70a. The control device 70a has substantially the same configuration as the control device 70 in the first embodiment, except that it does not have an attitude acquisition unit 703 (see FIG. 2) and has a model generation unit 704a.

[0122] <Example of functional configuration of model generation unit 704a> The functional configuration of the model generation unit 704a will be described with reference to Fig. 14 to Fig. 16. Fig. 14 is a block diagram showing an example of the functional configuration of the model generation unit 704a. Fig. 15 is a diagram explaining an example of a first loss according to this embodiment. Fig. 16 is a diagram explaining an example of a second loss according to this embodiment.

[0123] 14, the model generation unit 704a includes an acquisition unit 41a, a geometric loss calculation unit 48, and a model update unit 46a. Except for this, the model generation unit 704a has substantially the same functional configuration as the model generation unit 704 according to the first embodiment.

[0124] The acquisition unit 41a outputs the first image 30, the second image 40, and the third image 50 acquired from the first camera 3 and the second camera 4 via the shooting control unit 702 in Fig. 13 to the rendering unit 43, the corresponding pixel identification unit 44, and the geometric loss calculation unit 48, respectively. Except for this, the acquisition unit 41a has almost the same functions as the acquisition unit 41 shown in Fig. 2.

[0125] The geometric loss calculation unit 48 calculates a geometric loss (for example, the first loss L) for optimizing the parameters of the object shape NN and the object color NN, and the information on the postures of the first camera 3 and the second camera 4. SR (θ, τ) or second loss L ST Calculate (θ, τ).

[0126] 15 and 16, like FIG. 9, show a point x in a three-dimensional scene 55, a function S(x;θ) of an object shape NN having a parameter θ, a first pixel Pa in a first image 30, a second pixel Pb in a second image 40, etc.

[0127] Since a set of points x that satisfy the function S(x;θ) = 0 represents the surface of an object, the intersection x of a ray from a camera with an orientation o with the surface of the object can be found by equation (1), as described above. The ray is a ray that passes from the center of the camera through pixel P in the camera view, and its line-of-sight vector is v. Equation (1) can be found by using the implicit derivative of equation S(x;θ) = 0, where o = o(τ), t = t(θ, τ), v = v(τ), and x = o + t · vt, where θ is the current parameter and τ is the camera parameter that represents the current camera orientation (in this embodiment, τ is the optimization target). Note that instead of using equation (1), multiple sample points x may be set on the ray, S(x;θ) may be calculated for each sample point x, and a sample point x that satisfies S(x;θ) ≈ 0 may be found as the intersection of the ray with the surface of the object.

[0128] 1st loss L SR is information about the loss obtained by projecting a point x on the surface of an object (function S(x;θ)=0) of a corresponding pixel in a captured image onto another captured image (i.e., the corresponding pixel in another captured image). Therefore, the first loss L SR 15, the first coordinate xa and the second coordinate xb are obtained by the corresponding pixel specifying unit 44 and the corresponding coordinate calculating unit 45 in the same manner as described with reference to FIG.

[0129] The geometric loss calculation unit 48 calculates a first point Q(xb, τa) obtained by projecting the second coordinate xb onto the first camera 3, and a second point Q(xa, τb) obtained by projecting the first coordinate xa onto the second camera 4. Note that τa and τb are camera parameters of the first camera 3 and the second camera 4, respectively. The geometric loss calculation unit 48 calculates the sum of the distance |Pa-Q(xb, τa)| and the distance |Pb-Q(xa, τb)| as the first loss LSR The first loss L SR may be the sum of the distances for all corresponding pixels.

[0130] 2nd loss L ST is information about the loss obtained using the shape of the object calculated by triangulation from the first image 30 and the second image 40. Therefore, the second loss L ST 16, the first coordinate xa and the second coordinate xb are obtained by the corresponding pixel specifying unit 44 and the corresponding coordinate calculating unit 45 in the same manner as described with reference to FIG.

[0131] The geometric loss calculation unit 48 calculates the intersection of rays Ra(τ) and Rb(τ) passing through corresponding pixels pa and pb in the images captured by the first camera 3 and the second camera 4, respectively, using triangulation. Here, the rays passing through corresponding pixels pa and pb, respectively, may not intersect due to an attitude error between the first camera 3 and the second camera 4. For this reason, the geometric loss calculation unit 48 calculates a point ya on the ray Ra(τ) that is closest to the ray Rb(τ), and a point yb on the ray Rb(τ) that is closest to the ray Ra(τ). Finally, the geometric loss calculation unit 48 calculates a second loss L that evaluates the distances between the four points (for example, the sum of the distance |xa-ya|, the distance |ya-yb|, and the distance |yb-xb|) so that the four points (xa, xb, ya, yb) are positioned in the same position. ST Calculate the second loss L ST may be the sum of the distances for all corresponding pixels.

[0132] In FIG. 14, the geometric loss calculation unit 48 calculates the first loss L SR and second loss L ST are output to the model update unit 46a.

[0133] The model update unit 46a further updates the parameters θ and φ of the object shape NN and the object color NN, and the camera parameters τa and τb, respectively, using at least the first point Q(xb, τa) and the second point Q(xa, τb). The initial values ​​of the camera parameters may be acquired by a known method (for example, a method of calculation based on the appearance of a predetermined marker included in the captured image, or a method of calculation based on feature point extraction from the captured image, such as SfM (Structure from Motion)). In the example shown in this specification, the model update unit 46a updates the loss L R (θ) and the first loss L SR and the second loss L ST Using these, the loss function La(θ, φ, τ) is calculated using the following equation. La(θ, φ, τ)=L R +λ SR L SR +λ ST L ST λ SR , λ ST is a predetermined coefficient. Note that, in this embodiment as well, a regularization constraint of the SDF may be added to the loss function La(θ, φ, τ) as in the first embodiment.

[0134] The model update unit 46a updates the parameters of the object shape NN and object color NN, and the camera parameters τa and τb by backpropagation based on the loss function La(θ, φ, τ). The function of the model update unit 46a is almost the same as the function of the model generation unit 704 in the first embodiment, except that the loss function La(θ, φ, τ) is used instead of the loss function L(θ, φ, τ).

[0135] <Example of processing by the control device 70a> Fig. 17 is a flowchart showing an example of processing by the control device 70a. For example, the control device 70a starts the processing of Fig. 17 when it receives an operation input of a reconstruction start instruction from a user via its operation unit. Before the processing of Fig. 17 starts, an object 10 is placed on the upper surface 2A of the support table 2 in Fig. 1. The position where the object 10 is placed is preferably the center of rotation of the support table 2.

[0136] This embodiment differs from the first embodiment in Fig. 10 in that external parameters such as posture information of each of the first camera 3 and the second camera 4 relative to the object 10 are not acquired in advance. On the other hand, the processes of steps S171 to S183 are the same as the processes of steps S91 to S103 in Fig. 10, and therefore redundant explanations will be omitted here.

[0137] In step S184, the control device 70a calculates the first loss L for optimizing the information about the postures of the first camera 3 and the second camera 4 by the geometric loss calculation unit 48 of the model generation unit 704a. SR and second loss L ST After completing the calculation, the geometric loss calculation unit 48 outputs the calculation result to the model update unit 46a of the model generation unit 704a.

[0138] Subsequently, in step S185, the control device 70a causes the model update unit 46a of the model generation unit 704a to calculate the loss function La(θ, φ, τ). After the calculation is completed, the model update unit 46a determines whether the loss function La(θ, φ, τ) is less than a predetermined value.

[0139] If it is determined in step S185 that the difference is not less than the predetermined value (step S185, NO), in step S186 (update step), the control device 70a calculates the loss function La(θ, φ, τ) for each shooting direction using the model update unit 46a of the model generation unit 704a. The model update unit 46a updates each parameter of the object shape NN and the object color NN using the backpropagation algorithm based on this loss function La(θ, φ, τ). Thereafter, the control device 70a performs the processes from step S181 onwards again. The control device 70a repeats the update of each NN using the model update unit 46a until a predetermined criterion is met.

[0140] On the other hand, if it is determined in step S185 that the difference is less than the predetermined value (step S185, YES), in step S187 (output step), the control device 70a causes the model output unit 47 of the model generation unit 704a to generate a 3D model (3D mesh and texture map) based on the object shape NN and object color NN finally obtained as a result of the update, and outputs it as reconstruction information. When the processing of step S187 is completed, this processing flow ends.

[0141] In this way, the 3D reconstruction system 100a simultaneously optimizes the object shape NN and information relating to the orientations of the first camera 3 and the second camera 4, and performs 3D reconstruction. As a result, in this embodiment, even if there is no known information relating to the orientations of the first camera 3 and the second camera 4, it is possible to optimize the information relating to the orientations of the first camera 3 and the second camera 4. As a result, this embodiment can provide a 3D reconstruction system and a 3D reconstruction method that can increase the reconstruction accuracy of an object.

[0142] In this embodiment, the geometric loss calculation unit 48 calculates the second loss L ST Without calculating the first loss L SR The model update unit 46a may calculate only the loss L R and distance loss L D and the first loss L SRThe loss function La(θ, φ, τ) may be calculated using the following equation. Furthermore, the model update unit 46a may update only the object shape NN without updating the object color NN. Furthermore, the 3D reconstruction system 100a may reconstruct only the object shape without reconstructing the object color. These also have the effect of providing a 3D reconstruction method and a 3D reconstruction system 100 that can improve the reconstruction accuracy of the object 10 even when there is no known information about the orientations of the first camera 3 and the second camera 4.

[0143] <Example of hardware configuration in the above-described embodiment> Some or all of the devices (3D reconstruction systems 100 and 100a) in the above-described embodiments may be configured as hardware, or may be configured as software (program) information processing executed by a CPU, GPU (Graphics Processing Unit), or the like. In the case of software information processing, software that realizes at least some of the functions of each device in the above-described embodiments may be stored in a non-transitory storage medium (non-transitory computer-readable medium) such as a CD-ROM (Compact Disc-Read Only Memory) or USB (Universal Serial Bus) memory, and the software information processing may be executed by reading it into a computer. The software may also be downloaded via a communication network. Furthermore, all or part of the software processing may be implemented in a circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), thereby executing the software information processing by hardware.

[0144] The storage medium that stores the software may be a removable medium such as an optical disk, or a fixed medium such as a hard disk, memory, etc. The storage medium may be provided inside the computer (main storage device, auxiliary storage device, etc.) or outside the computer.

[0145] 18 is a block diagram showing an example of the hardware configuration of each device (3D reconstruction systems 100 and 100a) in the above-mentioned embodiment. Each device may be realized as a computer 7 including, for example, a processor 71, a main storage device 72 (memory), an auxiliary storage device 73 (memory), a network interface 74, and a device interface 75, which are connected via a bus 76.

[0146] Although the computer 7 in FIG. 18 includes one of each component, it may also include multiple of the same component. Also, while FIG. 18 shows one computer 7, the software may be installed on multiple computers, and each of the multiple computers may execute the same or different parts of the software. In this case, a distributed computing configuration may be used in which each computer communicates via a network interface 74 or the like to execute processing. That is, each device in the above-described embodiment (the three-dimensional reconstruction systems 100 and 100a) may be configured as a system in which one or more computers execute instructions stored in one or more storage devices to realize its functions. Furthermore, it may also be configured such that information transmitted from a terminal is processed by one or more computers provided on a cloud, and the processing results are transmitted to the terminal.

[0147] The various calculations of each device (3D reconstruction systems 100 and 100a) in the above-described embodiments may be executed in parallel using one or more processors, or using multiple computers via a network. Furthermore, the various calculations may be distributed to multiple processing cores in a processor and executed in parallel. Furthermore, some or all of the processes, means, etc. disclosed herein may be realized by at least one of a processor and a storage device provided on a cloud that can communicate with computer 7 via a network. Thus, each device in the above-described embodiments may be implemented in the form of parallel computing using one or more computers.

[0148] The processor 71 may be an electronic circuit (processing circuit, processing circuitry, CPU, GPU, FPGA, ASIC, etc.) that at least controls a computer or performs calculations. The processor 71 may be a general-purpose processor, a dedicated processing circuit designed to perform a specific calculation, or a semiconductor device that includes both a general-purpose processor and a dedicated processing circuit. The processor 71 may also include an optical circuit or a calculation function based on quantum computing.

[0149] The processor 71 may perform arithmetic processing based on data or software input from each device or the like configured inside the computer 7, and may output the calculation results or control signals to each device or the like. The processor 71 may control each component constituting the computer 7 by executing the OS (Operating System) of the computer 7, applications, etc.

[0150] Each device (3D reconstruction systems 100 and 100a) in the above-described embodiments may be realized by one or more processors 71. Here, the processor 71 may refer to one or more electronic circuits arranged on one chip, or may refer to one or more electronic circuits arranged on two or more chips or two or more devices. When multiple electronic circuits are used, the electronic circuits may communicate with each other via wire or wirelessly.

[0151] The main memory device 72 may store instructions to be executed by the processor 71, various data, etc., and information stored in the main memory device 72 may be read by the processor 71. The auxiliary memory device 73 is a memory device other than the main memory device 72. Note that these memory devices refer to any electronic component capable of storing electronic information, and may be semiconductor memory. The semiconductor memory may be either volatile memory or non-volatile memory. The memory device for saving various data, etc. in each device (the three-dimensional reconstruction systems 100 and 100a) in the above-described embodiments may be realized by the main memory device 72 or the auxiliary memory device 73, or may be realized by an internal memory built into the processor 71.

[0152] When each device (3D reconstruction systems 100 and 100a) in the above-described embodiments is configured with at least one storage device (memory) and at least one processor connected (coupled) to this at least one storage device, at least one processor may be connected to one storage device. Also, at least one storage device may be connected to one processor. Also, a configuration in which at least one processor among multiple processors is connected to at least one storage device among multiple storage devices may be included. Also, this configuration may be realized by storage devices and processors included in multiple computers. Furthermore, a configuration in which a storage device is integrated with a processor (for example, a cache memory including an L1 cache and an L2 cache) may be included.

[0153] The network interface 74 is an interface for connecting to the communication network 8 wirelessly or via a wire. The network interface 74 may be an appropriate interface, such as one that conforms to an existing communication standard. Information may be exchanged with an external device 9A connected via the communication network 8 via the network interface 74. The communication network 8 may be any one of a WAN (Wide Area Network), a LAN (Local Area Network), a PAN (Personal Area Network), etc., or a combination thereof, as long as information is exchanged between the computer 7 and the external device 9A. An example of a WAN is the Internet, an example of a LAN is IEEE802.11 or Ethernet (registered trademark), and an example of a PAN is Bluetooth (registered trademark) or NFC (Near Field Communication), etc.

[0154] The device interface 75 is an interface such as a USB that directly connects to the external device 9B.

[0155] The external device 9A is a device connected to the computer 7 via a network. The external device 9B is a device connected directly to the computer 7.

[0156] For example, the external device 9A or the external device 9B may be an input device. The input device may be a device such as a camera, a microphone, a motion capture device, various sensors, a keyboard, a mouse, or a touch panel, and provides acquired information to the computer 7. Alternatively, the external device 9A or the external device 9B may be a device equipped with an input unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.

[0157] Furthermore, the external device 9A or the external device 9B may be, for example, an output device. The output device may be, for example, a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) panel, or a speaker that outputs sound or the like. Alternatively, the output device may be a device including an output unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.

[0158] Furthermore, the external device 9A or the external device 9B may be a storage device (memory). For example, the external device 9A may be a network storage or the like, and the external device 9B may be a storage such as an HDD.

[0159] Furthermore, the external device 9A or the external device 9B may be a device having some of the functions of the components of each device (the three-dimensional reconstruction systems 100 and 100a) in the above-described embodiments. That is, the computer 7 may transmit some or all of the processing results to the external device 9A or the external device 9B, or may receive some or all of the processing results from the external device 9A or the external device 9B.

[0160] In this specification (including the claims), when the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used, it includes any of a, b, c, ab, ac, bc, or abc. It may also include multiple instances of any element, such as aa, abb, aabbcc, etc. Furthermore, it also includes the addition of elements other than the enumerated elements (a, b, and c), such as having d, as in abcd.

[0161] In this specification (including claims), when expressions such as "using data as input / based on / according to / in response to data" (including similar expressions) are used, unless otherwise specified, this includes cases where the data itself is used, or where data that has been processed in some way (e.g., data with noise added, normalized data, features extracted from data, intermediate representations of data, etc.) is used. Furthermore, when a statement is made that a result is obtained "using data as input / based on / according to / in response to data" (including similar expressions), this includes cases where the result is obtained based solely on the data, or where the result is influenced by other data, factors, conditions, and / or states other than the data itself, unless otherwise specified. Furthermore, when a statement is made that "data is output" (including similar expressions), this includes cases where the data itself is used as output, or where data that has been processed in some way (e.g., data with noise added, normalized data, features extracted from data, intermediate representations of various data, etc.) is used as output, unless otherwise specified.

[0162] When the terms "connected" and "coupled" are used in this specification (including the claims), they are intended as open-ended terms that encompass any of direct connection / coupling, indirect connection / coupling, electrically connection / coupling, communicatively connection / coupling, functionally connection / coupling, and physically connection / coupling. These terms should be interpreted appropriately according to the context in which they are used, but any form of connection / coupling that is not intentionally or naturally excluded should be interpreted as being included in these terms without limitation.

[0163] In this specification (including the claims), the expression "A configured to B" may include the physical structure of element A having a configuration capable of performing operation B, and the permanent or temporary setting / configuration of element A being configured / set to actually perform operation B. For example, if element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and is configured to actually perform operation B by setting a permanent or temporary program (instruction). Also, if element A is a dedicated processor, dedicated arithmetic circuit, etc., it is sufficient that the circuit structure, etc. of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached.

[0164] Whenever words implying containing or possessing (e.g., "comprising / including," "having," etc.) are used in this specification (including the claims), they are intended to be open-ended terms that include the inclusion or possession of things other than the object designated by the object of the term. When the object of such words implying containing or possessing does not specify a quantity or suggests a singular number (e.g., expressions using the articles "a" or "an"), the expression should be construed as not being limited to a specific number.

[0165] In this specification (including the claims), even if expressions such as "one or more" and "at least one" are used in some places and expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") should be interpreted as not necessarily being limited to a specific number.

[0166] In this specification, when a particular advantage / result is described as being obtained with respect to a particular configuration of an embodiment, it should be understood that the same advantage / result can also be obtained with one or more other embodiments having the same configuration, unless otherwise stated. However, it should be understood that the presence or absence of the effect generally depends on various factors, conditions, and / or circumstances, and that the effect is not necessarily obtained with the configuration. The effect is merely obtained by the configuration described in the embodiment when various factors, conditions, and / or circumstances are satisfied, and the effect does not necessarily occur in a claimed invention that defines the same or a similar configuration.

[0167] In this specification (including claims), when multiple pieces of hardware perform a predetermined process, the pieces of hardware may cooperate to perform the predetermined process, or some of the hardware may perform all of the predetermined process. Furthermore, some of the hardware may perform part of the predetermined process, and other hardware may perform the rest of the predetermined process. In this specification (including claims), when an expression such as "one or more pieces of hardware perform a first process, and the one or more pieces of hardware perform a second process" (including similar expressions) is used, the hardware performing the first process and the hardware performing the second process may be the same or different. In other words, it is sufficient that the hardware performing the first process and the hardware performing the second process are included in the one or more pieces of hardware. Note that hardware may include electronic circuits, devices including electronic circuits, etc.

[0168] In this specification (including the claims), when multiple storage devices (memories) store data, each of the multiple storage devices may store only a portion of the data, or may store the entire data. Also, a configuration in which only some of the multiple storage devices store data may be included.

[0169] Although the embodiments of the present disclosure have been described in detail above, the present disclosure is not limited to the individual embodiments described above. Various additions, modifications, substitutions, partial deletions, etc. are possible within the scope of the conceptual idea and spirit of the present invention, which is derived from the content defined in the claims and their equivalents. For example, when numerical values ​​or formulas are used in the above-described embodiments, they are shown for illustrative purposes and do not limit the scope of the present disclosure. Furthermore, the order of each operation shown in the embodiments is also illustrative and does not limit the scope of the present disclosure.

[0170] The loss function is not limited to those including the losses described in the above-described embodiments, but may include a loss calculated using a first coordinate and a second coordinate corresponding to a first pixel in an image from a first viewpoint and a second pixel in an image from a second viewpoint, respectively, which correspond to the same point on the actual object surface onto which the structured light is projected. For example, the distance loss L D , and the first loss L according to the second embodiment SR and second loss L ST is the loss L that makes the SDF of the intersection zero. sdf (θ) may be replaced by the loss L sdf (θ) is expressed by the following equation. L sdf (θ)=|ya-yb|+|S(ya;θ)|+|S(yb;θ)| In addition, loss L sdf may be the sum of the distances for all corresponding pixels.

[0171] The above loss L sdfWhen (θ) is used, the 3D reconstruction system according to the embodiment calculates the point ya on the ray Ra(τ) that is closest to the ray Rb(τ) and the point yb on the ray Rb(τ) that is closest to the ray Ra(τ) in Fig. 16. Finally, the 3D reconstruction system according to the embodiment may evaluate the distance between the two points (ya, yb) so that the two points are positioned at the same position, and may calculate the absolute values ​​of the SDFs at the points ya and yb as losses so that the SDFs at these points become 0.

[0172] In the above-described embodiments, a signed distance field (SDF) has been described as an example of data representing the shape of the object 10 (data for reconstructing the three-dimensional shape of the object). However, neural radiance fields (NeRFs) may be employed as another example of data representing the shape of the object 10. NeRFs can be approximately represented by a neural network that receives the coordinates of a point x in three-dimensional space and a line-of-sight vector v of a viewpoint (ray) viewing the point x, and outputs the color (radiance) and opacity (volume density) of the point x as viewed from the viewpoint. This corresponds to the object shape NN and object color NN in the above-described embodiments. Using this property, a sampling point xa on a ray from a first viewpoint passing through a first corresponding pixel where the opacity of the sampling point x significantly changes (changes from almost zero to a predetermined threshold or more indicating the presence of an object) relative to adjacent sampling points may be regarded as a point on the object surface (first coordinate corresponding to the first corresponding pixel). Similarly, a point xb (second coordinate corresponding to the second corresponding pixel) may be calculated for a ray from a second viewpoint passing through a second corresponding pixel. Then, the distance loss or geometric loss is calculated using the first and second coordinates, and the calculated distance loss or geometric loss and the loss L R The parameters of the neural networks (object shape NN and object color NN) may be updated based on the above.

[0173] For example, aspects of the present invention are as follows. <1> The three-dimensional reconstruction method includes acquiring images of an actual object onto which structured light is projected, from at least two viewpoints including a first viewpoint and a second viewpoint; identifying a first pixel in the image from the first viewpoint and a second pixel in the image from the second viewpoint, which correspond to an identical point on the surface of the actual object onto which the structured light is projected, based on a projection pattern of the structured light; calculating a first coordinate corresponding to the first pixel and a second coordinate corresponding to the second pixel on the surface of the object reconstructed by the object shape neural network, based on current parameters in the object shape neural network that reconstructs the three-dimensional shape of the object; and updating parameters of the object shape neural network using at least the first coordinate and the second coordinate. <2> acquiring images of the actual object from at least two viewpoints where the structured light is not projected, calculating a color of the object represented by the object color neural network based on current parameters in the object color neural network that represent the color of the object, and updating the parameters of the object color neural network using information about the calculated color of the object; <1> 3. A three-dimensional reconstruction method according to claim 1. <3> The structured light includes a plurality of projection patterns. <1> or the above <2> 3. A three-dimensional reconstruction method according to claim 1. <4> the structured light includes a plurality of linear patterns each extending in a first direction and arranged side by side in a second direction perpendicular to the first direction, and a plurality of linear patterns each extending in the second direction and arranged side by side in the first direction, <1> From the above <3> The three-dimensional reconstruction method according to any one of the above items. <5> acquiring an image of the real object at the first viewpoint using a first camera; acquiring an image of the real object at the second viewpoint using a second camera; calculating a first point obtained by projecting the second coordinates onto the first camera and a second point obtained by projecting the first coordinates onto the second camera; and further updating parameters of the object shape neural network using at least the first point and the second point. <1> 3. A three-dimensional reconstruction method according to claim 1. <6> acquiring an image of the real object at the first viewpoint using a first camera, acquiring an image of the real object at the second viewpoint using a second camera, calculating a first point obtained by projecting the second coordinates onto the first camera and a second point obtained by projecting the first coordinates onto the second camera, and further updating parameters of the object color neural network using at least the first point and the second point; <2> 3. A three-dimensional reconstruction method according to claim 1. <7> a camera that acquires images of an actual object onto which the structured light is projected, from at least two viewpoints including a first viewpoint and a second viewpoint; a corresponding pixel specifying unit that specifies, based on a projection pattern of the structured light, a first pixel in the image from the first viewpoint and a second pixel in the image from the second viewpoint, which correspond to an identical point on a surface of the actual object onto which the structured light is projected; a corresponding coordinate calculation unit that calculates, based on current parameters in an object shape neural network that reconstructs a three-dimensional shape of the object, a first coordinate corresponding to the first pixel and a second coordinate corresponding to the second pixel on the surface of the object reconstructed by the object shape neural network; and a model update unit that updates parameters of the object shape neural network by using at least the first coordinate and the second coordinate.

Claims

1. acquiring images of an actual object onto which structured light is projected, from at least two viewpoints including a first viewpoint and a second viewpoint; Identifying a first pixel in the first viewpoint image and a second pixel in the second viewpoint image, which correspond to the same point on the actual object surface onto which the structured light is projected, based on the projection pattern of the structured light; calculating a first coordinate corresponding to the first pixel and a second coordinate corresponding to the second pixel on the surface of the object reconstructed by the object shape neural network based on current parameters in the object shape neural network that reconstructs the three-dimensional shape of the object; A three-dimensional reconstruction method, comprising: updating parameters of the object shape neural network using at least the first coordinate and the second coordinate.

2. acquiring an image of the real object without the structured light projected thereon for at least one viewpoint; further calculating a color of the object based on current parameters in the object color neural network representing the color of the object and current parameters in the object shape neural network; The three-dimensional reconstruction method according to claim 1 , further comprising updating parameters of the object shape neural network using at least information about the first coordinates, the second coordinates, and the calculated color of the object.

3. acquiring an image of the real object without the structured light projected thereon for at least one viewpoint; further calculating a color of the object represented by the object color neural network based on current parameters in the object color neural network representing the color of the object; 2. The three-dimensional reconstruction method according to claim 1, further comprising updating parameters of said object color neural network using information about the calculated color of said object.

4. acquiring an image of the real object without the structured light projected thereon for at least one viewpoint; further calculating a color of the object based on current parameters in the object color neural network representing the color of the object and current parameters in the object shape neural network; 2. The three-dimensional reconstruction method according to claim 1, further comprising updating parameters of the object shape neural network and the object color neural network using at least information about the first coordinates, the second coordinates, and the calculated color of the object.

5. The three-dimensional reconstruction method according to claim 1 , wherein the structured light includes a plurality of projection patterns.

6. The structured light is a plurality of linear patterns each extending in a first direction and arranged side by side in a second direction perpendicular to the first direction; 5. The three-dimensional reconstruction method according to claim 1, further comprising: a plurality of linear patterns each extending in the second direction and arranged side by side in the first direction.

7. acquiring an image of the real object at the first viewpoint with a first camera; acquiring an image of the real object at the second viewpoint with a second camera; calculating a first point obtained by projecting the second coordinates onto the first camera and a second point obtained by projecting the first coordinates onto the second camera; The three-dimensional reconstruction method according to claim 1 , further comprising the step of updating parameters of the object shape neural network using at least the first point and the second point.

8. acquiring an image of the real object at the first viewpoint with a first camera; acquiring an image of the real object at the second viewpoint with a second camera; calculating a first point obtained by projecting the second coordinates onto the first camera and a second point obtained by projecting the first coordinates onto the second camera; 4. The three-dimensional reconstruction method according to claim 3, further comprising updating parameters of the object color neural network using at least the first point and the second point.

9. acquiring images of an actual object onto which structured light is projected, from at least two viewpoints including a first viewpoint and a second viewpoint; Identifying a first pixel in the first viewpoint image and a second pixel in the second viewpoint image, which correspond to the same point on the actual object surface onto which the structured light is projected, based on the projection pattern of the structured light; A three-dimensional reconstruction method that updates parameters of the object shape neural network so that, with respect to the three-dimensional shape of the object reconstructed by the object shape neural network, a point on the surface of the object corresponding to the first pixel coincides with a point on the surface of the object corresponding to the second pixel.

10. a projection unit that projects structured light; a camera that acquires images of the actual object onto which the structured light is projected from at least two viewpoints including a first viewpoint and a second viewpoint; a corresponding pixel specifying unit that specifies, based on the projection pattern of the structured light, a first pixel in the first viewpoint image and a second pixel in the second viewpoint image that correspond to the same point on the actual object surface onto which the structured light is projected; a corresponding coordinate calculation unit that calculates, based on current parameters in an object shape neural network that reconstructs the three-dimensional shape of the object, a first coordinate corresponding to the first pixel and a second coordinate corresponding to the second pixel on the surface of the object reconstructed by the object shape neural network; a model update unit that updates parameters of the object shape neural network using at least the first coordinates and the second coordinates.