Three-dimensional reconstruction method and three-dimensional reconstruction system
The three-dimensional reconstruction system uses patterned background members and machine learning to distinguish and reconstruct objects with the same color as the background, enhancing reconstruction accuracy and versatility.
Patent Information
- Application Number
- JP2022081861
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-08-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Conventional 3D reconstruction methods struggle to reconstruct objects with portions that have the same color as the background, limiting the variety of objects that can be reconstructed.
A three-dimensional reconstruction system and method that uses cameras positioned to capture images of an object from different directions, with a transparent support base and a background member having a pattern, allowing the system to distinguish between object and background regions, and employs machine learning to generate a three-dimensional model.
Enables the reconstruction of a variety of objects by effectively separating object and background regions, improving reconstruction accuracy and overcoming limitations of conventional methods.
Smart Images

Figure 2025124948000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a three-dimensional reconstruction method and a three-dimensional reconstruction system. [Background technology]
[0002] Conventionally, a 3D reconstruction method is known that generates a 3D model of an object based on multiple images obtained by capturing the object from different directions using a camera. Such a 3D reconstruction method is used in 3DCG (3 Dimensional Computer Graphics) and the like.
[0003] There is also a technology for performing three-dimensional reconstruction by volume rendering using a neural network. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Peng Wang et al., "35th Conference on Neural Information Processing Systems", 16,December,2021, Internet<URL:https: / / arxiv.org / pdf / 2106.10689.pdf> ,internet<URL:https: / / lingjie0206.github.io / papers / NeuS / index.htm> Summary of the Invention [Problem to be solved by the invention]
[0005] However, conventional 3D reconstruction methods may be limited in the objects that can be reconstructed because they cannot reconstruct portions of an object that have the same color as the background. Therefore, there is a demand for a 3D reconstruction system that can reconstruct a variety of objects.
[0006] The present disclosure aims to provide a three-dimensional reconstruction system capable of reconstructing a variety of objects. [Means for solving the problem]
[0007] A three-dimensional reconstruction method according to one aspect of an embodiment of the present disclosure is a method for three-dimensionally reconstructing an object based on multiple images obtained by photographing the object from different directions, wherein a background image region of the object included in each of the multiple images has a pattern.
[0008] Similarly, a three-dimensional reconstruction system according to one aspect of an embodiment of the present disclosure includes a camera capable of photographing an object from different directions, and a control device configured to output reconstruction information of the object based on multiple images taken by the camera, wherein a background image area of the object included in each of the multiple images has a pattern. [Effects of the Invention]
[0009] According to the present disclosure, a three-dimensional reconstruction system capable of reconstructing a variety of objects can be provided. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a diagram illustrating an example of the overall configuration of a three-dimensional reconstruction system according to an embodiment. [Figure 2] 1A and 1B are diagrams illustrating examples of captured images including a background image region according to an embodiment. [Figure 3] FIG. 2 is a block diagram illustrating an example of a functional configuration of a model generation unit according to the embodiment. [Figure 4] FIG. 1 is a diagram illustrating an example of a three-dimensional scene. [Figure 5] 4 is a flowchart illustrating an example of processing by a control device according to the embodiment. [Figure 6] FIG. 10 is a flowchart of an example of processing by a rendering unit of the control device according to the embodiment. [Figure 7] FIG. 10 is a flowchart of an example of processing by a model generation unit of the control device according to the embodiment. [Figure 8] FIG. 2 is a block diagram of an example of a hardware configuration of a control device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. To facilitate understanding of the description, the same components in the drawings will be denoted by the same reference numerals as much as possible, and duplicate descriptions will be omitted where appropriate.
[0012] In each figure, the x, y, and z directions are perpendicular to each other. The z direction is the normal direction of the mounting surface and is typically the vertical direction. The positive z direction is referred to as the upper side, and the negative z direction is referred to as the lower side. The x and y directions are the extension directions of the mounting surface and are typically horizontal.
[0013] [Embodiment] <Configuration example of 3D reconstruction system 100> The configuration of a three-dimensional reconstruction system 100 according to this embodiment will be described with reference to Figures 1 and 2. Figure 1 is a diagram illustrating an example of the overall configuration of the three-dimensional reconstruction system 100. Figure 2 is a diagram illustrating an example of a captured image including a background image region.
[0014] As shown in FIG. 1, the three-dimensional reconstruction system 100 includes a support base 2, upper cameras 3A and 3B, lower cameras 4A and 4B, a rotating unit 5, a control device 70, lighting units 8A and 8B, and a background member 9.
[0015] The three-dimensional reconstruction system 100 generates a three-dimensional model of the object 10 based on a plurality of images of the object 10 taken from different directions by the upper camera 3A, the upper camera 3B, the lower camera 4A, and the lower camera 4B. The object 10 may be, for example, a vessel-shaped object such as a cup as shown in Fig. 1, or an object of any shape that is sized to be placed inside the outer edge of the support table 2 when viewed from above.
[0016] In this specification, the term "object" includes both an object that exists in reality and is the subject of three-dimensional reconstruction, and an object reconstructed based on multiple images of the object. Whether "object" refers to an object that exists in reality or a reconstructed object can be appropriately distinguished depending on the context. For example, when describing the handling of an object in real space, such as "placing an object on a support base," "object" refers to an object that exists in reality. On the other hand, when describing the handling of an object in virtual space, such as "storing an object" or "reconstructing an object," "object" refers to a reconstructed object that exists in reality.
[0017] The support table 2 is a base on which the object 10 is placed and rotated on its upper surface 2A. The support table 2 is a transparent plate such as an acrylic plate. Because the support table 2 is a transparent plate, it is possible to photograph the object 10 placed on the upper surface 2A from the lower surface 2B, which is the back side of the upper surface 2A, as shown in FIG. 2. By using a transparent plate for the support table 2, the three-dimensional reconstruction system 100 can photograph the object 10 from all directions, thereby obtaining three-dimensional reconstruction information of the object 10 without any information loss.
[0018] Each of the upper camera 3A, the upper camera 3B, the lower camera 4A, and the lower camera 4B is an example of a camera capable of photographing an object 10 placed on the upper surface 2A from different directions. The upper cameras 3A and 3B are installed toward the upper surface 2A of the support base 2, and can be directed toward the object 10 from a diagonally upward direction relative to the upper surface 2A to photograph the object 10. The upper cameras 3A and 3B are installed at different inclination angles. The number of cameras may be one or more.
[0019] The lower cameras 4A and 4B are installed toward the lower surface 2B of the support base 2 and can be directed toward the object 10 from a diagonally downward direction relative to the upper surface 2A to photograph the object 10. The lower cameras 4A and 4B are installed at different inclination angles.
[0020] Each of the upper camera 3A, upper camera 3B, lower camera 4A, and lower camera 4B can capture an RGB image including the colors R (Red), G (Green), and B (Blue) of the object 10. In the following, the upper cameras 3A and 3B will be collectively referred to as the "upper camera 3," and the lower cameras 4A and 4B will be collectively referred to as the "lower camera 4." The positions of the upper camera 3 and lower camera 4 may be rotated.
[0021] In FIG. 1, the dashed arrow extending from upper camera 3A represents a line-of-sight vector v, which is the direction of a line of sight from upper camera 3A.
[0022] In this embodiment, the term "camera" refers to an element that can capture an RGB image of the object 10. This "camera" encompasses the entire camera device, sensors such as a CMOS sensor and a depth sensor that are built into the camera device, and sensors that are used independently.
[0023] Each of the upper camera 3A, the upper camera 3B, the lower camera 4A, and the lower camera 4B has internal parameters and external parameters. The internal parameters include information related to the distortion of the lens provided in each camera. In this embodiment, the internal parameters are known based on simulation results, etc. The external parameters include attitude information of each camera, etc. The camera attitude includes the relative attitude of the camera with respect to a predetermined reference and the absolute attitude in the world coordinate system. The camera attitude also corresponds to the tilt of the optical axis of the optical system, such as the lens, included in each camera. The external parameters are known based on simulation results, etc., or can be calculated by any method for each reconstruction operation by the 3D reconstruction system 100. In this embodiment, a case where the external parameters are known based on simulation results, etc. is illustrated.
[0024] Rotation unit 5 rotates object 10 by rotating support base 2. For example, rotation unit 5 rotates support base 2 in the direction of arrow 50. Each of upper camera 3A, upper camera 3B, lower camera 4A, and lower camera 4B photographs object 10 at each of a plurality of rotation angles set by rotation unit 5, thereby being able to photograph images of object 10 from different directions.
[0025] A known power transmission system can be used for the mechanism of the rotating unit 5. For example, the rotating unit 5 has a motor and a gear mechanism. The rotating unit 5 may be configured so that the driving force of the motor is transmitted to the rotation shaft of the support base 2 via the gear mechanism. Alternatively, the rotating unit 5 may be configured so that a driving force is applied to the outer edge of the support base 2 to rotate the support base 2.
[0026] The control device 70 controls the operation of the 3D reconstruction system 100. In this embodiment, the control device 70 is configured to output reconstruction information of the object 10 based on a plurality of RGB images taken by the upper camera 3 and the lower camera 4.
[0027] Specifically, the control device 70 controls the photographing of the object 10 by the upper camera 3 and the lower camera 4. The control device 70 also reconstructs the object 10 by generating a three-dimensional model of the object 10 based on the photographed image of the object 10. The control device 70 has a photographing control unit 701, a posture acquisition unit 702, and a model generation unit 703 as functions related to these.
[0028] The control device 70 can realize each of the above functions by software (CPU: Central Processing Unit) or electric circuits, or can realize each of these functions by multiple electric circuits or multiple pieces of software. Furthermore, the control device 70 may realize each of the above functions by distributed processing with components other than the control device 70.
[0029] The photographing control unit 701 controls the operations of the rotating unit 5, the upper camera 3, and the lower camera 4 so that the upper camera 3 and the lower camera 4 photograph the object 10 and acquire multiple images during rotation by the rotating unit 5. The photographing control unit 701 may also control the lighting 8A and 8B described below.
[0030] The attitude acquisition unit 702 can acquire attitude information of the upper camera 3A, the upper camera 3B, the lower camera 4A, and the lower camera 4B that has been acquired in advance and stored in memory.
[0031] The model generation unit 703 reconstructs the object 10 by generating a three-dimensional model of the object 10 based on multiple images of the object 10 captured by the upper camera 3 and the lower camera 4 and the orientation information of the upper camera 3 and the lower camera 4 acquired by the orientation acquisition unit 702. The model generation unit 703 outputs the generated three-dimensional model of the object 10 as reconstruction information of the object 10.
[0032] For example, the model generation unit 703 reconstructs the object 10 by a machine learning method based on the error between an image obtained by rendering three-dimensional information and a captured image. However, the reconstruction processing method by the model generation unit 703 is not limited to this method. For example, the model generation unit 703 can also reconstruct the object 10 by using a method in which a model shape is generated from a depth image and a texture is generated from an RGB image, a so-called photogrammetry method, or the like.
[0033] Lighting 8A is arranged toward the upper surface 2A of the support base 2. Lighting 8B is arranged toward the lower surface 2B of the support base 2. Lighting 8A and 8B each illuminate the object 10. Lighting 8A and 8B are arranged according to the installation positions of the upper camera 3 and the lower camera 4 so that there is no shadow on the surface of the object 10 in the images captured by the upper camera 3 and the lower camera 4.
[0034] Background member 9 is a member that is installed so as to form the background of object 10 when photographed by upper camera 3 and lower camera 4. There are no particular limitations on the material of background member 9, and it can be made of materials that appropriately include resin materials, metal materials, etc. Furthermore, background member 9 is, for example, a plate-shaped member, but is not limited to a plate-shaped member and may be a member of any shape, such as a columnar member.
[0035] The background member 9 has a pattern on its surface. The term "pattern" here refers to a non-solid color. "Solid color" refers to a single color with no pattern. In this embodiment, if a reproducible pattern is not recognized as an image within an image region corresponding to the background member 9 in the captured image, the background member 9 can be said to be solid color. Conversely, if a reproducible pattern is recognized as an image within an image region corresponding to the background member 9 in the captured image, the background member 9 can be said to have a pattern on its surface. An example of a pattern is a black-and-white dot pattern in which relatively bright white dots are arranged on a relatively dark black background. However, there are no particular restrictions on the type of pattern or the colors that make up the pattern. Examples of patterns include stripes, grids, dots, and other patterns with distinctive shapes, geometric patterns, polka dots, and other patterns with distinctive color variations, such as monochromatic gradations and spots, as well as other patterns not exemplified here. Furthermore, the pattern does not necessarily have to be periodic.
[0036] The multiple images captured by the upper camera 3 and the lower camera 4 are obtained by capturing images of the object 10 with the same patterned background member 9 as its backdrop. The multiple images are also obtained by capturing images of the object 10 rotated with the background member 9 as its backdrop from different directions.
[0037] As shown in FIG. 2, the captured image 60 obtained by each of the upper camera 3 and the lower camera 4 includes an object image region 61 and a background image region 62. The background image region 62 corresponds to the portion of the background member 9 that is not occluded by the object 10 and is captured by the camera. When the object 10 is rotated by the rotation unit 5, the capturing direction of the object 10 by each of the upper camera 3 and the lower camera 4 changes. The object image region 61 changes in accordance with the change in the position and orientation of the object 10 relative to each camera that accompanies the change in capturing direction. On the other hand, the background image region 62 does not move even when the object 10 is rotated by the rotation unit 5, and therefore remains the same regardless of the capturing direction of each camera. In other words, the background image region 62 of the object 10 included in each of the multiple captured images 60 has the same pattern.
[0038] The pattern of the background image area 62 is the pattern of the background member 9 photographed.
[0039] Each of the upper camera 3 and the lower camera 4 can also see through the support base 2 to capture an image of the object 10, thereby obtaining an object image area 61. In addition, each of the upper camera 3 and the lower camera 4 can also see through the support base 2 to capture an image of the background member 9, thereby obtaining a background image area 62.
[0040] If the background image region 62 is solid colored corresponding to a solid background component, the object cannot be properly reconstructed because portions of the object that are the same color as the background color cannot be distinguished from the background. This can limit the objects that can be reconstructed using conventional 3D reconstruction methods. In the 3D reconstruction system 100, the background image region 62 has a pattern, which includes multiple colors, so no portion of the object 10 has the exact same color as the background. This allows the 3D reconstruction system 100 to reconstruct a variety of objects 10 without being subject to reconstruction limitations related to the background color.
[0041] In this embodiment, cases in which the multiple captured images include the background image region 62 include cases in which a pattern is formed on the background member 9 itself, cases in which a pattern is projected onto a plain background member 9 using a projector or the like, and cases in which the background member 9 displays a pattern. When a pattern is formed, this includes cases in which the pattern is printed, processed, recorded, or embroidered. In any case, the background member 9 has a pattern. The background member onto which the pattern is projected using a projector or the like is a wall, floor, or the like of a building. When a pattern is displayed, the background member 9 is, for example, a display device such as a liquid crystal display or an organic EL (Electro Luminescence) display.
[0042] When a projected pattern is used, an existing wall or the like can be used as the background member 9, which is advantageous in that it is possible to simplify the configuration of the three-dimensional reconstruction system 100 or to save the installation space of the three-dimensional reconstruction system 100. On the other hand, when a pattern is formed on the background member 9 itself, or when the background member 9 displays a pattern, it is advantageous in that it is possible to obtain the background image area 62 even if there is no wall or floor suitable for projecting a pattern.
[0043] <Example of functional configuration of model generation unit 703> The functional configuration of the model generation unit 703 will be described with reference to Fig. 3 and Fig. 4. Fig. 3 is a block diagram showing an example of the functional configuration of the model generation unit 703. Fig. 4 is a diagram showing an example of a three-dimensional scene 40.
[0044] 3, the model generation unit 703 includes an object data storage unit 31, a background data storage unit 32, a rendering unit 33, a model update unit 34, and a model output unit 35. The model generation unit 703 performs three-dimensional reconstruction of the object 10 based on data representing the shape and color of the object 10, which is separated from data representing the background of the object 10.
[0045] The object data storage unit 31 stores data representing the shape and color of the object 10 .
[0046] 4 shows a three-dimensional scene 40 including an upper camera 3A, an object 10, a line-of-sight vector v, and a camera view 41. Note that although the upper camera 3A is shown here as an example, it can be replaced with an upper camera 3B, a lower camera 4A, and a lower camera 4B.
[0047] An example of data representing the shape of the object 10 is a signed distance field (SDF), which represents a field formed corresponding to each location in a three-dimensional scene. Each point included in the SDF contains information about the signed distance from that point to the nearest surface of the object 10. The distance d in FIG. 4 represents this signed distance. The signed distance takes a negative (-) value for points inside the object 10 and a positive (+) value for points outside the object 10. The zero isosurface of the SDF represents the surface of the object 10.
[0048] The line of sight vector v corresponds to the line of sight from the camera and is also called a ray. The camera view 41 is an image generated (rendered) as an image seen from the upper camera 3A. The color of a pixel at position P in the camera view 41 is determined by the color of light (point color) emitted from each point in the three-dimensional scene on the line of sight vector v from the camera that passes through that pixel. The color of this light can be calculated by a weighted sum, which is an accumulation of color information at each three-dimensional coordinate on the line of sight vector v.
[0049] The sampling points Sm represent sampling points of pixel colors on the gaze direction vector v. A weight is assigned to each sampling point Sm. The weight increases as the sampling point Sm approaches the surface of the object 10. The weight assigned to a sampling point Sm may be expressed as a function of the SDF calculated for that sampling point Sm. The area Ar indicates the area where the weight increases.
[0050] The SDF may be approximately represented as a neural field by a neural network. For example, when a point x (three-dimensional coordinates: x=(x, y, z)) in a three-dimensional scene 40 is input, this neural network outputs a signed distance d from the point x to the surface of the object 10. In this embodiment, this neural network is called an object shape NN (Neural Network). The object shape NN can represent the zero isosurface of the SDF, i.e., the object surface, and therefore corresponds to data representing the shape of the object 10.
[0051] An example of data representing the color of the object 10 is a vector field (color field) representing the color (e.g., RGB values) of each point x in the three-dimensional scene 40, conditioned by the line of sight. This vector field may be approximately represented as a neural field by a neural network. For example, when a point x in the three-dimensional scene and a line of sight vector v for that point x are input, this neural network outputs the color c of the pixel corresponding to that point x. In this embodiment, this neural network is referred to as an object color NN. The object color NN corresponds to data representing the color of the object 10.
[0052] The background data storage unit 32 stores data representing the color of the background. This data is prepared for each of the upper camera 3 and the lower camera 4, and represents the background color of each pixel in the camera view 41. This data may be approximately represented as a neural field by a neural network. When the position P (two-dimensional coordinates) of a pixel in the camera view 41 is input, this neural network outputs the background color c_bg of that pixel. In this embodiment, this neural network is referred to as a background color NN. The background color NN corresponds to data representing the background of the object 10.
[0053] The background color NN is stored in the background data storage unit 32, and the object shape NN and object color NN are stored in the object data storage unit 31. As a result, the background color NN is treated as data (information) separate from the object shape NN and object color NN.
[0054] The rendering unit 33 generates a camera view 41 using the object shape NN and object color NN stored in the object data storage unit 31 and the background color NN stored in the background data storage unit 32. In this embodiment, the camera view 41 may be generated for each pixel according to a volume rendering method, for example, using the method of the NeuS paper (see, for example, Non-Patent Document 1).
[0055] The model update unit 34 calculates the error (loss) between the camera view 41 generated by the rendering unit 33 and the actually captured image for each of the multiple camera views 41. The error can be calculated, for example, by accumulating the color differences between corresponding pixels for all pixels between the camera view 41 and the captured image. However, the definition of the error is not limited to the above-mentioned accumulation of the color differences between pixels for all pixels, and other definitions may be used. Furthermore, the method of calculating the error is not limited to the above-mentioned method of accumulating the color differences between pixels for all pixels, and other methods may be used.
[0056] The model update unit 34 updates the parameters of the object shape NN, object color NN, and background color NN using an error backpropagation method based on this error. The model update unit 34 repeats updating each NN until a predetermined criterion is met. As a result of updating by the model update unit 34, the object shape NN and object color NN finally obtained represent the three-dimensional shape and color of the actual object 10, i.e., a three-dimensional model. The predetermined criterion may be, for example, until the error falls below a predetermined value. Alternatively, the predetermined criterion may be, for example, until a predetermined number of updates have been performed.
[0057] The model update unit 34 stores the three-dimensional model obtained as a result of the update as reconstruction information in the memory of the control device 70. The model generation unit 703 can also output the reconstruction information to an external device such as a display device or a PC (Personal Computer).
[0058] The model output unit 35 uses the data stored in the object data storage unit 31 to generate three-dimensional mesh data representing the shape of the three-dimensional model of the object 10 and a texture map to be applied to this three-dimensional mesh data. The model output unit 35 converts these into a data format viewable by a three-dimensional viewer and outputs them.
[0059] In this embodiment, the finally obtained object shape NN represents an SDF, and the model output unit 35 generates a mesh that approximates the zero isosurface of this SDF. The model output unit 35 can convert the zero isosurface of the SDF into mesh data using, for example, the marching cubes method or other known methods. Furthermore, for each polygon constituting the mesh, the model output unit 35 inputs the polygon's three-dimensional coordinates and the inverse vector of the normal vector of the SDF's zero isosurface at these three-dimensional coordinates to the object color NN. The model output unit 35 obtains the polygon color as output from the object color NN. The model output unit 35 can generate a texture map for the mesh using the obtained polygon color.
[0060] The method of acquiring the polygon colors is not limited to acquiring one color per polygon. For example, the method of acquiring the polygon colors may be to acquire one color for each of a plurality of unit areas obtained by further dividing one polygon. In this case, the spatial resolution can be improved compared to acquiring one color per polygon. For example, when the reconstructed object 10 is displayed as an image, the unit area corresponds to one pixel.
[0061] <Example of processing by the control device 70> Fig. 5 is a flowchart showing an example of processing by the control device 70. For example, the control device 70 starts the processing of Fig. 5 when it receives an operation input of a reconstruction start instruction from a user via its operation unit. Before the processing of Fig. 5 starts, an object 10 is placed on the upper surface 2A of the support table 2. The position where the object 10 is placed is preferably the center of rotation of the support table 2. It is also assumed that the external parameters of the upper camera 3 and the lower camera 4 relative to the object 10 have been acquired in advance.
[0062] First, in step S51, the control device 70 causes the imaging control unit 701 to drive the rotation unit 5 and make the support base 2 start rotating.
[0063] Next, in step S52 (photographing step), the control device 70 causes the photographing control unit 701 to operate the upper camera 3 and the lower camera 4 to photograph the rotating object 10 from both above and below, and acquire RGB images from each of the upper camera 3 and the lower camera 4. The photographing control unit 701 causes the upper camera 3 and the lower camera 4 to photograph at the same timing when the object 10 is rotating. This allows the photographing control unit 701 to concurrently acquire RGB images from each camera when the object 10 is at a position of any rotation angle. However, the photographing by the upper camera 3 and the lower camera 4 does not necessarily have to be at the same timing, and may be intentionally shifted.
[0064] In step S52, the photographing control unit 701 continues the photographing operation to acquire images while the rotation unit 5 rotates the object 10 once. This allows each of the upper camera 3 and the lower camera 4 to photograph the object 10 from directions tilted at different angles relative to the top surface 2A and from multiple directions along the direction of rotation. Each camera can acquire an RGB image in each photographing direction. In other words, the photographing control unit 701 can photograph the object 10 simultaneously from different directions using multiple cameras. The photographing control unit 701 outputs the photographed images to each of the model generation units 703.
[0065] Subsequently, in step S53, the control device 70 causes the imaging control unit 701 to stop the rotation unit 5, thereby stopping the rotation of the support table 2.
[0066] Subsequently, in step S54 (camera view generation step), the control device 70 causes the model generation unit 703 to generate a camera view 41 for each shooting direction.
[0067] Next, in step S55, the control device 70 determines, by the rendering unit 33 of the model generation unit 703, whether or not the error between the camera view 41 and the captured image is less than a predetermined value.
[0068] If it is determined in step S55 that the error is not less than the predetermined value (step S55, NO), the control device 70 calculates the error between the camera view 41 and the captured image for each camera view 41 in each shooting direction using the model update unit 34 of the model generation unit 703 in step S56 (update step). The model update unit 34 updates the parameters of the object shape NN, object color NN, and background color NN using the error backpropagation method based on this error. The control device 70 then performs the processes from step S54 onwards again. The control device 70 repeats the update of each NN using the model update unit 34 until the predetermined criterion is met.
[0069] On the other hand, if it is determined in step S55 that the difference is less than the predetermined value (step S55, YES), in step S57 (output step), the control device 70 generates a three-dimensional model (three-dimensional mesh and texture map) using the model output unit 35 of the model generation unit 703 based on the object shape NN and object color NN finally obtained as a result of the update, and outputs it as reconstruction information. When the processing of step S57 is completed, this processing flow ends.
[0070] <Example of Generation Process of Camera View 41 by Rendering Unit 33> 6 is a flowchart showing an example of a process for generating a camera view 41 by the rendering unit 33 of the model generation unit 703. For example, in the process for generating a camera view 41 by the model generation unit 703, the rendering unit 33 starts the process of FIG.
[0071] First, in step S61, the rendering unit 33 acquires the background color c_bg of the pixel at the position P based on the data stored in the background data storage unit 32.
[0072] Subsequently, in step S62, the rendering unit 33 determines a line of sight vector v passing through the pixel at the selected position P.
[0073] Subsequently, in step S63, the rendering unit 33 obtains the three-dimensional coordinates of each sample point x_i on the determined line-of-sight direction vector v.
[0074] Subsequently, in step S64, the rendering unit 33 obtains the color c_obj of each sample point x_i based on the line-of-sight vector v, the three-dimensional coordinates of each sample point x_i, and the object color NN.
[0075] Subsequently, in step S65, the rendering unit 33 acquires information on the signed distance d_i from each sample point x_i to the surface of the object 10 based on the three-dimensional coordinates of each sample point x_i and the object shape NN.
[0076] Next, in step S66, the rendering unit 33 calculates the color weight w_i and background color weight w_bg of each sample point x_i based on the signed distance d_i of each sample point x_i.
[0077] Next, in step S67, the rendering unit 33 calculates the weighted sum of the color of each sample point and the background color based on the background color c_bg of the pixel, the color c_obj of each sample point x_i, and the color weight w_i of each sample point x_i and the background color weight w_bg, and obtains the (final) color c of the pixel. In this case, the weights are normalized so that the sum is 1. Also, it is designed in advance that the weight becomes larger when the signed distance d_i is closer to 0, that is, closer to the surface of the object 10.
[0078] Next, in step S68, rendering unit 33 outputs color c of the pixel at position P in camera view 41. When the processing of step S68 is completed, the processing of acquiring color c of one pixel at position P in camera view 41 ends.
[0079] By performing the above processing for all pixels of the camera view 41, the rendering unit 33 can generate the camera view 41.
[0080] <Example of a three-dimensional model generation process by the model output unit 35> Fig. 7 is a flowchart showing an example of a process for generating a three-dimensional model (three-dimensional mesh and texture map) by the model output unit 35. The model output unit 35 starts the process in Fig. 7 when it is determined in step S55 of Fig. 5 that the value is less than a predetermined value (step S55, YES). Note that this flowchart describes an example in which one color is calculated for one polygon, but as described above, it is also possible to calculate colors for each of multiple locations within one polygon.
[0081] First, in step S71, the model output unit 35 generates a three-dimensional mesh that approximates the shape of the object 10 based on the object shape NN.
[0082] Next, in step S72, the model output unit 35 calculates, for each polygon constituting the three-dimensional mesh, the inverse vector u of the normal vector of the polygon at the three-dimensional coordinate q that represents the position of the polygon. Note that the model output unit 35 may, for example, calculate the three-dimensional coordinate corresponding to the position of the centroid of the polygon and acquire the coordinate as the three-dimensional coordinate q of the polygon, or may acquire another coordinate value relating to the position of the polygon in three-dimensional space as the three-dimensional coordinate of the polygon.
[0083] Subsequently, in step S73, the model output unit 35 acquires the three-dimensional coordinate q of each polygon, for example, by the calculation described above. Note that the order of execution of steps S72 and S73 may be reversed, or steps S72 and S73 may be executed in parallel.
[0084] Next, in step S74, the model output unit 35 calculates and obtains the color c_p of each polygon based on the three-dimensional coordinates q of the polygon, the inverse vector u, and the object color NN.
[0085] Subsequently, in step S75, the model output unit 35 outputs the polygon color c_p acquired for each polygon in step S74. That is, in this step, the model output unit 35 generates a texture map to be applied to the three-dimensional mesh generated in step S71.
[0086] <Actions and Effects of the 3D Reconstruction System 100> As described above, the three-dimensional reconstruction system 100 includes the upper camera 3 and the lower camera 4 (cameras) that can capture images of the object 10 from different directions, and the control device 70 that is configured to output reconstruction information of the object 10 based on a plurality of images 60 (a plurality of images) captured by the upper camera 3 and the lower camera 4. The background image region 62 of the object 10 included in each of the plurality of captured images 60 has a pattern.
[0087] For example, if the background image region 62 corresponds to a solid color, portions of the object that are the same color as the background cannot be distinguished from the background, making it impossible to properly reconstruct the object. This can limit the objects that can be reconstructed using conventional three-dimensional reconstruction methods. In the three-dimensional reconstruction system according to this embodiment, the background image region 62 has a pattern, which includes multiple colors, so no portion of the object 10 has the exact same color as the background. This allows the object 10 to be reconstructed without being subject to reconstruction limitations related to the background color, making it possible to provide a three-dimensional reconstruction system 100 capable of reconstructing a variety of objects 10.
[0088] In this embodiment, each of the multiple captured images 60 includes an object image region 61 that changes depending on the capturing direction, and a background image region 62 that remains the same regardless of the capturing direction. This allows the 3D reconstruction system 100 to distinguish between the object image region 61 and the background image region 62 in the captured image 60, and to 3D reconstruct the object 10.
[0089] This embodiment also includes a three-dimensional reconstruction method. For example, the three-dimensional reconstruction method is a three-dimensional reconstruction method of an object 10 based on a plurality of captured images 60 (a plurality of images) obtained by photographing the object 10 from different directions, and the background image region 62 of the object 10 included in each of the plurality of captured images 60 has a pattern. The background image region 62 of the object 10 included in each of the plurality of captured images 60 has the same pattern. The plurality of captured images 60 are obtained by photographing the object 10 with the same background member 9 having a pattern behind it. The plurality of captured images 60 are also obtained by photographing the object 10 rotated with the background member 9 behind it from different directions. These three-dimensional reconstruction methods can also achieve the same effects as the three-dimensional reconstruction system 100.
[0090] Furthermore, in this embodiment, the three-dimensional reconstruction of the object 10 is performed based on data representing the shape and color of the object 10, which is treated as data separate from data representing the background of the object 10. In this embodiment, the object 10 is three-dimensionally reconstructed using machine learning, and therefore the reconstruction accuracy of the object 10 can be improved.
[0091] In the present embodiment, the case where the internal parameters and external parameters of each of the upper camera 3 and the lower camera 4 are known has been exemplified, but the internal parameters and external parameters of each camera may be acquired based on a plurality of captured images 60 when capturing an image of the object 10. When acquiring the external parameters of each camera based on a plurality of captured images 60, markers placed on the support base 2 may be used.
[0092] In this embodiment, a method using a background color NN is exemplified, but the present invention is not limited to this. For example, the 3D reconstruction system 100 can use images of the background member 9 captured by the upper camera 3 and the lower camera 4 in FIG. 1 in a state where the object 10 is not present as data representing the color of the background. In this case, the background data storage unit 32 stores the images of the background member 9 captured by the upper camera 3 and the lower camera 4 in a state where the object 10 is not present as data representing the color of the background.
[0093] Using an image of the background member 9 captured in advance is advantageous in that the reconstruction process can be simplified. On the other hand, using the background color NN is advantageous in that it is not necessary to capture an image of the background member 9 in advance, and therefore the capturing operation can be simplified.
[0094] In this embodiment, at least two of the multiple images may be obtained by capturing the object 10 with the background member 9 as its backside while the object 10 is rotated at different angles.
[0095] In this embodiment, the pattern of the background image area 62 may be a black and white design.
[0096] In this embodiment, the pattern of the background image region 62 may be a pattern in which white dots are arranged on a black background.
[0097] As described above, the three-dimensional reconstruction system 100a can reconstruct the object 10 by generating a three-dimensional model of the object 10. The effects of the three-dimensional reconstruction system 100a are the same as those of the three-dimensional reconstruction system 100 according to this embodiment.
[0098] <Example of hardware configuration in the above-described embodiment> Some or all of the devices (3D reconstruction systems 100 and 100a) in the above-described embodiments may be configured as hardware, or may be configured as software (program) information processing executed by a CPU, GPU (Graphics Processing Unit), or the like. In the case of software information processing, software that realizes at least some of the functions of each device in the above-described embodiments may be stored in a non-transitory storage medium (non-transitory computer-readable medium) such as a CD-ROM (Compact Disc-Read Only Memory) or USB (Universal Serial Bus) memory, and the software information processing may be executed by reading it into a computer. The software may also be downloaded via a communication network. Furthermore, all or part of the software processing may be implemented in a circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), thereby executing the software information processing by hardware.
[0099] The storage medium that stores the software may be a removable medium such as an optical disk, or a fixed medium such as a hard disk, memory, etc. The storage medium may be provided inside the computer (main storage device, auxiliary storage device, etc.) or outside the computer.
[0100] 8 is a block diagram showing an example of the hardware configuration of each device (3D reconstruction systems 100 and 100a) in the above-mentioned embodiment. Each device may be realized as a computer 7 including, for example, a processor 71, a main storage device 72 (memory), an auxiliary storage device 73 (memory), a network interface 74, and a device interface 75, which are connected via a bus 76.
[0101] Although the computer 7 in FIG. 8 includes one of each component, it may also include multiple of the same component. Although FIG. 8 shows one computer 7, the software may be installed on multiple computers, and each of the multiple computers may execute the same or different parts of the software. In this case, a distributed computing configuration may be used in which each computer communicates via a network interface 74 or the like to execute processing. That is, each device in the above-described embodiment (the three-dimensional reconstruction systems 100 and 100a) may be configured as a system in which one or more computers execute instructions stored in one or more storage devices to realize its functions. Furthermore, the system may be configured such that information transmitted from a terminal is processed by one or more computers provided on a cloud, and the processing results are transmitted to the terminal.
[0102] The various calculations of each device (3D reconstruction systems 100 and 100a) in the above-described embodiments may be executed in parallel using one or more processors, or using multiple computers via a network. Furthermore, the various calculations may be distributed to multiple processing cores in a processor and executed in parallel. Furthermore, some or all of the processes, means, etc. disclosed herein may be realized by at least one of a processor and a storage device provided on a cloud that can communicate with computer 7 via a network. Thus, each device in the above-described embodiments may be implemented in the form of parallel computing using one or more computers.
[0103] The processor 71 may be an electronic circuit (processing circuit, processing circuitry, CPU, GPU, FPGA, ASIC, etc.) that at least controls a computer or performs calculations. The processor 71 may be a general-purpose processor, a dedicated processing circuit designed to perform a specific calculation, or a semiconductor device that includes both a general-purpose processor and a dedicated processing circuit. The processor 71 may also include an optical circuit or a calculation function based on quantum computing.
[0104] The processor 71 may perform arithmetic processing based on data or software input from each device or the like configured inside the computer 7, and may output the calculation results or control signals to each device or the like. The processor 71 may control each component constituting the computer 7 by executing the OS (Operating System) of the computer 7, applications, etc.
[0105] Each device (3D reconstruction systems 100 and 100a) in the above-described embodiments may be realized by one or more processors 71. Here, the processor 71 may refer to one or more electronic circuits arranged on one chip, or may refer to one or more electronic circuits arranged on two or more chips or two or more devices. When multiple electronic circuits are used, the electronic circuits may communicate with each other via wire or wirelessly.
[0106] The main memory device 72 may store instructions to be executed by the processor 71, various data, etc., and information stored in the main memory device 72 may be read by the processor 71. The auxiliary memory device 73 is a memory device other than the main memory device 72. Note that these memory devices refer to any electronic component capable of storing electronic information, and may be semiconductor memory. The semiconductor memory may be either volatile memory or non-volatile memory. The memory device for saving various data, etc. in each device (the three-dimensional reconstruction systems 100 and 100a) in the above-described embodiments may be realized by the main memory device 72 or the auxiliary memory device 73, or may be realized by an internal memory built into the processor 71.
[0107] When each device (3D reconstruction systems 100 and 100a) in the above-described embodiments is configured with at least one storage device (memory) and at least one processor connected (coupled) to this at least one storage device, at least one processor may be connected to one storage device. Also, at least one storage device may be connected to one processor. Also, a configuration in which at least one processor among multiple processors is connected to at least one storage device among multiple storage devices may be included. Also, this configuration may be realized by storage devices and processors included in multiple computers. Furthermore, a configuration in which a storage device is integrated with a processor (for example, a cache memory including an L1 cache and an L2 cache) may be included.
[0108] The network interface 74 is an interface for connecting to the communication network 8 wirelessly or via a wire. The network interface 74 may be an appropriate interface, such as one that conforms to an existing communication standard. Information may be exchanged with an external device 9A connected via the communication network 8 via the network interface 74. The communication network 8 may be any one of a WAN (Wide Area Network), a LAN (Local Area Network), a PAN (Personal Area Network), etc., or a combination thereof, as long as information is exchanged between the computer 7 and the external device 9A. An example of a WAN is the Internet, an example of a LAN is IEEE802.11 or Ethernet (registered trademark), and an example of a PAN is Bluetooth (registered trademark) or NFC (Near Field Communication), etc.
[0109] The device interface 75 is an interface such as a USB that directly connects to the external device 9B.
[0110] The external device 9A is a device connected to the computer 7 via a network. The external device 9B is a device connected directly to the computer 7.
[0111] For example, the external device 9A or the external device 9B may be an input device. The input device may be a device such as a camera, a microphone, a motion capture device, various sensors, a keyboard, a mouse, or a touch panel, and provides acquired information to the computer 7. Alternatively, the external device 9A or the external device 9B may be a device equipped with an input unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.
[0112] Furthermore, the external device 9A or the external device 9B may be, for example, an output device. The output device may be, for example, a display device such as an LCD (Liquid Crystal Display) or an organic EL (Electro Luminescence) panel, or a speaker that outputs sound or the like. Alternatively, the output device may be a device including an output unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.
[0113] Furthermore, the external device 9A or the external device 9B may be a storage device (memory). For example, the external device 9A may be a network storage or the like, and the external device 9B may be a storage such as an HDD.
[0114] Furthermore, the external device 9A or the external device 9B may be a device having some of the functions of the components of each device (the three-dimensional reconstruction systems 100 and 100a) in the above-described embodiments. That is, the computer 7 may transmit some or all of the processing results to the external device 9A or the external device 9B, or may receive some or all of the processing results from the external device 9A or the external device 9B.
[0115] In this specification (including the claims), when the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used, it includes any of a, b, c, ab, ac, bc, or abc. It may also include multiple instances of any element, such as aa, abb, aabbcc, etc. Furthermore, it also includes the addition of elements other than the enumerated elements (a, b, and c), such as having d, as in abcd.
[0116] In this specification (including claims), when expressions such as "using data as input / based on / according to / in response to data" (including similar expressions) are used, unless otherwise specified, this includes cases where the data itself is used, or where data that has been processed in some way (e.g., data with noise added, normalized data, features extracted from data, intermediate representations of data, etc.) is used. Furthermore, when a statement is made that a result is obtained "using data as input / based on / according to / in response to data" (including similar expressions), this includes cases where the result is obtained based solely on the data, or where the result is influenced by other data, factors, conditions, and / or states other than the data itself, unless otherwise specified. Furthermore, when a statement is made that "data is output" (including similar expressions), this includes cases where the data itself is used as output, or where data that has been processed in some way (e.g., data with noise added, normalized data, features extracted from data, intermediate representations of various data, etc.) is used as output, unless otherwise specified.
[0117] When the terms "connected" and "coupled" are used in this specification (including the claims), they are intended as open-ended terms that encompass any of direct connection / coupling, indirect connection / coupling, electrically connection / coupling, communicatively connection / coupling, functionally connection / coupling, and physically connection / coupling. These terms should be interpreted appropriately according to the context in which they are used, but any form of connection / coupling that is not intentionally or naturally excluded should be interpreted as being included in these terms without limitation.
[0118] In this specification (including the claims), the expression "A configured to B" may include the physical structure of element A having a configuration capable of performing operation B, and the permanent or temporary setting / configuration of element A being configured / set to actually perform operation B. For example, if element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and is configured to actually perform operation B by setting a permanent or temporary program (instruction). Also, if element A is a dedicated processor, dedicated arithmetic circuit, etc., it is sufficient that the circuit structure, etc. of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached.
[0119] Whenever words implying containing or possessing (e.g., "comprising / including," "having," etc.) are used in this specification (including the claims), they are intended to be open-ended terms that include the inclusion or possession of things other than the object designated by the object of the term. When the object of such words implying containing or possessing does not specify a quantity or suggests a singular number (e.g., expressions using the articles "a" or "an"), the expression should be construed as not being limited to a specific number.
[0120] In this specification (including the claims), even if expressions such as "one or more" and "at least one" are used in some places and expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") should be interpreted as not necessarily being limited to a specific number.
[0121] In this specification, when a particular advantage / result is described as being obtained with respect to a particular configuration of an embodiment, it should be understood that the same advantage / result can also be obtained with one or more other embodiments having the same configuration, unless otherwise stated. However, it should be understood that the presence or absence of the effect generally depends on various factors, conditions, and / or circumstances, and that the effect is not necessarily obtained with the configuration. The effect is merely obtained by the configuration described in the embodiment when various factors, conditions, and / or circumstances are satisfied, and the effect does not necessarily occur in a claimed invention that defines the same or a similar configuration.
[0122] In this specification (including claims), when multiple pieces of hardware perform a predetermined process, the pieces of hardware may cooperate to perform the predetermined process, or some of the hardware may perform all of the predetermined process. Furthermore, some of the hardware may perform part of the predetermined process, and other hardware may perform the rest of the predetermined process. In this specification (including claims), when an expression such as "one or more pieces of hardware perform a first process, and the one or more pieces of hardware perform a second process" (including similar expressions) is used, the hardware performing the first process and the hardware performing the second process may be the same or different. In other words, it is sufficient that the hardware performing the first process and the hardware performing the second process are included in the one or more pieces of hardware. Note that hardware may include electronic circuits, devices including electronic circuits, etc.
[0123] In this specification (including the claims), when multiple storage devices (memories) store data, each of the multiple storage devices may store only a portion of the data, or may store the entire data. Also, a configuration in which only some of the multiple storage devices store data may be included.
[0124] Although the embodiments of the present disclosure have been described in detail above, the present disclosure is not limited to the individual embodiments described above. Various additions, modifications, substitutions, partial deletions, etc. are possible within the scope of the conceptual idea and spirit of the present invention, which is derived from the content defined in the claims and their equivalents. For example, when numerical values or formulas are used in the above-described embodiments, they are shown for illustrative purposes and do not limit the scope of the present disclosure. Furthermore, the order of each operation shown in the embodiments is also illustrative and does not limit the scope of the present disclosure.
Claims
1. 1. A method for three-dimensionally reconstructing an object based on a plurality of images obtained by photographing the object from different directions, comprising: A three-dimensional reconstruction method, wherein a background image region of the object included in each of the plurality of images has a pattern.
2. The three-dimensional reconstruction method according to claim 1 , wherein background image regions of the object included in each of the plurality of images have the same pattern.
3. The three-dimensional reconstruction method according to claim 1 , wherein the plurality of images are obtained by photographing the object against the same background member having a pattern.
4. The three-dimensional reconstruction method according to claim 3 , wherein the plurality of images are obtained by photographing the object rotated with the background member fixed behind it.
5. The three-dimensional reconstruction method according to claim 4 , wherein at least two of the plurality of images are obtained by capturing the object with the background member as a backdrop while the object is rotated at different angles.
6. The three-dimensional reconstruction method according to claim 1 , wherein the pattern is a black and white design.
7. 7. The three-dimensional reconstruction method according to claim 6, wherein the pattern is a design in which white dots are arranged on a black background.
8. 6. The three-dimensional reconstruction method according to claim 1, wherein the three-dimensional reconstruction of the object is performed based on data representing the shape and color of the object, the data being treated as separate data from data representing a background of the object.
9. A camera that can capture an object from different directions, a control device configured to output reconstruction information of the object based on a plurality of images taken by the camera; A three-dimensional reconstruction system, wherein a background image region of the object included in each of the plurality of images has a pattern.
10. The three-dimensional reconstruction system according to claim 9 , wherein each of the plurality of images includes an object image region that changes depending on the imaging direction, and the background image region that remains the same regardless of the imaging direction.