Method of computer vision and system

The method employs neural networks to process estimated shapes from multiple viewpoints, sample surface points, and render three-dimensional reconstructions, effectively addressing the challenges of inaccurate shape estimation in existing photometric stereo methods by producing accurate and realistic 3D reconstructions.

JP2025077980AActive Publication Date: 2025-05-19KK TOSHIBA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024122420
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-07
Filing Date
2024-07-29
Publication Date
2025-05-19
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

Existing methods for three-dimensional reconstruction of objects using photometric stereo face challenges due to non-linear dependencies on reflection types and material properties, leading to inaccurate shape estimations, especially for objects without color texture or geometric features.

Method used

A computer vision method involving the use of neural networks to process estimated shapes of objects from multiple line-of-sight directions, sampling points along these directions to identify surface points, and rendering these points to generate a three-dimensional reconstruction, with the neural network being trained on defects to improve accuracy.

Benefits of technology

This approach enables the generation of accurate and realistic three-dimensional reconstructions of objects, even those without color texture or geometric features, by performing efficient surface optimization and learning material properties, thus overcoming limitations of existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025077980000001_ABST
    Figure 2025077980000001_ABST
Patent Text Reader

Abstract

To provide a method of computer vision for generating three-dimensional reconfiguration of an object.SOLUTION: A method comprises the steps of: receiving at least one estimated shape of an object; providing the at least one estimated shape of the object generated on the basis of capture of one or more images from the corresponding visual line direction by using a plurality of point light sources as input to a first neural network; sampling one or more points along the corresponding visual line direction by using the first neural network to identify the plurality of points on the surface of the object; rendering each of the plurality of identified points to generate three-dimensional reconfiguration of the object; and generating at least one defect to train the first neural network at least partially on the basis of the at least one defect.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to methods and systems for computer vision. In particular, the present disclosure relates to performing three-dimensional reconstruction of objects.

Background Art

[0002] Computer vision tasks include tasks such as acquisition, analysis, and processing of image data. Some computer vision systems can perform three-dimensional reconstruction of objects by performing these tasks. However, photometric stereo is a classical problem in computer vision. Photometric stereo is a technique for estimating local geometric features such as the surface normal of an object and the depth of an object. Photometric stereo utilizes the relationship between the orientation of the surface (e.g., with respect to a light source and an observer) on an object and the amount of light reflected by the surface.

[0003] Photometric stereo can involve extracting the three-dimensional shape of an object from the light reflected from the object. This can often be difficult, especially due to its non-linear dependence on the type of reflection and material properties (e.g., the material properties of the object). Some existing methods use single-view photometric stereo to estimate the shape of an object. The three-dimensional shape estimated using such methods is not always accurate. Recent advancements have been made towards three-dimensional reconstruction of objects, but the quality and practicality of the estimated shapes have not been satisfactory. Therefore, accurate three-dimensional reconstruction of objects is needed. In addition, there is an unmet need for accurate three-dimensional reconstruction of objects and / or scenes that do not have color texture or geometric features.

Summary of the Invention

[0004] In one aspect, a computer vision method for generating a three-dimensional reconstruction of an object is provided, the method comprising Receiving at least one estimated shape of an object, the at least one estimated shape being generated based on capturing one or more images from a corresponding line-of-sight direction using a plurality of point light sources, Providing the at least one estimated shape of the object as an input to a first neural network, the first neural network being initialized to calculate a representation of the object, Sampling one or more points along the corresponding line-of-sight direction using the first neural network to identify a plurality of points on the surface of the object, Rendering each of the identified plurality of points to generate a three-dimensional reconstruction of the object, wherein rendering generates at least one defect and the first neural network is trained at least partially based on the at least one defect.

[0005] In one example, the representation is a surface representation of the object. In some variations, the surface representation of the object is a height map. Sampling one or more points may include fitting the height map to at least one estimated shape of the object. In some variations, the surface representation of the object is a neural function, where the neural function is represented by the first neural network. In some variations, the at least one estimated shape of the object comprises at least one estimated surface of the object, wherein the first neural network maps points on a plane to heights from the surface of the object. In some variations, the first neural network may map points on the surface to the surface depth of the object.

[0006] In another example, the representation is a three-dimensional geometric representation of an object. In some variations, at least one estimated shape of the object comprises at least one estimated three-dimensional geometry of the object, where the first neural network maps points in three-dimensional space to a signed distance field of points on the object. The method may further comprise calculating a surface representation of the object from the calculated three-dimensional geometric representation of the object. In some variations, the three-dimensional geometric representation of the object is a neural function, where the neural function is represented by the first neural network.

[0007] The method may further comprise performing backpropagation of at least one loss and adjusting one or more weights of the first neural network.

[0008] In one example, rendering is implemented by a second neural network. The method may further comprise performing backpropagation of at least one loss and adjusting one or more weights of the second neural network. The second neural network is further configured to learn the material of the object.

[0009] In one example, rendering may further comprise rendering at least one of a surface normal and an intensity at each of the identified plurality of points. In one example, at least one estimated shape is generated using a third neural network.

[0010] In one example, receiving at least one estimated shape may comprise receiving a first estimated shape and a second estimated shape, the first estimated shape being generated based on a first set of a plurality of images captured from a first line-of-sight direction, and the second estimated shape being generated based on a second set of a plurality of images captured from a second line-of-sight direction. The method may further comprise concatenating the first estimated shape and the second estimated shape to generate a concatenated estimated shape, and providing the concatenated estimated shape as an input to a first neural network.

[0011] In an example where the representation of the object is a surface representation of the object, the method may further comprise calculating at least one surface normal of the object based on a partial derivative of one or more weights of the first neural network and at least one estimated shape of the object. Rendering may include rendering at least one surface normal calculated at each of the identified plurality of points.

[0012] In one example, sampling one or more points along a corresponding line-of-sight direction includes determining a plurality of points at the intersection of incident light rays from an image capture device that captures one or more images and the surface of the object, and sampling the determined plurality of points to identify a plurality of points on the surface of the object, where the incident light rays pass through the center of each pixel.

[0013] The method may also comprise determining a rendering defect based on a difference between a three-dimensional reconstruction of the object and the actual object.

[0014] A carrier medium may be provided that carries computer-readable instructions adapted to cause a computer to execute the method described above.

[0015] In one aspect, a system for three-dimensional reconstruction of an object may comprise an interface and a processor. The interface has an image input and is configured to receive one or more images of an object from corresponding line-of-sight directions captured using a plurality of point light sources. The processor receives at least one estimated shape of the object, where the at least one estimated shape is generated based on one or more images of the object. provides the at least one estimated shape of the object as an input to a first neural network, where the first neural network is initialized to compute a representation of the object. uses the first neural network to sample one or more points along the corresponding line-of-sight direction and identify a plurality of points on the surface of the object. is configured to render each of the identified plurality of points and generate a three-dimensional reconstruction of the object, where rendering generates at least one defect and the first neural network is trained at least partially based on the at least one defect.

[0016] Next, a system and method by way of non-limiting examples will be described with reference to the accompanying drawings.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

[0018] This specification discloses methods and systems for computer vision. More particularly, methods and systems for computer vision for three-dimensional reconstruction of objects are disclosed herein. As used herein, the term "object" may refer to what is being imaged. However, it will be readily understood that the term "object" may refer to multiple objects, scenes, combinations thereof, and the like. As used herein, "accurate three-dimensional reconstruction" may refer to an accurate reconstruction of the shape of an object. For example, "accurate three-dimensional reconstruction" may refer to a realistic reconstruction of the shape of an object.

[0019] Accurately reconstructing the three-dimensional (3D) structure of an object can be difficult. Such challenges often stem from problems related to photometric stereo. Some existing methods utilize multiple lines of sight (e.g., by using multiple cameras to capture images from multiple lines of sight) to generate a 3D reconstruction of the object. For example, some existing methods involve image-based models that extract the depth between the object and the cameras of the image capture device. The color texture at each pixel can be projected according to the extracted depth to generate a 3D reconstruction of the object. However, texture depends on light reflection, and thus, a realistic 3D reconstruction of the object may not always be possible. Additionally, the color texture depends on the material properties of the object. That is, the same image of an object may appear substantially different from different lines of sight.

[0020] Some existing methods combine 3D scanning techniques such as, for example, laser scanners and structured light scanners. However, these approaches may require merging data from different sensors. Merging data can lead to several problems, such as loss of accuracy, limited resolution of the 3D reconstruction, and noisy and incomplete image data (e.g., scans) of shiny and metallic objects. Additionally, merging is often performed at the per-pixel level, which can be a difficult task.

[0021] Until recently, methods that utilize multiple lines of sight did not include learnable components (e.g., neural networks). Such methods often lead to errors. Recent advancements have been made in utilizing multiple lines of sight by including learnable components. However, such learnable components can be configured to learn the textured volume of an object. Therefore, these methods are not efficient and are not adjusted to perform surface optimization procedures. More specifically, the techniques described herein include learnable components for performing surface parameterization. This enables the techniques described herein to implement another neural network to render a 3D reconstruction of an object and backpropagate the deficiencies from the data, thereby making the techniques described herein efficient. This is particularly possible because, unlike existing methods with learnable components, the neural surface is completely smooth and differentiable. Additionally, such methods are often not feasible when implemented for applications such as robotic interaction, conveyor belt scanning, etc. In particular, implementing these methods may not be feasible due to the long capture time of images and precise camera pose calibration. Furthermore, such methods are more vulnerable, less concise, and significantly more complex than the techniques disclosed herein.

[0022] The techniques disclosed herein can generate an accurate (e.g., realistic) 3D reconstruction of an object that has no color texture or geometric features. In contrast to existing methods, the techniques disclosed herein enable better adjusted and efficient surface optimization procedures.

[0023] In one example, the techniques disclosed herein may comprise receiving at least one estimated shape of an object. The estimated shape may be generated based on the capture of one or more images from corresponding line-of-sight directions using a plurality of point light sources. For example, an image of the object may be captured in a controlled manner to generate an estimated shape of the object. For example, a plurality of point light sources (e.g., LEDs) may be used in a calibrated manner. The images may be captured using a binocular illuminance difference setting. In other words, the image capture device may include two cameras for capturing an image of the object under controlled illumination conditions (e.g., using a plurality of point light sources). Capturing the images in such a controlled and calibrated manner enables the techniques disclosed herein to generate an accurate estimated shape of an object with a general reflectivity.

[0024] At least one estimated shape of the object is provided as an input to the first neural network. In some variations, the first neural network is a deep neural network. The first neural network can be initialized to compute a representation of the object (e.g., a surface representation of the object and / or a 3D geometric representation of the object). More specifically, the techniques disclosed herein formulate the task of generating a 3D reconstruction of an object as the learning of a differentiable surface representation and texture representation, and / or a 3D geometric representation and texture representation. This can minimize the discrepancy between local geometric features (e.g., surface normals) estimated for two viewing directions using varying illumination conditions. This can also minimize the discrepancy between the rendered surface intensity (e.g., rendered as discussed below) and the observed image. As an example, the surface of the object can be represented as a neural height map where the height of points on the surface of the object can be computed using the first neural network. Representing the surface as a neural height map can enable a better tuned and more efficient surface optimization procedure compared to existing methods that utilize volume-based representations. In some variations, the first neural network can be initialized to compute a 3D geometric representation of the object. For example, the geometry of the object can be represented as a neural signed distance field where the signed distance field of points on the geometry of the object can be computed using the first neural network. In such variations, the surface representation of the object can be derived from the 3D geometric representation of the object.

[0025] The technology disclosed in this specification may further comprise using a first neural network to sample one or more points along corresponding lines of sight to identify a plurality of points on the surface of an object. Each of the identified plurality of points may be rendered to generate a three-dimensional reconstruction of the object (e.g., by rendering surface normals, surface depths, a three-dimensional representation of the object, combinations thereof, etc.). Rendering the identified plurality of points may result in one or more deficiencies, as further described herein. In one example, these deficiencies may be backpropagated to the first neural network. The first neural network may be trained based on these deficiencies. In some examples, the identified plurality of points may be rendered by implementing a second neural network. This second neural network may be configured to learn the material of the object (e.g., based on the rendered intensity deficiencies). In other words, instead of predicting the individual intensity for each point, a second neural network is used to render the identified plurality of points based on the learned material and albedo parameters of the object. Since the second neural network learns the material of the object based on robust intensity deficiencies, the technology disclosed herein is less susceptible to the effects of intensity overfitting than direct intensity matching deficiencies. In particular, robust intensity deficiencies preserve the peak of the reflection lobe and are thus less susceptible to the effects of intensity overfitting and limitations of the renderer quality than direct intensity matching deficiencies.

[0026] Thus, in some examples, the techniques described herein include estimating a base shape for each line of sight of an object by using an illuminance difference stereo normal estimated for each line of sight, using the estimated base shape for each line of sight of the object to initialize a neural height map network (and / or a neural signed distance field network) derived from the estimated normal and depth for each pixel, and subsequently performing a robust fitting of the initialized neural height map to the image intensity and the estimated normal information.

[0027] The techniques disclosed herein can be utilized for various applications. For example, the techniques disclosed herein can be used for accurate quality control measurements of shiny objects. As another example, the techniques disclosed herein can also be used for depth perception in robotic applications. As yet another example, the techniques disclosed herein can be used for biometric three-dimensional feature extraction (e.g., fingerprint patterns).

[0028] FIG. 1 shows a schematic diagram of a system that can capture three-dimensional (3D) image data of an object and be used to reconstruct the object.

[0029] The 3D image data of object 10 is captured using device 11. Further details of device 11 are provided with reference to FIG. 2. The 3D image captured by device 11 may be provided to computing device 12, where it is processed. In some variations, computing device 12 may include an interface for receiving the 3D image captured by device 11. In some examples, device 11 may be communicatively coupled to computing device 12 via a network (e.g., the Internet, a local area network (LAN), a wide area network (WAN), etc.). Some non-limiting examples of computing device 12 include computers (e.g., desktops, personal computers, laptops, etc.), tablets and e-readers (e.g., Apple iPad®, Samsung Galaxy® Tab, Microsoft Surface®, Amazon Kindle®, etc.), mobile devices and smartphones (e.g., Apple iPhone®, Samsung Galaxy®, Google Pixel®, etc.), and the like.

[0030] Computing device 12 includes at least one processor. The processor can be any suitable processing device configured to launch and / or execute a set of instructions or code, and may include one or more data processors, image processors, graphics processing units, digital signal processors, and / or central processing units. The processor can be, for example, a general-purpose processor, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or the like.

[0031] When executed, the processor causes the processor to: (1) generate at least one estimated shape of an object based on the capture of one or more images from one or more viewing directions using a plurality of point light sources; (2) initialize a first neural network to calculate a representation of the object (e.g., a surface representation and / or a three-dimensional geometric representation); (3) use at least one estimated shape as an input to the first neural network to calculate a representation of the object (e.g., a surface representation and / or a three-dimensional geometric representation); (4) use the first neural network to sample one or more points along one or more viewing directions to identify a plurality of points on the surface of the object; or (5) render each of the identified plurality of points to generate a 3D reconstruction of the object (e.g., rendering may be based on implementing a second neural network). The instructions may include one or more of the above operations that can be performed.

[0032] FIG. 2 shows an exemplary arrangement of the apparatus 11 of FIG. 1 that can be used for binocular illuminance difference stereo. More specifically, the systems and methods described herein use image data (e.g., an image of an object) captured in a controlled manner from two different viewing directions to generate an estimated shape of the object. FIG. 2 shows an exemplary variant of a controlled setting for capturing image data for generating an estimated shape.

[0033] FIG. 2 shows a mount that holds a first camera 21 and a second camera 22. The mount further holds a plurality of light sources 23 in a fixed relationship with each other and with the first camera 21 and the second camera 22. This arrangement of the apparatus 20 allows the first camera 21, the second camera 22, and the plurality of light sources 23 to move together while maintaining a constant separation from each other.

[0034] In FIG. 2, the light source 23 is provided so as to surround the cameras 21 and 22. However, it will be readily understood that the light source 23 and the cameras 21 and 22 can be provided in any suitable arrangement. The cameras 21 and 22 are used together with the light source 23 to obtain binocular illuminance difference stereo data of the object 10. The individual light sources of the plurality of light sources 23 can be sequentially activated so that the cameras 21 and 22 can capture binocular illuminance difference stereo data.

[0035] In a specific arrangement of the device 11, an FLEA3.2 megapixel camera equipped with an 8 mm lens can be used as the cameras 21 and 22. The cameras 21 and 22 may have an 8 mm lens and can be firmly attached to a printed circuit board. The printed circuit board may further include 15 white bright LEDs, which are used as the light source 23, arranged on the same plane as the image plane, and provided around the cameras 21 and 22 at intervals of up to 6.5 centimeters.

[0036] The device 11 can be used to capture a binocular illuminance difference stereo image of the object 10. The object 10 is located in front of the cameras 21 and 22 within the depth range of the cameras 21 and 22. The depth range, also called the depth of field (DOF), is a term used to indicate the range of distances in focus on the object 10 between each of the cameras 21 and 22 and the object 10. If the object 10 is outside the respective depth ranges of the cameras 21 and 22, too close to or too far from the cameras 21 and 22, the object will be out of focus and details cannot be resolved. For example, the cameras 21 and 22 equipped with 8 mm lenses may have a depth range between 5 and 30 centimeters.

[0037] The light sources 23 are individually and sequentially activated to enable each of the cameras 21 and 22 to capture illuminance difference stereo data of the object under different illumination conditions. This is achieved by switching on only one light source at a time. For example, the device 11 may comprise 15 LEDs, and thus a set of 15 illuminance difference stereo images can be captured by each of the cameras 21 and 22.

[0038] In illuminance difference stereo, each of the light sources 23 is individually activated with respect to the object 10 to capture illuminance difference stereo data of the object under different illumination conditions. This is achieved by switching on only one light source at a time. Then, the amount of light reflected onto each pixel of the cameras 21 and 22 from a known surface or object is measured.

[0039] Figure 3 shows an example of a binocular illuminance difference stereo data capture setup described in Figure 2. The mount 30 holds the cameras 31 and 32. The mount 30 also holds a plurality of point light sources (e.g., LEDs) fixed as described in Figure 2. The cameras 31 and 32 capture images of the object 10 under different illumination conditions from each other.

[0040] FIG. 4 is a flowchart showing an overview of an exemplary method 400 for generating a 3D reconstruction of an object. At step 401, method 400 may include receiving at least an estimated shape of the object. The estimated shape of the object may be generated based on the capture of one or more images from corresponding line-of-sight directions using a plurality of point light sources. For example, the (image) data capture settings shown in FIGS. 2 and 3 may be used to capture one or more images. In particular, the images may be captured using two cameras (e.g., camera 21 and camera 22) from two line-of-sight directions. A plurality of point light sources (e.g., bright LEDs) (e.g., a plurality of point light sources 23) may be used to capture one or more images in a controlled manner under various lighting conditions. For example, if the data capture settings include 15 point light sources, each point light source may be switched on one at a time. Each time a point light source is switched on, a first image may be captured by a first camera (e.g., camera 21) from a first line-of-sight direction, and a second image may be captured by a second camera (e.g., camera 22) from a second line-of-sight direction. Thus, using 15 point light sources, 15 images of the object may be captured by the first camera (e.g., camera 21) from the first line-of-sight direction, and 15 images of the object may be captured by the second camera (e.g., camera 22) from the second line-of-sight direction.

[0041] These images can be used, for example, to estimate local geometric features such as surface normals. The estimated local geometric features can be used to generate at least one estimated shape of the object. The local geometric features can be estimated in any suitable way. As an example, the local geometric features can be determined by implementing an approach based on a convolutional neural network (CNN) described in A cnn based approach for the point-light photometric stereo problem, Int. Journal of Computer Vision, 2022, by Fotios Logothetis, Roberto Mecca, Ignas Budvytis, and Roberto Cipolla (referred to herein as the "PX-Net approach"). In this approach, a CNN can be used to compute a normal map for each line of sight. This approach provides a general proximity field network for obtaining high-quality normal maps for calibrated proximity (and telephoto) field illuminance difference stereo setups. Additionally, an initial estimate of the geometry (as a depth map for each view) can also be obtained using this approach. The normal map is fused onto a unified surface by using normal guidance and feedback from one or more images captured in step 401. For example, in this approach, a per-pixel CNN is trained to predict the surface normal from reflectance samples. Depth can be computed by integrating the normal field to iteratively estimate the light direction and attenuation that can be used to compensate the input image for computing the reflectance samples of the next iteration.

[0042] Although a PX-Net approach is described to generate the estimated shape, it will be readily understood that any suitable approach, e.g., using an image capture setting including a point light source such as the image capture setting described herein (e.g., the setting in FIG. 2 and FIG. 3), may be used to generate the estimated shape of the object. For example, the estimated shape may be generated using a computer-aided design (CAD) model. In an example where two or more estimated shapes are generated (e.g., the PX-Net approach), the estimated shapes may be concatenated together to generate a single estimated shape. For example, the estimated shapes from two or more viewing directions may be concatenated together by fusing local geometric features into a unified surface.

[0043] In step 403, the method includes providing at least one estimated shape (received in step 401) as an input to a first neural network. The first neural network may be initialized to compute a representation (e.g., a surface representation and / or a three-dimensional geometric representation) of the object. In some examples, the first neural network may be initialized to compute a surface representation of the object. In other examples, the first neural network may be initialized to compute a three-dimensional geometric representation of the object (e.g., a 3D geometric representation such as a signed distance field). In other words, the surface representation and / or the three-dimensional geometric representation may be a learned representation. The estimated shape (e.g., an estimated surface and / or an estimated three-dimensional geometry) may be provided as an input to a learnable component (e.g., the first neural network). Thus, the surface representation and / or the three-dimensional geometric representation may be defined as a function of the estimated shape, where the function may be a learnable function configured to optimize its weights.

[0044] In an example where the first neural network is initialized to compute a surface representation, the surface of the object is represented by a continuous depth map z s =F(x s ,y s) can be expressed as, where the subscript s indicates surface coordinates. More specifically, the first neural network can map a point (e.g., any point) on a plane (e.g., the x, y plane) to the surface depth of an object. As used herein, surface depth and / or surface height can mean the height from a point on the surface of an object. In other words, surface depth and / or surface height can be the vertical distance from a point on the surface of an object. Thus, in step 403, the surface depth of a plurality of points on the object can be expressed as a neural function. The neural function is expressed by the first neural network. In the case of the image capture settings (i.e., binocular stereo cameras) shown in FIGS. 2 and 3, this coordinate system can be selected as the average between two cameras (i.e., a modified stereo system). To convert the surface coordinates (e.g., indicated by the subscript s) to the original camera coordinates (e.g., indicated by the subscript c), the rotation translation R c t c can be used. For example, it is as follows.

[0045] [Number]

[0046] The unknown function F is a first neural network, the purpose of which is to optimize its weights. In some variants, the surface normal of an object can be calculated based on the partial derivatives of one or more weights of the first neural network and at least one estimated shape. For example, in some variants, the first neural network may comprise a SIREN architecture (e.g., described in Implicit neural representations with periodic activation functions, by Vincent Sitzmann, Julien N.P.Martel, Alexander W.Bergman, David B.Lindell, and Gordon Wetzstein, In Proc.NeurIPS, 2020), where the SIREN architecture is a multi-layer perceptron with sine wave activation functions. Such an architecture can ensure that the surface is infinitely differentiable, and thus, the normal integration problem can correspond to solving a partial differential equation. In other words, the surface normal can be calculated as follows.

[0047] [Number]

[0048] As seen in Equation 2, the proportionality can be due to the unit amplitude constraint of the normal vector. Furthermore, the partial derivatives can be calculated using automatic differentiation and are thus functions of the network weights. In some variants, the first neural network can be trained using backpropagation (as described below) based on rendering deficiencies. The rendering deficiencies can be determined using one or more images captured at step 401. In some variants, the surface representation can be a height map of the object (e.g., a map of surface depth).

[0049] In an example where the first neural network is initialized to compute a three-dimensional geometric representation, the three-dimensional geometric representation of an object can be represented as a signed distance field of the object. For example, the geometric representation of an object can be parameterized as the zero-th level set of an implicit function (e.g., points in space where the implicit function is zero - exactly zero, and / or approximated to zero using the marching cubes method, ray sampling, and / or averaging procedures described in equations 7, 8, and 9 herein). The geometric meaning of the signed distance field is to assign a scalar value d to each point [x w ,y w ,z w T in space (where a negative sign indicates a point inside the surface).

[0050]

Number

[0051] The unknown function F is the first neural network, and its purpose is to optimize its weights. The subscript w indicates world coordinates. More specifically, the first neural network can map a point (e.g., any point) in three-dimensional space (e.g., x, y, z space) to the signed distance field of a point on the object. In other words, the signed distance fields of multiple points on the object can be represented as a neural function. The neural function is represented by the first neural network. In the case of the image capture setup (i.e., binocular stereo cameras) shown in Figures 2 and 3, this coordinate system can be selected as the average between the two cameras (i.e., a modified stereo system). To convert world coordinates (indicated, for example, by the subscript w) to the original camera coordinates (indicated, for example, by the subscript c), a rotation and translation R c t c can be used. For example, as follows.

[0052] ​ [Mathematics]

[0053] A signed distance field of any geometry can always be continuous, differentiable, and can satisfy the following unit amplitude gradient eiconal equation.

[0054] [Mathematics]

[0055] For a surface point s where F(x s , y s , z s ) = 0, the surface normal can be the gradient of the signed distance field.

[0056] [Mathematics]

[0057] The above equation provides a direct relationship between the signed distance field and the surface normal. Thus, in these examples, the first neural network can be trained based on rendering defects (described below) from one or more images captured in step 401 (e.g., using backpropagation). After the weights of the first neural network are optimized, the calculated signed distance field may be sampled at a regular grid of points, and the triangular mesh surface can be restored using the standard marching cubes method. In this way, the surface representation of the object can be calculated from the 3D geometric representation of the object.

[0058] Figure 5 shows an exemplary implementation of steps 401 and 403 of method 400 described in FIG. 4. As can be seen in FIG. 5, images 58a and 58b show images captured from a pair of cameras using the image capture settings described herein (e.g., the image capture settings shown in FIGS. 2 and 3). In other words, image 58a can be the first image captured by the first camera (e.g., camera 21 in FIG. 2) from the first line-of-sight direction. Image 58b can be the second image captured by the second camera (e.g., camera 22 in FIG. 2) from the second line-of-sight direction. These images 58a and 58b can be used to generate an estimated shape of the object. For example, 62a and 62b show the normal maps for each line of sight (e.g., the normal maps for each of the first line-of-sight direction and the second line-of-sight direction). 64a and 64b show the estimated shape of the object for each line of sight (e.g., the estimated shape for each of the first line-of-sight direction and the second line-of-sight direction). These estimated shapes 64a and 64b for each line of sight can be received by a processor (e.g., the processor included in computing device 12 in FIG. 1). The estimated shapes 64a, 64b for each line of sight can be provided to a CNN (e.g., the first neural network). The first neural network can be randomly initialized as shown at 68. At 70, the first neural network can be initialized based on the estimated shapes 64a and 64b for each line of sight and the estimated normals 62a and 62b for each line of sight. As described above, in some examples, the first neural network is configured to compute a surface representation of the object. For example, the first neural network can compute a height map of the object. The first neural network can be trained to optimize its weights. In some variations, the first neural network can be trained (e.g., at 72) based on a comparison between the neural height map generated at 70 and the first image 58a and the second image 58b captured by the image capture settings.

[0059] Referring back to FIG. 4, at step 405, the method includes using a first neural network to sample one or more points along a corresponding line of sight direction to identify a plurality of points on the surface of the object. For example, sampling can include determining one or more points at the intersection of the incident light rays from the cameras (e.g., cameras 21 and 22 that capture the image at step 401) passing through the center of each pixel and the surface representation calculated at step 403. These determined points can be sampled to identify a plurality of points on the surface of the object.

[0060] In an example where the first neural network is initialized to calculate the surface representation at step 403, the projected depth z in Equation 1 c is unknown, so an exact transformation between the surface coordinate system and the camera coordinate system may not be possible. To overcome this, method 400 can consider tentative depth samples z ci 1, z c 2, z c 3,... to generate a point i along the line of sight direction v as {vz c}. The method can include applying coordinate transformation equation 1 to transform the generated points into the world coordinate system {[x ri , y ri , z ri}. The learnable function F can be queried at the position {[x si = F[x ri , y ri} to obtain the surface depth estimate {Z ri , y ri}. The set of surface depth estimates can be reduced as a weighted average using the average weights calculated through the disparities z e and z si and z ri to obtain the "actual" depth estimate z

[0061]

Number

[0062] Here, SM represents the softmax operator (e.g., exponent divided by the sum of exponents). The scaling factor -f can be used to obtain the minimum value of the correct scale (where the softmax of negative values is softmin) (e.g., to convert between millimeters within the sampling interval and the normalized unit). The surface normal n e can be estimated using similar softmax averaging. The surface normal n e is then projected back into camera coordinates using the transposed rotation R T ·n e (e.g., to compare with the original normal and apply the deficiencies (described below)).

[0063] In the case of multiple intersections of a ray with a surface (i.e., occlusion), the softmax distribution can become broad. Method 400 may detect these points and discard them (e.g., from the perspective of deficiency application described below).

[0064] In an example where a first neural network is initialized to calculate a 3D geometric representation in step 403, the signed distance field can be optimized by identifying the intersections of each ray from each camera (e.g., camera 21 and camera 22) pixel with the surface representation calculated in step 403. To achieve this, method 400 generates a point i along the line of sight direction v as {vz ci} by considering tentative depth samples z c 1, z c 2, z c 3,... The method may include applying coordinate transformation formula 4 to convert the generated points into the world coordinate system r i = {[x ri , y ri , z ri}. The learnable function F is the signed distance field value dr iTo obtain these, queries can be made at these positions. In some variations, the signed distance field value dr i can be converted to a differentiable volume permeability α using Equation 8 (where β is a trainable parameter and f is a scaling factor for converting between millimeters within the sampling interval and the normalized unit).

[0065]

Number

[0066] Thus, the intersection point p can be calculated as the weighted sum of the ray samples p = Σ i (w i r i ), where the weight w i of the sample corresponds to the accumulated opacity. That is,

[0067]

Number

[0068] The surface normal and albedo can be calculated in a similar way to the position p and re-projected into camera coordinates using the transposed rotation R T ·n e .

[0069] In step 407, method 500 includes rendering each point identified in step 405 to generate a 3D reconstruction of the object. Rendering each point identified in step 405 can result in one or more deficiencies, as further described herein. The points identified in step 405 can be rendered into any suitable representation of the object (e.g., surface normal, surface depth, 3D representation of the object, combinations thereof, etc.).

[0070] In some examples, rendering may be implemented by a second neural network. The second neural network may be trained for one or more deficiencies generated due to rendering the identified points. In other words, the second neural network may be trained to render each pixel based on the plurality of points identified in step 405. Since the second neural network is trained to render pixels, the second neural network may implicitly learn the material of the object. In some examples, the second neural network may also learn other characteristics of the object based on the deficiencies for which the second neural network continues to be trained. For example, the second neural network may learn characteristics such as the surface heat of the object, the grayscale representation of the object, the RGB representation of the object, etc. Thus, the second neural network may be configured to learn the material of the object. In some variations, the second neural network may be trained based on robust intensity deficiencies (described in further detail below), thereby learning the material of the object. The second neural network may be different from the first neural network described above.

[0071] To render the light intensity, first, the proximity field illumination vector l m and the light attenuation a m can be calculated for the point light source m. A proximity illumination model (such as the model described in Near Field Photometric Stereo with Point Light Sources, SIAM Journal on Imaging Sciences, 2014, by R. Mecca, A. Wetzler, A. Bruckstein, and R. Kimmel) can be implemented to do so. The illumination vector for each surface point (e.g., the points identified in step 405) p can be calculated as pl m =s m -p, where s mrepresents the position of the point light source m. The attenuation coefficient can be calculated as follows.

[0072] [Number]

[0073] Here,

[0074] [Number]

[0075] is the illumination direction, and

[0076] [Number]

[0077] is the operating direction of the point light source, and μ m is the angular deficit coefficient.

[0078] Next, the total intensity i m is, i m = s m · a m · BRDF(n, l m , v, mat) can be calculated as. s mis a "soft" indicator variable that is 0 in the case of a shadow point and 1 otherwise. To approximate the BRDF (i.e., bidirectional reflectance distribution function), a multilayer perceptron with 3×128 hidden layers having gelu activation and relu activation in the final layer (when m≧0) can be used. In some variants, the material vector mat can be 10-dimensional (e.g., based on the number of parameters of the BRDF). For example, this can include a single albedo excluding anisotropy. In some variants, the multilayer perceptron renderer can be trained on synthetic data from the PX-Net approach described above. The trained renderer should approximate the Disney shader, but robustness to noise is added (e.g., by data augmentation using synthetic data from the PX-Net approach). The surface point can be calculated as a weighted sum over the incident light rays (as described above), but the rendering can be calculated for only a single sample.

[0079] To estimate the shadow, each surface point p (e.g., the point identified in step 405) can be traced back to the point light source m by following the direction of the illumination vector

[0080]

Number

[0081] calculated above. For example, for each ray, 8 samples h can be obtained starting 3mm from the start and every 1.5mm. The depth of the surface representation (e.g., height map) may be queried, and for example, the difference d h =z h -z s can be calculated. If at least one of the differences is negative, it can indicate the presence of a shadow. Its differentiable approximation can be calculated as SM(-sigmoid(d h ))

[0082] Figure 6 shows an exemplary implementation of steps 405 and 407 of method 400 of FIG. 4. The intersection of the incident light rays from the cameras (e.g., camera 21 and camera 22) passing through the center of each pixel and the estimated surface representation is z r,1 , z r,2 , z r,3 can be represented by, and each surface value is z s,1 , z s,2 , z s,3 , and these are used to calculate the intersection point z e (Equation 9). Then, the surface point (the point z e’ is shown to avoid clutter) is traced back to the light source passing through the points z h,4 , z h,5 , z h,6 , and the surface is queried again to obtain z s,4 , z s,5 , z s,6 and calculate the shadow using the above equation. In this way, the surface normal, depth, and shadow can be estimated from the calculated surface representation of the object.

[0083] Method 400 may also include determining one or more deficiencies (e.g., angular normal deficiency, rendering deficiency, robust rendering deficiency, normal deficiency, depth deficiency, depth average deficiency, combination deficiency, iconal deficiency, free space regularization term, combinations thereof, etc.) based on the 3D reconstruction of the object. These deficiencies can be used to train the first neural network and / or the second neural network described herein. More specifically, these deficiencies can be used to optimize the performance of the first neural network and / or the second neural network described herein. In some examples, these deficiencies can be used to perform backpropagation to adjust the weights of the first neural network and / or the second neural network described herein.

[0084] Angular Normal Deficiency - Since the methods and systems described herein integrate surface normals into the surface, it may be important to calculate the angular deficiency. Thus, for the network normal n n and the surface normal n s the angular deficiency can be calculated as follows.

[0085]

Equation

[0086] The angle is measured in degrees. Additionally, the deficiency at each pixel can be weighted by max(n n ·v, 0). This can minimize the effect of the predicted normal on slanted surface points which are usually not very accurate.

[0087] Rendering Deficiency - The error in the rendered intensity (for all point lights m) can be calculated as follows.

[0088]

Equation

[0089] Robust Rendering Deficiency - The peaks (across all point lights) for both the rendered intensity and the actual intensity may have to be aligned. This property is albedo invariant and should be approximately material invariant (for all reflective materials, the peak of the reflection lobe should be aligned with the half vector). Thus, for the rendered intensity i r,m and the true intensity i t,m the robust rendering deficiency is as follows.

[0090]

Equation

[0091] The normal regularization term - The normal defect is extremely disadvantageous for oblique points, so it may be necessary to include the following regularization term.

[0092]

Number

[0093] Here, n w,z is the z-component of the surface normal (in world coordinates) and is encouraged to approach -1 (adding +1 ensures that the defect is always positive).

[0094] Depth defect - This can only be used for the initialization of the first neural network (e.g., SIREN architecture) described in this specification. Surface z s and target surface z t are considered (since there is no ground truth available yet, the target surface is considered).

[0095]

Number

[0096] Depth average defect - Even under mathematically perfect conditions, the integration constant is always ambiguous, so numerical integration is inherently inappropriate. Stereo set theory eliminates this problem (in principle, there is a unique surface that coincides with both lines of sight). However, numerical instability remains. To overcome this problem, an average depth regularization term can be used. Therefore, the defect with respect to the average depth can be added to the world coordinates as follows.

[0097]

Number

[0098] The average may be calculated for each mini-batch of data points, and z 0 is the average in the world coordinate system, and λ z = 0.1 is the regularization weight.

[0099] Combination Deficiency - When the above deficiencies are combined, the total deficiency is expressed as follows:

[0100]

Number

[0101] In some variants, empirically, α = 1.0, β = 1000, γ = 100, δ = 0.1, η = 0.1 can be set. Also, during the initialization of the first neural network (e.g., neural height map), φ can be set to φ = 1. However, the final major stage of training may not include depth deficiency.

[0102] In an example where the first neural network is initialized to calculate a three - dimensional geometric representation in step 403, the following deficiencies can also be calculated to train the first neural network and / or the second neural network described herein.

[0103] Iconal Deficiency - For all rays, to apply Equation 5, the iconal deficiency L e can be applied.

[0104]

Number

[0105] Free - space Regularization Term - To encourage the network to recover the minimum surface, the signed distance field values with positive signs are encouraged to increase and are suppressed from decreasing according to the following equation.

[0106]

Number

[0107] This deficiency can also be applied to all ray samples. Training In an example where the first neural network is initialized to calculate a surface representation in step 403, the first neural network can be trained in the following manner. The surface sampling procedure described in step 405 of FIG. 4 may require that the surface representation be properly initialized since the softmax distribution need not be overly broad. Thus, an initial estimate of the surface representation (e.g., depth map) can be unified in a single coordinate system (e.g., world coordinate system). A learnable function (e.g., SIREN function) can be pre-trained using normal and depth deficits. The initial depth maps may be inconsistent, and the neural network can converge at best to an estimate of their average. A data augmentation of ±1 mm on world coordinate points can be applied to smooth the initial surface (e.g., instead of generating two copies of the surface representation by the first neural network). The pre-initialization stage may be trained over 10 epochs, the main stage may be trained over 250 epochs, and the total training time is about 1 day per object.

[0108] In an example where the first neural network is initialized to calculate a 3D geometric representation in step 403, the first neural network can be trained in the following manner. Initialization stage - To accelerate convergence, the first neural network can be initialized with an initial surface estimate. To achieve this, an initial surface estimate for each line of sight can be calculated. This may be projected into the world coordinate system and fused with Poisson reconstruction to obtain an initial surface representation. Then, the distance transform of the surface can be calculated in a regular grid of 128×128×128 points P g around the object. (That is, for each of these grid points, the minimum distance to the surface point is the signed distance field estimate d(P g) can be calculated using brute - force search to have). The network can be trained to directly match the signed - distance fields at these points. That is, (using a simple absolute - value loss) F(P g ) = d(P g ). Dataset The methods and systems described herein were evaluated using two datasets, namely, the DiLiGenT dataset and the LUCES dataset.

[0109] DiLiGenT Dataset - The DiLiGeNT dataset contains five objects, namely, a bear, a Buddha statue, a cow, a pot2, and a reading, captured from 20 viewpoints using 96 light images at a resolution of 612×512. The objects are approximately 1.5 m away from all cameras (turntable - capture setup), and a camera focal length of 50 mm approximates the orthographic viewpoints. The selected viewpoints were viewpoints 3 and 4.

[0110] LUCES Stereo Dataset - This dataset was generated using the image - capture setup described herein (e.g., in FIGS. 2 and 3). In this example, the LUCES stereo dataset contains a subset of six objects (a bell, a rabbit, a camel, an owl, a queen, a mouse) out of the original 14 objects. The image - capture device (e.g., device 11) includes a 2x - Flea3 FL3 - U3 - 32S2C - CS1 / 2.8” Color USB3.0 Camera pointgray camera (2080×1552 px) and 15 LED lights as shown in FIG. 3. The camera lens was 8 mm. The objects were located 15 - 20 cm away from the camera (in contrast to DiLiGeNT) such that the proximity - illumination effect was prominent.

[0111] To measure performance, since a stereo pair of lines of sight is used to calculate the 3D reconstruction, the complete object Hausdorff distance may not be beneficial. Thus, to make a fair comparison, the ground truth (e.g., through back-projection of the ground truth depth map) can be calculated. The Hausdorff distance can be calculated from the ground truth to the reconstructed object. It can ensure that the metric is fair regardless of the reconstructed surface sampling (e.g., since the neural surface can be arbitrarily sampled). Thus, an error map can be calculated and shown, and different approaches can be compared.

[0112] In one example, the methods and systems disclosed herein were implemented by disabling various configurations of rendering and robust rendering defects. The surface normals estimated using the PX-Net approach were used to generate the estimated shape and perform surface fitting. The quantitative results are shown in Table 1.

[0113]

Table 1

[0114] As can be seen in Table 1, using both rendering defects and robust rendering defects significantly improves with respect to the use of normals consistent with the PS-Nerf approach. Using rendering shows significantly better performance for the bear and leading objects, while robust defects are good for the cow object. The better performance for the cow object can be explained by the non-uniform albedo that can prevent the rendering defect-based network from achieving better accuracy.

[0115] Figure 7 shows the evaluation of the methods and systems disclosed herein with respect to the DiLiGenT dataset. The results from the methods disclosed herein are compared to the single-viewpoint illuminance difference stereo approach (Logothetis et al.) and the PS-Nerf volumetric multi-viewpoint illuminance difference stereo approach. For each of these methods, the predicted shape and the predicted error map are visualized in Figure 7. As can be seen in Figure 7, the method disclosed herein is superior to the single-viewpoint method. The error map of the method disclosed herein looks relatively good, even for the pot and Buddha object where the shape is erroneously haloed, because the visualized error is calculated by finding the average PS-Nerf of the closest distance from the ground truth mesh to the predicted mesh. This emphasizes the strength of the learnable surface representation (e.g., neural height map) over the volumetric approach (e.g., PS-Nerf), because such halos of copies of parts of the object are not possible for the volumetric approach (e.g., PS-Nerf). Additionally, as can be seen in Figure 7, the PS-Nerf approach itself tries to prevent creating copies of parts of the object due to its extremely sparse supervision (e.g., pot 2 and reading object reference).

[0116] Figure 8 shows plots of the mean and median shape errors of the objects from the DiLiGenT dataset. These plots show the mean and median shape errors for all objects and the mean of the corresponding metrics across all objects. As can be seen in Figure 8, most of the median shape error metric is less than 1 mm.

[0117] Figure 9 shows the evaluation of the methods and systems disclosed herein for the LUCES stereo dataset. The first three rows of Figure 9 show the cropped image, the corresponding normals estimated from the PX-Net approach, and the ground truth shape. The last two rows show the shape predicted by the methods and systems disclosed herein and the corresponding error map. Similarly, Table 2 shows the quantitative evaluation of the methods and systems disclosed herein for the LUCES stereo dataset.

[0118]

Table 2

[0119] Table 2 also includes the results after scaled rigid body iterative closest point alignment (row 2). This is shown because the camera image planes in LUCES stereo are in a parallel configuration. This makes the task of recovering the true depth values an ambiguous one. After scaling, as seen in Table 2, the error for all objects except the bell is less than 1 mm.

[0120] Therefore, the techniques disclosed herein can generate accurate 3D reconstructions of objects from extremely sparse viewpoints. The techniques disclosed herein are far better than the single-viewpoint illuminance difference stereo approach. Additionally, the techniques disclosed herein do not suffer from the duplication of parts of the object that would occur for volume approaches such as PS-Nerf.

[0121] Although some embodiments of the present invention have been described, these embodiments are presented solely by way of example and are not intended to limit the scope of the invention. In fact, the novel devices and methods described herein may be embodied in various other forms, and furthermore, various omissions, replacements, and changes may be made in the forms of the devices, methods, and products described herein without departing from the spirit of the present invention. The appended claims and their equivalents are intended to cover such forms or modifications as fall within the scope and spirit of the present invention.

Claims

1. 1. A computer vision method for generating a three-dimensional reconstruction of an object, comprising: receiving at least one estimated shape of an object, the at least one estimated shape being generated based on capturing one or more images from corresponding line of sight directions using a plurality of point light sources; providing the at least one estimated shape of the object as an input to a first neural network, the first neural network being initialized to compute a representation of the object; sampling one or more points along the corresponding gaze direction using the first neural network to identify a plurality of points on a surface of the object; and rendering each of the identified plurality of points to generate a three-dimensional reconstruction of the object, wherein the rendering generates at least one defect, and the first neural network is trained based at least in part on the at least one defect.

2. The computer vision method of claim 1 , wherein the representation is a surface representation of the object.

3. The method of computer vision of claim 1 , wherein the representation is a three-dimensional geometric representation of the object.

4. 2. The method of claim 1, further comprising performing backpropagation of the at least one deficit to adjust one or more weights of the first neural network.

5. The computer vision method of claim 1 , wherein the rendering is implemented by a second neural network.

6. 6. The method of claim 5, further comprising performing backpropagation of the at least one deficit to adjust one or more weights of the second neural network.

7. The method of computer vision of claim 5 , wherein the second neural network is further configured to learn materials of the object.

8. The method of claim 1 , wherein the rendering comprises rendering at least one of a surface normal and an intensity at each point of the identified plurality of points.

9. The method of computer vision of claim 1 , wherein the at least one estimated shape is generated using a third neural network.

10. 2. The computer vision method of claim 1, wherein receiving the at least one estimated shape includes receiving a first estimated shape and a second estimated shape, the first estimated shape being generated based on a first set of a plurality of images captured from a first line of sight direction and the second estimated shape being generated based on a second set of a plurality of images captured from a second line of sight direction.

11. concatenating the first estimated shape and the second estimated shape to generate a concatenated estimated shape; providing the concatenated estimated shape as an input to the first neural network; The computer vision method of claim 10 further comprising:

12. 3. The method of computer vision of claim 2, further comprising: calculating at least one surface normal of the object based on partial derivatives of one or more weights of the first neural network and the at least one estimated shape of the object.

13. The method of computer vision of claim 12 , wherein the rendering comprises rendering the calculated at least one surface normal at each point of the identified plurality of points.

14. 2. The computer vision method of claim 1, wherein sampling the one or more points along the corresponding line of sight direction comprises determining a plurality of points at intersections of an incident ray from an image capture device that captures the one or more images with the surface of the object, and sampling the determined plurality of points to identify the plurality of points on the surface of the object, wherein the incident ray passes through a center of each pixel.

15. The method of computer vision of claim 1 , further comprising determining a rendering defect based on a difference between the three-dimensional reconstruction of the object and an actual object.

16. The computer vision method of claim 2 , wherein the surface representation of the object is a height map.

17. 20. The method of claim 16, wherein sampling the one or more points comprises fitting the height map to the at least one estimated shape of the object.

18. The method of computer vision of claim 2 , wherein the surface representation of the object is a neural function, wherein the neural function is represented by the first neural network.

19. 3. The method of computer vision of claim 2, wherein the at least one estimated shape of the object includes at least one estimated surface of the object, and the first neural network maps points on a plane to heights above the surface of the object.

20. The method of computer vision of claim 2 , wherein the first neural network maps points on a surface to a surface depth of the object.

21. 4. The method of computer vision of claim 3, wherein the at least one estimated shape of the object includes at least one estimated three-dimensional geometry of the object, and the first neural network maps points in three-dimensional space to a signed distance field of points on the object.

22. The method of computer vision of claim 3 , further comprising: computing a surface representation of the object from the computed three-dimensional geometric representation of the object.

23. The method of computer vision of claim 3 , wherein the three-dimensional geometric representation of the object is a neural function, wherein the neural function is represented by the first neural network.

24. A carrier medium carrying computer readable instructions adapted to cause a computer to perform the method of claim 1.

25. 1. A system for three-dimensional reconstruction of an object, comprising: an interface; and a processor; The interface has an image input and is configured to receive one or more images of an object from corresponding line of sight directions captured using a plurality of point light sources; The processor, receiving at least one estimated shape of the object, the at least one estimated shape being generated based on the one or more images of the object; providing the at least one estimated shape of the object as an input to a first neural network, the first neural network being initialized to compute a representation of the object; sampling one or more points along the corresponding gaze direction using the first neural network to identify a plurality of points on a surface of the object; and rendering each of the identified plurality of points to generate a three-dimensional reconstruction of the object, wherein the rendering generates at least one defect, and the first neural network is trained based at least in part on the at least one defect.

Citation Information

Patent Citations

  • Computer vision method and system

    JP2022032937A

  • Depth Estimation

    JP2022519194A