Grid reconstruction using data driven priors

By using a data-driven prior approach and employing a variational autoencoder to learn mesh prior information, combined with image observation and geometric constraints, the problem of insufficient reconstruction accuracy and efficiency in multi-view stereo vision technology is solved, and more efficient 3D model generation is achieved.

CN111445581BActive Publication Date: 2026-01-13NVIDIA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201911293872.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-12-19
Filing Date
2019-12-16
Publication Date
2026-01-13
Estimated Expiration
2040-06-18

AI Technical Summary

Technical Problem

Existing multi-view stereo vision technology struggles to effectively utilize the stereo correspondence between images when reconstructing 3D surfaces, resulting in insufficient reconstruction accuracy and efficiency.

Method used

A data-driven prior approach is adopted, which generates a machine learning model through a training engine, uses a variational autoencoder to learn prior information related to the mesh, and combines image observation and geometric constraints to refine the mesh reconstruction process and generate a more accurate 3D model.

Benefits of technology

It improves the accuracy and efficiency of 3D reconstruction, can better fit the shape of the object, reduces the dependence on global pose estimation, and enhances the adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111445581B_ABST
    Figure CN111445581B_ABST
Patent Text Reader

Abstract

A grid reconstruction using data driven priors is disclosed. One embodiment of a method includes predicting one or more three-dimensional (3D) mesh representations based on a plurality of digital images, wherein the one or more 3D mesh representations are refined by minimizing at least one discrepancy between the one or more 3D mesh representations and the plurality of digital images.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Multi-view stereo vision (MVS) technology involves constructing a three-dimensional (3D) surface from multiple overlapping two-dimensional (2D) images of an object. This technology can estimate the most probable 3D shape from 2D images based on assumptions related to texture, viewpoint, lighting, and / or other conditions associated with the captured images. Given a set of images of an object and corresponding assumptions, MVS uses stereo correspondences between the images to reconstruct the 3D geometry of the scene captured by the images. Attached Figure Description

[0002] Therefore, the above-described features of the various embodiments can be understood in more detail. A more specific description of the inventive concept briefly outlined above can be obtained by referring to the various embodiments, some of which are illustrated in the accompanying drawings. However, it should be noted that the drawings only show typical embodiments of the inventive concept and should not be considered as limiting its scope in any way, and other equivalent embodiments exist.

[0003] Figure 1 This is a block diagram illustrating a system configured to implement one or more aspects of various embodiments.

[0004] Figure 2 According to various embodiments Figure 1 A more detailed explanation of the training and execution engines.

[0005] Figure 3 This is a flowchart of method steps for performing mesh reconstruction using data-driven priors, according to various embodiments.

[0006] Figure 4 This is a flowchart of method steps for training a machine learning model to learn a data-driven grid prior, according to various embodiments.

[0007] Figure 5 It is a block diagram of a computer system configured to implement one or more aspects of various embodiments.

[0008] Figure 6 According to various embodiments Figure 5 A block diagram of the parallel processing units (PPUs) included in the parallel processing subsystem.

[0009] Figure 7 According to various embodiments Figure 6 A block diagram of the General Processing Cluster (GPC) included in the Parallel Processing Unit (PPU);

[0010] Figure 8 These are block diagrams of exemplary system-on-chip (SoC) according to various embodiments. Detailed Implementation

[0011] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of the various embodiments. However, it will be apparent to those skilled in the art that the inventive concept can be practiced without one or more of these specific details.

[0012] System Overview

[0013] Figure 1 A computing device 100 configured to implement one or more aspects of various embodiments is described. In one embodiment, the computing device 100 may be a desktop computer, laptop computer, smartphone, personal digital assistant (PDA), tablet computer, or any other type of computing device configured to receive input, process data, and selectively display images, and is adapted to practice one or more embodiments. The computing device 100 is configured to run a training engine 122 and an execution engine 124 residing in memory 116. It should be noted that the computing device described herein is illustrative, and any other technically feasible configuration falls within the scope of this disclosure.

[0014] In one embodiment, the computing device 100 includes, but is not limited to, an interconnect (bus) 112 connecting one or more processing units 102, an input / output (I / O) device interface 104 coupled to one or more input / output (I / O) devices 108, memory 116, storage 114, and a network interface 106. The one or more processing units 102 can be any suitable processor implemented as a central processing unit (CPU), graphics processing unit (GPU), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), artificial intelligence (AI) accelerator, any other type of processing unit, or a combination of different processing units (such as a CPU configured to operate in conjunction with a GPU). In general, the one or more processing units 102 can be any technically feasible hardware unit capable of processing data and / or executing software applications. Furthermore, in the context of this disclosure, the computing elements shown in the computing device 100 can correspond to a physical computing system (e.g., a system in a data center) or can be virtual computing instances performed in a computing cloud.

[0015] In one embodiment, I / O device 108 includes devices capable of providing input, such as a keyboard, mouse, touchscreen, etc., and devices capable of providing output, such as a display device. Furthermore, I / O device 108 may include devices capable of receiving input and providing output, such as a touchscreen, a Universal Serial Bus (USB) port, etc. I / O device 108 may be configured to receive various types of input from end users of computing device 100 (e.g., designers) and provide various types of output to end users of computing device 100, such as displayed digital images, digital video, or text. In some embodiments, one or more of I / O devices 108 are configured to couple computing device 100 to network 110.

[0016] In one embodiment, network 110 is any technically feasible type of communication network that allows data exchange between computing device 100 and external entities or devices, such as web servers or other networked computing devices. For example, network 110 may include a wide area network (WAN), a local area network (LAN), a wireless (WiFi) network, and / or the Internet.

[0017] In one embodiment, memory 114 includes non-volatile memory for applications and data, which may include fixed or removable disk drives, flash memory devices, and CD-ROMs, DVD-ROMs, Blu-ray, HD-DVDs, or other magnetic, optical, or solid-state storage devices. Training engine 122 and execution engine 124 may be stored in memory 114 and loaded into memory 116 during execution.

[0018] In one embodiment, memory 116 includes a random access memory (RAM) module, flash memory cells, or any other type of storage cell or a combination thereof. One or more processing units 102, I / O device interfaces 104, and network interfaces 106 are configured to read data from memory 116 and write data to memory 116. Memory 116 includes various software programs that can be executed by one or more processors 102 and application data associated with those software programs, including a training engine 122 and an execution engine 124.

[0019] In one embodiment, training engine 122 generates one or more machine learning models for performing mesh reconstruction using data-driven priors. Each machine learning model may learn priors related to vertices, edges, corners, faces, polygons, surfaces, shapes, and / or other properties of the two-dimensional (2D) and / or three-dimensional (3D) mesh. For example, each machine learning model may include a variational autoencoder (VAE) that learns to convert an input mesh into latent vectors and reconstruct the mesh from the latent vectors.

[0020] In one embodiment, execution engine 124 uses a mesh prior learned by a machine learning model to perform mesh reconstruction. Continuing the example above, execution engine 124 can input initial values ​​of the latent vectors into the VAE's decoder to produce an initial estimate of the mesh. Execution engine 124 can refine the mesh by selecting subsequent values ​​of the latent vectors based on geometric constraints associated with a set of image observations of the object. In various embodiments, subsequent values ​​of the latent vectors are selected to minimize the error between the mesh and the image observations, thus allowing the mesh to approximate the shape of the object. Training engine 122 and execution engine 124 are referenced below. Figure 2 Further detailed description.

[0021] Mesh Reconstruction Using Data-Driven Priors

[0022] Figure 2 According to various embodiments Figure 1 A more detailed description of the training engine 122 and execution engine 124 is provided below. In the illustrated embodiment, the training engine 122 creates a machine learning model that learns to reconstruct a set of training meshes 208 by learning and / or encoding priors associated with the training meshes 208. For example, the machine learning model can learn from the training meshes 208 several “basic types” related to vertices, edges, corners, faces, triangles, polygons, surfaces, shapes, and / or other properties of 2D and / or 3D meshes.

[0023] In one embodiment, execution engine 124 uses a machine learning model to perform inverse rendering of an object 260 captured in a set of images 258 into a corresponding mesh 216. For example, execution engine 124 can use a machine learning model to estimate the 3D mesh 216 of the object from multiple views and / or multiple lighting conditions of the 2D images 258 capturing the object 260. Thus, images 258 can represent ground truth image observations of the object 260.

[0024] The machine learning model created by training engine 122 can include any technically feasible form of machine learning model. For example, the machine learning model can include recurrent neural networks (RNNs), convolutional neural networks (CNNs), deep neural networks (DNNs), deep convolutional networks (DCNs), deep belief networks (DBNs), restricted Boltzmann machines (RBMs), long short-term memory (LSTM) units, gated recurrent units (GRUs), generative adversarial networks (GANs), self-organizing maps (SOMs), and / or other types of artificial neural networks or components of artificial neural networks. In another example, the machine learning model can include functionality to perform clustering, principal component analysis (PCA), latent semantic analysis (LSA), Word2vec, and / or other unsupervised learning techniques. In a third example, the machine learning model can include regression models, support vector machines, decision trees, random forests, gradient boosting trees, naive Bayes classifiers, Bayesian networks, hierarchical models, and / or ensemble models.

[0025] In some embodiments, the machine learning model created by the training engine 122 includes a VAE 200, which includes an encoder 202 and a decoder 206. For example, the training engine 122 may input a 2D and / or 3D training mesh 208 into the encoder 202, and the encoder 202 may "encode" the training mesh 208 into latent vectors 204 (e.g., scalars, geometric vectors, tensors, and / or other geometric objects) in a latent space of a lower dimension than the training mesh 208. The decoder 206 in the VAE 200 may "decode" the latent vectors 204 into a reconstructed mesh 210 of a higher dimension that substantially reproduces the training mesh 208.

[0026] In various embodiments, the training mesh 208 may include a set of 2D and / or 3D points, representations of edges between point pairs, and / or representations of triangles, polygons, and / or other shapes formed by points and edges. Edges and / or shapes in the training mesh 208 may be encoded according to the order in which the points are arranged. For example, a 2D mesh may be parameterized as an ordered set of points, where each adjacent pair of points in the ordered set is connected by an edge between the points. In another example, a 3D mesh may be parameterized using one or more triangular strips, triangular sectors, and / or other representations of triangles within the mesh.

[0027] In one or more embodiments, the training engine 122 configures the VAE 200 to encode the training mesh 208 in a manner unaffected by the parameterization of the training mesh 208 input to the encoder 202. In these embodiments, the training engine 122 inputs multiple ordered sets of points in each training mesh into the encoder 202 and normalizes the output of the encoder 202 from the multiple ordered sets of points. Therefore, the training engine 122 can decouple the parameterization of the training mesh 208 from the latent space of the VAE 200 to which the training mesh 208 is encoded.

[0028] In one embodiment, the ordered set of points input to encoder 202 may include all possible sorts of points in the training grid, a randomly selected subset of possible sorts of points in the training grid, and / or any other combination of sorts of points in the training grid. For example, training engine 122 can generate different ordered sets of points in a 2D grid of a polygon by selecting different vertices of the polygon and / or different points along the edges between the vertices as starting points in the ordered set. In another example, training engine 122 can generate two different sorts of points in a 2D grid from the same starting point by traversing the grid in clockwise and counterclockwise directions.

[0029] In one embodiment, training engine 122 trains different encoders not in VAE 200 to generate intermediate representations of ordered sets of points and / or other parameterizations of training mesh 208. In one embodiment, training engine 122 averages and / or otherwise aggregates the intermediate representations into standardized values ​​in the latent space, such that the standardized values ​​define the mean and standard deviation of the samples taken from the output of encoder 202.

[0030] In one embodiment, training engine 122 inputs the output sampled with normalized values ​​into decoder 206 to obtain a reconstructed mesh 210 as the output of decoder 206. Training engine 122 also calculates the minimum error between each training mesh and its corresponding reconstructed mesh, specifically the minimum error between the ordering of points in the training mesh input to encoder 202 and the ordering of points from the output of decoder 206. For example, training engine 122 may calculate the minimum error based on the intersection of points in the training mesh and points in the reconstructed mesh, the chamfer distance between the training mesh and the reconstructed mesh, the earthmover's distance, and / or other measures of distance, similarity, and / or difference between points, edges, shapes, and / or other attributes of the training mesh and the reconstructed mesh.

[0031] In one embodiment, training engine 122 uses the error between each training grid and the corresponding reconstructed grid to update the parameters of encoder 202 and decoder 206. For example, training engine 122 may backpropagate the error across layers and / or parameters of decoder 206 so that decoder 206 learns to decode the reconstructed grid from the sampled output of encoder 202. Training engine 122 may also backpropagate the error across layers and / or parameters of encoder 202 so that encoder 202 learns to produce consistent latent vector representations (e.g., mean and standard deviation vectors) for different orderings of points in the corresponding training grids.

[0032] In one embodiment, training engine 122 trains decoder 206 to generate reconstructed mesh 210 from training mesh 208 at different resolutions until the desired mesh resolution is reached. For example, decoder 206 may include multiple neural network layers, each producing a reconstructed mesh with more vertices than the previous layer (e.g., by adding vertices to the edges and / or centers of polygon faces in the mesh from the previous layer). In one embodiment, during training of decoder 206, training engine 122 may compare the coarse reconstructed mesh produced by the first layer of decoder 206 with the corresponding training mesh and update the parameters of the first layer to reduce the error between the vertices of the reconstructed mesh and the corresponding vertices of the training mesh. Training engine 122 may repeat this process for subsequent layers of decoder 206 that upsample the reconstructed mesh from the previous layer until the output layer of decoder 206 and / or the desired reconstructed mesh resolution is reached.

[0033] By performing coarse-to-fine auto-encoding on the training mesh 208, the training engine 122 can allow the coarse mesh from the first layer of the decoder 206 to be defined within a given reference coordinate frame, and subsequent layers of the decoder 206 to refine the coarse mesh further, defining it in relevant terms relative to the coarse mesh. Thus, the refinement, and in turn, of most of the reconstructed mesh can be independent of the object's reference frame or global pose. This allows the decoder 206 to more easily learn changes in the reference coordinate system between meshes, and / or potentially omit global pose estimation during the generation of the reconstructed mesh 210.

[0034] In the illustrated embodiment, execution engine 124 executes one or more parts of VAE 200 to generate a mesh 216 of object 260 from latent vector values ​​212 of vectors in the latent space of VAE 200 and an image 258 of object 260. For example, execution engine 124 may input latent vector values ​​212 into decoder 206 to generate mesh 216 as a decoded representation of vectors.

[0035] More specifically, in one embodiment, execution engine 124 uses fixed parameters to modify latent vector values ​​212 input to decoder 206 to generate mesh 216 so that mesh 216 reproduces the object 260 captured in a set of images 258. In one or more embodiments, images 258 are associated with a set of assumptions that allow mesh 216 to be reconstructed based on images 258. For example, images 258 may be associated with a known camera pose that generates multiple views of object 260, multiple lighting conditions that allow object 260 to be captured under different lighting conditions, and / or a Lambertian surface on object 260 with isotropic brightness (i.e., a surface with uniform brightness from any direction).

[0036] In one embodiment, execution engine 124 initially selects the latent vector value 212 based on one or more criteria. For example, execution engine 124 may randomly select the initial latent vector value 212 and / or set the initial latent vector value 212 as the default latent vector value associated with decoder 206. In another example, execution engine 124 may set the initial latent vector value 212 based on sparse features extracted from 2D image 258 (e.g., features extracted using scale-invariant feature transform (SIFT), corner detection, and / or other image feature detection techniques). Execution engine 124 may generate the initial latent vector value 212 using sparse features by estimating the 3D positions of the sparse features based on assumptions associated with image 258, input these positions into encoder 202, and obtain the initial latent vector value 212 based on the output of encoder 202.

[0037] In one embodiment, execution engine 124 uses assumptions associated with image 258 to determine geometric constraints 222 associated with object 260, which may include constraints related to the shape and / or geometry of object 260. Execution engine 124 also refines mesh 216 by updating latent vector values ​​212 based on mesh 216 and geometric constraints 222.

[0038] In one embodiment, execution engine 124 uses geometric constraints 222 to compute one or more errors associated with mesh 216 and updates latent vector values ​​212 based on gradients associated with the errors. For example, execution engine 124 can project an image backward from a first camera onto mesh 216 using known camera pose and / or lighting conditions associated with image 258, and project the distorted image back onto a second camera. Execution engine 124 can compute photometric and / or contour errors between the distorted image and corresponding image observations of object 260 from image 258 (e.g., images of object 260 taken from the second camera position and / or under the same lighting conditions used to generate the distorted image). Execution engine 124 can use gradients associated with the photometric and / or contour errors to modify latent vector values ​​212 so that mesh 216 better approximates the shape of object 260. The execution engine 124 can also repeat the process to iteratively update the latent vector values ​​212 to search the latent space associated with the VAE 200, and use the decoder 206 to generate a more accurate mesh 216 from the latent vector values ​​212 until the mesh 216 is a substantially accurate representation of the object 260.

[0039] In one embodiment, execution engine 124 generates a distorted image of grid 216 by projecting light rays from a first camera onto grid 216, selecting the nearest intersection point of the light rays on grid 216, and projecting another light ray from that intersection point onto a second camera, for comparison with a ground reality image 258 of object 260. In this process, the faces of grid 216 can be represented as binary indicator functions in the image, where a face is "visible" in the image at the location where a light ray from the corresponding camera intersects with that face, and is otherwise invisible. Because the indicator function is binary and includes abrupt changes from "visible" pixels of that face to "invisible" pixels at other locations, the image may be non-differentiable, which could interfere with the generation of gradients used to update the latent vector values ​​212.

[0040] In one or more embodiments, execution engine 124 generates an image of mesh 216 using a differentiable index function that smooths the transition between the visible and invisible portions of the faces in mesh 216. For example, a face in mesh 216 may have an index function that displays the face as "visible" regardless of where light from the camera intersects the face, while locations where light from the camera does not intersect the face are invisible, and the transition between visible and invisible is smooth but rapid near the boundary between the two.

[0041] When generating the grid 216 based on a single latent vector value 212 input into the decoder, the decoder 206 may need to learn the prior of the entire object 260. Therefore, the execution engine 124 is only able to generate the grid 216 as an accurate reconstruction of the object 260 if the object 260 belongs to the distribution of the training grid 208 used to generate the VAE 200.

[0042] In one or more embodiments, execution engine 124 reduces the number and / or complexity of mesh priors learned by decoder 206 by dividing mesh 216 into multiple smaller meshlets 218 representing different parts of mesh 218. For example, execution engine 124 may generate one or more polygons, surfaces, and / or shapes in mesh 216 for each meshlet. Each meshlet may be defined using a more fundamental prior than mesh 216, which allows the prior to be generalized to a wider range of objects and / or shapes. Each meshlet may also contain different parts of meshlet 218, and meshlets 218 may be combined to generate meshlet 218.

[0043] Similar to generating a mesh 216 from a single latent vector value 212 input to decoder 206, in one embodiment, execution engine 124 inputs multiple latent vector values ​​214 to decoder 206 to generate multiple corresponding small meshes 218 as the output of decoder 206. Execution engine 124 can also similarly update the latent vector values ​​214 to enforce a subset of the geometric constraints 222 associated with individual small meshes 218 and / or reduce the error between the small meshes 218 and corresponding portions of object 260. For example, execution engine 124 can map a small mesh to a portion of object 260 and use an image 258 containing that portion to determine the geometric constraints 222 associated with that portion. Execution engine 124 can use the geometric constraints 222 to generate an error between the small meshes and image 258 and perform gradient descent on the latent vector values ​​used to generate the meshes based on this error.

[0044] In the illustrated embodiment, the small mesh 218 is also associated with a pose 220 that projects the small mesh 218 back to the mesh 216. For example, the shape and / or geometry of each small mesh can be generated from a point in the latent space of the VAE 200. Each small mesh may also be associated with one or more custom poses 220 (e.g., one or more rotations, translations, and / or other transformations) that align the global pose of the image 258 and / or the mesh 216 with the standard pose of the small mesh learned by the decoder 206.

[0045] In one embodiment, execution engine 124 uses gradients and / or errors associated with geometric constraints 222 to update the latent vector values ​​214 and pose 220 of a small mesh 218, which translates the small mesh 218 into its corresponding position and / or orientation within mesh 216. For example, object 260 may include a rectangular surface with four right angles. During the reconstruction of mesh 216 of object 260, execution engine 124 can use geometric constraints 222 to identify values ​​in the latent space associated with VAE 200 (which generates small meshes with right angles) and learn four different poses 220 that map the small meshes to the four corners of the rectangular surface.

[0046] In one embodiment, execution engine 124 iteratively increases the resolution of the small mesh 218 to satisfy geometric constraints 222. For example, execution engine 124 may use a first layer of decoder 206 to generate an initial coarse mesh 216 that matches a low-resolution image 258 of object 260. Execution engine 124 may use subsequent layers of decoder 206 to upsample the coarse mesh 216 into increasingly smaller and more numerous small meshes 218. At each layer of decoder 206, execution engine 124 may perform gradient descent relative to the coarse mesh pair for generating the latent vector values ​​214 and pose 220 of the small mesh 218 until the desired mesh 216 and / or small mesh resolution is reached.

[0047] In one embodiment, execution engine 124 reconstructs mesh 216 from small mesh 218 and corresponding pose 220, and stores the reconstructed mesh 216 in association with object 260. For example, execution engine 124 may apply pose 220 to the corresponding small mesh 218 to map points in the standard pose of small mesh 218 to their positions in mesh 216. Execution engine 124 may then render the completed mesh 216 and / or store the points in mesh 216 under the identifier of object 260, along with the image 258 of object 260, and / or with other metadata of object 260.

[0048] Figure 3 This is a flowchart of method steps for performing mesh reconstruction using data-driven priors, according to various embodiments. Although combined... Figure 1 and Figure 2 The system describes the method steps, but those skilled in the art will understand that any system configured to perform the method steps in any order will fall within the scope of this disclosure.

[0049] As shown in the figure, training engine 122 generates a 302 machine learning model as a VAE, which learns and reconstructs a set of training grids input into the VAE, as described below. Figure 4Further detailed description. In various embodiments, the training mesh may include 2D and / or 3D meshes and / or small meshes.

[0050] Next, execution engine 124 executes a 304 machine learning model to generate a grid of objects from the first values ​​of vectors in the latent space. For example, execution engine 124 can choose the first value of the vector as a random value, a default value, a value based on a set of image observations of the object, and / or a value based on sparse features extracted from the image observations. Execution engine 124 can then input the first value into a decoder in the VAE and obtain the grid as the decoder output. The similarity between the grid and the object may be influenced by the criteria used to select the first value of the vector (e.g., the first value of the vector reflecting sparse features in the image observations is likely to produce a grid more similar to the object than a randomly selected first value of the vector).

[0051] Execution engine 124 refines the mesh of object 306 by selecting a second value of a vector in the latent space based on one or more geometric constraints associated with image observations of the object. For example, execution engine 124 can use known camera pose and / or lighting conditions under which image observations are made to project and / or post-project a distorted image of the mesh. Execution engine 124 can also compute the error between the distorted image of the mesh and the corresponding image observation (e.g., an image observation made under the same camera pose and / or lighting conditions). Execution engine 124 can then update the value of the vector using the gradient associated with this error and the constant parameters of the machine learning model.

[0052] In another example, execution engine 124 can project and / or divide the mesh into a set of smaller meshes representing different parts of the mesh. For each smaller mesh, execution engine 124 can choose values ​​of vectors in the latent space to learn a priori information about a portion of the mesh represented by the smaller mesh. Execution engine 124 can also learn custom poses for the smaller meshes that map their positions and / or orientations within the mesh and / or align the global pose of the image observations with the standard poses of the smaller meshes learned by the machine learning model. Execution engine 124 can then reconstruct the mesh from the set of smaller meshes by mapping points in the smaller meshes to the mesh's reference frame using the custom poses of the smaller meshes.

[0053] Finally, execution engine 124 stores the 308 grid in association with the object. For example, execution engine 124 may store points in grid 216 under the object's identifier, store them along with the object's image, and / or store them along with other metadata of the object.

[0054] Figure 4 This is a flowchart of method steps for training a machine learning model to learn a data-driven grid prior, according to various embodiments. Although combined... Figure 1 and Figure 2 The system describes the method steps, but those skilled in the art will understand that any system configured to perform the method steps in any order falls within the scope of this disclosure.

[0055] As shown in the figure, the training engine 122 applies the encoder 402 to multiple orderings of points in the training grid to generate intermediate representations of multiple orderings of points. For example, the training engine 122 can choose a random subset of all possible point orderings in the training grid as input to the encoder. The encoder can then "encode" each point ordering into a vector in the latent space.

[0056] Next, the training engine 122 averages the intermediate representations into standardized values ​​of vectors in the latent space. These standardized values ​​can define the mean vector and standard deviation vector output by the encoder in the VAE.

[0057] Then, training engine 122 trains the decoder in the 406VAE to reconstruct the training mesh from the normalized values. Training engine 122 may also optionally train the encoder in the 408VAE to generate normalized values ​​from each sort of points in the training mesh. For example, training engine 122 may calculate the error between the reconstructed mesh and the training mesh as the minimum error between the sorting of points in the training mesh input to the encoder and the sorting of points output from the decoder. This error may be based on the intersection of points between the training mesh and the reconstructed mesh, a distance metric between the training mesh and the reconstructed mesh, and / or other metrics of similarity or dissimilarity between the training mesh and the reconstructed mesh. Training engine 122 may backpropagate this error across layers and / or parameters of the decoder so that the decoder learns to reconstruct the training mesh based on the normalized values. Training engine 122 may also backpropagate this error between layers and / or parameters of the encoder so that the encoder learns to output normalized values ​​for different sortings of points in the corresponding training mesh.

[0058] Example hardware architecture

[0059] Figure 5 This is a block diagram of a computer system 500 configured to implement one or more aspects of various embodiments. In some embodiments, the computer system 500 is a server operating in a data center or cloud computing environment, providing scalable computing resources as a service over a network. In some embodiments, the computer system 500 implements... Figure 1 The functions of the computing device 100 shown.

[0060] In various embodiments, the computer system 500 includes, but is not limited to, a central processing unit (CPU) 502 and a system memory 504 coupled to a parallel processing subsystem 512 via a memory bridge 505 and a communication path 513. The memory bridge 505 is further coupled to an I / O (input / output) bridge 507 via a communication path 506, and the I / O bridge 507 is coupled to a switch 516.

[0061] In one embodiment, I / O bridge 507 is configured to receive user input from an optional input device 508 (such as a keyboard or mouse) and forward the input to CPU 502 for processing via communication path 506 and memory bridge 505. In some embodiments, computer system 500 may be a server in a cloud computing environment. In such embodiments, computer system 500 may not have input device 508. Instead, computer system 500 may receive equivalent input by receiving commands in the form of messages sent over the network and received via network adapter 518. In one embodiment, switch 516 is configured to provide connectivity between I / O bridge 507 and other components of computer system 500, such as network adapter 518 and various plug-in cards 520 and 521.

[0062] In one embodiment, I / O bridge 507 is coupled to system disk 514, which may be configured to store content, applications, and data for use by CPU 502 and parallel processing subsystem 512. In one embodiment, system disk 514 provides non-volatile storage for applications and data and may include fixed or removable hard disk drives, flash memory devices, and CD-ROMs (compressed disc read-only memory), DVD-ROMs (Digital Universal Disc-ROMs), Blu-ray discs, HD-DVDs (High Definition DVDs), or other magnetic, optical, or solid-state storage devices. In various embodiments, other components such as universal serial bus or other port connections, compressed disc drives, digital universal disc drives, film recording devices, etc., may also be connected to I / O bridge 507.

[0063] In various embodiments, memory bridge 505 may be a northbridge chip, and I / O bridge 507 may be a southbridge chip. Furthermore, communication paths 506 and 513, as well as other communication paths within the computer system 500, may be implemented using any technically suitable protocol, including but not limited to AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.

[0064] In some embodiments, the parallel processing subsystem 512 includes a graphics subsystem that transmits pixels to an optional display device 510, which may be any conventional cathode ray tube, liquid crystal display, light-emitting diode display, or similar device. In this embodiment, the parallel processing subsystem 512 includes circuitry optimized for graphics and video processing, including, for example, video output circuitry. (See below for further details.) Figure 6 and Figure 7 In more detail, such circuitry may be contained across one or more parallel processing units (PPUs, also known as parallel processors) included in the parallel processing subsystem 512.

[0065] In other embodiments, the parallel processing subsystem 512 includes circuitry optimized for general and / or computational processing. Similarly, such circuitry may be included across one or more PPUs included in the parallel processing subsystem 512, which are configured to perform such general and / or computational operations. In other embodiments, one or more PPUs included in the parallel processing subsystem 512 may be configured to perform graphics processing, general processing, and computational processing operations. System memory 504 includes at least one device driver configured to manage the processing operations of one or more PPUs in the parallel processing subsystem 512.

[0066] In various embodiments, the parallel processing subsystem 512 can be coupled with... Figure 5 It can be integrated with one or more other components to form a single system. For example, the parallel processing subsystem 512 can be integrated with the CPU 502 and other interconnect circuitry on a single chip to form a system-on-a-chip (SoC).

[0067] In one embodiment, CPU 502 is the main processor of computer system 500, controlling and coordinating the operation of other system components. In one embodiment, CPU 502 issues commands to control the operation of PPUs. In some embodiments, communication path 513 is a PCI Express link, in which dedicated channels, as known in the art, are allocated to each PPU. Other communication paths may also be used. The PPU advantageously implements a highly parallel processing architecture. The PPU can have any number of local parallel processing memories (PP memories).

[0068] It should be understood that the systems illustrated herein are illustrative, and variations and modifications are possible. The connection topology (including the number and arrangement of bridges, the number of CPUs 502, and the number of parallel processing subsystems 512) can be modified as needed. For example, in some embodiments, system memory 504 may be directly connected to CPU 502 instead of via memory bridge 505, with other devices communicating with system memory 504 via memory bridge 505 and CPU 502. In other embodiments, parallel processing subsystems 512 may be connected to I / O bridge 507 or directly to CPU 502 instead of via memory bridge 505. In other embodiments, I / O bridge 507 and memory bridge 505 may be integrated into a single chip rather than existing as one or more discrete devices. Finally, in some embodiments, Figure 5 One or more of the components shown may be absent. For example, switch 516, network adapter 518, and plug-in cards 520 and 521 can be removed and directly connected to I / O bridge 507.

[0069] Figure 6 Based on various embodiments, Figure 5 A block diagram of the parallel processing unit (PPU) 602 included in the parallel processing subsystem 512. As described above, although... Figure 6 A PPU 602 is described, but the parallel processing subsystem 512 may include any number of PPUs 602. As shown, the PPU 602 is coupled to a local parallel processing (PP) memory 604. The PPU 602 and PP memory 604 may be implemented using one or more integrated circuit devices, such as programmable processors, application-specific integrated circuits (ASICs), or memory devices, or in any other technically feasible manner.

[0070] In some embodiments, PPU 602 includes a graphics processing unit (GPU) configured to implement a graphics rendering pipeline to perform various operations related to generating pixel data based on graphics data provided by CPU 502 and / or system memory 504. When processing graphics data, PPU 604 can be used as graphics memory, storing one or more regular frame buffers, and possibly one or more other rendering targets if needed. Furthermore, PPU 604 can be used to store and update pixel data and transmit the final pixel data or display frame to an optional display device 510 for display. In some embodiments, PPU 602 can also be configured for general processing and computational operations. In some embodiments, computer system 500 may be a server in a cloud computing environment. In these embodiments, computer system 500 may not have a display device 510. Instead, computer system 500 can generate equivalent output information by sending commands in the form of messages over a network via network adapter 518.

[0071] In some embodiments, CPU 502 is the main processor of computer system 500, controlling and coordinating the operation of other system components. In one embodiment, CPU 502 issues commands to control the operation of PPU 602. In some embodiments, CPU 502 writes the command stream of PPU 602 into a data structure located in system memory 504, PP memory 604, or another storage location accessible to both CPU 502 and PPU 602. Figure 5 or Figure 6 (Not explicitly shown in the document). A pointer to a data structure is written to a command queue (also referred to herein as a push buffer) to initiate processing of the command stream in the data structure. In one embodiment, PPU 602 reads the command stream from the command queue and then executes the commands asynchronously relative to the operation of CPU 502. In embodiments that generate multiple push buffers, the scheduling of different push buffers can be controlled by the application specifying an execution prior for each push buffer via a device driver.

[0072] In one embodiment, PPU 602 includes an I / O (input / output) unit 605 that communicates with the rest of the computer system 500 via communication path 513 and memory bridge 505. In one embodiment, I / O unit 605 generates data packets (or other signals) for transmission on communication path 513 and also receives all incoming data packets (or other signals) from communication path 513, directing the incoming data packets to the appropriate components of PPU 602. For example, commands related to processing tasks may be directed to host interface 606, while commands related to memory operations (e.g., reading from or writing to PP memory 604) may be directed to crossbar switch unit 610. In one embodiment, host interface 606 reads each command queue and sends the command stream stored in the command queue to front end 612.

[0073] As described above Figure 5 The connection between PPU 602 and the rest of the computer system 500 may be different. In some embodiments, the parallel processing subsystem 512 (which includes at least one PPU 602) is implemented as an insert card that can be inserted into an expansion slot of the computer system 500. In other embodiments, PPU 602 may be integrated on a single chip with a bus bridge, such as memory bridge 505 or I / O bridge 507. Similarly, in other embodiments, some or all of the components of PPU 602 may be included together with CPU 502 in a single integrated circuit or system-on-a-chip (SoC).

[0074] In one embodiment, front-end 612 sends processing tasks received from host interface 606 to a work allocation unit (not shown) within task / work unit 607. In one embodiment, the work allocation unit receives pointers to processing tasks, which are encoded as Task Metadata (TMDs) and stored in memory. Pointers to TMDs are included in a command stream, which is stored as a command queue and received by front-end unit 612 from host interface 606. A processing task that can be encoded as a TMD includes an index associated with the data to be processed, as well as state parameters and commands defining how the data should be processed. For example, state parameters and commands may define a procedure to be executed on the data. Also, for example, a TMD may specify the number and configuration of a set of Cooperative Thread Arrays (CTAs). Typically, each TMD corresponds to one task. Task / work unit 607 receives tasks from front-end 612 and ensures that GPC 608 is configured to an active state before the processing task specified by each TMD is initiated. Priors may also be specified for each TMD used to schedule the execution of processing tasks. Processing tasks may also be received from processing cluster array 630. Optionally, the TMD may include parameters that control whether the TMD is added to the head or tail of the list of processing tasks (or to a list of pointers to processing tasks), thus providing another layer of control over execution priors.

[0075] In one embodiment, the PPU 602 implements a highly parallel processing architecture based on a processing cluster array 630, which includes a set of C general-purpose processing clusters (GPCs) 608, where C ≥ 1. Each GPC 608 is capable of executing a large number (e.g., hundreds or thousands) of threads simultaneously, where each thread is an instance of a program. In various applications, different GPCs 608 can be allocated to handle different types of programs or perform different types of computations. The allocation of GPCs 608 can vary depending on the workload generated by each type of program or computation.

[0076] In one embodiment, the memory interface 614 includes a set of D partition units 615, where D ≥ 1. Each partition unit 615 is coupled to one or more dynamic random access memories (DRAMs) 620 residing in the PP memory 604. In some embodiments, the number of partition units 615 is equal to the number of DRAMs 620, and each partition unit 615 is coupled to a different DRAM 620. In other embodiments, the number of partition units 615 may differ from the number of DRAMs 620. Those skilled in the art will understand that the DRAMs 620 can be replaced with any other technically suitable storage device. In operation, various rendering targets (such as texture maps and framebuffers) can be stored on the DRAMs 620, allowing the partition units 615 to write portions of each rendering target in parallel, thereby efficiently utilizing the available bandwidth of the PP memory 604.

[0077] In one embodiment, a given GPC 608 can process data to be written to any DRAM 620 within the PP memory 604. In one embodiment, a crossbar switch unit 610 is configured to route the output of each GPC 608 to the input of any partition unit 615 or any other GPC 608 for further processing. GPCs 608 communicate with the memory interface 614 via the crossbar switch unit 610 to read from or write to the respective DRAMs 620. In some embodiments, the crossbar switch unit 610 is connected to I / O units 605 and also to the PP memory 604 via the memory interface 614, thereby enabling processing cores in different GPCs 608 to communicate with system memory 504 or other memory not native to the PPU 602. Figure 6 In some embodiments, the crossbar switch unit 610 is directly connected to the I / O unit 605. In various embodiments, the crossbar switch unit 610 may use virtual channels to separate traffic flows between the GPC 608 and the partition unit 615.

[0078] In one embodiment, the GPC 608 can be programmed to perform processing tasks associated with various applications, including but not limited to linear and nonlinear data transformations, filtering of video and / or audio data, modeling operations (e.g., applying physical laws to determine the position, velocity, and other properties of an object), image rendering operations (e.g., tessellation shaders, vertex shaders, geometry shaders, and / or pixel / fragment shader programs), general computational operations, etc. In operation, the PPU 602 is configured to transfer data from system memory 504 and / or PP memory 604 to one or more on-chip memory units, process the data, and write the resulting data back to system memory 504 and / or PP memory 604. The resulting data can then be accessed by other system components (including CPU 502, another PPU 602 in parallel processing subsystem 512, or another parallel processing subsystem 512 in computer system 500).

[0079] In one embodiment, the parallel processing subsystem 512 may include any number of PPUs 602. For example, multiple PPUs 602 may be provided on a single insert card, or multiple insert cards may be connected to the communication path 513, or one or more PPUs 602 may be integrated into a bridge chip. The PPUs 602 in a multi-PPU system may be the same or different from each other. For example, different PPUs 602 may have different numbers of processing cores and / or different numbers of PP memories 604. In implementations with multiple PPUs 602, these PPUs can operate in parallel to process data with a higher throughput than possible with a single PPU 602. Systems containing one or more PPUs 602 can be implemented with various configurations and form factors, including but not limited to desktops, laptops, handheld personal computers or other handheld devices, servers, workstations, game consoles, embedded systems, etc.

[0080] Figure 7 Based on various embodiments, Figure 6 A block diagram of the General Processing Cluster (GPC) included in the Parallel Processing Unit (PPU) 602 is shown. As shown, the GPC 608 includes, but is not limited to, a pipeline manager 705, one or more texture units 715, a pre-raster operation unit 725, a job allocation crossbar switch 730, and an L1.5 cache 735.

[0081] In one embodiment, the GPC 608 can be configured to execute a large number of threads in parallel to perform graphics processing, general processing, and / or computational operations. As used herein, a “thread” refers to an instance of a specific program executed on a specific input dataset. In some embodiments, Single Instruction, Multiple Data (SIMD) instruction issuing techniques are used to support the parallel execution of a large number of threads without providing multiple independent instruction units. In other embodiments, Single Instruction, Multiple Threading (SIMT) techniques are used to support the parallel execution of a large number of typically synchronous threads using a common instruction unit configured to issue instructions to a set of processing engines in the GPC 608. Unlike SIMD execution mechanisms, where all processing engines typically execute the same instructions, SIMT execution allows different threads to more easily follow different execution paths with a given program. Those skilled in the art will understand that SIMD processing mechanisms represent a subset of the functionality of SIMT processing mechanisms.

[0082] In one embodiment, the operation of GPC 608 is controlled via pipeline manager 705, which assigns processing tasks received from work assignment units (not shown) in task / work units 607 to one or more streaming multiprocessors (SMs) 710. Pipeline manager 705 can also be configured to control work assignment crossbar switch 730 by specifying the destination of the processed data output from SM 710.

[0083] In various embodiments, the GPC 608 includes a set of M SM 710s, where M ≥ 1. Furthermore, each SM 710 includes a set of functional execution units (not shown), such as execution units and load-store units. Processing operations specific to any functional execution unit can be pipelined, allowing new instructions to be issued for execution before previous instructions have completed. Any combination of functional execution units in a given SM 710 can be provided. In various embodiments, the functional execution units can be configured to support a variety of different operations, including integer and floating-point arithmetic (e.g., addition and multiplication), comparison operations, Boolean operations (AND, OR, XOR), bit shifting, and computation of various algebraic functions (e.g., plane interpolation and trigonometric functions, exponential and logarithmic functions, etc.). Advantageously, the same functional execution units can be configured to perform different operations.

[0084] In various embodiments, each SM 710 includes multiple processing cores. In one embodiment, the SM 710 includes a large number (e.g., 128) of different processing cores. Each core may include a fully piped, single-precision, double-precision, and / or mixed-precision processing unit including floating-point arithmetic logic units and integer arithmetic logic units. In one embodiment, the floating-point arithmetic logic units implement the IEEE 754-2008 standard for floating-point arithmetic. In one embodiment, the core includes 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.

[0085] In one embodiment, a tensor core is configured to perform matrix operations, and in another embodiment, one or more tensor cores are contained within a core. Specifically, the tensor core is configured to perform deep learning matrix arithmetic, such as convolution operations used for neural network training and inference. In one embodiment, each tensor core performs operations on a 4×4 matrix and performs matrix multiplication and accumulation operations D = A×B + C, where A, B, C, and D are 4×4 matrices.

[0086] In one embodiment, matrix multiplication inputs A and B are 16-bit floating-point matrices, while accumulation matrices C and D can be either 16-bit or 32-bit floating-point matrices. The Tensor Core performs operations on the 16-bit floating-point input data using 32-bit floating-point accumulation. The 16-bit floating-point multiplication requires 64 operations to obtain a full-precision product, which is then added to other intermediate products using 32-bit floating-point addition to obtain a 4×4×4 matrix multiplication. In practice, the Tensor Core is used to perform operations on larger two-dimensional or higher-dimensional matrices constructed from these smaller elements. APIs (such as the CUDA 9 C++ API) expose specialized matrix payloads, matrix multiplication and accumulation, and matrix storage operations to efficiently utilize the Tensor Core from CUDA-C++ programs. At the CUDA level, the thread bundle-level interface assumes a 16×16 matrix that spans all 32 threads of the thread bundle.

[0087] Neural networks heavily rely on matrix mathematical operations, and complex multi-layered networks require significant floating-point performance and bandwidth to improve efficiency and speed. Employing thousands of processing cores optimized for matrix mathematical operations and delivering performance ranging from tens to hundreds of TFLOPS in various embodiments, the SM 710 provides a computing platform capable of delivering the performance required for deep neural network-based artificial intelligence and machine learning applications.

[0088] In various embodiments, each SM 710 may also include multiple Special Function Units (SFUs) that perform special functions (e.g., attribute evaluation, inverse square root, etc.). In one embodiment, an SFU may include a tree traversal unit configured to traverse a hierarchical tree data structure. In one embodiment, an SFU may include a texture unit configured to perform texture map filtering operations. The texture unit is configured to load a texture map (e.g., a two-dimensional texture pixel array) from memory and sample the texture map to produce sampled texture values ​​for use in a shading procedure executed by the SM. In various embodiments, each SM 710 also includes multiple Load / Store Units (LSUs) that implement load and store operations between shared memory / L1 cache and a register file within the SM 710.

[0089] In one embodiment, each SM 710 is configured to process one or more thread groups. As used herein, a "thread group" or "thread warp" refers to a group of threads that execute the same program simultaneously on different input data, wherein one thread in the group is assigned to a different execution unit in the SM 710. A thread group may include fewer threads than the number of execution units in the SM 710, in which case some executions may be idle during a cycle while the thread group is being processed. A thread group may also include more threads than the number of execution units in the SM 710, in which case processing may occur in consecutive clock cycles. Since each SM 710 can support up to G thread groups simultaneously, up to G*M thread groups can be executed in the GPC 608 at any given time.

[0090] Furthermore, in one embodiment, multiple related thread groups may be active simultaneously (at different execution stages) within the SM 710. This collection of thread groups is referred to herein as a “cooperative thread array” (“CTA”) or “thread array”. The size of a particular CTA is equal to m*k, where k is the number of threads executing concurrently in the thread group, which is typically an integer multiple of the number of execution units in the SM 710, and m is the number of concurrently active thread groups in the SM 710. In some embodiments, a single SM 710 may support multiple CTAs simultaneously, where the granularity of these CTAs is the granularity of work assigned to the SM 710.

[0091] In one embodiment, each SM 710 includes a Level 1 (L1) cache, or uses space in a corresponding L1 cache outside the SM 710 to support load and store operations performed by the execution unit. Each SM 710 can also access a Level 2 (L2) cache (not shown) shared among all GPCs 608 in the PPU 602. The L2 cache can be used to transfer data between threads. Finally, the SM 710 can also access off-chip “global” memory, which may include PP memory 604 and / or system memory 504. It should be understood that any memory outside the PPU 602 can be used as global memory. Furthermore, as Figure 7 As shown, a Level 1.5 (L1.5) cache 735 may be included in the GPC 608 and configured to receive and store data requested from memory by the SM 710 via the memory interface 614. Such data may include, but is not limited to, instructions, uniform data, and constant data. In embodiments where the GPC 608 has multiple SMs 710, the SMs 710 may advantageously share common instructions and data cached in the L1.5 cache 735.

[0092] In one embodiment, each GPC 608 may have an associated memory management unit (MMU) 720, which is configured to map virtual addresses to physical addresses. In various embodiments, the MMU 720 may reside within the GPC 608 or the memory interface 614. The MMU 720 includes a set of page table entries (PTEs) for mapping virtual addresses to physical addresses of tiles or memory pages, and optional cache line indexes. The MMU 720 may include an address translation lookaside buffer (TLB) or a cache that may reside within the SM 710, one or more L1 caches, or the GPC 608.

[0093] In one embodiment, in a graphics and computing application, the GPC 608 can be configured to couple each SM 710 to a texture unit 715 to perform texture mapping operations, such as determining texture sampling locations, reading texture data, and filtering texture data.

[0094] In one embodiment, each SM 710 sends a processed task to a work assignment crossbar switch 730 to provide the processed task to another GPC 608 for further processing, or stores the processed task in an L2 cache (not shown), parallel processing memory 604, or system memory 504 via the crossbar switch unit 610. Furthermore, a pre-raster operation (preROP) unit 725 is configured to receive data from SM 710, direct the data to one or more raster operation (ROP) units within partition unit 615, perform color blending optimization, organize pixel color data, and perform address translation.

[0095] It should be understood that the architecture described herein is illustrative and subject to change and modification. Furthermore, the GPC 608 can contain any number of processing units, such as SM 710, texture units 715, or preROP units 725. Additionally, when combined with… Figure 6 The PPU 602 may include any number of GPCs 608, which are configured to be functionally similar to each other so that their execution behavior is independent of which GPCs 608 receive a specific processing task. Furthermore, each GPC 608 operates independently of the other GPCs 608 in the PPU 602 to execute tasks for one or more applications.

[0096] Figure 8This is a block diagram of an exemplary system-on-chip (SoC) integrated circuit 800 according to various embodiments. The SoC integrated circuit 800 includes one or more application processors 802 (e.g., CPUs), one or more graphics processors 804 (e.g., GPUs), one or more image processors 806, and / or one or more video processors 808. The SoC integrated circuit 800 also includes peripheral or bus components, such as those implementing Universal Serial Bus (USB), Universal Asynchronous Receiver / Transmitter (UART), Serial Peripheral Interface (SPI), Secure Digital Input / Output (SDIO), and Inter-IC Voice (IoV). 2 S) and / or inter-integrated circuits (I 2 C) Serial interface controller 814. The SoC integrated circuit 800 also includes a display device 818 coupled to a display interface 820 (e.g., High Definition Multimedia Interface (HDMI) and / or Mobile Industrial Processor Interface (MIPI)). The SoC integrated circuit 800 also includes a flash memory subsystem 824 providing storage on the integrated circuit, and a memory controller 822 providing a memory interface for accessing the memory device.

[0097] In one or more embodiments, the SoC integrated circuit 800 is implemented using one or more types of integrated circuit components. For example, the SoC integrated circuit 800 may include one or more processor cores for application processor 802 and / or graphics processor 804. Additional functionality related to serial interface controller 814, display device 818, display interface 820, image processor 806, video processor 808, AI acceleration, machine vision, and / or other dedicated tasks may be provided by application-specific integrated circuits (ASICs), application-specific standard components (ASSPs), field-programmable gate arrays (FPGAs), and / or other types of custom components.

[0098] In summary, the disclosed technique uses a VAE to learn priors associated with the mesh and / or mesh portions. The decoder in the VAE generates an additional mesh of the object by searching for latent vector values ​​in the VAE's latent space, thereby reducing errors between the mesh and image observations of the object. Each mesh can also be subdivided into smaller meshes, and the shape and pose of the smaller meshes can be individually refined to meet geometric constraints associated with the image observations before the smaller meshes are combined back into the mesh.

[0099] One technical advantage of the disclosed technique is that the mesh prior learned by the VAE is used to perform inverse rendering of the object by enforcing geometric constraints associated with the image observations of the object. Another technical advantage of the disclosed technique includes reduced complexity in estimating the global pose of the object by performing coarse-to-fine mesh rendering. A third technical advantage of the disclosed technique is improved generalizability and reduced complexity of the reconstructed mesh by dividing the mesh into smaller meshes. Therefore, the disclosed technique provides technical improvements to machine learning models, computer systems, applications, and / or techniques for performing mesh reconstruction.

[0100] 1. In some embodiments, the processor includes logic for predicting one or more three-dimensional (3D) mesh representations based on a plurality of digital images, wherein the one or more 3D mesh representations are refined by minimizing the differences between the one or more 3D mesh representations and the plurality of digital images.

[0101] 2. The processor as described in Clause 1, wherein predicting the one or more 3D mesh representations based on the plurality of digital images comprises: executing a machine learning model to generate a mesh of an object from a first value in a latent space; and refining the mesh of the object by selecting a second value in the latent space based on one or more geometric constraints associated with the plurality of digital images of the object.

[0102] 3. The processor as described in Clauses 1-2, wherein refining the mesh of the object comprises: selecting a first value; calculating the error between the mesh of the object and the plurality of digital images; and performing gradient descent on the first value using parameters of the machine learning model to reduce the error.

[0103] 4. The processor as described in Clauses 1-3, wherein selecting the first value comprises at least one of: randomizing the first value, selecting the first value based on a set of image observations, and initializing the first value based on sparse features extracted from the set of image observations.

[0104] 5. The processor as described in Items 1-4, wherein refining the mesh of the object comprises: dividing the mesh into a set of smaller meshes; for each of the smaller meshes in the set of smaller meshes, selecting the second value in the latent space to learn a priori of a portion of the mesh represented by the smaller meshes; and reconstructing the mesh from the set of smaller meshes.

[0105] 6. The processor as described in Items 1-5, wherein refining the mesh of the object further comprises: for each of the set of small meshes, learning a custom pose of the small mesh, the custom pose aligning the global pose of the image with the canonical pose of the small mesh.

[0106] 7. The processor as described in Clauses 1-6, wherein refining the mesh of the object further comprises: iteratively increasing the resolution of the set of smaller meshes to satisfy one or more geometric constraints.

[0107] 8. The processor as described in Items 1-7, wherein the logic further generates the machine learning model as a variational autoencoder to reconstruct a set of training grids input to the variational autoencoder.

[0108] 9. The processor as described in Items 1-8, wherein generating the machine learning model comprises: for each of the set of training grids, sorting and aggregating a plurality of points in the training grids into normalized values ​​in the latent space; and training a decoder in the variational autoencoder to reconstruct the training grids from the normalized values.

[0109] 10. The processor as described in Clauses 1-9, wherein aggregating the plurality of point sorts in the training grid into the normalized value comprises: applying an encoder to the plurality of point sorts to generate an intermediate representation of the plurality of point sorts; and averaging the intermediate representations into the normalized value.

[0110] 11. In some embodiments, one or more three-dimensional (3D) mesh representations are predicted based on a plurality of digital images, wherein the one or more 3D mesh representations are refined by minimizing at least one difference between the one or more 3D mesh representations and the plurality of digital images.

[0111] 12. The method of claim 11, wherein predicting the one or more 3D mesh representations based on the plurality of digital images comprises: executing a machine learning model to generate a mesh of an object from a first value in a latent space; and refining the mesh of the object by selecting a second value in the latent space based on one or more geometric constraints associated with the plurality of digital images of the object.

[0112] 13. The method of terms 11-12, further comprising generating the machine learning model as a variational autoencoder to reconstruct a set of training grids input to the variational autoencoder.

[0113] 14. The method of terms 11-13, wherein generating the machine learning model comprises: for each of the set of training grids, sorting and aggregating a plurality of points in the training grids into normalized values ​​in the latent space; and training a decoder in the variational autoencoder to reconstruct the training grids from the normalized values.

[0114] 15. The method of terms 11-14, wherein aggregating the plurality of points sorted in the training grid into the normalized value comprises: applying an encoder to the plurality of points sorted to generate an intermediate representation of the plurality of points sorted; and averaging the intermediate representation into the normalized value.

[0115] 16. In some embodiments, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to at least: execute a machine learning model to generate a mesh of an object from a first value in a latent space; refine the mesh of the object by selecting a second value in the latent space based on one or more geometric constraints associated with a set of image observations of the object; and store the mesh in association with the object.

[0116] 17. The non-transitory computer-readable medium as described in Clause 16, wherein refining the grid of the object comprises: selecting the first value; and performing gradient descent on the first value using parameters of the machine learning model to reduce the error between the grid and the set of image observations.

[0117] 18. The non-transitory computer-readable medium as described in Clauses 16-17, wherein the error includes at least one of photometric error and profile error.

[0118] 19. The non-transitory computer-readable medium as described in Clauses 16-18, wherein refining the mesh of the object comprises: dividing the mesh into a set of smaller meshes; for each of the smaller meshes in the set, selecting a second value in the latent space to learn a priori of a portion of the mesh represented by the smaller meshes; and reconstructing the mesh from the set of smaller meshes.

[0119] 20. The non-transitory computer-readable medium as described in Clauses 16-19, wherein refining the mesh of the object further comprises: iteratively increasing the resolution of the set of smaller meshes to satisfy the one or more geometric constraints.

[0120] Any combination of claim elements recited in any claim and / or any element described in any way in this application falls within the scope of this embodiment and protection.

[0121] The various embodiments have been presented for illustrative purposes, but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.

[0122] Aspects of this embodiment may be embodied as a system, method, or computer program product. Therefore, various aspects of this disclosure may take the form of a completely hardware embodiment, a completely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware (generally referred to herein as a "module" or "system"). Furthermore, any hardware and / or software technology, process, function, component, engine, module, or system described in this disclosure may be implemented as a circuit or a set of circuits. Additionally, aspects of this disclosure may take the form of a computer program product embodied in one or more computer-readable media, wherein the computer-readable medium has computer-readable program code embodied thereon.

[0123] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium includes, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination thereof. More specific examples (not an exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the context of this document, a computer-readable storage medium can be any tangible medium that can contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0124] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine. When the instructions are executed by the processor of the computer or other programmable data processing apparatus, the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams can be implemented. Such processors can be, but are not limited to, general-purpose processors, special-purpose processors, application-specific processors, or field-programmable gate arrays.

[0125] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a portion of a module, segment, or code, including one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative embodiments, the functions indicated in the blocks may occur in the order indicated in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, may be implemented by a dedicated hardware-based system that performs the specified function or action or a combination of dedicated hardware and computer instructions.

[0126] While the foregoing description is directed to embodiments of this disclosure, other and further embodiments of this disclosure may be designed without departing from its essential scope, the scope of which is determined by the claims.

Claims

1. A method comprising: predicting one or more three-dimensional (3D) mesh representations based on one or more two-dimensional (2D) images, wherein the predicting comprises: executing a machine learning model to produce a mesh of an object from a first value of a vector in a latent space, wherein the first value is based on sparse features of the one or more features of the one or more 2D images; and subsequently refining the mesh of the object by minimizing at least one difference between the one or more 3D mesh representations and the one or more 2D images using a second value of the vector in the latent space, wherein the second value is selected based on one or more geometric constraints associated with the one or more 2D images.

2. The method of claim 1, further comprising generating the machine learning model as a variational autoencoder to reconstruct a set of training meshes input to the variational autoencoder.

3. The method of claim 2, wherein generating the machine learning model comprises: for each training mesh in the set of training meshes, aggregating a plurality of rankings of points in the training mesh into a normalized value in the latent space; and training a decoder in the variational autoencoder to reconstruct the training mesh from the normalized value.

4. The method of claim 3, wherein aggregating the plurality of rankings of points in the training mesh into the normalized value comprises: applying an encoder to the plurality of rankings of points to generate an intermediate representation of the plurality of rankings of points; and averaging the intermediate representation into the normalized value.

5. A processor comprising: logic configured to predict one or more three-dimensional (3D) mesh representations based on one or more two-dimensional (2D) images, wherein the predicting comprises: executing a machine learning model to produce a mesh of an object from a first value of a vector in a latent space, wherein the first value is based on sparse features of the one or more features of the one or more 2D images; and subsequently refining the mesh of the object by minimizing at least one difference between the one or more 3D mesh representations and the one or more 2D images using a second value of the vector in the latent space, wherein the second value is selected based on one or more geometric constraints associated with the one or more 2D images.

6. The processor of claim 5, wherein refining the mesh of the object comprises: selecting the first value; computing an error between the mesh of the object and the one or more 2D images; and performing gradient descent on the first value with parameters of the machine learning model to reduce the error. at least one of randomizing the first value, selecting the first value based on a set of image observations, and initializing the first value based on sparse features extracted from the set of image observations.

7. The processor of claim 6, wherein selecting the first value comprises:

8. The processor of claim 5, wherein refining the mesh of the object comprises: dividing the mesh into a set of small meshes; ​ For each small grid in the group of small grids, select the second value in the latent space to learn a priori information about a portion of the grid represented by the small grid; and Reconstruct the mesh from the group of small meshes.

9. The processor of claim 8, wherein refining the mesh of the object further comprises: For each small grid in the group of small grids, a custom pose of the small grid is learned, which aligns the global pose of one or more 2D images with the canonical pose of the small grid.

10. The processor of claim 8, wherein refining the mesh of the object further comprises: The resolution of the group of small meshes is iteratively increased to satisfy one or more geometric constraints.

11. The processor of claim 5, wherein the logic further generates the machine learning model as a variational autoencoder to reconstruct the training grid set input to the variational autoencoder.

12. The processor of claim 11, wherein generating the machine learning model comprises: For each training grid in the training grid group, multiple rankings of points in the training grid are aggregated into standardized values ​​in the latent space; as well as The decoder is trained in the variational autoencoder to reconstruct the training grid from the normalized values.

13. The processor of claim 12, wherein aggregating the plurality of sorts of the points in the training grid into the standardized value comprises: The encoder is applied to multiple sorts of the points to generate intermediate representations of the multiple sorts of the points; as well as The intermediate representation is averaged to the standardized value.

14. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to at least: Predict one or more three-dimensional (3D) mesh representations based on one or more two-dimensional (2D) images, wherein the prediction includes: Execute a machine learning model to generate a grid of objects from a first value of a vector in a latent space, wherein the first value is based on sparse features from one or more features of one or more 2D images; as well as The object's mesh is then refined using a second value of the vector in the latent space by minimizing at least one difference between one or more 3D mesh representations and the one or more 2D images, wherein the second value is selected based on one or more geometric constraints associated with the one or more 2D images.

15. The non-transitory computer-readable medium of claim 14, wherein refining the mesh of the object comprises: Select the first value; as well as Gradient descent is performed on the first value using the parameters of the machine learning model to reduce the error between the grid and the image observation group.

16. The non-transitory computer-readable medium of claim 15, wherein the error includes at least one of photometric error and profile error.

17. The non-transitory computer-readable medium of claim 14, wherein refining the mesh of the object comprises: Divide the grid into groups of smaller grids; For each small grid in the group of small grids, the second value in the latent space is selected to learn a priori information about a portion of the grid represented by the small grid; and Reconstruct the mesh from the group of small meshes.

18. The non-transitory computer-readable medium of claim 17, wherein refining the mesh of the object further comprises: The resolution of the group of small meshes is iteratively increased to satisfy one or more geometric constraints.

Citation Information

Patent Citations

  • Mesh reconstruction from heterogeneous sources of data

    US20150146971A1