Reconstructing 3D objects from videos
By using the time consistency and texture invariance constraints of 3D object construction neural network system, the 3D structure of non-rigid objects can be reconstructed from unlabeled videos, which solves the problem of reconstructing non-rigid objects in existing technologies and achieves efficient and reliable 3D reconstruction effects.
Patent Information
- Application Number
- CN202110864471.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-31
- Filing Date
- 2021-07-29
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2041-07-29
AI Technical Summary
Existing technologies have difficulty in effectively reconstructing the 3D structure of non-rigid objects in computer vision, especially objects such as animals captured in the wild. The restricted environment and limited annotations make it difficult to promote the methods.
A neural network system is constructed using 3D objects, and the temporal consistency and texture invariance constraints of the objects are exploited to reconstruct the 3D representation of non-rigid objects from unlabeled videos. The neural network is trained through self-supervised regularization and adaptive techniques to predict the 3D shape and texture of the objects.
It achieves efficient and reliable recovery of the temporally consistent 3D structure of non-rigid objects from videos, improving the accuracy and stability of 3D reconstruction in natural environments.
Smart Images

Figure CN114092665B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to three-dimensional (3D) object reconstruction, and in particular, to techniques for constructing 3D objects from video. Background Art
[0002] When we humans try to understand an image of an object, such as a duck, we immediately recognize "duck." We also immediately perceive and imagine what the duck would look like from other viewpoints, its shape in the 3D world. Furthermore, when we see the duck in a video, its 3D structure and deformation become even more apparent to us. Our ability to perceive the 3D structure of an object actively contributes to our rich understanding of the object.
[0003] Although 3D perception is easy for humans, 3D reconstruction of deformable objects remains a very challenging problem in computer vision, especially for objects in the wild. For learning-based algorithms, the bottleneck is the lack of supervision available for training. It is challenging to collect 3D annotations (such as 3D shape and camera pose) without limiting the domain (e.g., rigid objects, human bodies, and faces) for which 3D annotations can be captured in a constrained environment. However, conventional methods in limited domains do not generalize well to non-rigid objects captured in natural environments (e.g., animals). Due to the restricted environment and limited annotations, it is very difficult to generalize conventional methods to 3D construction of non-rigid objects (e.g., animals) from images and videos captured in the wild. There is a need to solve these problems and / or other problems associated with the existing technology. Summary of the Invention
[0004] The 3D object reconstruction neural network system learns to predict 3D representations of objects from videos containing the objects. Objects in the video maintain temporal consistency, having consistent shape and texture across multiple frames. The temporal consistency of the objects is exploited to reconstruct dynamic 3D representations of the objects from unlabeled videos. Texture, identity shape, and part correspondence invariance constraints can be applied to fine-tune the neural network system. The reconstruction technique generalizes well, especially for non-rigid objects, and the neural network system can reason in real time.
[0005] Methods, computer-readable media, and systems for constructing a 3D representation of an object from a video are disclosed. In one embodiment, a neural network model receives a video including images of an object captured from a camera pose and predicts a 3D shape representation of the object for a first image in the image based on a set of learned shape bases. The neural network model also predicts a texture flow for the first image and maps pixels from the first image to a texture space according to the texture flow to produce a texture image, wherein the transfer of the texture image to the 3D shape representation constructs a 3D object corresponding to the object in the first image. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The present system and method for reconstructing a three-dimensional (3D) object from a video are described in detail below with reference to the accompanying drawings, wherein:
[0007] Figure 1A A block diagram of an example 3D object reconstruction system suitable for implementing some embodiments of the present disclosure is shown.
[0008] Figure 1B A method for using a computer suitable for implementing some embodiments of the present disclosure is shown. Figure 1A Flowchart of a method by which the system shown in FIG. 1 reconstructs a 3D representation of an object.
[0009] Figure 1C A conceptual diagram illustrating a temporal consistency constraint according to an embodiment is shown.
[0010] Figure 1D A method for applying self-supervised adaptation to a Figure 1A A flow chart of the method of the system is shown in FIG.
[0011] Figure 2A shows some embodiments suitable for implementing the present disclosure, Figure 1A Block diagram of an example training configuration for the 3D object construction system shown in .
[0012] Figure 2B A conceptual diagram illustrating the use of temporal invariance to facilitate part correspondence according to an embodiment.
[0013] Figure 2C A conceptual diagram illustrating training using annotation reprojection suitable for implementing some embodiments of the present disclosure is shown.
[0014] Figure 2D A method for training a computer suitable for implementing some embodiments of the present disclosure is shown. Figure 1A Flowchart of the method of the 3D object construction system shown in .
[0015] Figure 3 An image and a reconstructed object according to an embodiment are shown.
[0016] Figure 4 An example parallel processing unit suitable for implementing some embodiments of the present disclosure is shown.
[0017] Figure 5A is suitable for use in implementing some embodiments of the present disclosure Figure 4 Conceptual diagram of the processing system implemented by the PPU.
[0018] Figure 5B An exemplary system is shown in which the various architecture and / or functionality of various previous embodiments may be implemented.
[0019] Figure 6A is suitable for implementing some embodiments of the present disclosure, Figure 4 Conceptual diagram of the graphics processing pipeline implemented by the PPU.
[0020] Figure 6B An exemplary game streaming system suitable for implementing some embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0021] The task of 3D reconstruction requires simultaneously recovering the 3D shape, texture, and camera pose of an object from a 2D image. Due to the inherent ambiguity in correctly estimating both shape and camera pose, this task is highly ill-posed. However, a 3D object reconstruction neural network system can learn to predict 3D representations of objects from videos.
[0022] In an embodiment, a 3D object construction neural network system is trained to reconstruct a temporally consistent 3D mesh of a deformable object instance from a video. In an embodiment, the video includes real animals in a natural environment. Prior to inference, the neural network system is trained to jointly predict the image's shape, texture, and camera pose using a collection of single-view images of the same class for class-specific 3D reconstruction. A first example class may include, but is not limited to, birds (including ducks). A second example class may be a horse. Typically, classes include animals with similar structures, such as animals within a single species. The neural network can be trained without requiring annotated 3D meshes, 2D keypoints, or camera poses for each video frame.
[0023] Then, at inference time, a self-supervised regularization term that exploits the temporal consistency of object instances is used to adapt the neural network system over time to enforce that all reconstructed meshes of an object share a common texture map, underlying (identity) shape, and part. As a result of the adaptive refinement, the neural network system recovers temporally consistent and reliable 3D structure from videos of non-rigid objects, including those of animals captured in the wild—a challenging task that has rarely been solved.
[0024] Figure 1AA block diagram of a 3D object construction system 100 according to an embodiment is shown. It should be understood that this and other arrangements described herein are set forth as examples only. In addition to or in place of those arrangements and elements shown, other arrangements and elements (e.g., machines, interfaces, functions, sequences, functional groupings, etc.) may be used, and some elements may be omitted altogether. Further, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities may be performed by hardware, firmware, and / or software. For example, the various functions may be performed by a processor executing instructions stored in a memory. Furthermore, it will be understood by those skilled in the art that any system that performs the operations of the 3D object construction system 100 is within the scope and spirit of embodiments of the present invention.
[0025] The 3D object construction system 100 includes a neural network model that includes at least an encoder 105, a shape decoder 115, and a motion decoder 120. The encoder 105 extracts features 110 from each frame (eg, image) in a video. Figure 1A , an input image 102 including an object and a predicted 3D object 104 output by a 3D object construction system 100 is shown. Features 110 are then processed by various decoders to predict identity shapes, motion (offsets), textures, and cameras (not shown). In an embodiment, a shape decoder 115 outputs an identity shape 116 that represents a basic shape of the same class (e.g., a duck, a flying bird, a fat bird, a standing bird, etc.). In an embodiment, the identity shape 116 is defined by a 3D mesh of vertices defining faces of the mesh surface. In an embodiment, meshes of a particular class are deformed from a predefined sphere and have the same number of vertices / faces. The motion decoder 120 predicts an offset relative to each vertex in the identity shape 116. For each video frame, the offsets define a shape deformation applied to the identity shape 116 and appear as movement over time.
[0026] In contrast to conventional 3D reconstruction techniques, the predicted shapes are not limited to symmetrical shapes. The assumption of symmetry does not apply to most non-rigid animals, e.g. a bird tilting its head or a horse walking. This is particularly important for recovering dynamic meshes in sequence, e.g. when a bird rotates its head the 3D shape is no longer mirror symmetrical. Therefore, the assumption of symmetry can be removed and the constructed mesh is allowed to fit more complex, non-rigid poses. Simply removing the symmetry assumption for the predicted vertex offsets results in too many degrees of freedom for shape deformation. To account for undesired deformations, the shape decoder 115 learns N b A collection of shape bases The basis or identity shape 116 is calculated as a weighted combination of shape bases, denoted as the basis shape V base Compared with a single grid template, the basic shape V base It is more powerful in capturing the identity of the object and relieves the shape decoder 115 from predicting large motion deformations (e.g., deforming a standing bird template into a flying bird).
[0027] Different shape basis sets are learned during training by clustering the constructed meshes, where the meshes in each set share similar shapes and the basis shape is their average shape. is predicted by the shape decoder 115 and used to combine corresponding meshes in the shape basis set to produce the identity shape 116. The identity shape 116 can be calculated as:
[0028]
[0029] The motion decoder 120 predicts the offset relative to each vertex in the identity shape 116 Offset 118 encodes the asymmetric, non-rigid motion of the object, defining a deformation for each vertex in identity shape 116 .
[0030]
[0031] The offsets 118 are applied to the vertices of the identity shape 116 by the 3D mesh construction unit 112 to construct a predicted 3D shape representation 108 (eg, a wireframe or mesh) of the object. The predicted 3D shape representation 108 is Figure 1A is shown as a checkerboard for visualization purposes, where the different faces defined by the vertices in the mesh are colored black or white.
[0032] The texture decoder 125 receives the features 110 and predicts the texture image 106 for each frame of the video. The texture decoder 125 predicts the texture flow of each image based on the features 110. The texture stream maps pixels from the image to a UV texture space to produce a texture image 106. The 3D mesh texture unit 122 can then use a predefined UV mapping function to map the texture image 106 from the texture space to the 3D shape representation 108. Applying the texture image 106 to the 3D shape representation 108 produces a 3D object 104 (e.g., a textured mesh). The 3D object 104 is represented by a |V| vertex. |F| side and having a height H uv and width W uv UV texture image The UV texture space provides a parameterization that is invariant to object deformation. Therefore, the image texture of an image in a video should be constant or invariant to deformation over time.
[0033] By enforcing the predicted values of the texture image 106 to be consistent across different frames of the video in UV texture space, the neural network model can be regularized to generate coherent reconstructions over time during inference. Temporal invariance can be used as a self-supervisory signal to tune the neural network model. In an embodiment, during inference, self-supervised regularization is used to adapt the 3D object construction system 100 based on shape invariance and texture invariance. Specifically, objects in the video maintain a consistent shape and consistent texture across multiple frames. The self-supervised adaptation technique exploits the temporal consistency of objects to construct dynamic 3D objects from unlabeled video. As further described herein, texture image, identity shape, and part correspondence invariance constraints can be applied to train and / or fine-tune the neural network model.
[0034] Now, more illustrative information about various optional architectures and features that can implement the aforementioned framework will be described based on the user's expectations. It should be strongly noted that the following information is described for illustrative purposes and should not be interpreted as limiting in any way. Any of the following features can be optionally combined with or without excluding the other features described.
[0035] Figure 1B A method for using the Figure 1A 100 is a flowchart of a method 130 for constructing a 3D representation of an object using the 3D object reconstruction system 100 shown in . Each block of the method 130 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, the various functions can be performed by a processor executing instructions stored in a memory. The method 130 can also be embodied as computer-usable instructions stored on a computer storage medium. To name a few, the method 130 can be provided by a standalone application, a service, or a hosted service (standalone or in combination with another hosted service), or a plug-in to another product. Additionally, as an example, regarding Figure 1A Method 130 is described with reference to a system. However, this method may additionally or alternatively be performed by any system or any combination of systems, including but not limited to the systems described herein. Furthermore, one of ordinary skill in the art will understand that any system that performs method 130 is within the scope and spirit of embodiments of the present invention.
[0036] At step 135, the 3D object construction system 100 receives a video including images of an object captured from a camera pose. In an embodiment, the video is unlabeled. In an embodiment, the object is a non-rigid animal. In an embodiment, the video is captured in the "wild."
[0037] At step 140, the 3D object construction system 100 predicts a 3D shape representation of an object in a first image of the image based on the set of learned shape basis. In an embodiment, the 3D object construction system 100 predicts an identity shape from a set of learned shapes. In an embodiment, the identity shape is calculated using equation (1). In an embodiment, the identity shape is calculated as the sum of the component shapes included in the learned shape basis set, and each component shape is scaled accordingly using a coefficient generated by the neural network model. In an embodiment, a shape offset (e.g., a non-rigid motion deformation) is calculated and applied to the vertices of the identity shape to predict a 3D shape representation of the object. In an embodiment, the 3D shape representation is a mesh of vertices defining a face.
[0038] At step 145, the 3D object construction system 100 predicts a texture flow for the first image. At step 147, the 3D object construction system 100 maps pixels from the first image to a texture space based on the texture flow to generate a texture image corresponding to the first image. In an embodiment, for a particular image, the texture image is transferred to a 3D shape representation to generate a 3D object corresponding to the object in the first image. In an embodiment, steps 140, 145, and 147 are repeated for each image in the video. In an embodiment, the 3D construction system 100 also predicts a camera pose based on features extracted from the video. When a 3D object is rendered based on the camera pose for each frame of the video, the rendered object appears as the object in the frame.
[0039] Figure 1C A conceptual diagram of a temporal consistency constraint according to an embodiment is shown. The neural network model 150 includes an encoder 105, a shape decoder 115, and a texture decoder 125. In an embodiment, to enforce texture invariance, the reconstructed texture images of any pair of frames (e.g., images) in a video sequence are exchanged. Figure 1C As shown, frames 132 and 136 are processed by a neural network model 150 to predict corresponding identity shapes 142 and 146 . Figure 1C The identity shapes 142 and 146 shown in FIG are each posed according to a corresponding camera pose θ for a corresponding frame, where In an embodiment, the camera pose represents a perspective transform. Aside from the different camera poses, identity shapes 142 and 146 are identical. The offsets predicted by neural network model 150 for frames 132 and 136 are applied to the corresponding identity shapes 142 and 146 to produce 3D shape representations 152 and 156.
[0040] The shape model consists of an identity or base shape V base and the offset or deformation term △V is expressed as, where the identity shape V basecorresponds intuitively to the “identity” of the instance (e.g., duck or flying bird, etc.). During online adaptation, the neural network model 150 is trained to predict a consistent V via an exchange loss function. base , to maintain the identity shape over time. Given two randomly sampled frames I i and I j , using the original offset _ΔV i and ΔV j Shape your identity and The swap and transformation are:
[0041]
[0042] where θ i and θ j is the camera pose, ΔV i and ΔV j is motion deformation, and S i and S j are the object silhouettes (masks) for frames i and j, respectively. In an embodiment, Represents a differentiable renderer for rendering a 3D representation or 3D object as a 2D outline, such as In an embodiment, Represents a differentiable renderer for rendering a textured mesh into an RGB image, such as (Note that the mesh plane F is omitted for the sake of brevity.) In an embodiment, Represents the projection of 3D point v to image space, such as niou(·,·) denotes the negative intersection over union (IoU) objective. In an embodiment, the foreground mask for the contour is obtained by a segmentation model trained with the ground truth foreground mask. The loss function L s May be used to enforce consistency between identity shapes 142 and 146 .
[0043] The neural network model 150 also predicts texture images 154 and 158 for frames 132 and 136, respectively. Based on the observation that the texture of an object mapped to UV space should be invariant to deformation and remain constant over time, a texture invariance constraint can be used to encourage consistent texture reconstruction from all frames. However, fully aggregating UV texture maps from all frames can result in blurry video-level texture images. Instead, as with shape identity, texture consistency can be enforced between random pairs of frames via an exchange loss. Given two randomly sampled frames I i and I j , texture image and is swapped and reconstructed with the original mesh Vi and V j The combination is:
[0044]
[0045] where dist(·,·) is a perceptual metric. The swapping technique enforces the consistency of texture images across frames, thereby improving Figure 1A The accuracy of the predicted texture image 106, 3D shape representation 108 and 3D object 104 in . For example, the loss function L t to enforce consistency between texture images 154 and 158. During inference, neural network model 150 can be fine-tuned on a specific video, where the invariance constraints are enforced by equations (3) and (4). In an embodiment, self-supervision is used to fine-tune neural network model 150. Fine-tuning can improve the performance of neural network model 150 when domain differences in video quality, such as lighting conditions, lead to inconsistent 3D mesh reconstruction (because each frame of the video is processed independently).
[0046] Figure 1D A method for applying self-supervised adaptation to a Figure 1A Flowchart of method 148 of the system shown in FIG. A neural network model 150 is refined by exploiting redundancy in a time series as a form of self-supervision to improve the construction of dynamic non-rigid objects. Although method 148 is described in the context of a neural network model, method 148 may also be performed by a program, custom circuitry, or a combination of custom circuitry and a program. For example, method 148 may be performed by a GPU, a CPU, or any processor capable of implementing a neural network model. Furthermore, one of ordinary skill in the art will understand that any system that performs method 148 is within the scope and spirit of embodiments of the present invention.
[0047] At step 155, neural network model 150 receives a sequence of frames of a video including an object. At step 160, a silhouette and camera pose are obtained. In an embodiment, the silhouette is ground truth data corresponding to the video. In an embodiment, the camera pose is predicted by neural network model 150. Steps 165 and 170 can be performed in parallel or sequentially. At step 165, neural network model 150 predicts a texture image for each frame. At step 170, neural network model 150 predicts an identity shape and offset for each frame.
[0048] At step 175, a loss function is calculated based on texture invariance and shape identity invariance. In an embodiment, the loss function is an exchange loss function. In an embodiment, the loss function is a combination of equations (3) and (4).
[0049] At step 180, if the error is reduced according to the loss function, then at step 190, refinement of the neural network model 150 is completed. Otherwise, at step 185, the parameters of the neural network model 150 are updated by back-propagating the loss function before the method 148 returns to step 155. In an embodiment, parameters of one or more of the shape decoder 115, the motion decoder 120, and the texture decoder 125 are updated.
[0050] In an embodiment, a texture image predicted for a first image is transferred to a second 3D shape representation predicted for a second image in an image in a video to produce a first 3D object. The first 3D object is projected according to a first camera pose associated with the first image to produce a first projected 3D object. The second texture image predicted for the second image is transferred to the 3D shape representation predicted for the first image to produce a second 3D object. The second 3D object is projected according to a second camera pose associated with the second image to produce a second projected 3D object, and parameters of the neural network model 150 are updated to promote consistency between the first projected 3D object and the second projected 3D object. In an embodiment, differences between overlapping portions of the first and second images are reduced during training.
[0051] In an embodiment, a first non-rigid motion deformation predicted for a first image is applied to a first identity shape predicted for a second image in an image to produce a first 3D shape representation. The first 3D shape representation is projected according to a first camera pose associated with the first image to produce a first projected 3D object. A second non-rigid motion deformation predicted for the second image is applied to a second identity shape predicted for the first image to produce a second 3D shape representation. The second 3D shape representation is projected according to a second camera pose associated with the second image to produce a second projected 3D object, and parameters of the neural network model 150 are updated to promote consistency between the first projected 3D object and the second projected 3D object.
[0052] Figure 2A The method for Figure 1A A block diagram of the training configuration of the 3D object construction system 100 is shown in FIG. Figure 1A In addition to the boxes shown in , the training framework also includes a camera pose unit 225, a differentiable renderer 222, and a loss function 215. In an embodiment, the camera pose unit 225 is also included in the neural network model 150 and / or the 3D object construction system 100. The camera pose unit 225 receives the features and predicts the camera pose (position and orientation) θ of the identity shape based on the features. The shape decoder 115, the motion decoder 120, the texture decoder 125, and the camera pose unit 225 can be configured to jointly predict the identity shape, offset, texture image, and camera pose.
[0053] The differentiable renderer 222 receives the 3D mesh and the texture image and renders the 3D mesh according to the predicted camera pose provided by the camera pose unit 225. In an embodiment, the texture image is transferred to the 3D shape representation by the 3D mesh texture unit 122 to construct the 3D object. The textured 3D object is then projected by the differentiable renderer 222 according to the camera pose to produce a rendered image.
[0054] During training, the rendered image can be compared to the video frame and / or annotations. The loss function 215 can adjust the parameters of the shape decoder 115, the motion decoder 120, the texture decoder 125, and / or the camera pose unit 225 to reduce the difference between the rendered image and the video frame. In some embodiments, the loss function 215 receives ground truth data, such as object outlines and / or key point annotations, for reducing the difference between the rendered image and the video frame. In an embodiment, training is performed in a self-supervised manner using object outlines extracted from class-specific images. The training configuration can be used to train the 3D object construction system 100 using both individual input images and videos. In an embodiment, training can be performed using various techniques (e.g., supervision, self-supervision, and semi-supervision using individual images and / or videos). For example, the 3D object construction system 100 can be trained to learn a shape basis set using a semi-supervised technique with a single-view input image, and then fine-tuned via a self-supervised technique using an unlabeled video of a specific object.
[0055] Conventional techniques use annotated 2D object keypoints and class-level template shapes or contours for training. However, scaling learning with 2D annotations to hundreds of thousands of images is nontrivial and can also limit the generalization ability of trained neural network models to new domains. For example, conventional 3D construction neural network models trained on single-view images often produce unstable and erratic predictions of video data. This is expected due to temporal perturbations. Therefore, using temporal signals in videos should provide advantages rather than disadvantages.
[0056] A balance can be achieved between model generalization and specialization. In an embodiment, an image-based neural network model is trained on a set of images, while at test time, the neural network model within the 3D object construction system 100 is adapted or fine-tuned to an input video including a specific identity object. During training at test time, no labels are provided for the video. Therefore, a self-supervised objective is introduced that enables continuous improvement of the neural network model. As previously described, the UV texture space provides a parameterization that is invariant to object deformation. When mapped from 2D via the predicted texture flow, the object parts of an instance of an object should be constant. Therefore, in addition to promoting temporal consistency between texture images and identity shapes predicted for different frames, temporal consistency between object parts in the UV texture space can also be promoted. In particular, an invariance loss can be used to train the camera pose unit 225 to predict the camera position θ. Using the constraint of temporal consistency, the recovered identity shape and camera pose can be significantly stabilized and adapted based on the video processed during fine-tuning and / or test time.
[0057] Figure 2B A conceptual diagram of using temporal invariance to enforce part correspondences according to an embodiment is shown. The video includes a bird object that the 3D object construction system 100 can be trained to build. Conventional techniques such as unsupervised video correspondence (UVC) can be used to automatically apply a (random) pattern to the parts of an object in an input video frame to generate a propagated part map for each frame. The UVC model learns an affinity matrix that captures pixel-level correspondences between video frames. The UVC model can be used to propagate any annotations (e.g., segmentation labels, keypoints, part labels, etc.) from annotated keyframes to other unannotated frames. Part correspondences are generated within a video clip by "painting" a set of random regions on the object on the first frame and propagating the part paintings to the rest of the video using the UVC model. Any particular part of an object (e.g., a wing) can be painted as a single region or multiple regions. In other words, the painting does not provide a semantic definition. As Figure 2B As shown, the painting appears as vertical stripes of different colors within the outline of the object visible in the propagated site map 232 .
[0058] In an embodiment, two strategies can be adopted to obtain accurate part propagation of object parts through the UVC model. First, the parameters in the 3D object construction system 100 can be fine-tuned on a sliding window instead of all video frames. Each sliding window can include N w = 50 consecutive frames, and the sliding step is set to N s = 10. The 3D object construction system 100 can use the frames in each sliding window for N t= 40 iterations for adjustment. Second, instead of "painting" a random part onto the first frame and propagating the painted part sequentially to the remaining frames in the window, the random part can be painted onto the middle frame in the window (i.e., the Nth frame). w / 2 frames), and the painted parts can be propagated backward to the first frame in the window and forward to the last frame. This strategy can improve the propagation quality by reducing the propagation range to half the window size. Within each sliding window, the consistency of the UV texture image, part UV map, and identity shape is promoted for all frames.
[0059] A sequence of propagated part maps associated with video frames, such as the propagated part map 232 associated with frame 230, is processed by the 3D object construction system 100 in a test configuration. Processing of the propagated part map 232 is described, however, Figure 2B Additional propagated part maps shown in FIG or additional frames in the sequence can be processed in a similar manner to produce additional rendered images. In an embodiment, the part map is propagated across the object in multiple frames to produce the propagated part map. In an embodiment, the number of frames is included in the sliding window.
[0060] The propagated part map for each frame is mapped into UV texture space using the predicted texture flow 234 to produce a part UV map 236. In an embodiment, the propagated part map is mapped into texture space according to the predicted texture flow for the frame to produce a part map in texture space. In an embodiment, the part maps are aggregated to produce a video-level part UV map 235. By aggregating the part UV maps (i.e., averaging), the noise in each individual part UV map is minimized. The video-level part UV map 235 of the object depicted in the video to be constructed in 3D is shared by all frames in the video. Thus, for each frame, the video-level part UV map 235 is wrapped onto the predicted 3D shape representation 237. Aggregation of the predicted parts in UV texture space facilitates learning of the camera pose by the 3D object construction system 100.
[0061] In an embodiment, the predicted 3D shape representations for the frame are rendered according to the associated camera pose (not shown), and the video-level part map is transferred (e.g., wrapped) onto each of the 3D shape representations to produce a rendered image. The wrapped 3D shape representation 237 can be rendered by the differentiable renderer 222 according to the predicted camera to produce a rendered image 238.
[0062] For each frame, the difference between the parts rendered back to 2D space and the parts propagated is penalized. In an embodiment, the loss function 215 can be configured to update the parameters of the neural network model 150 based on this difference. In an embodiment, the parameters are updated to promote consistency between the rendered image and the propagated part map.
[0063] Since the propagated part maps are typically smooth and continuous in time, the loss implicitly regularizes the 3D object construction system 100 to predict a consistent camera pose and object shape over time. In an embodiment, rather than minimizing the difference between the rendered part map (e.g., a rendered image, such as rendered image 238) and the propagated part map of a frame, it is more robust to penalize the geometric distance between the projection of the vertex assigned to each part and the 2D point sampled from the corresponding part. The Chamfer loss can be calculated as:
[0064]
[0065] where N f is the number of frames in the video, N p =6 is the number of sites, and is the vertex assigned to part i. Chamfer distance is used to calculate the loss because the vertex projection Not strictly one-to-one correspondence to the sampled 2D points
[0066] The input samples of the propagated part map are compared with the projected samples of the image rendered based on the predicted video-level part UV map, 3D representation and camera pose to calculate the Chamfer loss. The Chamfer loss reduces the error from the camera pose estimate. Alternatively, more samples (e.g., pixels) within the color part of the propagated part map 232 can be used as ground truth, and the predicted 3D shape representation 237 can be wrapped with the video-level part UV map 235 and rendered according to the camera pose to produce a rendered image 238 for comparison with the ground truth color samples.
[0067] In one embodiment, the colored parts are used to supervise the estimated camera pose without rendering the image. The input video frame can be sampled within each part of the part propagation and then compared with the predicted samples on a 3D representation (e.g., a mesh) projected according to the predicted camera pose (without rendering).
[0068] Another technique for training uses the propagation of ground truth annotations (such as keypoints) from the input image through the predicted texture image to the rendered image. A loss can be calculated based on the ground truth annotations and the rendered annotations. The ground truth annotations enable the establishment of correspondences across different instances in a set of shape bases. For example, the beak or wing tip is marked as a keypoint in different images of a bird, and different birds with similar shapes are clustered together to form a set of shape bases.
[0069] Figure 2C A conceptual diagram illustrating training using annotation reprojection, according to an embodiment, is shown. Although the technique is described for annotations as keypoints, those skilled in the art will recognize that other types of annotations can be used with the technique. When using weak supervision to train the 3D object construction system 100, 2D keypoints can be provided as ground truth annotations for semantically related instances of an object. For example, annotated frame 240 includes multiple keypoints, such as tail keypoint 248 for the tip of a bird's tail.
[0070] When projecting 2D keypoints onto a 3D representation (e.g., a mesh surface), the same semantic keypoints of different object instances should be matched to the same face on the mesh surface. Conventionally, to model the mapping between the 3D mesh surface and the 2D keypoints, an affinity matrix is learned that describes the probability of each 2D keypoint being mapped to each vertex on the 3D representation. The probability map is a heat map, such as heat map 242.
[0071] The affinity matrix is shared across all instances and is independent of individual shape variations. Conventional methods are suboptimal because: (i) mesh vertices are a subset of discrete points on a continuous mesh surface, and thus the weighted combination of mesh vertices defined by the affinity matrix may not lie on the mesh surface, resulting in an inaccurate mapping of 2D keypoints. (ii) The mapping from image space to the mesh surface is described by the affinity matrix. In contrast, because the mapping from image space to the mesh surface is already modeled by the texture flow, learning both the affinity matrix and the texture flow independently may be redundant.
[0072] Therefore, the texture flow is reused to map the 2D keypoints from each image to the mesh surface. For example, the 2D tail keypoint 248 is mapped to the 3D tail keypoint 246 on the 3D representation. First, each 2D keypoint in the annotated frame is mapped to a UV texture space that is independent of deformation. For example, the texture flow I predicted for the annotated frame 240 can be used ... flowThe keypoints in the annotated frame 240 are mapped to UV texture space to generate an annotation map in texture space, a keypoint UV map 244. Ideally, each semantic keypoint from different instances of an object should be mapped to the same point in UV space. In practice, this may not be true due to inaccurate texture flow prediction. In order to accurately map each keypoint to UV space, a canonical keypoint UV map 245 can be calculated by: (i) mapping the keypoint heatmap of each instance to UV space via the predicted texture flow, and (ii) aggregating the keypoint UV maps across all instances to eliminate outliers caused by incorrect texture flow predictions. The keypoint UV maps are aggregated (e.g., averaged) to produce the canonical keypoint UV map 245.
[0073] The canonical keypoint UV map 245 is then transferred to the 3D shape representation predicted for the frame to produce an annotated 3D shape representation. The annotated 3D shape representation can then be projected according to the associated camera pose to produce a projected annotation of the frame. In an embodiment, the projection includes rendering. In an embodiment, the keypoint reprojection is accomplished by: (i) warping the canonical keypoint UV map to each individual predicted mesh surface to produce 3D keypoints; (ii) reprojecting these 3D keypoints via the predicted camera pose. Project back to this 2D space to produce reprojected 2D keypoints; (iii) compare the reprojected 2D keypoints to the ground truth keypoints in 2D For example, the canonical keypoint UV map 245 is mapped to a 3D representation to produce a 3D keypoint 246 including a tail keypoint 250. The 3D keypoint 246 is reprojected according to the predicted camera pose and compared with the ground truth keypoints in the annotated frame 240. In an embodiment, the parameters of the neural network model 150 are updated to reduce the discrepancy and promote consistency between the projected keypoints and the ground truth keypoints.
[0074] Given a 3D correspondence (expressed as each 2D semantic keypoint of ), the keypoint reprojection loss enforces the projection of the former to be consistent with the latter via the following equation:
[0075]
[0076] where N k is the number of keypoints. Evaluation of the keypoint reprojection loss function implicitly reveals the correctness of both the predicted shape and camera pose of the mesh reconstruction algorithm, especially for objects without 3D ground truth annotations. The annotated reprojection enables weakly supervised training of the 3D object construction system 100.
[0077] A bottleneck of conventional image-based 3D mesh reconstruction methods is the assumption that the predicted shape is symmetrical. This assumption does not apply to most non-rigid animals, such as birds tilting their heads or horses walking. Therefore, the assumption of symmetry can be ignored and the reconstructed mesh representation can be allowed to fit more complex non-rigid poses via as rigid as possible (ARAP) constraints. The ARAP constraint is an additional loss target that can be used for self-supervised training of the 3D object construction system 100. By construction, the identity shape is smooth. However, the application of the offset predicted by the motion decoder 120 can produce discontinuities in the 3D mesh representation. The ARAP constraint is used to ensure that the edge length of the 3D mesh is maintained even when the 3D mesh is rotated. ARAP is a self-supervised regularization that maintains rigidity and can be used for individual input images and video sequences.
[0078] Without any pose-dependent regularization, the predicted motion deformation ΔV often results in erroneous random deformations and spikes on the surface of the 3D mesh, which does not faithfully describe the motion of non-rigid objects. The ARAP constraint promotes the rigidity of local transformations and the preservation of local mesh structure. The ARAP constraint is formulated as an objective that ensures that the predicted shape V is derived from the predicted base shape V by the following equation base Local rigid transformation of :
[0079]
[0080] in represents the adjacent vertices of vertex i, w ij and R i are the co-tangent weights and the best approximation rotation matrix, respectively. As another constraint that does not require any labels, ARAP can be enforced with video input during test-time training to improve shape prediction.
[0081] In an embodiment, based on the ARAP loss function L arap The parameters of the neural network model 150 are updated to reduce discontinuities in the predicted 3D shape representation. Non-rigid motion deformations of the 3D shape representation are predicted for each frame and applied to the identity shape predicted for each frame to produce a 3D shape representation of the object. The ARAP loss function can be evaluated based on the rotation difference between the identity shapes and the difference between the 3D shape representations.
[0082] In an embodiment, two image-based 3D object construction methods are used to train the neural network model 150: (i) a weakly supervised method (i.e., provided with object outlines and 2D annotations), and (ii) a self-supervised method where only object outlines are available. The trained image-based neural network model 150 is then adapted to videos. For example, in an embodiment, the trained neural network model 150 is adapted to videos of birds and zebras in the wild.
[0083] Figure 2D The method for training according to the embodiment is shown Figure 1A Flowchart of method 255 of 3D object construction system 100 shown in FIG. At step 260, 3D object construction system 100 is trained to learn a set of shape basis from a single-view image. In an embodiment, the objectives used for single-view construction, alone or in combination, include (i) foreground mask loss: a negative intersection-over-union objective between rendered and ground-truth contours; (ii) foreground RGB texture loss: a perceptual metric between the rendered RGB image and the input RGB image; (iii) mesh smoothness: a Laplacian constraint for promoting smooth mesh reconstruction; (iv) keypoint reprojection loss; and (v) ARAP constraint.
[0084] At step 270, the 3D object construction system 100 is trained on the video using self-supervision. Because it is feasible to predict the segmentation mask using a pre-trained segmentation model, the predicted foreground mask can be used to calculate the foreground mask loss. In embodiments, objectives used for video use, either alone or in combination, may include a foreground RGB texture loss and a mesh smoothing objective. In embodiments, an ARAP constraint may also be used.
[0085] At step 275 , the 3D object construction system 100 is fine-tuned to construct a specific 3D object using one or more of the invariance constraints of texture, identity shape, and part correspondence.
[0086] Figure 33D representation and object 305 are shown. The 2D bird object shown in images 300, 310, and 320 is 3D constructed using the 3D object construction system 100, which includes a camera pose unit 225 trained on single-view images to produce 3D representation and object 305, trained on single-view images and video (but without invariance constraints) to produce 3D representation and object 315, and trained on single-view images and video (with invariance constraints) to produce 3D representation and object 325. Note that the object shape, camera pose, and texture are less reliably predicted for 3D representation and object 315 than for 325. Similarly, the object shape, camera pose, and texture are less reliably predicted for 3D representation and object 305 than for 315. In summary, the 3D object construction algorithm recovers temporally consistent and reliable 3D structure from videos of non-rigid objects, including videos of animals captured in the wild.
[0087] The 3D object construction technique does not require a predefined template object mesh, annotations of 3D objects, 2D annotations, or camera poses for video frames. The video-based 3D object construction system 100 can be refined via self-supervised online adaptation for any incoming test video. First, a category-specific 3D construction neural network model 150 is learned from a collection of single-view images of the same category. The 3D object construction system 100 jointly predicts the shape, texture, and camera pose of the object in the image. Then, at inference time, the neural network model is fine-tuned over time by the object-specific test videos using a self-supervised regularization term that exploits the temporal consistency of object instances to enforce that all reconstructed meshes share a common texture map, basic shapes, and parts.
[0088] The 3D object construction technology can be used for content creation, such as generating 3D characters for games, movies, and 3D printing. Because the 3D character is generated from video, the content can also include the character's motion, as predicted based on the video. Compared to conventional solutions, the 3D object construction technology does not rely on a predefined parametric mesh (for example, a human object defined by a fixed number of vertices). The 3D object construction system also generalizes well, especially for non-rigid objects.
[0089] Parallel processing architecture
[0090] Figure 4 A parallel processing unit (PPU) 400 is shown according to an embodiment. The PPU 400 may be used to implement the 3D object construction system 100. The PPU 400 may be used to implement one or more of the encoder 105, the shape decoder 115, the motion decoder 120, the texture decoder 125, the 3D mesh construction unit 130, and the 3D mesh texture unit 122 within the server / client system 100.
[0091] In an embodiment, the PPU 400 is a multi-threaded processor implemented on one or more integrated circuit devices. The PPU 400 is a latency-hiding architecture designed for processing many threads in parallel. A thread (i.e., an execution thread) is an instance of an instruction set configured to be executed by the PPU 400. In one embodiment, the PPU 400 is a graphics processing unit (GPU) that is configured to implement a graphics rendering pipeline for processing three-dimensional (3D) graphics data to generate two-dimensional (2D) image data for display on a display device. In other embodiments, the PPU 400 can be used to perform general-purpose computations. Although an exemplary parallel processor is provided herein for illustrative purposes, it should be specifically noted that the processor is described for illustrative purposes only and any processor can be used to supplement and / or replace the processor.
[0092] One or more PPUs 400 can be configured to accelerate thousands of high-performance computing (HPC), data center, cloud computing, and machine learning applications. PPUs 400 can be configured to accelerate numerous deep learning systems and applications, including autonomous vehicle platforms, deep learning, high-precision speech, image, and text recognition systems, intelligent video analysis, molecular simulations, drug discovery, disease diagnosis, weather forecasting, big data analytics, astronomy, molecular dynamics simulations, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations, among others.
[0093] like Figure 4 As shown, the PPU 400 includes an input / output (I / O) unit 405, a front-end unit 415, a scheduler unit 420, a work distribution unit 425, a hub 430, a crossbar switch (Xbar) 470, one or more general processing clusters (GPCs) 450, and one or more memory partitioning units 480. The PPU 400 can be connected to a host processor or other PPUs 400 via one or more high-speed NVLink 410 interconnects. The PPU 400 can be connected to a host processor or other peripheral devices via interconnect 402. The PPU 400 can also be connected to a local memory 404 including multiple memory devices. In one embodiment, the local memory can include multiple dynamic random access memory (DRAM) devices. The DRAM devices can be configured as a high-bandwidth memory (HBM) subsystem, in which multiple DRAM dies are stacked within each device.
[0094] The NVLink 410 interconnect enables the system to scale and include one or more PPUs 400 in conjunction with one or more CPUs, supporting cache coherency between the PPU 400 and the CPU, and CPU mastering. Data and / or commands can be sent by the NVLink 410 through the hub 430 to or from other units of the PPU 400, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). Figure 5B NVLink 410 is described in more detail.
[0095] I / O unit 405 is configured to send and receive communications (e.g., commands, data, etc.) from a host processor (not shown) via interconnect 402. I / O unit 405 can communicate with the host processor directly via interconnect 402, or through one or more intermediary devices (such as a memory bridge). In one embodiment, I / O unit 405 can communicate with one or more other processors (e.g., one or more PPUs 400) via interconnect 402. In one embodiment, I / O unit 405 implements a Peripheral Component Interconnect Express (PCIe) interface for communicating over a PCIe bus, and interconnect 402 is a PCIe bus. In alternative embodiments, I / O unit 405 can implement other types of known interfaces for communicating with external devices.
[0096] I / O unit 405 decodes data packets received via interconnect 402. In one embodiment, the data packets represent commands configured to cause PPU 400 to perform various operations. I / O unit 405 sends the decoded commands to various other units of PPU 400 as specified by the commands. For example, some commands may be sent to front-end unit 415. Other commands may be sent to hub 430 or other units of PPU 400, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). In other words, I / O unit 405 is configured to route communications between and among the various logical units of PPU 400.
[0097] In one embodiment, a program executed by a host processor encodes a command stream in a buffer that provides a workload to the PPU 400 for processing. The workload may include many instructions and data to be processed by those instructions. A buffer is an area of memory that is accessible (e.g., read / write) by both the host processor and the PPU 400. For example, the I / O unit 405 may be configured to access a buffer in system memory connected to the interconnect 402 via a memory request transmitted over the interconnect 402. In one embodiment, the host processor writes a command stream into the buffer and then sends a pointer to the start of the command stream to the PPU 400. The front end unit 415 receives pointers to one or more command streams. The front end unit 415 manages the one or more streams, reads commands from the streams, and forwards the commands to the various units of the PPU 400.
[0098] Front-end unit 415 is coupled to scheduler unit 420, which configures various GPCs 450 to process tasks defined by one or more streams. Scheduler unit 420 is configured to track state information related to the various tasks managed by scheduler unit 420. The state may indicate which GPC 450 a task is assigned to, whether the task is active or inactive, the priority associated with the task, and the like. Scheduler unit 420 manages the execution of multiple tasks on one or more GPCs 450.
[0099] Scheduler unit 420 is coupled to work distribution unit 425, which is configured to dispatch tasks for execution on GPCs 450. Work distribution unit 425 can track a number of scheduled tasks received from scheduler unit 420. In one embodiment, work distribution unit 425 manages a pending task pool and an active task pool for each GPC 450. When a GPC 450 completes execution of a task, the task is evicted from the active task pool of GPC 450, and one of the other tasks from the pending task pool is selected and scheduled for execution on GPC 450. If an active task on a GPC 450 has become idle, for example, while waiting for a data dependency to be resolved, the active task can be evicted from GPC 450 and returned to the pending task pool, while another task in the pending task pool is selected and scheduled for execution on GPC 450.
[0100] In one embodiment, a host processor executes a driver kernel that implements an application programming interface (API) that enables one or more applications executing on the host processor to schedule operations for execution on PPU 400. In one embodiment, multiple computing applications are executed simultaneously by PPU 400, and PPU 400 provides isolation, quality of service (QoS), and independent address spaces for the multiple computing applications. An application can generate instructions (e.g., API calls) that cause the driver kernel to generate one or more tasks for execution by PPU 400. The driver kernel outputs the tasks to one or more streams being processed by PPU 400. Each task can include one or more related groups of threads, referred to herein as warps. In one embodiment, a warp includes 32 related threads that can execute in parallel. Collaborating threads can refer to multiple threads that include instructions to execute a task and can exchange data through shared memory. The task can be assigned to one or more processing units within GPC 450, and the instructions are scheduled for execution by at least one warp.
[0101] Work distribution unit 425 communicates with one or more GPCs 450 via XBar 470. XBar 470 is an interconnect network that couples many of the units of PPU 400 to other units of PPU 400. For example, XBar 470 can be configured to couple work distribution unit 425 to a particular GPC 450. Although not explicitly shown, one or more other units of PPU 400 can also be connected to XBar 470 via hub 430.
[0102] Tasks are managed by a scheduler unit 420 and dispatched to GPCs 450 by a work distribution unit 425. A GPC 450 is configured to process tasks and generate results. The results can be consumed by other tasks within the GPC 450, routed to a different GPC 450 via an XBar 470, or stored in memory 404. The results can be written to memory 404 via a memory partition unit 480, which implements a memory interface for reading data from and writing data to memory 404. The results can be sent to another PPU 400 or CPU via NVLink 410. In an embodiment, a PPU 400 includes a number U of memory partition units 480, which is equal to the number of separate and distinct memory devices coupled to the memory 404 of the PPU 400. Each GPC 450 may include a memory management unit to provide virtual to physical address translation, memory protection, and arbitration of memory requests. In one embodiment, the memory management unit provides one or more translation lookaside buffers (TLBs) for performing translations of virtual addresses to physical addresses in memory 404 .
[0103] In one embodiment, the memory partition unit 480 includes a raster operations (ROP) unit, a level 2 (L2) cache, and a memory interface coupled to the memory 404. The memory interface can implement a 32-, 64-, 128-, 1024-bit data bus, etc., for high-speed data transfer. The PPU 400 can be connected to up to Y memory devices, such as a high-bandwidth memory stack or graphics double data rate, version 5, synchronous dynamic random access memory, or other types of persistent storage. In an embodiment, the memory interface implements an HBM2 memory interface, and Y is equal to half of U. In an embodiment, the HBM2 memory stack is located on the same physical package as the PPU 400, providing significant power and area savings compared to conventional GDDR5 SDRAM systems. In an embodiment, each HBM2 stack includes four memory dies and Y is equal to 4, where each HBM2 stack includes two 128-bit channels per die (8 channels total) and a data bus width of 1024 bits.
[0104] In an embodiment, memory 404 supports single-error correction, double-error detection (SECDED) error correction code (ECC) to protect data. ECC provides higher reliability for computing applications that are sensitive to data corruption. Reliability is particularly important in large-scale cluster computing environments where PPU 400 processes very large data sets and / or runs applications for extended periods of time.
[0105] In one embodiment, the PPU 400 implements a multi-level memory hierarchy. In one embodiment, the memory partitioning unit 480 supports unified memory to provide a single unified virtual address space for the CPU and PPU 400 memory, thereby enabling data sharing between virtual memory systems. In one embodiment, the frequency of PPU 400 accesses to memory located on other processors is tracked to ensure that memory pages are moved to the physical memory of the PPU 400 where pages are accessed more frequently. In one embodiment, NVLink 410 supports address translation services that allow the PPU 400 to directly access the CPU's page tables and provide full access to the CPU's memory by the PPU 400.
[0106] In an embodiment, the copy engine transfers data between multiple PPUs 400 or between a PPU 400 and a CPU. The copy engine may generate a page fault for an address that is not mapped to a page table. The memory partition unit 480 may then service the page fault, map the address to a page table, and then the copy engine may perform the transfer. In conventional systems, for multiple copy engines operating between multiple processors, memory is pinned (e.g., non-pageable), significantly reducing available memory. In the event of a hardware page fault, the address can be passed to the copy engine without concern for memory page residency, and the copy process is transparent.
[0107] Data from memory 404 or other system memory can be retrieved by memory partition unit 480 and stored in L2 cache 460, which is located on-chip and shared between various GPCs 450. As shown, each memory partition unit 480 includes a portion of the L2 cache associated with the corresponding memory 404. Lower-level caches can then be implemented in different units within GPC 450. For example, each processing unit in GPC 450 can implement a level 1 (L1) cache. The L1 cache is a dedicated memory dedicated to a specific processing unit. L2 cache 460 is coupled to memory interface 470 and XBar 470, and data from the L2 cache can be retrieved and stored in each of the L1 caches for processing.
[0108] In one embodiment, the processing units within each GPC 450 implement a SIMD (single instruction, multiple data) architecture, in which each thread in a group of threads (e.g., a warp) is configured to process different data sets based on the same instruction set. All threads in a group of threads execute the same instructions. In another embodiment, the processing units implement a SIMT (single instruction, multiple thread) architecture, in which each thread in a group of threads is configured to process different data sets based on the same instruction set, but in which individual threads in the group of threads are allowed to diverge during execution. In one embodiment, a program counter, call stack, and execution state are maintained for each warp, thereby enabling concurrency between warps and serial execution within the warp when threads within the warp diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, thereby enabling equal concurrency between all threads, within a warp, and between warps. When execution state is maintained for each individual thread, threads executing the same instruction can converge and execute in parallel for maximum efficiency.
[0109] Cooperative Groups is a programming model for organizing groups of communicating threads that allows developers to express the granularity at which threads are communicating, enabling the expression of richer and more efficient decompositions of parallelism. The cooperative launch API supports synchronization between thread blocks to execute parallel algorithms. The conventional programming model provides a single simple construct for synchronizing cooperating threads: a barrier across all threads of a thread block (e.g., the syncthreads() function). However, programmers often want to define thread groups at a granularity smaller than the thread block granularity and synchronize within the defined group, enabling higher performance, design flexibility, and software reuse in the form of a collective group-wide function interface.
[0110] Cooperative Groups enable programmers to explicitly define thread groups at sub-block (e.g., as small as a single thread) and multi-block granularity and perform collective operations, such as synchronization, on threads in a cooperative group. The programming model supports clean composition across software boundaries so that libraries and utility functions can safely synchronize in their local environment without making assumptions about convergence. Cooperative Group primitives enable new patterns of cooperative parallelism, including producer-consumer parallelism, opportunistic parallelism, and global synchronization across the entire grid of thread blocks.
[0111] Each processing unit includes a large number (e.g., 128, etc.) of different processing cores (e.g., functional units), which can be fully pipelined, single-precision, double-precision, and / or mixed-precision, and include a floating-point arithmetic logic unit (FLU) and an integer arithmetic logic unit (ILU). In one embodiment, the FLU implements the IEEE 754-2008 standard for floating-point operations. In one embodiment, the core includes 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.
[0112] Tensor Cores are configured to perform matrix operations. Specifically, Tensor Cores are configured to perform deep learning matrix operations, such as convolution operations used for neural network training and inference. In one embodiment, each Tensor Core operates on a 4×4 matrix and performs a matrix multiplication and accumulation operation D=A×B+C, where A, B, C, and D are 4×4 matrices.
[0113] In one embodiment, the matrix multiplication inputs A and B are 16-bit floating point matrices, while the accumulation matrices C and D can be 16-bit floating point or 32-bit floating point matrices. The tensor cores operate on 16-bit floating point input data as well as 32-bit floating point accumulations. The 16-bit floating point multiplication requires 64 operations to produce a full-precision product, which is then accumulated using 32-bit floating point additions with other intermediate products of the 4×4×4 matrix multiplication. In practice, tensor cores are used to perform larger two-dimensional or higher-dimensional matrix operations built from these smaller elements. APIs (such as the CUDA 9 C++ API) expose specialized matrix load, matrix multiplication and accumulation, and matrix store operations to efficiently use tensor cores from CUDA-C++ programs. At the CUDA level, the warp-level interface assumes that the 16×16 size matrix spans all 32 threads of the warp.
[0114] Each processing unit may also include M special function units (SFUs) that perform special functions (e.g., attribute evaluation, reciprocal square root, etc.). In one embodiment, the SFUs may include a tree traversal unit that is configured to traverse a hierarchical tree data structure. In one embodiment, the SFUs may include a texture unit configured to perform texture map filtering operations. In one embodiment, the texture unit is configured to load a texture map (e.g., a 2D array of texture pixels) from memory 404 and sample the texture map to generate sampled texture values for use in a shader program executed by the processing unit. In one embodiment, the texture map is stored in a shared memory that may include or contain an L1 cache. The texture unit implements texture operations, such as filtering operations using mip maps (i.e., texture maps of different levels of detail). In one embodiment, each processing unit includes two texture units.
[0115] Each processing unit also includes N load-store units (LSUs), which implement load and store operations between the shared memory and the register file. Each processing unit includes an interconnect network that connects each core to the register file and connects the LSUs to the register file and the shared memory. In one embodiment, the interconnect network is a crossbar switch that can be configured to connect any core to any register in the register file and to connect the LSUs to memory locations in the register file and the shared memory.
[0116] Shared memory is an on-chip memory array that allows data storage and communication between processing units and between threads within a processing unit. In one embodiment, shared memory includes 128KB of storage capacity and is in the path from each processing unit to memory partition unit 480. Shared memory can be used for cache reads and writes. One or more of shared memory, L1 cache, L2 cache, and memory 404 is a backing store.
[0117] Combining data cache and shared memory functionality into a single memory block provides the best overall performance for both types of memory access. This capacity can be used by programs as a cache when the shared memory is not in use. For example, if the shared memory is configured to use half of its capacity, texture and load / store operations can use the remaining capacity. This integration within the shared memory enables the shared memory to function as a high-throughput pipeline for streaming data, providing both high-bandwidth and low-latency access to frequently reused data.
[0118] When configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics processing. Specifically, the fixed-function graphics processing units are bypassed, creating a simpler programming model. In a general-purpose parallel computing configuration, the work distribution unit 425 assigns and distributes thread blocks directly to processing elements within the GPC 450. The threads execute the same program, using unique thread IDs in the computation to ensure that each thread generates a unique result, using the processing element to execute the program and perform the computation, using shared memory to communicate between threads, and using the LSU to read and write to global memory through the shared memory and memory partition unit 480. When configured for general-purpose parallel computing, the processing element can also write commands that the scheduler unit 420 can use to start new work on the processing element.
[0119] The PPUs 430 may each include and / or be configured to perform the functions of one or more processing cores and / or components thereof, such as a tensor core (TC), a tensor processing unit (TPU), a pixel vision core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application-specific integrated circuit (ASIC), a floating point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI) or a peripheral component interconnect express (PCIe) element, etc.
[0120] The PPU 400 may be included in a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smartphone (e.g., wireless, handheld device), a personal digital assistant (PDA), a digital camera, a vehicle, a head-mounted display, a handheld electronic device, etc. In one embodiment, the PPU 400 is included on a single semiconductor substrate. In another embodiment, the PPU 400 is included on a system-on-chip (SoC) along with one or more other devices (such as an additional PPU 400, a memory 404, a reduced instruction set computer (RISC) CPU, a memory management unit (MMU), a digital-to-analog converter (DAC), etc.).
[0121] In one embodiment, PPU 400 may be included on a graphics card that includes one or more memory devices. The graphics card may be configured to interface with a PCIe slot on a desktop computer's motherboard. In another embodiment, PPU 400 may be an integrated graphics processing unit (iGPU) or parallel processor included in a chipset on the motherboard.
[0122] Exemplary Computing System
[0123] Systems with multiple GPUs and CPUs are being used across various industries as developers expose and exploit greater parallelism in applications such as artificial intelligence computing. High-performance GPU-accelerated systems with tens to thousands of computing nodes are deployed in data centers, research institutions, and supercomputers to solve larger problems. As the number of processing devices within high-performance systems increases, communication and data transmission mechanisms need to scale to support this increased bandwidth.
[0124] Figure 5A According to one embodiment, the Figure 4 Conceptual diagram of a processing system 500 implemented by a PPU 400. The exemplary system 500 may be configured to implement Figure 1B The method 130 shown in FIG. Figure 1D The method shown in 148 and / or Figure 2D The processing system 500 includes a CPU 530, a switch 510, and multiple PPUs 400 and corresponding memories 404. NVLink 410 provides a high-speed communication link between each PPU 400. Figure 5A 402 connections, but the number of connections connected to each PPU 400 and CPU 530 may vary. Switch 510 interfaces between interconnect 402 and CPU 530. PPU 400, memory 404, and NVLink 410 may be located on a single semiconductor platform to form a parallel processing module 525. In one embodiment, switch 510 supports two or more protocols that interface between various different connections and / or links.
[0125] In another embodiment (not shown), NVLink 410 provides one or more high-speed communication links between each PPU 400 and the CPU 530, and the switch 510 interfaces between the interconnect 402 and each PPU 400. The PPUs 400, memory 404, and interconnect 402 may be located on a single semiconductor platform to form a parallel processing module 525. In yet another embodiment (not shown), the interconnect 402 provides one or more communication links between each PPU 400 and the CPU 530, and the switch 510 interfaces between each PPU 400 using NVLink 410 to provide one or more high-speed communication links between the PPUs 400. In another embodiment (not shown), NVLink 410 provides one or more high-speed communication links between the PPUs 400 and the CPU 530 through the switch 510. In yet another embodiment (not shown), the interconnect 402 provides one or more communication links directly between each PPU 400. One or more NVLink 410 high-speed communication links may be implemented as a physical NVLink interconnect or as an on-chip or on-die interconnect using the same protocol as NVLink 410 .
[0126] In the context of this specification, a single semiconductor platform may refer to a unique, single semiconductor-based integrated circuit fabricated on a die or chip. It should be noted that the term single semiconductor platform may also refer to a multi-chip module with increased connectivity that emulates on-chip operation and is substantially improved by utilizing conventional bus implementations. Of course, various circuits or devices may also be placed separately or in various combinations of semiconductor platforms, depending on the needs of the user. Alternatively, the parallel processing module 525 may be implemented as a circuit board substrate, and each of the PPU 400 and / or memory 404 may be a packaged device. In one embodiment, the CPU 530, switch 510, and parallel processing module 525 are located on a single semiconductor platform.
[0127] In one embodiment, the signaling rate of each NVLink 410 is 20 to 25 Gbit / s, and each PPU 400 includes six NVLink 410 interfaces (e.g., Figure 5A As shown, each PPU 400 includes five NVLink 410 interfaces. Each NVLink 410 provides a data transfer rate of 25 Gbit / s in each direction, with six links providing 300 Gbit / s. When the CPU 530 also includes one or more NVLink 410 interfaces, the NVLink 410 can be used exclusively for Figure 5A PPU to PPU communication shown, or some combination of PPU to PPU and PPU to CPU.
[0128] In one embodiment, NVLink 410 allows direct load / store / atomic access from the CPU 530 to the memory 404 of each PPU 400. In one embodiment, NVLink 410 supports coherency operations, allowing data read from memory 404 to be stored in the cache hierarchy of the CPU 530, reducing cache access latency for the CPU 530. In one embodiment, NVLink 410 includes support for Address Translation Services (ATS), allowing the PPU 400 to directly access page tables within the CPU 530. One or more NVLinks 410 can also be configured to operate in a low-power mode.
[0129] Figure 5B An exemplary system 565 is shown in which various architectures and / or functionalities of various previous embodiments may be implemented. The exemplary system 565 may be configured to implement Figure 1B The method 130 shown in FIG. Figure 1D The method shown in 148 and / or Figure 2D Method 255 shown in .
[0130] As shown, a system 565 is provided that includes at least one central processing unit 530 connected to a communication bus 575. The communication bus 575 can directly or indirectly couple one or more of the following devices: main memory 540, network interface 535, one or more CPUs 530, one or more display devices 545, one or more input devices 560, switch 510, and parallel processing system 525. The communication bus 575 can be implemented using any suitable protocol and can represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The communication bus 575 can include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, HyperTransport, and / or another type of bus or link. In some embodiments, there are direct connections between the components. For example, the CPU 530 can be directly connected to the main memory 540. Further, the CPU 530 can be directly connected to the parallel processing system 525. In the case where there is a direct or point-to-point connection between components, the communication bus 575 may include a PCIe link for performing the connection. In these examples, the PCI bus need not be included in the system 565.
[0131] although Figure 5B The various boxes are shown as being connected to the line via the communication bus 575, but this is not intended to be restrictive and is only for clarity. For example, in some embodiments, a presentation component such as a display device 545 can be considered to be an I / O component such as an input device 560 (e.g., if the display is a touch screen). As another example, the CPU 530 and / or the parallel processing system 525 may include a memory (e.g., in addition to the parallel processing system 525, the CPU 530 and / or other components, the main memory 540 can also represent a storage device). In other words, the computing device is merely illustrative. No distinction is made between categories such as "workstation", "server", "laptop computer", "desktop computer", "tablet computer", "client device", "mobile device", "handheld device", "game console", "electronic control unit (ECU)", "virtual reality system" and / or other device or system types, as all are considered within the scope of the computing device.
[0132] System 565 also includes a main memory 540. Control logic (software) and data are stored in main memory 540, which can take the form of various computer-readable media. Computer-readable media can be any available media that can be accessed by system 565. Computer-readable media can include volatile and non-volatile media, as well as removable and non-removable media. By way of example, and not limitation, computer-readable media can include computer storage media and communication media.
[0133] Computer storage media may include both volatile and nonvolatile media and / or both removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, main memory 540 may store computer-readable instructions (e.g., representing programs and / or program elements, such as an operating system). Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by system 565. As used herein, computer storage media does not include signals themselves.
[0134] Computer storage media can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal (such as a carrier wave or other transport mechanism) and include any information delivery media. The term "modulated data signal" may refer to a signal that has one or more of its characteristics set or changed in a manner that encodes information into the signal. By way of example, and not limitation, computer storage media may include wired media (such as a wired network or a direct wired connection) and wireless media (such as acoustic, RF, infrared, and other wireless media). Combinations of any of the above should also be included within the scope of computer-readable media.
[0135] The computer program, when executed, enables the system 565 to perform different functions. The CPU 530 may be configured to execute at least some of the computer-readable instructions to control one or more components of the system 565 to perform one or more of the methods and / or processes described herein. The CPUs 530 may each include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing multiple software threads simultaneously. The CPU 530 may include any type of processor and, depending on the type of system 565 implemented, may include different types of processors (e.g., a processor with fewer cores for mobile devices and a processor with more cores for servers). For example, depending on the type of system 565, the processor may be an Advanced RISC Machine (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). In addition to one or more microprocessors or supplementary coprocessors (such as math coprocessors), the system 565 may also include one or more CPUs 530.
[0136] In addition to or in lieu of the CPU 530, the parallel processing module 525 may also be configured to execute at least some of the computer-readable instructions to control one or more components of the system 565 to perform one or more of the methods and / or processes described herein. The parallel processing module 525 may be used by the system 565 to render graphics (e.g., 3D graphics) or perform general-purpose computations. For example, the parallel processing module 525 may be used for general-purpose computations on a GPU (GPGPU). In embodiments, one or more CPUs 530 and / or parallel processing modules 525 may perform any combination of methods, processes, and / or portions thereof, either separately or in combination.
[0137] System 565 also includes one or more input devices 560, a parallel processing system 525, and one or more display devices 545. The one or more display devices 545 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The one or more display devices 545 may receive data from other components (e.g., parallel processing system 525, CPU 530, etc.) and output the data (e.g., as images, video, sound, etc.).
[0138] The network interface 535 can enable the system 565 to be logically coupled to other devices, including input devices 560, one or more display devices 545, and / or other components, some of which can be built into (e.g., integrated into) the system 565. Illustrative input devices 560 include microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dishes, scanners, printers, wireless devices, and the like. The input devices 560 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological input generated by the user. In some cases, the input can be transmitted to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, gesture recognition on and near the screen, air gestures, head and eye tracking, and touch recognition (as described in more detail below) associated with the display of the system 565. The system 565 can include a depth camera for gesture detection and recognition, such as a stereo camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof. Additionally, system 565 may include an accelerometer or gyroscope that enables detection of motion (e.g., as part of an inertial measurement unit (IMU)). In some examples, system 565 may use the output of the accelerometer or gyroscope to render immersive augmented or virtual reality.
[0139] Furthermore, the system 565 can be coupled to a network (e.g., a telecommunications network, a local area network (LAN), a wireless network, a wide area network (WAN) (such as the Internet), a peer-to-peer network, a cable network, etc.) for communication purposes via the network interface 535. The system 565 can be included in a distributed network and / or cloud computing environment.
[0140] The network interface 535 may include one or more receivers, transmitters, and / or transceivers that enable the system 565 to communicate with other computing devices via an electronic communication network (including wired and / or wireless communications). The network interface 535 may include components and functionality for enabling communication via any of a number of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., via Ethernet or InfiniBand communications), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.
[0141] The system 565 may also include an auxiliary storage device (not shown). The auxiliary storage device 610 includes, for example, a hard drive and / or a removable storage drive, representing a floppy disk drive, a tape drive, a compact disk drive, a digital versatile disk (DVD) drive, a recording device, or a universal serial bus (USB) flash memory. The removable storage drive reads from and / or writes to the removable storage unit in a known manner. The system 565 may also include a hardwired power supply, a battery power supply, or a combination thereof (not shown). The power supply can provide power to the system 565 to enable the components of the system 565 to operate.
[0142] Each of the aforementioned modules and / or devices may even be located on a single semiconductor platform to form system 565. Alternatively, different modules may also be located individually or in different combinations on the semiconductor platform according to the needs of the user.
[0143] Although various embodiments have been described above, it should be understood that these embodiments are presented by way of example only and not limitation. Thus, the breadth and scope of the preferred embodiment should not be limited by any of the above exemplary embodiments, but should be defined only in accordance with the appended claims and their equivalents.
[0144] Machine Learning
[0145] Deep neural networks (DNNs) developed on processors such as the PPU 400 are already being used in a variety of use cases: from self-driving cars to faster drug development, from automatic image captioning in online image databases to intelligent real-time language translation in video chat applications. Deep learning is a technology that models the neural learning process of the human brain, constantly learning, getting smarter, and delivering more accurate results faster over time. A child, initially taught by an adult to correctly identify and classify various shapes, eventually becomes able to recognize shapes without any tutoring. Similarly, deep learning or neural learning systems need to be trained in object recognition and classification in order to become smarter and more efficient at recognizing basic objects, occluded objects, and so on, while also assigning context to objects.
[0146] At the simplest level, neurons in the human brain examine the various inputs they receive, assign a level of importance to each of these inputs, and pass outputs to other neurons for processing. An artificial neuron, or perceptron, is the most basic model of a neural network. In one example, a perceptron can receive one or more inputs representing various features of the object it is being trained to recognize and classify, and each of these features is assigned a weight based on its importance in defining the object's shape.
[0147] Deep neural network (DNN) models consist of multiple layers of connected nodes (e.g., perceptrons, Boltzmann machines, radial basis functions, convolutional layers, etc.), which can be trained with large amounts of input data to solve complex problems quickly and accurately. In one example, the first layer of a DNN model breaks down an input image of a car into its components and looks for basic patterns (such as lines and angles). The second layer assembles the lines to find higher-level patterns, such as wheels, windshields, and mirrors. The next layer identifies the type of vehicle, and the final layers generate labels for the input image that identify the model of a specific car brand.
[0148] Once trained, a DNN can be deployed and used to recognize and classify objects or patterns in a process called inference. Examples of inference (the process by which a DNN extracts useful information from a given input) include recognizing handwritten digits on a check deposited in an ATM, identifying images of friends in photos, providing movie recommendations to over 50 million users, identifying and classifying different types of cars, pedestrians, and road hazards in self-driving cars, or translating human speech in real time.
[0149] During training, data flows through the DNN in a forward propagation phase until a prediction is produced, which indicates a label corresponding to the input. If the neural network does not correctly label an input, the error between the correct label and the predicted label is analyzed, and the weights are adjusted for each feature during a backpropagation phase until the DNN correctly labels that input and other inputs in the training dataset. Training complex neural networks requires a large amount of parallel computing performance, including floating-point multiplications and additions supported by the PPU 400. Inference, which is less computationally intensive than training, is a latency-sensitive process in which a trained neural network is applied to new inputs it has not seen before to classify images, translate speech, and generally reason about new information.
[0150] Neural networks rely heavily on matrix math operations, and complex, multi-layer networks require significant floating-point performance and bandwidth for efficiency and speed. With thousands of processing cores optimized for matrix math operations and delivering tens to hundreds of TFLOPS of performance, the PPU 400 is a computing platform capable of delivering the performance required for deep neural network-based artificial intelligence and machine learning applications.
[0151] In addition, images generated using one or more of the techniques disclosed herein can be used to train, test, or certify DNNs for recognizing objects and environments in the real world. Such images can include scenes of roads, factories, buildings, urban environments, rural environments, people, animals, and any other physical objects or real-world environments. Such images can be used to train, test, or certify DNNs employed in machines or robots to manipulate, process, or modify physical objects in the real world. Furthermore, such images can be used to train, test, or certify DNNs employed in autonomous vehicles to navigate and move vehicles in the real world. In addition, images generated using one or more of the techniques disclosed herein can be used to convey information to users of such machines, robots, and vehicles.
[0152] Graphics processing pipeline
[0153] In one embodiment, the PPU 400 includes a graphics processing unit (GPU). The PPU 400 is configured to receive commands specifying a shader for processing graphics data. Graphics data can be defined as a set of primitives, such as points, lines, triangles, quadrilaterals, triangle strips, etc. Typically, a primitive includes data specifying a plurality of vertices of the primitive (e.g., in a model space coordinate system) and attributes associated with each vertex of the primitive. The PPU 400 can be configured to process the primitives to generate a frame buffer (e.g., pixel data for each of the pixels of a display).
[0154] An application writes model data for a scene (e.g., a collection of vertices and attributes) into memory (such as system memory or memory 404). The model data defines each of the objects that may be visible on the display. The application then makes an API call to the driver kernel, requesting the model data to be rendered and displayed. The driver kernel reads the model data and writes commands to one or more streams to perform operations to process the model data. These commands may reference different shading programs to be implemented on the processing units of PPU 400, including one or more of vertex shading, hull shading, domain shading, geometry shading, and pixel shading. For example, one or more of the processing units may be configured to execute a vertex shading program that processes multiple vertices defined by the model data. In one embodiment, different processing units may be configured to execute different shading programs simultaneously. For example, a first subset of processing units may be configured to execute a vertex shading program, while a second subset of processing units may be configured to execute a pixel shading program. The first subset of processing units processes the vertex data to generate processed vertex data and writes the processed vertex data to L2 cache 460 and / or memory 404. After the processed vertex data is rasterized (e.g., converted from three-dimensional data to two-dimensional data in screen space) to generate fragment data, a second subset of processing units performs pixel shading to generate processed fragment data, which is then blended with other processed fragment data and written to a frame buffer in memory 404. The vertex shading program and the pixel shading program can be executed simultaneously, processing different data from the same scene in a pipelined manner, until all model data for the scene has been rendered to the frame buffer. The contents of the frame buffer are then transmitted to the display controller for display on the display device.
[0155] Figure 6A According to one embodiment, Figure 4 4. A conceptual diagram of a graphics processing pipeline 600 implemented by a PPU 400 of FIG. The graphics processing pipeline 600 is an abstract flow diagram of the processing steps implemented to generate a 2D computer-generated image from 3D geometric data. As is well known, pipeline architectures can perform long-latency operations more efficiently by dividing the operations into multiple stages, where the output of each stage is coupled to the input of the next consecutive stage. Thus, the graphics processing pipeline 600 receives input data 601 that is passed from one stage of the graphics processing pipeline 600 to the next stage to generate output data 602. In one embodiment, the graphics processing pipeline 600 may represent a graphics processing pipeline composed of API-defined graphics processing pipeline. Alternatively, graphics processing pipeline 600 can be implemented within the functional and architectural context of the previous figures and / or one or more of any subsequent figures.
[0156] like Figure 6AAs shown, graphics processing pipeline 600 includes a pipeline architecture comprising multiple stages. These stages include, but are not limited to, a data assembly stage 610, a vertex shading stage 620, a primitive assembly stage 630, a geometry shading stage 640, a viewport scale, cull, and clip (VSCC) stage 650, a rasterization stage 660, a fragment shading stage 670, and a raster operation stage 680. In one embodiment, input data 601 includes commands that configure a processing unit to implement the stages of graphics processing pipeline 600 and configure geometric primitives (e.g., points, lines, triangles, quads, triangle strips, or fans, etc.) to be processed by these stages. Output data 602 may include pixel data (i.e., color data), which is copied to a frame buffer or other type of surface data structure in memory.
[0157] The data assembly stage 610 receives input data 601, which specifies vertex data for high-level surfaces, primitives, etc. The data assembly stage 610 collects the vertex data in temporary storage or queues, such as by receiving a command from the host processor that includes a pointer to a buffer in memory and reading the vertex data from the buffer. The vertex data is then passed to the vertex shading stage 620 for processing.
[0158] The vertex shading stage 620 processes vertex data by executing a set of operations (e.g., a vertex shader or program) on each vertex at a time. A vertex may be specified, for example, as a 4-coordinate vector (e.g., ) associated with one or more vertex attributes (e.g., color, texture coordinates, surface normal, etc.).<x,y,z,w> ). The vertex shading stage 620 can manipulate various vertex attributes, such as position, color, texture coordinates, etc. In other words, the vertex shading stage 620 performs operations on the vertex coordinates or other vertex attributes associated with the vertex. These operations typically include lighting operations (e.g., modifying the color attribute of a vertex) and transformation operations (e.g., modifying the coordinate space of a vertex). For example, a vertex can be specified using coordinates in an object coordinate space, which is transformed by multiplying the coordinates by a matrix that converts the coordinates from the object coordinate space to world space or normalized-device-coordinate (NCD) space. The vertex shading stage 620 generates transformed vertex data that is passed to the primitive assembly stage 630.
[0159] The primitive assembly stage 630 collects the vertices output by the vertex shading stage 620 and groups the vertices into geometric primitives for processing by the geometry shading stage 640. For example, the primitive assembly stage 630 can be configured to group every three consecutive vertices into geometric primitives (e.g., triangles) for transmission to the geometry shading stage 640. In some embodiments, particular vertices can be reused for consecutive geometric primitives (e.g., two consecutive triangles in a triangle strip can share two vertices). The primitive assembly stage 630 transmits the geometric primitives (e.g., a collection of associated vertices) to the geometry shading stage 640.
[0160] The geometry shading stage 640 processes geometric primitives by executing a set of operations (e.g., geometry shaders or programs) on the geometric primitives. A tessellation operation can generate one or more geometric primitives from each geometric primitive. In other words, the geometry shading stage 640 can subdivide each geometric primitive into a finer mesh of two or more geometric primitives for processing by the rest of the graphics processing pipeline 600. The geometry shading stage 640 passes the geometric primitives to the viewport SCC stage 650.
[0161] In one embodiment, the graphics processing pipeline 600 may operate within a streaming multiprocessor and vertex shading stage 620, primitive assembly stage 630, geometry shading stage 640, fragment shading stage 670, and / or hardware / software associated therewith, and may perform processing operations sequentially. Once the sequential processing operations are completed, in one embodiment, the viewport SCC stage 650 may utilize the data. In one embodiment, primitive data processed by one or more stages in the graphics processing pipeline 600 may be written to a cache (e.g., an L1 cache, a vertex cache, etc.). In this case, in one embodiment, the viewport SCC stage 650 may access the data in the cache. In one embodiment, the viewport SCC stage 650 and the rasterization stage 660 are implemented as fixed function circuits.
[0162] The viewport SCC stage 650 performs viewport scaling, culling, and clipping of geometric primitives. Each surface being rendered is associated with an abstract camera position. The camera position represents the position of the viewer viewing the scene and defines a viewing cone that surrounds the objects of the scene. The viewing cone can include a viewing plane, a back plane, and four clipping planes. Any geometric primitives that are completely outside the viewing cone can be culled (e.g., discarded) because they will not contribute to the final rendered scene. Any geometric primitives that are partially inside the viewing cone and partially outside the viewing cone can be clipped (e.g., converted to new geometric primitives that are enclosed within the viewing cone). In addition, each geometric primitive can be scaled based on the depth of the viewing cone. All potentially visible geometric primitives are then transferred to the rasterization stage 660.
[0163] The rasterization stage 660 converts 3D geometric primitives into 2D fragments (e.g., capable of being used for display, etc.). The rasterization stage 660 can be configured to use the vertices of the geometric primitives to set a set of plane equations from which various attributes can be interpolated. The rasterization stage 660 can also calculate a coverage mask for multiple pixels, which indicates whether one or more sample positions of the pixel intercept the geometric primitive. In one embodiment, a z test can also be performed to determine whether the geometric primitive is occluded by other geometric primitives that have already been rasterized. The rasterization stage 660 generates fragment data (e.g., interpolated vertex attributes associated with a specific sample position for each covered pixel), which is passed to the fragment shading stage 670.
[0164] The fragment shading stage 670 processes the fragment data by executing a set of operations (e.g., a fragment shader or program) on each of the fragments. The fragment shading stage 670 can generate pixel data (e.g., color values) for the fragment, such as by performing lighting operations or sampling a texture map using the fragment's interpolated texture coordinates. The fragment shading stage 670 generates pixel data, which is sent to the raster operations stage 680.
[0165] The raster operations stage 680 may perform various operations on the pixel data, such as performing alpha tests, stencil tests, and blending the pixel data with other pixel data corresponding to other fragments associated with the pixel. When the raster operations stage 680 has completed processing the pixel data (e.g., output data 602), the pixel data may be written to a render object, such as a frame buffer, a color buffer, etc.
[0166] It should be appreciated that one or more additional stages may be included in graphics processing pipeline 600 in addition to or in place of one or more of the above-described stages. Various implementations of the abstract graphics processing pipeline may implement different stages. Furthermore, in some embodiments, one or more of the above-described stages may be excluded from the graphics processing pipeline (such as geometry shading stage 640). Other types of graphics processing pipelines are contemplated within the scope of the present disclosure. Furthermore, any stage of graphics processing pipeline 600 may be implemented by one or more dedicated hardware units within a graphics processor (such as PPU 400). Other stages of graphics processing pipeline 600 may be implemented by programmable hardware units (such as processing units within PPU 400).
[0167] The graphics processing pipeline 600 can be implemented via an application program executed by a host processor (such as a CPU). In one embodiment, a device driver can implement an application programming interface (API) that defines various functions that can be utilized by the application program to generate graphics data for display. A device driver is a software program that includes multiple instructions that control the operation of the PPU 400. The API provides an abstraction for programmers, allowing them to utilize specialized graphics hardware (such as the PPU 400) to generate graphics data without requiring them to utilize the specific instruction set of the PPU 400. An application program can include API calls that are routed to the device driver of the PPU 400. The device driver interprets the API calls and performs various operations in response to the API calls. In some cases, the device driver can perform operations by executing instructions on the CPU. In other cases, the device driver can perform operations at least in part by initiating operations on the PPU 400 using an input / output interface between the CPU and the PPU 400. In one embodiment, the device driver is configured to implement the graphics processing pipeline 600 using the hardware of the PPU 400.
[0168] Various programs may be executed within the PPU 400 to implement the various stages of the graphics processing pipeline 600. For example, a device driver may launch a kernel on the PPU 400 to execute the vertex shading stage 620 on one processing unit (or multiple processing units). The device driver (or the initial kernel executed by the PPU 400) may also launch other kernels on the PPU 400 to execute other stages of the graphics processing pipeline 600, such as the geometry shading stage 640 and the fragment shading stage 670. In addition, some of the stages of the graphics processing pipeline 600 may be implemented on fixed unit hardware, such as a rasterizer or data assembler implemented within the PPU 400. It should be appreciated that the results from one kernel may be processed by one or more intermediate fixed-function hardware units before being processed by subsequent kernels on the processing unit.
[0169] The image generated by applying one or more of the techniques disclosed herein can be displayed on a monitor or other display device. In some embodiments, the display device may be directly coupled to a system or processor that generates or renders the image. In other embodiments, the display device may be indirectly coupled to the system or processor, for example, via a network. Examples of such networks include the Internet, mobile telecommunications networks, WIFI networks, and any other wired and / or wireless networking systems. When the display device is indirectly coupled, the image generated by the system or processor can be streamed to the display device via the network. Such streaming allows, for example, a video game or other application that renders an image to be executed on a server, in a data center, or in a cloud-based computing environment, and the rendered image will be sent and displayed on one or more user devices (such as computers, video game consoles, smart phones, other mobile devices, etc.) that are physically separated from the server or data center. Therefore, the technology disclosed herein can be applied to enhanced streaming images and services that enhance streaming images, such as NVIDIA GeForce Now (GFN), Google Stadia, etc.
[0170] Sample game streaming system
[0171] Figure 6B is an example system diagram of a game streaming system 605 according to some embodiments of the present disclosure. Figure 6B Includes one or more game servers 603 (which may include Figure 5A The example processing system 500 and / or Figure 5B ), one or more client devices 604 (which may include components, features, and / or functionality similar to the exemplary system 565 of Figure 5A The example processing system 500 and / or Figure 5B 5 ) and one or more networks 606 (which may be similar to one or more networks described herein). In some embodiments of the present disclosure, system 605 may be implemented.
[0172] In system 605, for a game session, one or more client devices 604 may receive input data only in response to input to one or more input devices, send the input data to one or more game servers 603, receive encoded display data from the one or more game servers 603, and display the display data on a display 624. In this way, more computationally intensive calculations and processing are offloaded to one or more game servers 603 (e.g., rendering—particularly ray or path tracing—for one or more GPUs of one or more game servers 603 to perform graphical output for the game session). In other words, the game session is streamed from one or more game servers 603 to one or more client devices 604, thereby reducing the graphics processing and rendering requirements of one or more client devices 604.
[0173] For example, with respect to instantiation of a game session, client device 604 may display a frame of the game session on display 624 based on display data received from one or more game servers 603. Client device 604 may receive input from one of one or more input devices and generate input data in response. Client device 604 may send the input data to one or more game servers 603 via communication interface 621 and over one or more networks 606 (e.g., the Internet), and one or more game servers 603 may receive the input data via communication interface 618. The CPU may receive the input data, process the input data, and send the data to the GPU, causing the GPU to generate a rendering of the game session. For example, the input data may represent movement of a user's character in the game, firing a weapon, reloading, passing a ball, turning a vehicle, and the like. Rendering component 612 may render the game session (e.g., representing the results of the input data), and rendering capture component 614 may capture the rendering of the game session as display data (e.g., capturing image data of a rendered frame of the game session). The rendering of the game session may include ray or path tracing lighting and / or shadow effects calculated using one or more parallel processing units (such as GPUs), which may further employ one or more dedicated hardware accelerators or processing cores to perform ray or path tracing techniques on one or more game servers 603. The encoder 616 may then encode the display data to generate encoded display data, and the encoded display data may be sent to the client device 604 via the network 606 via the communication interface 618. The client device 604 may receive the encoded display data via the communication interface 621, and the decoder 622 may decode the encoded display data to generate display data. The client device 604 may then display the display data via the display 624.
[0174] Sample network environment
[0175] A network environment suitable for implementing embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be configured to: Figure 5A The processing system 500 and / or Figure 5B The exemplary system 565 may be implemented on one or more instances of the exemplary system 565 of the processing system 500—for example, each device may include similar components, features, and / or functionality of the exemplary system 565 and / or the processing system 500.
[0176] The components of the network environment can communicate with each other via one or more networks, which can be wired, wireless, or both. The network can include multiple networks or one of multiple networks. For example, the network can include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks (such as the Internet and / or the Public Switched Telephone Network (PSTN)), and / or one or more private networks. In the case where the network includes a wireless telecommunications network, components such as base stations, communication towers, or even access points (and other components) can provide wireless connections.
[0177] Compatible network environments may include one or more peer-to-peer network environments (in which case the server may not be included in the network environment) and one or more client-server network environments (in which case one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein with respect to one or more servers may be implemented on any number of client devices.
[0178] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more servers, which may include one or more core network servers and / or edge servers. The framework layer may include software supporting the software layer and / or a framework for one or more applications at the application layer. The software or applications may include network-based service software or applications, respectively. In an embodiment, one or more client devices may use network-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open source software network application framework, such as those frameworks that can use a distributed file system for large-scale data processing (e.g., "big data").
[0179] A cloud-based network environment can provide cloud computing and / or cloud storage that performs any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these different functions can be distributed across multiple locations from a central or core server (e.g., one or more data centers that can be distributed across a state, region, country, globe, etc.). If the connection to the user (e.g., client device) is relatively close to an edge server, the core server can assign at least a portion of the functionality to the edge server. A cloud-based network environment can be private (e.g., limited to a single organization), public (e.g., available to many organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0180] Client devices may include Figure 5A The example processing system 500 and / or Figure 5B At least some of the components, features, and functionality of the exemplary system 565 of FIG. By way of example and not limitation, the client device may be embodied as a personal computer (PC), a laptop, a mobile device, a smartphone, a tablet computer, a smartwatch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality head mounted display, a global positioning system (GPS) or device, a video player, a camera, a surveillance device or system, a vehicle, a boat, a spacecraft, a virtual machine, a drone, a robot, a handheld communication device, a hospital device, a gaming device or system, an entertainment system, a vehicle computer system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, an edge device, any combination of the depicted devices, or any other suitable device.
[0181] It should be noted that the techniques described herein can be embodied in executable instructions stored in a computer-readable medium for use by or in conjunction with a processor-based instruction execution machine, system, device, or apparatus. Those skilled in the art will appreciate that for some embodiments, different types of computer-readable media may be included for storing data. As used herein, "computer-readable medium" includes one or more of any suitable media for storing executable instructions of a computer program, such that an instruction execution machine, system, apparatus, or device can read (or obtain) instructions from the computer-readable medium and execute instructions for implementing the described embodiments. Suitable storage formats include one or more of electronic formats, magnetic formats, optical formats, and electromagnetic formats. A non-exhaustive list of conventional exemplary computer-readable media includes: portable computer disks; random access memory (RAM); read-only memory (ROM); erasable programmable read-only memory (EPROM); flash memory devices; and optical storage devices, including portable compact disks (CDs), portable digital video disks (DVDs), and the like.
[0182] It should be understood that the arrangement of the components shown in the drawings is for illustrative purposes and other arrangements are possible. For example, one or more of the elements described herein may be implemented as an electronic hardware assembly in whole or in part. Other elements may be implemented with software, hardware, or a combination of software and hardware. In addition, some or all of these other elements may be combined, some elements may be omitted completely, and additional components may be added while still implementing the functions described herein. Thus, the subject matter described herein may be embodied in many different variations, and all such variations are contemplated to be within the scope of the claims.
[0183] For ease of understanding the subject matter described herein, many aspects are described in terms of action sequences. Those skilled in the art will recognize that different actions can be performed by dedicated circuits or electronic circuits, by program instructions executed by one or more processors, or by a combination of the two. The description of any action sequence herein is not intended to imply that the described particular order for executing the sequence must be followed. Unless otherwise indicated herein or the context clearly contradicts, all methods described herein can be performed in any suitable order.
[0184] The use of the terms "a" and "the" and similar references in the context of describing the subject matter (particularly in the context of the following claims) should be interpreted to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by the context. The use of the term "at least one" followed by a list of one or more items (e.g., "at least one of A and B") should be interpreted to mean one item selected from the listed items (A or B) or any combination of two or more of the listed items (A and B), unless otherwise indicated herein or clearly contradicted by the context. In addition, the foregoing description is for illustrative purposes only and not for limiting purposes, as the scope of protection sought is defined by the claims set forth herein and any equivalents thereof. The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended only to better illustrate the subject matter and does not impose limitations on the scope of the subject matter, unless otherwise stated. The use of the term "based on" and other similar phrases indicating conditions that cause a result in the claims and written description is not intended to exclude any other conditions that cause the result. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention as claimed.
Claims
1. A computer-implemented method for constructing a three-dimensional (3D) representation of an object, comprising: receiving, by the neural network model, a video comprising an image of an object captured from a camera pose; predicting, by the neural network model, a 3D shape representation of the object for a first one of the images based on the learned set of shape basis; Predicting the texture flow of the first image by the neural network model; mapping pixels from the first image to a texture space according to the texture stream to produce a texture image, wherein transferring the texture image to the 3D shape representation constructs a 3D object corresponding to the object in the first image; propagating the part map across the object in a plurality of images to produce a propagated part map; mapping the propagated part map into the texture space according to respective texture flows predicted for the plurality of images to generate a part map in the texture space; as well as The part maps are aggregated to generate a video-level part map.
2. The computer-implemented method of claim 1 , further comprising: predicting a non-rigid motion deformation of the 3D shape representation of the first image; as well as The non-rigid motion deformation is applied to the identity shape to produce the 3D shape representation.
3. The computer-implemented method of claim 1 , further comprising: predicting a non-rigid motion deformation of the image; applying the non-rigid motion deformation to an identity shape predicted for the image to produce a 3D shape representation of the object; as well as A loss function is evaluated based on the rotational difference between the identity shapes and the difference between the 3D shape representations.
4. The computer-implemented method of claim 3 , further comprising: Parameters of the neural network model are updated based on the loss function to reduce discontinuities in the 3D shape representation.
5. The computer-implemented method of claim 1 , wherein the identity shape is computed as the sum of component shapes included in the set of learned shape bases, and each component shape is scaled by a coefficient generated by the neural network model. The computer-implemented method of claim 1 , wherein the 3D shape representation is a mesh of vertices defining a face. The computer-implemented method of claim 1 , wherein the images in the video are unlabeled.
8. The computer-implemented method of claim 1 , wherein the neural network model is further configured to predict the camera pose.
9. The computer-implemented method of claim 1 , further comprising: transferring the texture image onto the 3D shape representation to construct the 3D object; projecting the 3D object according to the camera pose to produce a rendered image; as well as Parameters of the neural network model are updated to reduce a difference between the rendered image and the first image.
10. The computer-implemented method of claim 1 , further comprising: transferring the texture image predicted for the first image to a second 3D shape representation predicted for a second one of the images to produce a first 3D object; projecting the first 3D object according to a first camera pose associated with the first image to produce a first projected 3D object; transferring a second texture image predicted for the second image to the 3D shape representation predicted for the first image to generate a second 3D object; projecting the second 3D object according to a second camera pose associated with the second image to produce a second projected 3D object; as well as Parameters of the neural network model are updated to promote consistency between the first projected 3D object and the second projected 3D object.
11. The computer-implemented method of claim 1 , further comprising: applying a first non-rigid motion deformation predicted for the first image to a first identity shape predicted for a second one of the images to produce a first 3D shape representation; projecting the first 3D shape representation according to a first camera pose associated with the first image to produce a first projected 3D object; applying a second non-rigid motion deformation predicted for the second image to a second identity shape predicted for the first image to produce a second 3D shape representation; projecting the second 3D shape representation according to a second camera pose associated with the second image to produce a second projected 3D object; and Parameters of the neural network model are updated to promote consistency between the first projected 3D object and the second projected 3D object.
12. The computer-implemented method of claim 1 , further comprising: rendering 3D shape representations predicted for the plurality of images according to associated camera poses, wherein the video-level part map is transferred to each of the 3D shape representations to produce a rendered image; as well as Parameters of the neural network model are updated to promote consistency between the rendered image and the propagated part map.
13. The computer-implemented method of claim 1 , propagating the part map comprising: A part pattern is applied to the object in a center image at the center of the plurality of images, and the part pattern is propagated from the center image to images before and after the center image in the video.
14. The computer-implemented method of claim 1 , wherein the images are each annotated, and the method further comprises: mapping the annotations into the texture space according to corresponding texture flows predicted for the image to generate an annotation map in the texture space; as well as The annotation maps are aggregated to generate a canonical annotation map for the video.
15. The computer-implemented method of claim 14, further comprising: transferring the canonical annotation map to a predicted 3D shape representation for the image to produce an annotated 3D shape representation; projecting the 3D shape representation of the annotation according to the associated camera pose to produce a projected annotation for the image; as well as Parameters of the neural network model are updated to promote consistency between the projected annotation and the annotation. The computer-implemented method of claim 14 , wherein the annotations are semantic keypoints.
17. The computer-implemented method of claim 1, wherein the object is a non-rigid animal.
18. The computer-implemented method of claim 1 , wherein the steps of receiving, predicting the 3D shape representation, predicting the texture stream, and mapping are performed in a data center or on a server in a cloud-based computing environment to construct the 3D object, and the 3D object is streamed to a user device.
19. The computer-implemented method of claim 1 , wherein the steps of receiving, predicting the 3D shape representation, predicting the texture flow, and mapping are performed to generate the 3D object for training, testing, or demonstrating a second neural network employed in a machine, robot, or autonomous vehicle.
20. A system comprising: A neural network model is configured to construct a three-dimensional 3D representation of an object through the following steps: receiving a video including an image of an object captured from a camera pose; predicting a 3D shape representation of the object for a first one of the images based on the learned set of shape basis; predicting a texture flow of the first image; mapping pixels from the first image to a texture space according to the texture stream to produce a texture image, wherein transferring the texture image to the 3D shape representation constructs a 3D object corresponding to the object in the first image; propagating the part map across the object in a plurality of images to produce a propagated part map; mapping the propagated part map into the texture space according to respective texture flows predicted for the plurality of images to generate a part map in the texture space; as well as The part maps are aggregated to generate a video-level part map.
21. A non-transitory computer readable medium storing computer instructions for constructing a three-dimensional (3D) representation of an object, the computer instructions, when executed by one or more processors, causing the one or more processors to perform the following steps: receiving, by the neural network model, a video comprising an image of an object captured from a camera pose; predicting, by the neural network model, a 3D shape representation of the object for a first one of the images based on the learned set of shape basis; Predicting the texture flow of the first image by the neural network model; mapping pixels from the first image to a texture space according to the texture stream to produce a texture image, wherein transferring the texture image to the 3D shape representation constructs a 3D object corresponding to the object in the first image; propagating a part map by the neural network model across the object in a plurality of images to produce a propagated part map; mapping, by the neural network model, the propagated part map into the texture space according to corresponding texture flows predicted for the plurality of images to generate a part map in the texture space; as well as The part maps are aggregated by the neural network model to generate a video-level part map.
Citation Information
Patent Citations
Image processing method, apparatus, and storage medium
US20190108646A1