System and method for generating splash-based differentiable two-dimensional renderings

Through the differentiable rendering method based on splash, the accuracy and efficiency of derivative calculation of three-dimensional polygon mesh at the occlusion boundary is solved, and efficient and accurate differentiable rendering is achieved.

CN116018617BActive Publication Date: 2025-05-13GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080104438.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-31
Publication Date
2025-05-13
Estimated Expiration
2040-08-31

AI Technical Summary

Technical Problem

In the prior art, when calculating the derivative of a three-dimensional polygon mesh, especially at the occlusion boundary, it is difficult to achieve accuracy and efficiency, making calculations expensive and difficult to achieve.

Method used

Using a differentiable rendering method based on splattering, a two-dimensional raster is generated by rasterizing a three-dimensional grid, a splatter is constructed using the coordinates of the pixels and the associated shading or texture data, and the updated color value of the pixel is determined based on the weighting of the subset of the splattering, and a two-dimensional differentiable rendering of the three-dimensional grid is generated.

Benefits of technology

Generating smooth derivatives at occlusion boundaries is achieved, significantly reducing the computing resources required to generate differentiable renderings, and improving computing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116018617B_ABST
    Figure CN116018617B_ABST
Patent Text Reader

Abstract

The systems and methods of the present disclosure relate to a method that may include obtaining a 3D mesh including polygons and texture / shading data. The method may include rasterizing the 3D mesh to obtain a 2D raster including pixels and coordinates associated with a subset of pixels, respectively. The method may include determining an initial color value for a subset of pixels based on the coordinates of the pixels and the associated shading / texture data. The method may include constructing a splash at the coordinates of the corresponding pixels. The method may include determining an updated color value for the corresponding pixel based on a weighting of the subset of splashes to generate a 2D rendering of the 3D mesh based on the coordinates of the pixels and the splashes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to differentiable rendering. More specifically, the present disclosure relates to splash-based differentiable rendering for accurate derivatives at occlusion boundaries. Background Art

[0002] In many industries, 3D polygonal meshes are the primary shape representation used for 3D modeling (e.g., graphics, computer vision, machine learning, etc.). As a result, there is growing interest in the ability to compute accurate derivatives of these meshes with respect to underlying scene parameters. However, a key difficulty in computing these derivatives is the discontinuity introduced by occlusion boundaries. Modern solutions to this problem (e.g., mesh processing, custom derivative calculations, etc.) have generally been found to be computationally expensive and / or difficult to implement. Summary of the invention

[0003] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.

[0004] An example aspect of the present disclosure relates to a computer-implemented method for efficiently generating a differentiable two-dimensional rendering of a three-dimensional model. The method may include obtaining, by a computing system including one or more computing devices, a three-dimensional mesh, the three-dimensional mesh including a plurality of polygons and at least one of associated texture data or associated shading data. The method may include rasterizing, by the computing system, the three-dimensional mesh to obtain a two-dimensional raster of the three-dimensional mesh, wherein the two-dimensional raster includes a plurality of pixels and a plurality of coordinates correspondingly associated with at least a subset of pixels of the plurality of pixels, wherein the coordinates of each pixel in the subset of pixels describe the position of the corresponding pixel relative to the vertex of the corresponding polygon in which the pixel is located in the plurality of polygons. The method may include determining, by the computing system, a corresponding initial color value for each pixel in the subset of pixels based at least in part on the coordinates of the pixels and at least one of the associated shading data or the associated texture data. The method may include constructing, by the computing system, a splash at the coordinates of the corresponding pixel for each pixel in the subset of pixels. The method may include determining, by a computing system, an updated color value for each pixel in a subset of pixels based on a weighting of each splash in a subset of splashes to generate a two-dimensional differentiable rendering of a three-dimensional grid, wherein, for each splash in the subset of splashes, the weighting of the corresponding splash is based at least in part on coordinates of the corresponding pixel and coordinates of the corresponding splash.

[0005] Another example aspect of the present disclosure relates to a computing system for efficiently generating a differentiable two-dimensional rendering of a three-dimensional model. The computing system may include one or more processors. The computing system may include one or more tangible non-transitory computer-readable media storing computer-readable instructions that, when executed by one or more processors, cause one or more processors to perform operations. The operations may include obtaining a three-dimensional mesh comprising a plurality of polygons and at least one of associated texture data or associated shading data. The operations may include rasterizing the three-dimensional mesh to obtain a two-dimensional raster of the three-dimensional mesh, wherein the two-dimensional raster comprises a plurality of pixels and a plurality of coordinates correspondingly associated with at least a subset of pixels of the plurality of pixels, wherein the coordinates of each pixel in the subset of pixels describe the position of the corresponding pixel relative to the vertex of the corresponding polygon in which the pixel is located in the plurality of polygons. The operations may include determining the corresponding initial color value of each pixel in the subset of pixels based at least in part on the coordinates of the pixels and at least one of the associated shading data or the associated texture data. The operations may include: for each pixel in the subset of pixels, constructing a splash at the coordinates of the corresponding pixel. The operation may include: for each pixel in the subset of pixels, determining an updated color value of the corresponding pixel based on a weight of each splash in the subset of splashes to generate a two-dimensional differentiable rendering of the three-dimensional grid, wherein, for each splash in the subset of splashes, the weight of the corresponding splash is at least partially based on the coordinates of the corresponding pixel and the coordinates of the corresponding splash.

[0006] Another example embodiment of the present disclosure relates to one or more tangible non-transitory computer-readable media storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations. The operations may include obtaining a three-dimensional mesh comprising a plurality of polygons and at least one of associated texture data or associated shading data. The operations may include rasterizing the three-dimensional mesh to obtain a two-dimensional raster of the three-dimensional mesh, wherein the two-dimensional raster comprises a plurality of pixels and a plurality of coordinates correspondingly associated with at least a subset of the plurality of pixels, wherein the coordinates of each pixel in the subset of pixels describe the position of the corresponding pixel relative to the vertex of the corresponding polygon in which the pixel is located in the plurality of polygons. The operations may include determining a corresponding initial color value for each pixel in the subset of pixels based at least in part on the coordinates of the pixel and at least one of the associated shading data or the associated texture data. The operations may include, for each pixel in the subset of pixels, constructing a splash at the coordinates of the corresponding pixel. The operation may include: for each pixel in the subset of pixels, determining an updated color value of the corresponding pixel based on a weight of each splash in the subset of splashes to generate a two-dimensional differentiable rendering of the three-dimensional grid, wherein, for each splash in the subset of splashes, the weight of the corresponding splash is at least partially based on the coordinates of the corresponding pixel and the coordinates of the corresponding splash.

[0007] Other aspects of the disclosure relate to various systems, apparatuses, non-transitory computer-readable media, user interfaces, and electronic devices.

[0008] These and other features, aspects and advantages of various embodiments of the present disclosure will be better understood with reference to the following description and appended claims.The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] A detailed discussion of embodiments for those of ordinary skill in the art is set forth in the specification with reference to the accompanying drawings, in which:

[0010] Figure 1A A block diagram depicting an example computing system that performs two-dimensional differentiable rendering according to an example embodiment of the present disclosure.

[0011] Figure 1B A block diagram depicting an example computing device that performs training of a machine learning model based on two-dimensional differentiable rendering according to an example embodiment of the present disclosure.

[0012] Figure 1C A block diagram depicting an example computing device that performs two-dimensional differentiable rendering according to an example embodiment of the present disclosure.

[0013] Figure 2 An illustrative example of splash construction for a one-dimensional line segment is depicted according to an example embodiment of the present disclosure.

[0014] Figure 3A A data flow diagram depicting an example method for generating a two-dimensional differentiable rendering according to an example embodiment of the present disclosure.

[0015] Figure 3B A data flow diagram depicting an example method for generating derivatives from a two-dimensional differentiable rendering to train a machine learning model according to an example embodiment of the present disclosure.

[0016] Figure 4 A flow chart depicting an example method for generating a two-dimensional differentiable rendering according to an example embodiment of the present disclosure.

[0017] Figure 5 A flowchart depicting an example method for generating derivatives at occlusion boundaries for two-dimensional differentiable rendering to train a machine learning model according to an example embodiment of the present disclosure.

[0018] Reference numerals repeated across multiple figures are intended to identify the same features in different embodiments. DETAILED DESCRIPTION

[0019] Overview

[0020] In general, the present disclosure relates to differentiable rendering. More specifically, the present disclosure relates to a splash-based forward rendering for a three-dimensional grid, which generates smooth derivatives near the occlusion boundary of the grid. As an example, a three-dimensional grid can be obtained, which includes a plurality of polygons and associated texture data and / or shading data. This three-dimensional grid can be rasterized to generate a two-dimensional grating of the three-dimensional grid (e.g., represented as a plurality of pixels, etc.). For a subset of pixels of the two-dimensional grating, coordinates (e.g., barycentric coordinates, etc.) describing the position of each pixel relative to the vertex of the corresponding polygon in which the pixel is located can be generated. Based on the coordinates of the pixel, an initial color value can be determined for the pixel based on the coordinates of each pixel and the associated shading and / or texture data (e.g., a pixel located within a polygon corresponding to yellow shading can have an initial color value of yellow, etc.). Once the color is determined for the subset of pixels, a splash (e.g., a small color area with smooth attenuation, etc.) can be constructed for the pixel at the corresponding coordinates of each pixel. An updated color value can be determined for each pixel in the pixel based on the weighting of the subset of splashes previously constructed. For example, a splash constructed for a yellow pixel can "spread" the color of the yellow pixel to pixels adjacent to the yellow pixel based on the location of the splash (e.g., at the yellow pixel) and the neighboring pixels (e.g., the splash will spread more strongly to a pixel that is one pixel away from the yellow pixel rather than a pixel that is five pixels away from the yellow pixel). By determining the updated color of each pixel in the subset of pixels, a two-dimensional differentiable rendering of the three-dimensional grid can be generated. Based on the two-dimensional differentiable rendering, derivatives can be generated for the splash (e.g., based on the weighting of the splash relative to the pixel, etc.). Therefore, in this way, a splash can be constructed and applied to a two-dimensional raster to generate a two-dimensional differentiable rendering of the three-dimensional grid, thereby providing a source of smooth derivatives for any point in the two-dimensional rendering (e.g., at an occlusion boundary, etc.).

[0021] More specifically, the computing system can obtain a three-dimensional mesh. The three-dimensional mesh can be or otherwise include a plurality of polygons and associated texture data and / or associated shading data. As an example, the three-dimensional mesh can be a mesh representation of an object, and the associated shading and / or texture data can indicate one or more colors of the polygons of the three-dimensional mesh (e.g., texture data of a character model mesh in a video game, etc.). In some embodiments, one or more machine learning techniques can be used to generate the three-dimensional mesh. As an example, the three-dimensional mesh can be generated by a machine learning model based on a two-dimensional input. For example, a machine learning model can process a two-dimensional image of a cup and can output a three-dimensional mesh representation of the cup.

[0022] As another example, a three-dimensional mesh may represent a pose and / or orientation adjustment to a first three-dimensional mesh. For example, a machine learning model may process a first three-dimensional mesh in a first pose and / or orientation and obtain a second three-dimensional mesh in a second pose and / or orientation different from the first pose and / or orientation. Thus, in some embodiments, the three-dimensional mesh and associated texture / shading data may be the output of a machine learning model, and thus the ability to generate smooth derivatives (e.g., at occlusion boundaries, etc.) for rendering of the three-dimensional mesh is necessary to optimize the machine learning model used to generate the three-dimensional mesh.

[0023] The computing system may rasterize the three-dimensional grid to obtain a two-dimensional raster of the three-dimensional grid. The two-dimensional raster may be or otherwise include a plurality of pixels and a plurality of coordinates associated with a subset of the plurality of pixels, respectively (e.g., by sampling the surface of the three-dimensional grid, etc.). These coordinates may describe the position of the pixel relative to the vertices of the polygon in which the pixel is located. As an example, a first coordinate may describe the position of the first pixel relative to the vertices of the first polygon in which the first pixel is located. A second coordinate may describe the position of the second pixel relative to the vertices of the second polygon in which the second pixel is located. In some embodiments, each of the plurality of coordinates may be or otherwise include a barycentric coordinate.

[0024] In some embodiments, the two-dimensional raster may further include a plurality of polygon identifiers for each pixel in the subset of pixels. The polygon identifier may be configured to identify, for each pixel, one or more polygons within which the corresponding pixel is located. As an example, a pixel may be located within two or more overlapping polygons in a plurality of polygons (e.g., because the polygons of the three-dimensional grid are represented in two dimensions, etc.). The polygon identifier of the pixel will identify that the pixel is located within two polygons. In addition, in some embodiments, when the pixel is located within two or more overlapping polygons, the polygon identifier may identify a front-facing polygon. As an example, the three-dimensional grid may represent a cube. The pixels of the two-dimensional raster of the three-dimensional grid may be located within a polygon on a first side of the cube and an interior of a polygon that is not visible (e.g., from two dimensions) on the opposite side of the cube. The polygon identifier may indicate that the polygon on the first side of the cube is a front-facing polygon. In this way, the polygon identifier may be used to ensure that the color applied to the pixel is associated with the polygon facing the viewpoint of the two-dimensional raster.

[0025] It should be noted that the computing system can rasterize the three-dimensional mesh using any type of rasterization scheme to obtain a two-dimensional raster. As an example, a conventional z-buffered graphics process (e.g., a depth buffer, etc.) can be used to obtain a two-dimensional raster (e.g., a plurality of coordinates, a polygon identifier, etc.). As another example, a three-dimensional ray casting process can be used to obtain a two-dimensional raster (e.g., a plurality of coordinates, a polygon identifier, etc.). Therefore, embodiments of the present disclosure can be used in a rendering and / or rasterization pipeline without the need to use a non-standard or unique rasterization scheme.

[0026] The computing system may determine an initial color value for each pixel in the subset of pixels. The initial color value may be based on the coordinates of the pixel and associated shading data and / or associated texture data. As an example, the polygon coordinates of a first pixel may indicate that the first pixel is located in a first polygon. Based on the shading and / or texture data, it may be determined that the polygon has an orange color where the pixel is located (e.g., relative to a vertex of the polygon, etc.). The first pixel may be colored orange based on the pixel coordinates and the shading and / or texture data.

[0027] More specifically, in some embodiments, the polygon identifier and coordinates of each pixel can allow perspective-corrected interpolation of vertex attributes (e.g., of a polygon of a three-dimensional mesh, etc.) to determine the color of the pixel. In some embodiments, determining the color of each pixel can include applying shading data and / or texture data to a two-dimensional raster using a shading and / or texturing scheme. For either application, any type of application scheme can be used. As an example, a deferred shading, image space application scheme can be used to apply shading data and / or texture data to determine the color of each pixel in a subset of pixels.

[0028] The computing system may construct a splash at the coordinates of each pixel in the corresponding pixel for each pixel in the subset of pixels. More specifically, each rasterized surface point of the two-dimensional grating may be converted into a splash centered at the coordinates of the corresponding pixel and colored by the corresponding coloring color C determined for the pixel. In some embodiments, the splash may be constructed as a small color area with smooth color decay that spreads out from the origin of the splash. As an example, the color of a splash with a bright red center position may decay smoothly as the distance from the center of the splash (e.g., a gradient from the edge of the splash to the center of the splash, etc.) further increases.

[0029] In order to preserve the derivatives of the vertices of the polygons of the three-dimensional mesh, the coordinates at which the splash is constructed (e.g., splash center points, etc.) can be calculated by interpolating the clip space vertex positions of the polygon in which the corresponding pixel is located based at least in part on the coordinates associated with the pixel (e.g., x, y, z, w homogeneous coordinates, barycentric coordinates, etc.). As an example, an h×w×4 buffer of per-pixel clip space positions V can be formed for each pixel. Perspective division and viewport transformation can be applied to the buffer to produce an h×w×2 screen space splash position buffer S. For example, if S is the position of a single splash, the weight of the splash at pixel p in the image is:

[0030]

[0031] where W is normalized by the sum of all the weights of the splash (see below). Then, the final color r at pixel p is p It can represent the adjacent pixels q∈N p Color c q The weighted sum of .

[0032] The computing system may determine an updated color value for each pixel in a subset of pixels to generate a two-dimensional differentiable rendering of a three-dimensional grid. The updated color value may be a weighted subset of the constructed splashes. The weighting of the corresponding splashes in the subset of the splashes of the pixel may be based at least in part on the coordinates of the corresponding pixel and the coordinates of the corresponding splashes. As an example, a first pixel may have a determined yellow color. The coordinates of the second pixel may be located a specific distance away from the first pixel, and the third pixel may be located at an even greater distance from the first pixel. Since the second pixel is located closer to the first pixel than the third pixel, the application of the first pixel splash (e.g., constructed at the coordinates of the first pixel, etc.) to the second pixel may be more heavily weighted than the application of the first pixel splash to the third pixel. Therefore, the weighting of the subset of splashes when determining the updated color value of the pixel may be based on the proximity of each splash to the pixel. By determining the updated color value of each splash in the subset of splashes, a differentiable two-dimensional rendering of a three-dimensional grid may be generated. Therefore, a differentiable 2D rendering (eg, generated by applying splatter or the like) can be used to find smooth derivatives at occlusion boundaries of the rendering.

[0033] It should be noted that in some embodiments, a subset of splashes may include splashes less than the number of constructed splashes. In addition, in some embodiments, a subset of splashes may be uniquely selected for each pixel in a subset of pixels. As an example, a subset of pixel splashes may include only splashes constructed for pixels adjacent to the pixel. As another example, a subset of pixel splashes may include only splashes constructed for pixels within a specific distance or region of the pixel (e.g., a 5x5 pixel region around the pixel, a specific distance as defined by pixel coordinates, etc.). Alternatively, in some embodiments, a subset of splashes may include each splash in the splashes constructed for the subset of pixels. As an example, in order to determine the updated color value of a single pixel, each splash in the constructed splashes may be weighted (e.g., a normalized weight from 0.0 to 1.0, etc.). For example, each splash of a pixel may be weighted based on the distance of the splash from the pixel, and the splash closest to the pixel may be weighted the most, while the splash farthest from the pixel may be weighted the least.

[0034] As an example, S may be the location of a single splash. P may be a pixel in a two-dimensional raster. The weight of a splash at a pixel p in a two-dimensional raster may be defined as:

[0035]

[0036] where W can be normalized by the sum of all the weights of the splash (see below). The updated color value r of pixel p is p It can be the adjacent pixel q∈N p The initial color value c of the coloring q The weighted sum of:

[0037]

[0038] As previously described, in some embodiments, a subset of neighbors (e.g., adjacent, a certain distance away, etc.) may be used. For example, a 3x3 pixel neighborhood may be used with σ=0.5, which may yield a normalization factor:

[0039]

[0040] It should be noted that, unlike point-based rendering, where the splashes themselves are representations of the surface of a three-dimensional mesh, the splashes of the present disclosure are sampled from the underlying mesh at an exact pixel rate. Therefore, the splash radius can be fixed in screen space, and the normalization factor for each splash can be constant, thus avoiding the need to create artificial gradients that are robust to per-pixel weight normalization.

[0041] In some embodiments, the computing system may generate one or more derivatives of one or more corresponding splashes of the two-dimensional rendering based on the two-dimensional differentiable rendering (e.g., at an occlusion boundary, etc.). More specifically, one or more derivatives may be generated for the weights of one or more splashes of the two-dimensional differentiable rendering relative to the vertex position (e.g., using an automatic differentiation function, etc.). Figure 2 The generation of the derivative of splash is discussed in more detail.

[0042] In some embodiments, a two-dimensional differentiable rendering of a three-dimensional mesh may be processed with a machine learning model to generate a machine learning output (e.g., a depiction of a mesh in a second pose / orientation, a second three-dimensional mesh, a mesh fit, a pose estimate, etc.).

[0043] In some embodiments, a loss function can be evaluated that evaluates the difference between the machine learning output and the training data associated with the three-dimensional mesh based at least in part on one or more derivatives. As an example, using a gradient descent algorithm, one or more derivatives can be evaluated to determine the difference between the machine learning output and the training data associated with the three-dimensional mesh (e.g., a true value associated with the three-dimensional mesh, etc.). Based on the loss function, one or more parameters of the machine learning model can be adjusted.

[0044] As an example, the machine learning model may be or otherwise include a machine learning pose estimation model, and the machine learning output may include image data depicting an entity represented by a three-dimensional mesh (e.g., an object, at least a portion of a human body, a surface, etc.). The machine learning output may be image data (e.g., a three-dimensional mesh, a two-dimensional differentiable rendering, etc.) depicting the entity (e.g., as represented by a three-dimensional mesh, etc.) having at least one of a second pose or a second orientation different from the first pose and / or orientation of the entity.

[0045] As another example, the machine learning model can be a machine learning three-dimensional mesh generation model (e.g., a model configured to fit a mesh to a depicted entity, generate a mesh from two-dimensional image data, etc.). The machine learning model can process the two-dimensional differentiable rendering to generate a second three-dimensional mesh. In this way, the two-dimensional differentiable rendering can allow any type or configuration of machine learning model to be trained by differentiation of the machine learning output of the model.

[0046] It should be noted that while the 2D differential rendering generated by the construction of the splatter may resemble a blurred or anti-aliased version of the 2D raster, simply rasterizing and then blurring (even with arbitrary supersampling) is not sufficient to generate non-zero derivatives. More specifically, rasterization alone typically produces zero derivatives with respect to vertices at any sampling rate, and thus the blurred result will lack the ability to generate non-zero derivatives (e.g., be differentiable, etc.).

[0047] The present disclosure provides many technical effects and benefits. As an example technical effect and benefit, the system and method of the present disclosure enable the generation of differentiable two-dimensional renderings of three-dimensional meshes that provide smooth derivatives at occlusion boundaries. As an example, modern differentiable rendering techniques are often computationally expensive and / or unable to produce accurate derivatives at the occlusion boundaries of the rendering. Therefore, the present disclosure provides a method for constructing splashes at the locations of pixels of two-dimensional differentiable renderings. Since the construction of the splashes is computationally efficient, the present disclosure provides a system and method for significantly reducing the amount of computing resources (e.g., power, memory, instruction cycles, storage space, etc.) required to generate differentiable renderings. In addition, the construction of the splashes provides accurate and simple derivatives at the occlusion boundaries of the two-dimensional rendering, thereby significantly optimizing the performance of operations that require differentiable renderings (e.g., machine learning model training, model alignment, pose estimation, mesh fitting, etc.). In this way, the system and method of the present disclosure can be used to generate two-dimensional differentiable renderings of three-dimensional meshes, which are more computationally efficient and more accurate than traditional methods.

[0048] Referring now to the drawings, example embodiments of the present disclosure will be discussed in further detail.

[0049] Example devices and systems

[0050] Figure 1A A block diagram of an example computing system 100 that performs two-dimensional differentiable rendering is depicted according to an example embodiment of the present disclosure. System 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 that are communicatively coupled via a network 180.

[0051] The user computing device 102 may be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0052] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be one processor or a plurality of processors operably connected. The memory 114 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 114 may store data 116 and instructions 118 executed by the processor 112 to cause the user computing device 102 to perform operations.

[0053] As an example, instructions 118 may be executed to perform operations for generating a two-dimensional differentiable rendering 124. More specifically, the user computing device 102 may obtain a three-dimensional mesh including a plurality of polygons and associated texture data and / or shading data (e.g., from one or more machine learning models 120, via network 180, from computing systems 130 and / or 150, etc.). The user computing device 102 may rasterize the three-dimensional mesh to generate a two-dimensional raster of the three-dimensional mesh. For a subset of pixels of the two-dimensional raster, the user computing device 102 may generate coordinates describing the position of each pixel relative to the vertex of the corresponding polygon in which the pixel is located. Based on the coordinates of the pixels, the user computing device 102 may determine an initial color value for each pixel based on the coordinates of the pixels and the associated shading and / or texture data. The user computing device 102 may construct a splash for each pixel at the corresponding coordinates of the pixel (e.g., based at least in part on the initial color value of the pixel, etc.). The user computing device 102 may determine an updated color value for each of the pixels based on a weighting of a subset of previously constructed splashes. By determining the updated color of each pixel in the subset of pixels, the user computing device 102 can generate a two-dimensional differentiable rendering 124 of the three-dimensional mesh. Based on the two-dimensional differentiable rendering, the user computing device 102 can generate derivatives for the splash (e.g., based on the weighting of the splash relative to the pixel, etc.). In this manner, the user computing device 102 can execute instructions 118 to perform operations configured to generate a two-dimensional differentiable rendering 124 from the three-dimensional mesh. Figure 3A and 3B The generation of the two-dimensional differentiable rendering 124 is discussed in more detail.

[0054] It should be noted that in some embodiments, the operations previously described with respect to instructions 118 may also be included in instructions 138 of server computing system 130 and instructions 158 at training computing system 150. Thus, according to example embodiments of the present disclosure, any computing device and / or system may be configured to execute instructions to perform operations configured to generate a two-dimensional differentiable rendering from a three-dimensional mesh.

[0055] In some implementations, the user computing device 102 may store or include one or more machine learning models 120. For example, the machine learning model 120 may be or may otherwise include various machine learning models, such as a neural network (e.g., a deep neural network) or other types of machine learning models, including nonlinear models and / or linear models. The neural network may include a feedforward neural network, a recursive neural network (e.g., a long-short term memory recursive neural network), a convolutional neural network, or other forms of neural networks.

[0056] In some implementations, one or more machine learning models 120 may be received from the server computing system 130 via the network 180, stored in the user computing device memory 114, and subsequently used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 may implement multiple parallel instances of a single machine learning model 120 (e.g., to perform parallel operations across multiple instances of the machine learning model 120).

[0057] More specifically, the machine learning model 120 may be or otherwise include one or more machine learning models 120 that are trained based on processing of two-dimensional differentiable renderings of three-dimensional meshes generated using example embodiments of the present disclosure. As an example, the machine learning model 120 may be a machine learning model (e.g., a pose estimation model, a mesh fitting model, a mesh alignment model, etc.) that is trained based at least in part on processing the two-dimensional differentiable rendering to obtain a machine learning output and then evaluating a loss function that evaluates the machine learning output and a true value using a gradient descent algorithm (e.g., at the user computing device 102, at the server computing system 130, at the training computing system 150, etc.). As another example, the machine learning model 120 may be a machine learning model that is trained at least in part by evaluating a loss function that evaluates the difference between a derivative generated at an occlusion boundary of the two-dimensional differentiable rendering and training data associated with the three-dimensional mesh. It should be noted that the machine learning model 120 may execute any function that can be evaluated based at least in part on a derivative of the two-dimensional differentiable rendering.

[0058] Additionally or alternatively, one or more machine learning models 140 may be included in or otherwise stored and implemented by a server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine learning model 140 may be implemented by the server computing system 140 as part of a web service (e.g., a pose estimation service, a three-dimensional mesh generation service, a pose estimation service, etc.). Thus, one or more machine learning models 120 may be stored and implemented at the user computing device 102, and / or one or more machine learning models 140 may be stored and implemented at the server computing system 130.

[0059] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other devices through which a user can provide user input.

[0060] The server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and may be one processor or a plurality of processors operably connected. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 executed by the processor 132 to cause the server computing system 130 to perform operations.

[0061] In some implementations, the server computing system 130 includes, or is otherwise implemented by, one or more server computing devices. Where the server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0062] As described above, server computing system 130 may store or otherwise include one or more machine learning models 140. For example, model 140 may be or may otherwise include a variety of machine learning models. Example machine learning models include neural networks or other multi-layer nonlinear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. The machine learning model may be the same machine learning model as machine learning model 120 and may be trained in the same manner (e.g., based on two-dimensional differentiable rendering 124, etc.).

[0063] User computing device 102 and / or server computing system 130 may train models 120 and / or 140 via interaction with training computing system 150 communicatively coupled via network 180. Training computing system 150 may be separate from server computing system 130 or may be part of server computing system 130.

[0064] The training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors that are operably connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some embodiments, the training computing system 150 includes one or more server computing devices, or is otherwise implemented by one or more server computing devices.

[0065] The training computing system 150 may include a model trainer 160 that trains the machine learning models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques such as, for example, back error propagation. For example, a loss function may be back propagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques may be used to iteratively update parameters through multiple training iterations.

[0066] In some implementations, performing back-error propagation may include performing truncated back-propagation through time.The model trainer 160 may perform a number of generalization techniques (eg, weight decay, dropout, etc.) to improve the generalization capabilities of the model being trained.

[0067] Specifically, the model trainer 160 can train the models 120 and / or 140 based on the set of training data 162. The training data 162 may include, for example, real-value data associated with a three-dimensional mesh. As an example, the training data 162 may include real-value data of an entity (e.g., an estimated pose of the entity and / or an estimated orientation of the entity, etc.) represented by a three-dimensional mesh (e.g., an object, at least a portion of a human body, a surface, etc.). As an example, the machine learning model (e.g., the machine learning model 120 and 140) may be or additionally include a machine learning pose estimation model. The machine learning pose estimation model 120 / 140 may receive the training data 162, which includes a first pose and / or a first orientation of an entity represented by a three-dimensional mesh, and may output image data (e.g., a second three-dimensional mesh, a two-dimensional rendering, etc.) depicting the entity in a second pose and / or a second orientation. The loss function may be evaluated by the training computing system 150, which evaluates the difference between the image data and the real-value data describing the first pose and / or orientation included in the training data 162.

[0068] In some implementations, if the user has consented, the training examples may be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 may be trained by the training computing system 150 on user-specific data received from the user computing device 102. In some cases, this process may be referred to as personalizing the model.

[0069] The model trainer 160 includes computer logic for providing the desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general purpose processor. For example, in some embodiments, the model trainer 160 includes a program file stored on a storage device, loaded into a memory, and executed by one or more processors. In other embodiments, the model trainer 160 includes one or more sets of computer executable instructions stored in a tangible computer readable storage medium such as a RAM hard disk or an optical or magnetic medium.

[0070] Network 180 may be any type of communications network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and may include any number of wired or wireless links. In general, communications through network 180 may be carried via any type of wired and / or wireless connection using a variety of communications protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).

[0071] Figure 1AThe diagram illustrates an example computing system that can be used to implement the present disclosure. Other computing systems may also be used. For example, in some embodiments, the user computing device 102 may include a model trainer 160 and a training data set 162. In such embodiments, the model 120 may be trained and used locally at the user computing device 102. In some of such embodiments, the user computing device 102 may implement the model trainer 160 to personalize the model 120 based on user-specific data.

[0072] Figure 1B A block diagram depicting an example computing device 10 for performing training of a machine learning model based on two-dimensional differentiable rendering according to an example embodiment of the present disclosure. The computing device 10 may be a user computing device or a server computing device.

[0073] Computing device 10 includes multiple applications (e.g., applications 1 to N). Each application contains its own machine learning library and machine learning model. For example, each application can include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, etc.

[0074] like Figure 1B As illustrated in FIG. 1 , each application can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some embodiments, each application can communicate with each device component using an API (e.g., a public API). In some embodiments, the API used by each application is specific to the application.

[0075] Figure 1C Depicted is a block diagram of an example computing device 50 that performs two-dimensional differentiable rendering according to an example embodiment of the present disclosure. The computing device 50 may be a user computing device or a server computing device.

[0076] The computing device 50 includes a plurality of applications (e.g., applications 1 to N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and the models stored therein) using an API (e.g., a public API across all applications).

[0077] The central intelligence layer includes multiple machine learning models. For example, Figure 1CAs illustrated in , a corresponding machine learning model (e.g., model) can be provided for each application, and the corresponding machine learning model is managed by the central intelligence layer. In other embodiments, two or more applications can share a single machine learning model. For example, in some embodiments, the central intelligence layer can provide a single model (e.g., a single model) for all applications. In some embodiments, the central intelligence layer is included in the operating system of the computing device 50 or is otherwise implemented by the operating system of the computing device 50.

[0078] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized repository for data of computing devices 50. Figure 1C As shown in , the central device data layer can communicate with multiple other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some embodiments, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0079] Example Model Layout

[0080] Figure 2 An illustrative example 200 of a splash construction for a one-dimensional line segment 201 is depicted according to an example embodiment of the present disclosure. It should be noted that Figure 2 An example construction of a splash 204 of a pixel 208 on a one-dimensional line segment 201 is depicted to more easily illustrate the construction of a splash of a two-dimensional raster. Thus, a one-dimensional line segment 201 may be similar to a two-dimensional polygon (e.g., of a three-dimensional grid, etc.), and the construction of a splash 204 of a one-dimensional line segment 201 may be the same or substantially similar to the construction of a splash of a two-dimensional polygon.

[0081] As an example, line segment 201 may include first vertex 202 and second vertex 206. For a pixel located at pixel position 208, a splash may be constructed at splash position 204. More specifically, barycentric coordinate α 212 and barycentric coordinate β 210 may be one-dimensional barycentric coordinates of splash position s204 along line segment 201. Thus, splash position s204 may be defined as:

[0082] s=αv1+βv2

[0083] Where α may represent barycentric coordinate 212, v1 may represent vertex 202, v2 may represent point 206, and β may represent barycentric coordinate 210. As previously described, the pixel at pixel location 208 may be updated (e.g., updated color value, etc.) based at least in part on the weight w(p) of the splash at splash location 204. More specifically, weight w(p) may represent the weight of the splash as applied to the pixel at pixel location 208, and may be weighted based at least in part on the difference between the location of splash location 204 and pixel location 208. Thus, when calculating weight w(p), substituting the above formula for s204 (and omitting normalization and σ factors for simplicity) yields:

[0084] w(p)=exp(-(p-αv1-βv2) 2 ).

[0085] After constructing the splash, a derivative may be generated from the splash based at least in part on the coordinates of the splash relative to the line segment and / or pixel (e.g., 210 and / or 212). Since the barycentric coordinates α 212 and β 210 are generally non-differentiable and may be considered constant, the derivative of the weight w(p) of the splash on a pixel with respect to v1 may be defined as:

[0086]

[0087] The exponential term and the barycentric coordinate α212 may be positive, so the sign of the derivative may be determined by the sign of p-αv1-βv2 or ps. Thus, increasing the distance of vertex v1202 along line segment 201 in a particular direction will increase the weight w(p) of the splash as applied to pixel 208.

[0088] As an example, if pixel position 208 is positioned to the right of splash 204 on line segment 201 (e.g., p>s, etc.), the weight w(p) of splash 204 relative to the pixel at position 208 can be increased. Similarly, if pixel position 208 is positioned to the right of splash 204 on line segment 201 (e.g., p<s, etc.), the weight w(p) of splash 204 relative to the pixel can be decreased. Derivatives can be generated relative to v2 in a similar manner. In this way, splash 204 can be constructed for pixel 208 and the splash can be used to generate derivatives based at least in part on the position of the splash.

[0089] Figure 3AA data flow diagram depicting an example method 300A for generating a two-dimensional differentiable rendering according to an example embodiment of the present disclosure. More specifically, a three-dimensional mesh 302 can be obtained. The three-dimensional mesh 302 can be or otherwise include a plurality of polygons 302A and associated texture data and / or associated shading data 302B. As an example, the three-dimensional mesh 302 can be a mesh representation of an object, and the associated shading and / or texture data 302B can indicate one or more colors of the polygons of the three-dimensional mesh 302 (e.g., texture data for a character model mesh in a video game, etc.). In some embodiments, the three-dimensional mesh 302 can be generated using one or more machine learning techniques. As an example, the three-dimensional mesh 302 can be generated based on inputs through a machine learning model (e.g., Figure 3B machine learning model 326, etc.) to generate the three-dimensional mesh 302.

[0090] The three-dimensional grid 302 can be rasterized by a rasterizer 304 (e.g., a component of a computing system, one or more operations performed by a computing system, etc.) to obtain a two-dimensional raster 306 of the three-dimensional grid 302. The two-dimensional raster 306 can be or otherwise include a plurality of pixels and a plurality of coordinates associated with a subset of the plurality of pixels 306, respectively. The coordinates can describe the position of the pixels relative to the vertices of the polygon 302A in which the pixels 306 are located. As an example, a first coordinate of the pixel / coordinate 306A can describe the position of the first pixel relative to the vertex of the first polygon in which the first pixel is located in the polygon 302A. A second coordinate of the pixel / coordinate 306A can describe the position of the second pixel relative to the vertex of the second polygon in which the second pixel is located in the polygon 302A. In some embodiments, each of the plurality of coordinates of the pixel / coordinate 306A can be or otherwise include a barycentric coordinate.

[0091] The two-dimensional raster 306 may also include a plurality of polygon identifiers 306B for each pixel in the subset of pixels of the pixels / coordinates 306A. The polygon identifiers 306B may be configured to identify, for each pixel 306A, one or more polygons 302A within which the corresponding pixel is located. As an example, a pixel in the pixels 306A may be located within two or more overlapping polygons in the plurality of polygons 302A (e.g., because the polygons of the three-dimensional grid are represented in two dimensions, etc.). The polygon identifier 306B for the pixel 306A will identify that the pixel 306A is located within two polygons 302A. Additionally, in some embodiments, when the pixel in the pixels 306A is located within two or more overlapping polygons of the polygons 302A, the polygon identifier 306B may identify the front-facing polygon 302A. As an example, the three-dimensional grid 302 may represent a cube. The pixels 302A of the two-dimensional raster 306 of the three-dimensional grid 302 may be located within the polygons 302A on a first side of the cube and the interior of the polygons 302A on the opposite side of the cube that are not visible (e.g., from two dimensions). Polygon identifier 306B may indicate that polygon 302A on the first side of the cube is a front-facing polygon. In this way, polygon identifier 306B may be used to ensure that the color applied to the pixel is associated with polygon 302A facing the viewpoint of two-dimensional raster 306.

[0092] It should be noted that the rasterizer 304 can utilize any type of rasterization scheme to rasterize the three-dimensional mesh to obtain a two-dimensional raster. As an example, the two-dimensional raster 306 (e.g., the plurality of pixels / coordinates 306A, the polygon identifier 306B, etc.) can be obtained by the rasterizer 304 using conventional z-buffered graphics processing (e.g., depth buffer, etc.). As another example, the two-dimensional raster 306 can be obtained by the rasterizer 304 using a three-dimensional ray casting process. Therefore, embodiments of the present disclosure can be used in a rendering and / or rasterization pipeline without the need to use a non-standard or unique rasterization scheme.

[0093] A color determiner 308 (e.g., a component of a computing system, one or more operations performed by a computing system, etc.) may determine an initial color value 310 for each pixel in a subset of pixels 306A. The initial color value 310 may be based on the coordinates of the pixel and the associated shading data and / or the associated texture data 302B. As an example, polygon coordinates 306A of a first pixel may indicate that the first pixel is located in a first polygon of polygons 302A. Based on the shading and / or texture data 302B, the color determiner 308 may determine that the polygon will have an orange color where the pixel is located (e.g., relative to a vertex of the polygon, etc.). The first pixel may be colored orange based on the pixel coordinates 306A and the associated shading and / or texture data 302B.

[0094] More specifically, in some embodiments, the polygon identifier 306B and coordinates of each pixel 306A may allow for perspective-corrected interpolation of vertex attributes by a color determiner 308 (e.g., of a polygon 302A of a three-dimensional mesh 302, etc.) to determine the color of the pixel. In some embodiments, determining the color of each pixel by the color determiner 308 may include applying shading data and / or texture data 302B to the two-dimensional raster 306 using a shading and / or texturing scheme. For either application, the color determiner 308 may use any type of application scheme. As an example, the color determiner 308 may apply the shading data and / or texture data 302B using a deferred shading, image space application scheme, and thereby determine an initial color value 310 for the pixel 306A.

[0095] A splat constructor 312 (e.g., a component of a computing system, one or more operations performed by a computing system, etc.) may construct a splat for each pixel in a subset of pixels 306A at the coordinates of each pixel in the corresponding pixel. More specifically, each rasterized surface point of two-dimensional raster 306 may be converted by splat constructor 312 to a splat centered at the coordinates of the corresponding pixel and colored by a corresponding shading color C of initial color value 310. To preserve derivatives of vertices of polygons 302A of three-dimensional mesh 302, coordinates at which the splat is constructed (e.g., splat center points, etc.) may be calculated by interpolating clip space vertex positions of the polygon in which the corresponding pixel is located based at least in part on coordinates associated with pixel 306A (e.g., x, y, z, w homogeneous coordinates, barycentric coordinates, etc.). As an example, an h×w×4 buffer of per-pixel clip space positions V may be formed for each pixel of pixels / coordinates 306A. The splash constructor 312 may apply perspective division and viewport transformation to the buffer to produce an h×w×2 screen-space splash position buffer S. For example, if s is the position of a single splash, the weight of the splash at pixel p in the image is:

[0096]

[0097] where W is normalized by the sum of all the weights of the splash (see below). Then, the final color r at pixel p is p is the adjacent pixel q∈N p Color c q The weighted sum of .

[0098] Based on the constructed splash, an updated color value may be determined for each pixel in the subset of pixels of pixel / coordinate 306A to generate a two-dimensional differentiable rendering 314 of the three-dimensional grid 302. The updated color value may be a weight based on a subset of splashes constructed by the splash constructor 312 (e.g., a subset of splashes located near a pixel, etc.). The weighting of the corresponding splashes in the subset of splashes of the pixel may be based at least in part on the coordinates of the corresponding pixel and the coordinates of the corresponding splash. As an example, the first pixel of pixel / coordinate 306A may have a determined yellow color. The coordinates of the second pixel of pixel / coordinate 306A may be located a specific distance away from the first pixel, and the third pixel of pixel / coordinate 306A may be located at an even greater distance from the first pixel. Since the second pixel is located closer to the first pixel than the third pixel, the application of the first pixel splash (e.g., constructed by the splash constructor 312 at the coordinates of the first pixel, etc.) to the second pixel may be more heavily weighted than the application of the first pixel splash to the third pixel. Thus, the weighting of the subset of splashes in determining the updated color value of the pixel can be based on the proximity of each of the splashes to the pixel. By determining the updated color value of each splash in the subset of splashes, a differentiable two-dimensional rendering 314 of the three-dimensional grid 302 can be generated. Thus, the differentiable two-dimensional rendering 314 (e.g., generated by applying splashes, etc.) can be used to find smooth derivatives at occlusion boundaries of the rendering.

[0099] The two-dimensional differential rendering 314 generated by the construction of the splash at the splash constructor 312 may resemble a blurred or anti-aliased version of the two-dimensional raster 306, but simply rasterizing and then blurring (e.g., with any supersampling) the two-dimensional raster 306 is not sufficient to generate non-zero derivatives. More specifically, rasterization alone typically produces zero derivatives with respect to vertices at any sampling rate, and thus the blurred result will lack the ability to generate non-zero derivatives (e.g., be differentiable, etc.).

[0100] Figure 3B A data flow diagram depicting an example method 300B for training a machine learning model using two-dimensional differentiable rendering according to an example embodiment of the present disclosure. The two-dimensional differentiable rendering 314 can be obtained together with the 3D mesh training data 303. The 3D mesh training data 303 can be or otherwise include real-valued data associated with some aspect of the three-dimensional mesh 302. As an example, the three-dimensional mesh 302 can be generated by a machine learning model based on two-dimensional input data (e.g., a model generation network, etc.). The 3D mesh training data 303 can include real-valued data describing a first pose and / or orientation of an entity depicted by the three-dimensional mesh 302. Therefore, the training data can describe any aspect of the three-dimensional mesh 302, one or more entities represented by the three-dimensional mesh 302, and / or any other aspect of the three-dimensional mesh 302.

[0101] One or more derivatives 318 may be generated for one or more corresponding splashes (e.g., at occlusion boundaries, etc.) of the two-dimensional rendering 314. More specifically, one or more derivatives 318 may be generated using an automatic differentiation function 316 based at least in part on weights of one or more corresponding splashes of the two-dimensional differentiable rendering 314 relative to vertex positions.

[0102] The two-dimensional differentiable rendering 314 of the three-dimensional mesh 302 may be processed with a machine learning model 316 to generate a machine learning output 318 (e.g., a depiction of the mesh 302 in a second pose / orientation, a second three-dimensional mesh, a mesh fit, a pose estimate, etc.).

[0103] A loss function 320 may be evaluated that evaluates the difference between the machine learning output 318 and the training data 303 associated with the three-dimensional mesh 302. As an example, the machine learning output 318 may be evaluated at least in part using a gradient descent algorithm (e.g., using the loss function 320, etc.) to determine the difference between the machine learning output 318 and the training data 303 associated with the three-dimensional mesh 302 (e.g., true values ​​associated with the three-dimensional mesh 302, etc.). Based on the loss function 320, one or more parameters of the machine learning model 316 may be adjusted via parameter adjustment 322. As an example, the machine learning output 318 may be generated by the machine learning model 326, and the loss function 320 may be used to evaluate the difference between the machine learning output 318 and the true value associated with the generated three-dimensional mesh 302 (e.g., the training data 303), and generate parameter adjustment 322 to optimize the machine learning model 316.

[0104] As an example, machine learning model 316 may be or otherwise include machine learning pose estimation model 316, and machine learning output 318 may include image data depicting an entity (e.g., an object, at least a portion of a human body, a surface, etc.) represented by three-dimensional mesh 302. Machine learning output 318 may be image data (e.g., a three-dimensional mesh, a two-dimensional differentiable rendering, etc.) depicting the entity (e.g., as represented by a three-dimensional mesh, etc.) having at least one of a second pose or a second orientation different from the first pose and / or orientation of the entity.

[0105] Example Method

[0106] Figure 4 Flowchart depicting an example method 400 for generating a two-dimensional differentiable rendering according to an example embodiment of the present disclosure. Although for purposes of illustration and discussion, Figure 4The steps are depicted as being performed in a particular order, but the method of the present disclosure is not limited to the particular illustrated order or arrangement. The various steps of the method 400 may be omitted, rearranged, combined and / or adapted in various ways without departing from the scope of the present disclosure.

[0107] At 402, a computing system may obtain a three-dimensional mesh comprising a plurality of polygons and associated texture and / or shading data. As an example, the three-dimensional mesh may be a mesh representation of an object, and the associated shading and / or texture data may indicate one or more colors of polygons of the three-dimensional mesh (e.g., texture data of a character model mesh in a video game, etc.). In some implementations, one or more machine learning techniques may be used to generate the three-dimensional mesh. As an example, the three-dimensional mesh may be generated by a machine learning model based on a two-dimensional input. For example, a machine learning model may process a two-dimensional image of a cup and may output a three-dimensional mesh representation of the cup.

[0108] As another example, a three-dimensional mesh may represent a pose and / or orientation adjustment to a first three-dimensional mesh. For example, a machine learning model may process a first three-dimensional mesh in a first pose and / or orientation and obtain a second three-dimensional mesh in a second pose and / or orientation different from the first pose and / or orientation. Thus, in some embodiments, the three-dimensional mesh and associated texture / shading data may be the output of a machine learning model, and thus the ability to generate smooth derivatives (e.g., at occlusion boundaries, etc.) for rendering of the three-dimensional mesh is necessary to optimize the machine learning model used to generate the three-dimensional mesh.

[0109] At 404, the computing system may rasterize the three-dimensional grid to obtain a two-dimensional raster of the three-dimensional grid. The two-dimensional raster may be or otherwise include a plurality of pixels and a plurality of coordinates associated with a subset of the plurality of pixels, respectively (e.g., by sampling the surface of the three-dimensional grid, etc.). The coordinates may describe the position of the pixel relative to the vertices of the polygon in which the pixel is located. As an example, a first coordinate may describe the position of the first pixel relative to the vertices of the first polygon in which the first pixel is located. A second coordinate may describe the position of the second pixel relative to the vertices of the second polygon in which the second pixel is located. In some embodiments, each of the plurality of coordinates may be or otherwise include a barycentric coordinate.

[0110] In some embodiments, the two-dimensional raster may further include a plurality of polygon identifiers for each pixel in the subset of pixels. The polygon identifier may be configured to identify, for each pixel, one or more polygons within which the corresponding pixel is located. As an example, a pixel may be located within two or more overlapping polygons in a plurality of polygons (e.g., because the polygons of the three-dimensional grid are represented in two dimensions, etc.). The polygon identifier of the pixel will identify that the pixel is located within two polygons. In addition, in some embodiments, when the pixel is located within two or more overlapping polygons, the polygon identifier may identify a front-facing polygon. As an example, the three-dimensional grid may represent a cube. The pixels of the two-dimensional raster of the three-dimensional grid may be located within a polygon on a first side of the cube and an interior of a polygon that is not visible (e.g., from two dimensions) on the opposite side of the cube. The polygon identifier may indicate that the polygon on the first side of the cube is a front-facing polygon. In this way, the polygon identifier may be used to ensure that the color applied to the pixel is associated with the polygon facing the viewpoint of the two-dimensional raster.

[0111] It should be noted that the computing system can rasterize the three-dimensional mesh using any type of rasterization scheme to obtain a two-dimensional raster. As an example, a conventional z-buffered graphics process (e.g., a depth buffer, etc.) can be used to obtain a two-dimensional raster (e.g., a plurality of coordinates, a polygon identifier, etc.). As another example, a three-dimensional ray casting process can be used to obtain a two-dimensional raster (e.g., a plurality of coordinates, a polygon identifier, etc.). Therefore, embodiments of the present disclosure can be used in a rendering and / or rasterization pipeline without the need to use a non-standard or unique rasterization scheme.

[0112] At 406, the computing system may determine an initial color value for each pixel in the subset of pixels. The initial color value may be based on the coordinates of the pixel and the associated shading data and / or the associated texture data. As an example, the polygon coordinates of a first pixel may indicate that the first pixel is located in a first polygon. Based on the shading and / or texture data, it may be determined that the polygon has an orange color where the pixel is located (e.g., relative to a vertex of the polygon, etc.). The first pixel may be colored orange based on the pixel coordinates and the shading and / or texture data.

[0113] More specifically, in some embodiments, the polygon identifier and coordinates of each pixel can allow perspective-corrected interpolation of vertex attributes (e.g., of a polygon of a three-dimensional mesh, etc.) to determine the color of the pixel. In some embodiments, determining the color of each pixel can include applying shading data and / or texture data to a two-dimensional raster using a shading and / or texturing scheme. For either application, any type of application scheme can be used. As an example, a deferred shading, image space application scheme can be used to apply shading data and / or texture data to determine the color of each pixel in a subset of pixels.

[0114] At 408, the computing system may construct a splash at the coordinates of each pixel in the corresponding pixel for each pixel in the subset of pixels. More specifically, each rasterized surface point of the two-dimensional raster may be converted to a splash centered at the coordinates of the corresponding pixel and colored by the corresponding shading color C determined for the pixel. In order to preserve the derivatives of the vertices of the polygons of the three-dimensional mesh, the coordinates at which the splash is constructed (e.g., splash center points, etc.) may be calculated by interpolating the clip space vertex positions of the polygons in which the corresponding pixel is located based at least in part on the coordinates associated with the pixel (e.g., x, y, z, w homogeneous coordinates, barycentric coordinates, etc.). As an example, an h×w×4 buffer of per-pixel clip space positions V may be formed for each pixel. Perspective division and viewport transformations may be applied to the buffer to produce an h×w×2 screen space splash position buffer S. For example, if S is the position of a single splash, the weight of the splash at pixel p in the image is:

[0115]

[0116] where W is normalized by the sum of all the weights of the splash (see below). Then, the final color r at pixel p is p It can represent the adjacent pixels q∈N p Color c q The weighted sum of .

[0117] At 410, the computing system may determine an updated color value for each pixel in the subset of pixels to generate a two-dimensional differentiable rendering of the three-dimensional grid. The updated color value may be a weighted subset of the constructed splashes. The weighting of the corresponding splashes in the subset of the splashes of the pixel may be based at least in part on the coordinates of the corresponding pixel and the coordinates of the corresponding splashes. As an example, a first pixel may have a determined yellow color. The coordinates of the second pixel may be located a specific distance away from the first pixel, and the third pixel may be located at an even greater distance from the first pixel. Since the second pixel is located closer to the first pixel than the third pixel, the application of the first pixel splash (e.g., constructed at the coordinates of the first pixel, etc.) to the second pixel may be more heavily weighted than the application of the first pixel splash to the third pixel. Therefore, the weighting of the subset of splashes when determining the updated color value of the pixel may be based on the proximity of each splash in the splash to the pixel. By determining the updated color value of each splash in the subset of splashes, a differentiable two-dimensional rendering of the three-dimensional grid may be generated. Therefore, a differentiable 2D rendering (eg, generated by applying splatter or the like) can be used to find smooth derivatives at occlusion boundaries of the rendering.

[0118] It should be noted that in some embodiments, a subset of splashes may include splashes less than the number of constructed splashes. In addition, in some embodiments, a subset of splashes may be uniquely selected for each pixel in a subset of pixels. As an example, a subset of pixel splashes may include only splashes constructed for pixels adjacent to the pixel. As another example, a subset of pixel splashes may include only splashes constructed for pixels within a specific distance or region of the pixel (e.g., a 5x5 pixel region around the pixel, a specific distance as defined by pixel coordinates, etc.). Alternatively, in some embodiments, a subset of splashes may include each splash in the splashes constructed for the subset of pixels. As an example, in order to determine the updated color value of a single pixel, each splash in the constructed splashes may be weighted (e.g., a normalized weight from 0.0 to 1.0, etc.). For example, each splash of a pixel may be weighted based on the distance of the splash from the pixel, and the splash closest to the pixel may be weighted the most, while the splash farthest from the pixel may be weighted the least.

[0119] As an example, s may be the location of a single splash. P may be a pixel in a two-dimensional raster. The weight of a splash at a pixel p in a two-dimensional raster may be defined as:

[0120]

[0121] where W can be normalized by the sum of all the weights of the splash (see below). The updated color value r of pixel P p It can be the adjacent pixel q∈N p The initial color value c of the coloring q The weighted sum of:

[0122]

[0123] As previously described, in some embodiments, a subset of neighbors (e.g., adjacent, a certain distance away, etc.) may be used. For example, a 3x3 pixel neighborhood may be used with σ=0.5, which may yield a normalization factor:

[0124]

[0125] Figure 5 A flowchart depicting an example method 500 for generating derivatives at occlusion boundaries of two-dimensional differentiable rendering to train a machine learning model according to an example embodiment of the present disclosure. Although for purposes of illustration and discussion, Figure 5 The steps are depicted as being performed in a particular order, but the method of the present disclosure is not limited to the particular illustrated order or arrangement. The various steps of the method 500 may be omitted, rearranged, combined and / or adapted in various ways without departing from the scope of the present disclosure.

[0126] At 502, the computing system may generate one or more derivatives of one or more corresponding splashes of the two-dimensional rendering based on the two-dimensional differentiable rendering (e.g., at an occlusion boundary, etc.). More specifically, one or more derivatives may be generated for the weights of the one or more splashes of the two-dimensional differentiable rendering with respect to the vertex position (e.g., using an automatic differentiation function, etc.). This may be as previously described with respect to Figure 4 The described generates a two-dimensional differentiable rendering.

[0127] At 504, the computing system may process the two-dimensional differentiable rendering of the three-dimensional mesh with the machine learning model to generate a machine learning output (e.g., a depiction of the mesh at a second pose / orientation, a second three-dimensional mesh, a mesh fit, a pose estimate, etc.).

[0128] At 506, the computing system may evaluate a loss function that evaluates a difference between the machine learning output and the training data associated with the three-dimensional mesh based at least in part on one or more derivatives. As an example, using a gradient descent algorithm, one or more derivatives may be evaluated to determine a difference between the machine learning output and the training data associated with the three-dimensional mesh (e.g., a true value associated with the three-dimensional mesh, etc.). Based on the loss function, one or more parameters of the machine learning model may be adjusted.

[0129] As an example, the machine learning model may be or otherwise include a machine learning pose estimation model, and the machine learning output may include image data depicting an entity represented by a three-dimensional mesh (e.g., an object, at least a portion of a human body, a surface, etc.). The machine learning output may be image data (e.g., a three-dimensional mesh, a two-dimensional differentiable rendering, etc.) depicting the entity (e.g., as represented by a three-dimensional mesh, etc.) having at least one of a second pose or a second orientation different from the first pose and / or orientation of the entity.

[0130] As another example, the machine learning model can be a machine learning three-dimensional mesh generation model (e.g., a model configured to fit a mesh to a depicted entity, generate a mesh from two-dimensional image data, etc.). The machine learning model can process the two-dimensional differentiable rendering to generate a second three-dimensional mesh. In this way, the two-dimensional differentiable rendering can allow any type or configuration of machine learning model to be trained by differentiation of the machine learning output of the model.

[0131] At 508, the computing system may adjust one or more parameters of the machine learning model based at least in part on the loss function (e.g., an evaluation of the loss function, etc.).

[0132] Additional public content

[0133] The techniques discussed herein refer to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for various possible configurations, combinations, and divisions of tasks and functions between and within components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0134] Although the subject matter has been described in detail with respect to various specific example embodiments of the subject matter, each example is provided by way of explanation rather than limitation of the present disclosure. Those skilled in the art can easily produce changes, variations, and equivalents to such embodiments after obtaining and understanding the foregoing. Therefore, the present disclosure does not exclude such modifications, variations, and / or additions to the subject matter, as will be apparent to those of ordinary skill in the art. For example, a feature illustrated or described as a part of an embodiment can be used together with another embodiment to produce yet another embodiment. Therefore, it is contemplated that the present disclosure covers such changes, variations, and equivalents.

Claims

1. A computer-implemented method for generating a differentiable two-dimensional rendering of a three-dimensional model, the method comprising: obtaining, by a computing system including one or more computing devices, a three-dimensional mesh comprising a plurality of polygons and at least one of associated texture data or associated shading data; rasterizing, by the computing system, the three-dimensional grid to obtain a two-dimensional raster of the three-dimensional grid, wherein the two-dimensional raster includes a plurality of pixels and a plurality of coordinates respectively associated with at least a subset of the plurality of pixels, wherein the coordinates of each pixel in the subset of pixels describe a position of the corresponding pixel relative to a vertex of a corresponding polygon of the plurality of polygons in which the pixel is located; determining, by the computing system, a respective initial color value for each pixel in the subset of pixels based at least in part on coordinates of the pixel and the at least one of the associated shading data or the associated texture data; For each pixel in the subset of pixels, constructing, by the computing system, a splash at the coordinates of the corresponding pixel; as well as For each pixel in the subset of pixels, the computing system determines an updated color value of the corresponding pixel based on a weighting of each splash in the subset of splashes to generate a two-dimensional differentiable rendering of the three-dimensional grid, wherein, for each splash in the subset of splashes, the weighting of the corresponding splash is based at least in part on the coordinates of the corresponding pixel and the coordinates of the corresponding splash.

2. The computer-implemented method of claim 1 , wherein the method further comprises: One or more derivatives of one or more corresponding splashes of the two-dimensional differentiable rendering are generated by the computing system based on the two-dimensional differentiable rendering. 3 . The computer-implemented method of claim 2 , wherein the one or more derivatives are generated using an automatic differentiation function.

4. The computer-implemented method of claim 2, wherein: generating each of the one or more derivatives based on respective coordinates of one or more pixels at which the one or more splashes are constructed; and The coordinates of each pixel in the subset of pixels include one or more barycentric coordinates.

5. The computer-implemented method of claim 2, wherein the method further comprises: processing, by the computing system, the two-dimensional rendering using a machine learning model to generate a machine learning output; evaluating, by the computing system, a loss function that evaluates a difference between the output and training data associated with the three-dimensional mesh based at least in part on the one or more derivatives; as well as One or more parameters of the machine learning model are adjusted by the computing system based at least in part on the loss function.

6. The computer-implemented method of claim 5, wherein: The two-dimensional rendering depicts the entity represented by the three-dimensional mesh; The training data includes real-valued data describing at least one of a first posture or a first orientation of the entity; The machine learning output includes image data depicting the entity in at least one of a second pose or a second orientation different from the first pose or the first orientation; and The machine learning model includes a machine learning posture estimation model.

7. The computer-implemented method of claim 5, wherein: The machine learning model includes a machine learning three-dimensional grid generation model; The machine learning output comprises a second three-dimensional mesh based at least in part on the two-dimensional rendering; and The training data includes real-valued data associated with the three-dimensional mesh.

8. The computer-implemented method of claim 6, wherein the three-dimensional mesh comprises a mesh representation of: Object; at least part of a human body; or surface.

9. The computer-implemented method of claim 5, wherein: The machine learning output is evaluated by the loss function based at least in part on a gradient descent algorithm.

10. The computer-implemented method of claim 1, wherein: The two-dimensional raster further comprises a respective subset of polygon identifiers, wherein each polygon identifier in the respective subset of polygon identifiers is configured to identify, for each pixel in the subset of pixels, one or more polygons within which the respective pixel is located; and The initial color value of each pixel in the subset of pixels is based at least in part on the polygon identifier of the corresponding pixel.

11. The computer-implemented method of claim 10, wherein: The pixel is located within two or more overlapping polygons of the plurality of polygons; and The polygon identifier is configured to identify a front polygon of the two or more overlapping polygons.

12. The computer-implemented method of any one of claims 1 to 11, wherein: For each splash in the subset of the splashes, coordinates of the corresponding splash are adjacent to coordinates of the corresponding pixel.

13. A computing system for generating a differentiable two-dimensional rendering of a three-dimensional model, comprising: one or more processors; as well as One or more tangible, non-transitory computer-readable media storing computer-readable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: obtaining a three-dimensional mesh comprising a plurality of polygons and at least one of associated texture data or associated shading data; rasterizing the three-dimensional grid to obtain a two-dimensional raster of the three-dimensional grid, wherein the two-dimensional raster comprises a plurality of pixels and a plurality of coordinates respectively associated with at least a subset of the plurality of pixels, wherein the coordinates of each pixel in the subset of pixels describe a position of the corresponding pixel relative to a vertex of a corresponding polygon of the plurality of polygons in which the pixel is located; determining a respective initial color value for each pixel in the subset of pixels based at least in part on coordinates of the pixel and the at least one of the associated shading data or the associated texture data; For each pixel in the subset of pixels, constructing a splash at the coordinates of the corresponding pixel; and For each pixel in the subset of pixels, an updated color value of the corresponding pixel is determined based on a weighting of each splash in the subset of splashes to generate a two-dimensional differentiable rendering of the three-dimensional grid, wherein, for each splash in the subset of splashes, the weighting of the corresponding splash is based at least in part on the coordinates of the corresponding pixel and the coordinates of the corresponding splash.

14. The computing system of claim 13, wherein: The operations further include: generating one or more derivatives of one or more corresponding splashes of the two-dimensional differentiable rendering based on the two-dimensional differentiable rendering; generating each of the one or more derivatives based on respective coordinates of one or more pixels at which the one or more splashes are constructed; and The coordinates of each pixel in the subset of pixels include one or more barycentric coordinates.

15. The computing system of claim 13, wherein the operations further comprise: processing the two-dimensional rendering using a machine learning model to generate a machine learning output; evaluating a loss function that evaluates a difference between the output and training data associated with the three-dimensional mesh based at least in part on the one or more derivatives; as well as One or more parameters of the machine learning model are adjusted based at least in part on the loss function.

16. The computing system of claim 15, wherein: The two-dimensional rendering depicts the entity represented by the three-dimensional mesh; The training data includes real-valued data describing at least one of a first posture or a first orientation of the entity; The machine learning output includes image data depicting the entity in at least one of a second pose or a second orientation different from the first pose or the first orientation; and The machine learning model includes a machine learning posture estimation model.

17. The computing system of claim 15, wherein: The machine learning model includes a machine learning three-dimensional grid generation model; The machine learning output comprises a second three-dimensional mesh based at least in part on the two-dimensional rendering; and The training data includes real-valued data associated with the three-dimensional mesh.

18. The computing system of claim 16, wherein the three-dimensional grid comprises a grid representation of: Object; at least part of a human body; or surface.

19. The computing system of any one of claims 13 to 18, wherein: The two-dimensional raster further comprises a respective subset of polygon identifiers, wherein each polygon identifier in the respective subset of polygon identifiers is configured to identify, for each pixel in the subset of pixels, one or more polygons within which the respective pixel is located; and The initial color value of each pixel in the subset of pixels is based at least in part on the polygon identifier of the corresponding pixel.

20. One or more tangible, non-transitory computer-readable media storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising: obtaining a three-dimensional mesh comprising a plurality of polygons and at least one of associated texture data or associated shading data; rasterizing the three-dimensional grid to obtain a two-dimensional raster of the three-dimensional grid, wherein the two-dimensional raster comprises a plurality of pixels and a plurality of coordinates respectively associated with at least a subset of the plurality of pixels, wherein the coordinates of each pixel in the subset of pixels describe a position of the corresponding pixel relative to a vertex of a corresponding polygon of the plurality of polygons in which the pixel is located; determining a respective initial color value for each pixel in the subset of pixels based at least in part on coordinates of the pixel and the at least one of the associated shading data or the associated texture data; For each pixel in the subset of pixels, constructing a splash at the coordinates of the corresponding pixel; as well as For each pixel in the subset of pixels, an updated color value of the corresponding pixel is determined based on a weighting of each splash in the subset of splashes to generate a two-dimensional differentiable rendering of the three-dimensional grid, wherein, for each splash in the subset of splashes, the weighting of the corresponding splash is based at least in part on the coordinates of the corresponding pixel and the coordinates of the corresponding splash.

Citation Information

Patent Citations

  • Continuous depth-ordered image compositing

    CN108604389A

  • Differentiable rendering pipeline for inverse graphics

    CN109472858A