Stylized animatable representation
Using training sample datasets and machine learning models, calculating grid mapping and proximity evaluations, solves the problem of personalized stylized avatar generation, and achieves efficient, personalized and low-cost avatar generation, reducing computing resource requirements and reducing the uncanny valley effect.
Patent Information
- Application Number
- CN202380080269.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-22
- Filing Date
- 2023-09-27
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art is difficult to create personalized stylized animated representation efficiently and scalably, especially when facing a large number of individual objects, and the use of photo-level avatars has high computational costs and uncanny valley problems.
Using a stylized representation generator, a model is formed by training the sample dataset, compute grid maps and evaluate proximity, select stylized animated representations from the population, and use machine learning models and grid mapping techniques to generate a personalized and stylized 3D model.
It realizes the generation of personalized stylized avatars more efficiently in computing, reduces the demand for computing resources, reduces the uncanny valley effect, and promotes human computer interaction.
Smart Images

Figure CN120283264A_ABST
Abstract
Description
Background Art
[0001] Avatars of humans or animals can be used in many applications, including but not limited to: video games, video conferencing, telepresence, virtual reality, augmented reality, virtual reality, etc.
[0002] To render an avatar of a human or animal to create an animation, a 3D model of the human or animal is typically used. The 3D model is an animatable representation. By controlling the pose of the 3D model (where the pose is the 3D position and orientation) and rendering a 2D image based on the posed 3D model (depicting the human or animal as an avatar), an animation can be created. Generally, the complexity of the 3D model is increased to increase the precision of control of the 3D model in order to depict fine details such as facial expressions, movements of fingers, and subtle movements of the body.
[0003] The embodiments described below are not limited to implementations that solve any or all of the disadvantages of known stylized animatable representations. Summary of the Invention
[0004] A simplified overview of the present disclosure is given below to provide a basic understanding to the reader. This overview is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Its sole purpose is to present a selection of concepts disclosed herein in a simplified form as a prelude to the more detailed description presented later.
[0005] In various examples, a stylized animatable representation is created for an object such as a specific human or animal. The representation is personalized because it has a likeness of the specific object. In some examples, the stylized animatable representation is a 3D rigged model.
[0006] In various examples, there is a method for calculating a stylized animatable representation of an object based on a population of stylized animatable representations. The method includes the steps of: accessing a ground truth representation of the object; using a model to calculate a mesh mapping, the model being formed using a dataset of training examples that pair ground truth representations of other objects with instances of the population. The method also includes applying the mesh mapping to the ground truth representation of the object to produce a target mesh; and selecting the stylized animatable representation from the population by evaluating the proximity of the target mesh to the instances of the population.
[0007] Many additional features will be more readily understood by reference to the following detailed description considered in conjunction with the accompanying drawings. Brief Description of the Drawings
[0008] The present specification will be better understood by reading the following detailed description in conjunction with the accompanying drawings, wherein:
[0009] Figure 1 shows a stylized representation generator deployed as a cloud service;
[0010] Figure 2 is a schematic diagram of a training example;
[0011] Figure 3 is a schematic diagram of a 2D image of an object and shows the stages of a process for generating a stylized avatar of the object;
[0012] Figure 4 is a flowchart of a method for calculating an animatable representation of a style of an object;
[0013] Figure 5 illustrates an exemplary computing-based device in which an embodiment of a stylized representation generator is implemented.
[0014] In the drawings, like reference numerals are used to represent like components. DETAILED DESCRIPTION
[0015] The detailed description provided below in conjunction with the accompanying drawings is intended as a description of the present example and is not intended to represent the only form in which the present example is constructed or utilized. The description sets forth the functions of the example and the sequence of operations for constructing and operating the example. However, the same or equivalent functions and sequences may be implemented by different examples.
[0016] The term "animatable representation" is used herein to refer to a 3D model of a person or animal, wherein the pose of the 3D model can be controlled using the parameters of the 3D model. In some examples, the animatable representation is articulated because it includes at least one joint, such as a neck joint, a shoulder joint, a hip joint, or other joints. In some examples, the animatable representation is rigged because it includes one or more controls, such as joints or groups of joints. In a non-limiting example, a group of joints is a skeleton or a part of a skeleton. In some examples, the animatable representation includes at least one blend shape, the at least one blend shape being a function representing facial muscle movements or other deformations, which may be non-articulated, such as the movement of hair or the appearance of wrinkles. In some examples, the animatable representation is a rigged 3D mesh or a rigged smooth surface 3D model.
[0017] As mentioned above, to render an avatar of a person or animal to create an animation, a 3D model of the person or animal is typically used. By controlling the pose of the 3D model (where the pose is 3D position and orientation) and rendering a 2D image based on the posed 3D model, an animation can be created. In the case where the 3D model depicts an object such as a person or animal, a photo-realistic 2D image can be rendered with high fidelity. However, near-real but not fully photo-real avatars suffer from the "uncanny valley" problem, whereby when viewing such synthetic images of a person or animal, users may experience a negative emotional response.
[0018] As mentioned above, avatars of persons or animals can be used in many applications, including but not limited to: video games, video conferencing, telepresence, virtual reality, augmented reality, virtual reality, etc. The inventors have recognized that in these applications, the purpose of using an avatar is to facilitate human-computer interaction and / or human communication. Having a near photo-real avatar may impede human-computer interaction (due to the uncanny valley problem). A near photo-real avatar is also computationally expensive to compute and uses a large amount of resources such as memory and power. Therefore, the inventors have recognized that a stylized avatar is beneficial for facilitating human-computer interaction and / or human communication. A stylized avatar is a semi-realistic depiction of a person or animal. A stylized avatar can be a depiction of a person or animal, where at least one feature of the person or animal is enlarged or reduced relative to another feature of the person or animal. A non-exhaustive list of features is: body stance, eye shape, nose shape, chin shape. Using a stylized avatar also enables savings in resources such as memory and power compared to using a photo-real avatar.
[0019] However, the inventors have found that it is very difficult to create an animatable representation of a person or animal that is stylized and also maintains the likeness of the individual in a scalable manner. Such an animatable representation is considered personalized because it retains the likeness of a particular object (person or animal). Since there may be millions of individual objects (for web-scale applications), it is a huge task to be able to compute a stylized animatable representation for each of these individuals. It is challenging to do so in a way that preserves identity or retains the likeness of the individual.
[0020] The inventors have developed a way to compute a stylized animatable representation of an object in a computationally efficient and scalable manner using training examples.
[0021] Figure 1 is a schematic diagram of a computer-implemented stylized representation generator 100. The stylized representation generator includes a model 108, at least one processor 104, and a memory 106.
[0022] InFigure 1 In the example, the stylized representation generator 100 is deployed at a computing entity communicating with the communication network 124. The communication network is the Internet, an intranet, or any other communication network. The stylized representation generator receives an input including a ground truth representation 118. The ground truth representation 118 is a 3D model of an object such as a person or an animal. In some cases, the ground truth representation 118 is calculated based on at least one 2D image of the object. In some cases, the ground truth representation 118 is obtained from the repository 126 via the communication network 124. The output of the stylized representation generator 100 is a 3D model of the object in a stylized form. The 3D model is sent to the downstream application 130 to use the renderer 102 to render a 2D image animating the avatar of the object. A non-exhaustive list of downstream applications is: video conferencing applications, video game applications, movie creation applications, telepresence applications, virtual assistants. In some cases, the downstream application 130 is deployed in an end-user computing device.
[0023] The output from the renderer 102 includes a 2D image that can be sent to the end-user device, for example, by one or more of the following: inserted into a virtual network camera stream for display at the smartphone 122, inserted into a video game controlled by the game console 110, displayed via the head-mounted computing device 114. The renderer 102 is any function for calculating a 2D image based on a 3D model, such as by using ray tracing, ray casting, neural rendering, rasterization, or otherwise. A non-exhaustive list of example renderers that can be used is: SolidWorks Visualize (trademark), Sunflow (trademark), LuxCoreRender (trademark), Unity (trademark).
[0024] As mentioned above, the stylized representation generator 100 includes a model 108. The model 108 is formed using a dataset of training examples that pair the ground truth representations of other objects with instances of ethnic groups. In some examples, the model is a trained machine learning model. In other examples, the model is used to calculate a mesh map using vertex displacements observed in multiple training examples. In some examples, the model is used to calculate at least one distortion, such as a 3D space distortion. The term "3D space distortion" is used to refer to a function from a position in 3D space to another position (or equivalently, a function from a position in 3D space to a displacement vector).
[0025] The stylized representation generator can access at least one population 120 of stylized representations. Each population includes a potentially infinite number of stylized representations, where each stylized representation is a rigged 3D model of a person or animal. In an example, an instance of a population is realized by varying the values of successive parameters of a rigged 3D model of a person or animal. Within a population, the stylized representations have the same style. For example, the style of a given population may include enlarged eyes; the style of another population may include the length of limbs relative to the specified proportion of the whole body. Within a population, each stylized representation is a rigged 3D model depicting a particular object (person or animal) where the object is real or synthetic. The stylized representations are computationally expensive to create and are pre-created and stored in a database or other storage device accessible to the stylized representation generator 100 via the communication network 124.
[0026] The stylized representation generator 100 is capable of accessing multiple sets of training examples 128 via the communication network 124. Each training example is a pair including a real item and a stylized version of the real item created by a human artist. The real item is a real image of an object (such as a photo or video frame) or a real 3D mesh model, or multiple images captured by an offline multi-camera capture device, which can be used to create a real 3D mesh model. The stylized version of the real item is a 3D mesh model where the human artist has manually set the values of the rigging parameters of the 3D mesh model. Each population has a set of training examples. Each set of training examples includes multiple training examples. In some embodiments, the number of training examples is set to be small, such as approximately 20. In some embodiments, each set has thousands of training examples. More details about the training examples are given in Figure 2 reference.
[0027] In one example, a user is wearing a head-mounted computing device 114 and participating in a video call. The remote participant in the video call is visible to the user as a hologram 112. The hologram is stylized and generated using a stylized representation generator 100. The remote participant has an account with the provider of the video call, and associated with the account is a stored true representation of the remote participant. In the example, a 2D image of the remote participant is obtained and used to compute a true representation including a 3D model of the remote participant. The true representation is stored in a repository 126. The video call provider determines the account of the remote participant and accesses the true representation from the repository 126. The true representation is sent to the stylized representation generator 100. The stylized representation generator 100 generates a stylized representation, which is a 3D-bound model. Then, the 3D-bound model is used by a downstream application 130 and a renderer 102 to create a hologram 112 for display by the head-mounted computing device 114. The hologram 112 facilitates the video call between the remote participant and the wearer of the head-mounted computing device.
[0028] The stylized representation generator 100 of the present disclosure operates in an unconventional manner to achieve an animatable representation that can maintain the style of a portrait of a person or an animal.
[0029] By using training examples 128 to compute a mesh mapping, the stylized representation generator 100 is capable of improving the functionality of a computing device by computing an animatable representation that facilitates human-computer interaction.
[0030] In Figure 1 an example, the stylized representation generator 100 is deployed at a computing entity remote from the end-user device or the downstream application. However, in some cases, the stylized representation generator 100 functionality is deployed at the end-user device or at the computing entity providing the downstream application 130. In other examples, the functionality of the stylized representation generator 100 is shared between the stylized representation generator 100 and at least one other computing device (such as the end-user computing devices 110, 114, 122 or the computing device providing the downstream application 130).
[0031] Alternatively or additionally, at least part of the functionality of the stylized representation generator 100 described herein is performed by one or more hardware logic components. By way of example, and not limitation, illustrative types of hardware logic components that may optionally be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and graphics processing units (GPUs).
[0032] Figure 2 is a schematic diagram of a training example that includes a 2D image 200 of an object and a corresponding stylized picture 204 of the object created by a human artist. The image 200 is a photograph or video frame of the object captured by, for example, a video camera, a web camera, a smart phone camera, or other digital camera. In Figure 2 the example, the image 200 of the object shows the head and shoulders of a female facing the camera. The female has long straight hair with a center parting line, is wearing a blue top, and has a neutral expression. The corresponding stylized picture 204 shows the head and shoulders of a female facing the camera, with wavy hair just above shoulder length, a center parting line, wearing a blue top, and having a neutral expression. The eyes in the stylized picture 204 are larger than the eyes in the image 200. Note that in Figure 2 the training example is shown as a black and white line drawing, while the actual training examples can include color digital images. Figure 2 is thus a schematic diagram of one of the training examples 128 that can be stored in Figure 1 the training example.
[0033] Figure 3 is a schematic diagram of a 2D image 300 of an object and shows a processing stage of the 2D image for generating a stylized avatar 310 of the object. There is a 2D image 300 of the object. The 2D image is not available in the training example 128. The 2D image is a digital photograph captured by an image capture device such as any one of the following: a web camera, a smart phone camera, a video camera, an RGB camera. The 2D image is used to create a true representation of the object, where the true representation is a 3D model, such as a rigged 3D model. In Figure 2 the true representation is schematically shown at 302, which is a black and white line drawing schematically showing the 3D model of the object. The true representation 302 truly depicts the object.
[0034] In the example, the 2D image is used to create the true representation using techniques to reconstruct a 3D model using dense landmarks. A trained machine learning model is used to predict the positions of the dense landmarks in the 2D image. Then, the 3D model is fitted to the predicted positions of the dense landmarks. The machine learning model is trained using synthetic training data that gives ground truth landmark annotations. Using this scheme only gives the benefit of a single 2D image of the object when the accuracy of the 3D model is high.
[0035] In another example, the 2D image is one of multiple 2D images of the object captured using a camera device. Since the position of the camera in the device is known and images are captured simultaneously from different viewpoints, it is possible to construct a 3D model of the object using the geometry and information about the parameters of the camera.
[0036] In various examples, the 3D model 302 is a 3D mesh model formed by a plurality of polygons.
[0037] In some but not all embodiments, the 3D model 302 is processed to change the number of polygons of the 3D mesh to a target number of polygons. The resulting 3D mesh 304 is schematically shown in Figure 3 In.
[0038] A mesh mapping is applied to the 3D model 302 (or 3D model 304). The mesh mapping is found using the model 108, as explained in more detail below. The result of applying the mesh mapping is schematically shown at 306 in Figure 3 and is referred to as the target mesh 306.
[0039] The target mesh 306 is used to select an instance 308 from the population 120 of stylized animatable representations. The target mesh 306 is less complex than the instances in the population 120, and thus, by selecting an instance from the population, a more powerful representation is obtained that is stylized and able to maintain the likeness or identity of the object. Instances from the population are more powerful because they can be animated with greater precision and detail than the target mesh 306. The selected instance 308 is a stylized animatable representation of the object, which can then be used by a downstream application 130 to render a 2D image 310 of the object in a stylized form. The 2D image 310 of the object in a stylized form retains the likeness of the object as depicted in the real 2D image 300.
[0040] Although Figure 3 the example in shows the object face facing forward with a neutral expression, the 2D image 310 can depict the object with different poses and / or different expressions. This is achieved by changing the pose parameters of the stylized animatable representation 308 before rendering.
[0041] Figure 4 is a flowchart of a method for calculating a stylized animatable representation of an object according to a population 120 of stylized animatable representations. In an example, Figure 4 the method of is by Figure 1It is performed by the stylized representation generator 100. Access the true representation 118 of the object. In an example, the true representation 118 is accessed via a communication network from a repository (such as a repository 126 of 3D mesh models of individual objects). In an example, the object has registered with the service and has provided its own 2D image, which has been used to calculate a 3D mesh model depicting the object. In another example, the true representation is received from another computing entity such as a telepresence service or an online game application.
[0042] The stylized representation generator 100 uses the model 108 to calculate 400 a mesh mapping. The model 108 is formed using a dataset of training examples 128 that pair the true representations of other objects with instances of the ethnic group 120. More details about the model 108 are given below.
[0043] The stylized representation generator 100 applies 402 the mesh mapping to the true representation of the object to produce a target mesh 404; that is, the target mesh 404 is the result of the application operation. In Figure 3 In, item 306 is an example of a target mesh with a smooth surface that is applied and drawn as a line drawing. The stylized representation generator 100 selects 406 an instance from the ethnic group 120 by evaluating the proximity of the target mesh 404 to the instances of the ethnic group 120. Thus, the operation of evaluating proximity can be considered part of the selection operation 406. Instances from the ethnic group are selected based on their proximity to the target mesh. Any suitable way of evaluating proximity is used, such as calculating a similarity metric or calculating an optimization. The selected instance is a stylized animatable representation 408 of the object that preserves the portrait of the object.
[0044] In various examples, evaluating the proximity of the target mesh to the instances of the ethnic group includes optimizing an energy function that includes a geometric residual term (landmark or vertex differences) and an identity prior term (penalizing extreme identity coefficients or impossible combinations of identity coefficients). In various examples, evaluating the proximity of the target mesh to the instances of the ethnic group includes fitting the target mesh to the ethnic group by the computational optimization now described:
[0045] Let β be the identity parameter of the target ethnic group.
[0046] Let M(β) be the personalized bound mesh in the target ethnic group.
[0047] Let T be the transformed (e.g., warped) target mesh.
[0048] Energy minimization algorithms and energy terms are used to solve for the optimal target population identity parameters for a given transformed (e.g., warped) target mesh. Levenberg Marquardt is one optimization algorithm that can be used, but other optimization algorithms such as gradient descent are also available. In an example, the energy terms include 3D data terms (at landmark vertices or all vertices) and identity priors (such as terms that minimize the L1 or L2 norm of beta or a Gaussian mixture model).
[0049] In an example, the 3D data term is E Landmarks :=∑ j∈L ||T j -M(β) j || 2 , which in words is: the energy at the landmarks of the target mesh equals the sum of the squares of the magnitudes of the differences between the vertices of the target mesh and the corresponding vertices of the personalized mesh in the target population at the vertices of the target mesh. The symbol L represents the set of landmark vertex indices (or all vertex indices). The symbol T j represents the j-th vertex of the target mesh. The symbol M(β) j represents the j-th vertex M(β).
[0050] In an example, the identity prior using the L1 norm is E IdentityL1 :=|β|1.
[0051] Figure 4 The process of is an efficient and effective way to compute a stylized animatable representation of an object. Figure 4 The process of is scalable because the process of computing 400 the mesh mapping and selecting 406 from the population is efficient. Since the mesh mapping is computed 400 using a model 108 formed using a data set of training examples 128 that pair real representations of other objects with instances of the population, the mesh mapping can apply the style of the population to the real representation 118. Computing the mesh mapping is computationally efficient and scalable. Due to the way the mesh mapping is computed using model 108, the target mesh retains some likeness of the object depicted in the real representation.
[0052] In some embodiments, model 108 is a trained machine learning model. Any machine learning model can be used, such as a convolutional neural network, a random decision forest, a support vector machine. The trained machine learning model is trained using supervised learning with training examples 128 that, in this embodiment, include tens of thousands or more examples, as in Figure 2As shown. Depending on the type of machine learning model used, any suitable supervised learning algorithm is used. Once trained, the machine learning model is used to take the ground truth representation 118 as input and predict the target mesh 404. In doing so, the trained machine learning model has effectively computed the mesh mapping 400 and applied 402 the mesh mapping to the ground truth representation 118 to produce the target mesh 404, although these processes are part of the overall trained machine learning model inference process and are not easily separable from that inference process. By using a machine learning model, the style can be accurately transferred from the training examples 128 to the ground truth representation 118. Using a machine learning model gives the benefit of generalization; that is, even if the ground truth representation is different from the training examples 128, the target mesh 404 is accurate.
[0053] In some embodiments, the model is used to compute the mesh mapping 400 using vertex displacements observed in multiple training examples. The multiple training examples are selected as the nearest neighbors of the ground truth representation from the dataset of training examples 128. In embodiments where the model is used to compute the mesh mapping 400 using vertex displacements, the number of training examples 128 can be approximately 20, and the number of nearest neighbors can be approximately 3. In some cases, the dataset includes only dozens of training examples. Unexpectedly, using such a small number of training examples and nearest neighbors is found to give good performance and is beneficial for scalability.
[0054] To compute the nearest neighbors, the ground truth representation 118 including the 3D mesh is compared with the ground truth 3D meshes of each training example 128. Any suitable similarity metric (such as the mesh-to-mesh distance or the distance between semantic representations) is used to compute the comparison. In some cases, the mesh-to-mesh distance compares each vertex with its corresponding vertex in another mesh. In some cases, the comparison is made only between a subset of the mesh vertices (such as those representing key points). Then, the training examples are ranked according to the similarity metric, and the top k are selected. In the example, k is 3. However, other values of k are available.
[0055] In some examples, retopology is computed before comparing the meshes. The retopology includes adjusting the mesh of the ground truth representation 118 to have the same number of vertices as each of the meshes in the training examples 128. In this way, there is a one-to-one mapping between the vertices in the ground truth representation and the vertices in the meshes of the training instances. In cases where the mesh mappings are already roughly aligned, no alignment operation is required.
[0056] In some examples, scale - invariant alignment is applied to the mesh of the ground truth representation 118; or, without re - topology, scale - invariant alignment is applied to the re - topologized mesh of the ground truth representation or directly to the mesh of the ground truth representation.
[0057] Once the top - k training examples are identified, per - vertex displacements are computed for each of the k training examples. The per - vertex displacement for a given training example is the displacement that must be applied to the 3D mesh of the ground truth representation in the training pair to reach the 3D mesh of the stylized representation in the training pair. Thus, there are k sets of per - vertex displacements. Aggregate the k sets of per - vertex displacements. The aggregation is any suitable aggregation, such as weighted average, median, mode, or other aggregations. In the case of a weighted average, the weights can take into account the similarity metric obtained when computing the nearest neighbors. The benefit of using a weighted average is that the influence of the k training examples can take into account the similarity between the ground truth representation and the training examples.
[0058] In another example, the k sets of per - vertex displacements are aggregated before applying the result of the aggregation to the target mesh.
[0059] The per - vertex displacement process has been found to be efficient and effective, especially since applying the mesh mapping 402 as per - vertex displacements can be performed in parallel for each vertex in the mesh, and the displacements are efficiently implemented in a computing device.
[0060] In some embodiments, the model is used to compute the mesh mapping 400 as one or more warps. Using warps is more powerful than using only per - vertex displacements, enabling the mesh mapping 400 to apply more styles in a concise manner. In an example, as explained above, the k nearest neighbors of the ground truth representation are computed. The k nearest neighbors are training examples. For each of the k nearest neighbors (which are training examples), a warp is computed that transforms the 3D mesh of the ground truth representation of the training example to the 3D mesh of the stylized representation of the training example. As a result, k warps are determined. In some cases, each of the k warps is applied to the ground truth representation 118 to produce k intermediate target meshes. Then, the k intermediate target meshes are aggregated to produce the target mesh 404. The aggregation method is any suitable aggregation, such as weighted average, median, mode, or other aggregations. The benefit of producing the aggregated k target meshes is that a warp kernel can be used to compute the k target meshes. Conversely, in some cases, aggregating warp kernels produces unexpected results because in some cases aggregating warp kernels may result in a loss of accuracy.
[0061] In other cases, k warps are aggregated to produce an aggregated warp, which is then applied to the ground truth representation 118 to produce the target mesh 404. Aggregating the k warps gives a viable result and leads to efficiency as there is no need to store k intermediate target meshes.
[0062] In embodiments where the model is used to compute the mesh mapping 400 using warps, the number of training examples 128 can be approximately 20, and the number of nearest neighbors can be approximately 3. In some cases, the dataset includes only dozens of training examples. Unexpectedly, using such a small number of training examples and nearest neighbors is found to give good performance and facilitate the scalability of using warps.
[0063] In various examples, the mesh mapping is a learned vertex displacement as now explained. Computation is found to be very efficient in these examples.
[0064] Let N be the number of training examples.
[0065] Let k be the number of nearest neighbor training examples.
[0066] Let j ∈ {0, …, N - 1} be the example index.
[0067] Let be the stylized warp function (also called spatial warp) that captures the style of example j.
[0068] Let d ∈ {x, y, z} be the spatial dimension.
[0069] Let be the stylized warp function that captures the style of example j along dimension d.
[0070] Let V be the number of vertices in the mesh.
[0071] Let be the i-th kernel center (vertex position) of the example input j.
[0072] Let M be the mesh to be warped.
[0073] Let M′ be the output mesh.
[0074] Let v = (x, y, z) be a vertex of M.
[0075] Let v′ = (x′, y′, z′) be a vertex of M′.
[0076] Let D j be the per-vertex displacement function that captures the style of example j.
[0077] In the case of only a single "training pair" (such as the average mesh from each ethnic group), the same per-vertex displacement is applied to all objects as follows:
[0078] M′ = M + D shared
[0079] In the case of using multiple nearest neighbors, the output grid is calculated as which in words is: the output grid is equal to the input grid plus the average of the neighbors of the displacement of the vertices.
[0080] Now describe an example where the mesh mapping is a distortion such as a spatial warp. In the example, the distortion function is one or more thin plate splines (TPS). Each TPS encodes the displacement along one spatial dimension and can be implemented as the sum of V "linear" kernels plus an overall affine term.
[0081] For all vertices v of the ground truth representation 118 mesh aligned with training example j, the spatial warp is applied as follows:
[0082] v′ = v + W j (v)
[0083] This is achieved by calculating the spatial warp for each of the three spatial dimensions x, y, and z as follows:
[0084] v′.x = v.x + W j,x (x, y, z)
[0085] v′.y = v.y + W j,y (x, y, z)
[0086] v′.z = v.z + W j,z (x, y, z)
[0087] As a result, the vertices v′ of the ground truth representation 118 mesh are stylized like training example j.
[0088] The thin plate spline calculation is expressed as follows:
[0089] For d ∈ {x, y, z}:
[0090]
[0091] The terms in Equation 1 are scalar coefficients calculated according to the input to MM ′ and Equation 1 in words is: the warp that captures the stylization of training example j in spatial dimension d is equal to the sum of the scalar coefficients calculated from the input to MM ′ (the 3D mesh of the training example), plus the sum of the learned weights of the thin plate spline for a given vertex index, the index into the artist example, and the spatial dimension multiplied by the magnitude of the displacement between vertices in each dimension on the vertex.
[0092] The individual warps can be calculated as follows:
[0093] M′ = M + W j (M) (2)
[0094] It is calculated per vertex. For each vertex v ∈ M, the corresponding warped vertex is v′ ∈ M′:
[0095] v′ = v + W j (v) (3)
[0096] Wherein, the three dimensions are warped individually:
[0097] v′.x = v.x + W j,x (x,y,z) (4)
[0098] v′.y = v.y + W j,y (x,y,z) (5)
[0099] v′.z = v.z + W j,z (x,y,z) (6)
[0100] The k-nearest neighbor warp calculation can be calculated as follows:
[0101] M′ = M + (1 / k)∑ j∈neighbors W j (M) (7)
[0102] Which in words is: the warped grid M ′ equals the grid M to be warped plus the average of the warps applied to each of the grids in the grid M for k training examples.
[0103] Figure 5 Illustrated are various components of an exemplary computing-based device 500 implemented as any form of computing and / or electronic device, and in some examples, embodiments of a stylized representation generator 100 are implemented.
[0104] The computing-based device 500 includes one or more processors 104, the one or more processors 104 being microprocessors, controllers, or any other suitable type of processor for processing computer-executable instructions to control the operation of the device to compute an animatable representation of a stylized version of an object. In some examples, such as in the case of using a system-on-chip architecture, the processor 104 includes one or more fixed function blocks (also referred to as accelerators) which are implemented in hardware (rather than software or firmware) Figure 3 and Figure 4Part of the method of any of the above. The functionality of the stylized representation generator 100 is deployed at the computing-based device 500. Platform software including an operating system 508 or any other suitable platform software is provided at the computing-based device to enable the application software 510 to execute on the device. The data repository 522 stores values of coefficients, results of nearest neighbor calculations, training examples, 3D mesh models, populations of stylized representations, rendered images, and other data.
[0105] Computer-executable instructions are provided using any computer-readable medium accessible by the computing-based device 500. Computer-readable media include, for example, computer storage media such as memory 106 and communication media. Computer storage media such as memory 106 includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, etc. Computer storage media includes but is not limited to: random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage devices, magnetic tape cartridges, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other non-transmission media for storing information accessible by a computing device. In contrast, communication media embodies computer-readable instructions, data structures, program modules, etc. in a modulated data signal such as a carrier wave or other transmission mechanism. As defined herein, computer storage media does not include communication media. Thus, computer storage media should not be construed as propagating signals themselves. Although the computer storage media (memory 106) is shown within the computing-based device 500, it should be appreciated that in some examples, the storage is distributed or remotely located and accessed via a network or other communication link (e.g., using the communication interface 516).
[0106] The computing-based device 500 optionally includes a capture device 518, such as a camera for capturing an image of an object. The computing-based device optionally includes a display device 520 for displaying an image rendered according to the stylized animatable representation. The display device 520 can also display other information such as training examples, results of nearest neighbor calculations, and other data.
[0107] Alternatively or in addition to the other examples described herein, examples include any combination of the following:
[0108] Clause A, A method for calculating a stylized animatable representation of an object based on a population of stylized animatable representations, the method comprising the steps of:
[0109] Access the true representation of the object;
[0110] Use a model to calculate a mesh mapping, the model being formed using a dataset of training examples that pair true representations of other objects with instances of the ethnic group;
[0111] Apply the mesh mapping to the true representation of the object to produce a target mesh; and
[0112] Select the stylized animatable representation from the ethnic group by evaluating the proximity of the target mesh to instances of the ethnic group. By using the mesh mapping, the target mesh can be effectively produced. Then, the target mesh is used to make a selection from the ethnic group, such as by finding an instance from the ethnic group that is closest to the target mesh according to a similarity metric. The ethnic group can have hundreds or thousands or more instances within the ethnic group. By calculating the target mesh and then making a selection from the ethnic group, a scalable way of creating the stylized representation is given, which scales for web-scale applications such as where personal, stylized avatars are to be created for millions of people.
[0113] Clause B. The method according to clause A, wherein the model is a trained machine learning model. By using a trained machine learning model, generalization according to the training examples is enabled; that is, even if the object is significantly different from the objects depicted in the training examples, the trained machine learning model can generalize according to the training examples and produce an appropriate mesh mapping that enables the portrait of the object to be maintained.
[0114] Clause C. The method according to clause A, wherein the model is used to calculate the mesh mapping using vertex displacements observed in a plurality of the training examples. By using vertex displacements, a highly effective way of creating the stylized representation is given, since the vertex displacements can be applied to the true representation in an efficient manner.
[0115] Clause D. The method according to clause C, wherein the plurality of the training examples are selected from the dataset as the nearest neighbors of the true representation. By using the nearest neighbors of the true representation, efficiency is increased since not all training examples have to be used. Additionally, the nearest neighbors can be calculated in an efficient manner.
[0116] Clause E. The method according to any one of the preceding clauses includes: before selecting the plurality of training examples, calculating a re-topology of the ground truth representation. By calculating the re-topology, the accuracy is improved because the application of the mesh mapping is facilitated. In an example, the re-topology adjusts the total number of vertices of the ground truth representation to match the total number of vertices of the mesh mapping.
[0117] Clause E1. The method according to any one of the preceding clauses includes: calculating the mesh mapping by deriving the mesh mapping according to a spatial distortion.
[0118] Clause F. The method according to any one of the preceding clauses, wherein the mesh mapping is calculated by calculating a distortion for each of the plurality of training examples and then aggregating the distortions, and wherein the plurality of training examples are selected from the dataset as the nearest neighbors of the ground truth representation. By using the distortions, an effective way of applying the mesh mapping is given, which is found to give good results in practice.
[0119] Clause G. The method according to any one of the preceding clauses, wherein the mesh mapping is calculated by calculating a separate distortion for each of the plurality of training examples selected from the dataset as the nearest neighbors of the ground truth representation. Using a plurality of separate distortions enables the use of information from more than one training example in an efficient manner.
[0120] Clause H. The method according to Clause G, wherein applying the mesh mapping includes: applying each of the separate distortions to the ground truth representation individually to obtain a plurality of distorted ground truth representations, and aggregating the plurality of distorted ground truth representations. By aggregating after the distortion, a kernel can be used during the distortion process to give an efficient and effective process.
[0121] Clause I. The method according to any one of the preceding clauses, wherein the dataset includes only dozens of training examples. Using only a dozen training examples is beneficial because obtaining training examples is time-consuming and expensive. Using a large number of training examples also makes the process more difficult to scale. It has been found that the method works well even for a small number of training examples, such as about 20 training examples.
[0122] Clause J. The method according to any one of the preceding clauses, wherein the stylized animatable representation is a rigged model. Using a rigged model facilitates fine-grained, accurately controllable animation.
[0123] Clause K. The method according to any one of the preceding clauses, wherein the truthful representation is a mesh model or an image of an object from which a mesh model is derived. By using the mesh model, the ability to calculate the mesh mapping is facilitated.
[0124] Clause L. The method according to any one of the preceding clauses, wherein the population of stylized animatable representations is a stylistically related population of stylized representations and includes representations of different objects. By using this type of population, it is possible to create stylized representations of different objects while using the same style.
[0125] Clause M. The method according to any one of the preceding clauses, wherein the object is at least a part of a human or an animal. By creating a stylized representation, an avatar of the whole human or animal is obtained for use in one or more downstream processes such as computer games, video conferencing, or the metaverse. In some cases, the stylized representation can be a part of a human or an animal, such as the head and shoulders or the face.
[0126] Clause N. The method according to any one of the preceding clauses, wherein evaluating the proximity of the target mesh to an instance of the population includes: optimizing an energy function including a geometric residual term and an identity prior term. By using optimization to evaluate proximity, a principled and accurate way of selecting the instance of the population is given.
[0127] Clause O. An apparatus for calculating a stylized animatable representation of an object according to a population of stylized animatable representations, the apparatus comprising:
[0128] At least one processor;
[0129] A memory that stores the truthful representation of the object and stores instructions that, when executed by the at least one processor:
[0130] Use a model to calculate a mesh mapping, the model being formed using a dataset of training examples that pair truthful representations of other objects with instances of the population;
[0131] Apply the mesh mapping to the truthful representation of the object to produce a target mesh; and
[0132] Select the stylized animatable representation from the population by evaluating the proximity of the target mesh to an instance of the population.
[0133] Clause P. A method for calculating a stylized animatable 3D representation of an object according to a population of stylized animatable 3D representations, the method comprising the steps of:
[0134] Access the truthful representation of the object;
[0135] Use a model to calculate a mesh mapping, the model being formed using a data set of training examples that pair ground truth representations of other objects with instances of the ethnic group;
[0136] Apply the mesh mapping to the ground truth representation of the object to produce a target 3D mesh; and
[0137] Select the stylized animatable representation from the ethnic group by evaluating the proximity of the target 3D mesh to instances of the ethnic group.
[0138] Clause Q. The method according to Clause P, including rendering an image according to the selected stylized animatable representation. The rendered image can be stored or used in downstream processes, such as being inserted into a video conferencing stream, a video game, or being displayed using a display device.
[0139] Clause R. The method according to Clause P or Clause Q, including: receiving a value of a parameter of the selected stylized animatable representation; applying the received value to the selected stylized animatable representation; and rendering an image according to the selected stylized animatable representation. In this way, detailed animation of an avatar is achieved in an efficient manner that can be implemented in real time.
[0140] Clause S. The method according to any one of Clauses P to R, including adding the image to a video. By adding the image to a video avatar, the video avatar or other content can be added to a game.
[0141] Clause T. The method according to any one of Clauses P to S, wherein the selected stylized animatable representation preserves the likeness of the object depicted in the ground truth representation. Preserving the likeness of the object is very useful as it facilitates human-computer interaction using an animated avatar.
[0142] The term 'computer' or 'computing-based device' is used herein to refer to any device having processing capabilities such that it can execute instructions. Those skilled in the art will recognize that such processing capabilities are incorporated into many different devices, and thus, the terms 'computer' and 'computing-based device' each include personal computers (PCs), servers, mobile phones (including smart phones), tablet computers, set-top boxes, media players, game consoles, personal digital assistants, wearable computers, and many other devices.
[0143] In some examples, the methods described herein are performed by software in machine-readable form on a tangible storage medium (e.g., in the form of a computer program comprising computer program code units) that are adapted to perform all operations of one or more of the methods described herein when the program is run on a computer, and wherein the computer program can be embodied on a computer-readable medium. The software is suitable for execution on a parallel processor or a serial processor such that the method operations can be performed in any suitable order or simultaneously.
[0144] Those skilled in the art will recognize that the storage devices used to store program instructions are optionally distributed across a network. For example, a remote computer can store an example of a process described as software. A local or terminal computer can access the remote computer and download some or all of the software to run the program. Alternatively, the local computer can download fragments of the software as needed, or execute some software instructions at the local terminal and at the remote computer (or computer network). Those skilled in the art will also recognize that all or part of the software instructions can be executed by special-purpose circuitry, such as a digital signal processor (DSP), a programmable logic array, etc., by using conventional techniques known to those skilled in the art.
[0145] As will be apparent to those skilled in the art, any range or device value given herein can be extended or altered without losing the desired effect.
[0146] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
[0147] It will be understood that the above benefits and advantages may relate to one embodiment or may relate to several embodiments. Embodiments are not limited to those that solve any or all of the stated problems or have any or all of the stated benefits and advantages. It should also be understood that references to "a" item refer to one or more of those items.
[0148] The operations of the methods described herein can be performed in any suitable order or, where appropriate, simultaneously. Additionally, individual blocks can be deleted from any of the methods without departing from the scope of the subject matter described herein. Aspects of any of the examples described above can be combined with aspects of any of the other examples described to form further examples without losing the desired effect.
[0149] The term "comprising" is used in this document to mean that the identified method blocks or elements are included, but such blocks or elements do not include an exclusive list, and the method or apparatus may contain additional blocks or elements.
[0150] It should be understood that the above description is given by way of example only, and various modifications can be made by those skilled in the art. The above specification, examples, and data provide a complete description of the structure and use of the exemplary embodiments. Although the various embodiments have been described above with a certain degree of particularity or with reference to one or more individual embodiments, many changes can be made to the disclosed embodiments by those skilled in the art without departing from the scope of this specification.
Claims
1. A method for calculating a stylized animatable representation of an object according to a population of stylized animatable representations, the method comprising: Accessing a ground truth representation of the object; Using a model to calculate a mesh mapping, the model being formed using a data set of training examples that pair ground truth representations of other objects with instances of the population; Applying the mesh mapping to the ground truth representation of the object to produce a target mesh; And Selecting the stylized animatable representation from the population by evaluating the proximity of the target mesh to instances of the population.
2. The method according to claim 1, wherein The model is a trained machine learning model.
3. The method according to claim 1, wherein, The model is used to calculate the mesh mapping using vertex displacements observed in a plurality of the training examples.
4. The method according to claim 3, wherein The plurality of the training examples are selected from the data set as the nearest neighbors of the ground truth representation.
5. The method according to claim 1, comprising: Prior to applying the mesh mapping, a re-topology of the ground truth representation is calculated.
6. The method according to claim 1, comprising: The mesh mapping is calculated by deriving the mesh mapping according to a spatial distortion.
7. The method according to claim 1, wherein The model is used to calculate the mesh mapping by calculating a distortion for each of a plurality of the training examples and then aggregating the distortions, and wherein the plurality of the training examples are selected from the data set as the nearest neighbors of the ground truth representation.
8. The method according to claim 1, wherein The mesh mapping is calculated by calculating a separate distortion for each of a plurality of the training examples selected from the data set as the nearest neighbors of the ground truth representation.
9. The method according to claim 7, wherein Applying the mesh mapping includes: separately applying each of the separate distortions to the ground truth representation to obtain a plurality of distorted ground truth representations, and aggregating the plurality of distorted ground truth representations.
10. The method according to claim 1, wherein, The data set includes only dozens of training examples.
11. The method according to claim 1, wherein, The stylized animatable representation is a rigging model.
12. The method according to claim 1, wherein, The ground truth representation is a mesh model or an image of the object from which a mesh model is derived.
13. The method according to claim 1, wherein, The population of stylized animatable representations is stylistically related stylized representations and includes representations of different objects.
14. The method according to claim 1, wherein, Evaluating the proximity of the target mesh to instances of the population includes: optimizing an energy function including a geometric residual term and an identity prior term.
15. An apparatus for calculating a stylized animatable representation of an object according to a population of stylized animatable representations, the apparatus comprising: A processor; A memory that stores the ground truth representation of the object and stores instructions that, when executed by the processor: Use a model to calculate a mesh mapping, the model being formed using a data set of training examples that pair ground truth representations of other objects with instances of the population; Apply the mesh mapping to the ground truth representation of the object to produce a target mesh; And Select the stylized animatable representation from the population by evaluating the proximity of the target mesh to instances of the population.