Markerless 3D motion capture using gaussian-based avatar reconstruction
The 3D Gaussian-based animatable avatar model addresses the limitations of existing markerless motion capture by aligning with RGB images for precise pose estimation and avatar creation, achieving high-fidelity motion capture and photorealistic avatars.
Patent Information
- Application Number
- PCT/EP2024/067801
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2026-01-02
AI Technical Summary
Existing markerless motion capture technologies face challenges in precisely fitting the body shape and appearance of a subject due to sparse keypoint detection, occlusions, and the inability to utilize color information effectively, leading to inaccurate 3D pose estimation and animatable avatar creation.
A method using a 3D Gaussian-based animatable avatar model that includes a skeleton, mesh, and Gaussian layers to align with RGB images, allowing for precise pixel-level pose estimation and simultaneous motion capture and avatar generation through iterative optimization and mesh densification.
Enables precise alignment of the 3D body model at a pixel level, resulting in improved photorealistic animatable avatars with accurate motion capture, capable of handling complex textures and non-flat geometries.
Smart Images

Figure EP2024067801_02012026_PF_FP_ABST
Abstract
Description
[0001] MARKERLESS 3D MOTION CAPTURE USING GAUSSIAN-BASED AVATAR RECONSTRUCTION
[0002] FIELD OF THE INVENTION
[0003] This invention relates to motion capture, particularly using animatable avatars as the body model, which may be used in computer vision applications.
[0004] BACKGROUND
[0005] The concept of markerless motion capture, which may also be referred to as 3D human pose estimation, is schematically illustrated in Figure 1. Markerless motion capture involves estimating the 3D body pose 102 (i.e. the 3D orientation of each articulation of a pre-defined skeleton model) of a human subject for each frame of a video comprising an RGB image 101 of the subject. This is done without requiring markers on the body of the human subject that are detected and used to estimate the 3D body pose of the subject.
[0006] Automated animatable avatar creation involves creating a 3D representation of a person that can be deformed to a given 3D body pose and rendered as an image from any camera viewpoint. The common approach in the industry is to use a mesh to represent avatars and obtaining photorealistic appearance and realistic motion involves tedious manual work from 3D artists.
[0007] Figure 2 schematically illustrates a known method for the estimation of a motion sequence performed by a subject in a multiview RGB video. In this example, frames from multiple view videos 201 are used to separately perform motion capture using human body keypoints at 202, to form a 3D motion sequence 203 for the subject, and for digital avatar creation 204. The 3D motion sequence is then applied to the digital avatar to form an animatable avatar that moves according to the determined 3D motion sequence.
[0008] Recently, automated data-driven approaches have emerged. Figure 3 shows a pipeline of a Human Gaussian Splatting (HuGS) approach (as described in Moreau et al., “Human Gaussian Splatting: Real-time rendering of Animatable Avatars'”, Conference on Computer Vision and Pattern Recognition 2024). Given a 3D body pose as input, the canonical representation (or T pose), shown at 301, is deformed and 3D Gaussians are rendered into RGB images for the different poses, as shown at 302. The deformation model can be trained using RGB images, as shown at 303.
[0009] Markerless motion capture is usually framed as an optimization problem where a human body model is aligned with features from an image. Figures 4(a)-4(c) show an overview of some different 3D human body models and how they can be used for 3D pose estimation.
[0010] In Figure 4(a), a skeleton model is used to estimate a body pose (i.e. a 3D rotation for each body joint). The most common procedure is to align keypoints that represent body joints. In this scenario, keypoints detected in the image are compared to the position of joints from a skeleton model. While this solution has proven to be robust to different scenarios, it presents several disadvantages. Only a sparse set of points is considered, neglecting most of the information present in the image. Furthermore, some keypoints might not be observed in the images due to occlusions, resulting in failures. Body joints (i.e. where articulations happen) are never truly observable in images, such that detecting them is an ambiguous problem where pixel-level precision is not possible to obtain. Alternatively, one can use a parametric mesh template as a body model, as illustrated in Figure 4(b) (see, for example, Loper et al., “SMPL: A Skinned Multi-Person Linear Model”, SIGGRAPH Asia 2015). In addition to the skeleton, a skin surface representation is given by a mesh. It enables to render a mask of the mesh shape for a given body pose and align it with a segmentation mask of the human in the image, providing more signal than keypoints only. This approach is more precise but still limited. Using a binary segmentation mask discards the color information in the image, which can be helpful to estimate the body pose.
[0011] Figure 4(c) shows a Gaussian-based avatar, as mentioned above, which can be animated based on body poses determined using any of the above body models.
[0012] Although the body shape provided by the mesh of these models can be adapted to the morphology of the person through shape parameters, it is not expressive enough to precisely fit the shape of the subject at a pixel- level, preventing RGB alignment with the image.
[0013] It is desirable to develop an approach that may overcome at least some of the above issues.
[0014] SUMMARY OF THE INVENHON
[0015] According to a first aspect, there is provided a method for training a model to estimate a 3D motion sequence of a subject, the method comprising: receiving multiple RGB images depicting a subject in movement; forming a candidate model representing a body of the subject or a part thereof, the candidate model comprising a set of primitives each shaped as 3D Gaussians and having a set of candidate parameters; forming one or more preliminary outputs using the candidate model, the or each preliminary output comprising a respective rasterized RGB image of the subject having a respective estimated 3D pose of the body of the subject or the part thereof; comparing the or each preliminary output with a corresponding RGB image of the received multiple RGB images having the respective estimated 3D pose of the body of the subject or the part thereof to determine a respective image similarity measurement for the or each preliminary output; and updating the set of candidate parameters in dependence on the image similarity measurements) to form a set of updated parameters.
[0016] By using 3D Gaussians in the candidate model, this can preserve the local appearance details when the model is deformed with a novel body pose. Comparing the preliminary outputs with RGB images and updating the model in dependence on a similarity measure can result in a more precise pose estimation.
[0017] The received multiple RGB images may be observations in the RGB space. The received multiple RGB images may be real images. The received multiple RGB images may be 2D RGB images. The multiple RGB images may be from one or more videos captured using one or more cameras depicting a subject performing a movement. This may allow the rendered images produced by the candidate model to be compared with the real RGB images to iteratively update the model.
[0018] The candidate model may be a model of the subject as a 3D Gaussian-based animatable avatar. Using a 3D Gaussian-based animatable avatar as the 3D model of a motion estimation system may allow the animatable avatar and motion estimation to be performed simultaneously, in one go.
[0019] Determining the respective similarity measurement for the or each preliminary output may comprise forming a respective set of RGB colour differences comprising, for each pixel of a respective preliminary output having a given 3D pose, a difference between an RGB value of that pixel and a corresponding pixel of the RGB image having the given estimated 3D pose. Using RGB alignment can enable to align the 3D body model at a pixel level, resulting in a very precise pose and body shape estimation. This may also allow for optimization of the estimated motion sequence such that the rendered images from a 3D Gaussian-based avatar align with observations in the RGB space.
[0020] The method may comprise forming the candidate model by defining one or more 3D Gaussians on the surface of a mesh representing the body of the subject or the part thereof, the mesh comprising multiple triangles. By defining the 3D Gaussians on the surface of a mesh, this can help to preserve the local appearance of details when the model is deformed with a novel body pose.
[0021] The method may further comprise forming a respective rasterized RGB image by deforming a canonical mesh using an estimated 3D pose corresponding to a selected RGB image of the multiple RGB images to form a posed mesh and rotating the posed mesh with a global transformation and / or a 3D translation offset that corresponds to the selected RGB image. This may allow the canonical mesh to be deformed so that a rasterized RGB image can be produced using the candidate version of the model that can be compared with real RGB images.
[0022] Each triangle of the mesh may contain a first 3D Gaussian. The shape of the first 3D Gaussian for a respective triangle may be computed from the vertices of a respective triangle as the Steiner ellipse of the respective triangle. The first 3D Gaussian may be defined on the surface of the mesh. By defining the first 3D Gaussian on the surface of a mesh, this can help to preserve the local appearance of details when the model is deformed with a novel body pose.
[0023] One or more triangles of the mesh may contain one or more further 3D Gaussians on the surface of the mesh. The one or more further 3D Gaussians may represent texture in the respective triangle. The inclusion of one or more texture Gaussians in each triangle, as well as a background Gaussian, can notably increase the image rendering of the Gaussian model, resulting in a higher quality avatar.
[0024] The one or more further 3D Gaussians may each have a learned position represented as barycentric coordinates of the triangle. This may help to preserve the local textures when the triangle shape changes under a novel body pose.
[0025] The position(s) of the one or more further 3D Gaussians may be computed by barycentric interpolation of the coordinates of the vertices of the respective triangle. This may help to preserve the local textures when the triangle shape changes under a novel body pose.
[0026] The one or more further 3D Gaussians may have one or more of a learned scale, a learned orientation, a learned opacity and a learned view-dependent colour. This may allow the parameters to be learned to improve the motion-capture ability of the model and the appearance of images generated by the model.
[0027] The number of further 3D Gaussians per triangle may be pre-determined or iteratively increased during optimization of the model parameters. If pre-determined, this may help to prevent a very large number of Gaussians that would increase the storage size of the model. If iteratively increased, this may help to allocate more Gaussians where more local details (such as complex textures) are observed.
[0028] The one or more 3D Gaussians may be allowed to shift along the direction normal to the triangle by a learned offset. This may help to represent non-flat geometries for body parts such as hair or clothes made out of fur that are difficult to represent with the mesh geometry only. The method may comprise performing the steps of any preceding claim until a predetermined level of convergence is reached. The candidate parameters may be updated to form the updated parameters using the gradient descent principle.
[0029] This may allow the model to be iteratively updated before deployment.
[0030] The method may comprise iteratively increasing the resolution of the mesh. By iteratively increasing the mesh resolution during the optimization process, this may enable the avatar geometry to be represented with a high-resolution mesh, enabling to fit details such as garment wrinkles with a better accuracy. Using a high-resolution mesh can result in a highly detailed avatar with improved photorealism that is suitable for image comparison using, for example, RGB alignment.
[0031] The method may comprise increasing the resolution of the mesh by: splitting an original triangle into three new triangles by introducing a new vertex initially located at the centre of the original triangle; and replacing the first 3D Gaussian from the original triangle with three new 3D Gaussians for each new triangle. This may allow the mesh to more precisely fit the body shape of the subject.
[0032] Each of the one or more further 3D Gaussians may be attributed to the closest new triangle to the respective further 3D Gaussian. This may allow the mesh resolution to be increased and the 3D Gaussians representing texture in the triangles to be assigned to appropriate new triangles.
[0033] The candidate model may comprise a skeleton layer for the subject. Together with the mesh layer and 3D Gaussians layer, this may form the model of the body of the subject or part thereof. This may allow the body shape of the subject to be more realistically represented.
[0034] The set of candidate parameters and the set of updated parameters may each comprise parameters for both body shape and the estimated motion sequence of the subject. This may allow the model to perform both motion estimation and avatar generation.
[0035] The entire pipeline may be differentiable, allowing to back-propagate the image reconstruction error from the similarity measurement to update both body shape and motion sequence parameters. This pipeline can also be combined with other keypoints and segmentation alignment methods.
[0036] The method may further comprise deploying the model having the set of updated parameters as a model for performing motion capture and generating an animatable avatar of the subject. This may allow the model to perform both motion estimation and avatar generation simultaneously.
[0037] According to a further aspect, there is provided one or more computer programs for instructing a computer comprising one or more processors to implement the method above.
[0038] According to a further aspect there is provided a data carrier storing in non-transitory form the one or more computer programs above.
[0039] According to another aspect, there is provided a device comprising one or more processors, the one or more processors being configured to implement a model trained according to the method of any preceding claim, the system being configured to receive multiple RGB images of a subject in movement and output an animatable avatar performing an estimated 3D motion sequence for the subject. After the optimization process, the learned parameters enable to deploy the animatable avatar and the motion sequence. The one or more processors may be configured to generate a video of the animatable avatar performing the estimated 3D motion sequence. This may allow the technique to be used in applications such as movies, video games and augmented and virtual reality.
[0040] The one or more processors may be configured to generate a video of the animatable avatar performing a 3D motion sequence imported from another motion capture operation. This may allow the approach to be used in motion transfer applications.
[0041] The one or more processors may be configured to render a video or image of the animatable avatar performing a movement of the 3D motion sequence observed from any camera pose. This may allow the approach to be used for novel view synthesis.
[0042] According to a further aspect, there is provided a device for training a model to estimate a 3D motion sequence of a subject, the device comprising one or more processors configured to: receive multiple RGB images depicting a subject in movement; form a candidate model representing a body of the subject or a part thereof, the candidate model comprising a set of primitives each shaped as 3D Gaussians and having a set of candidate parameters; form one or more preliminary outputs using the candidate model, the or each preliminary output comprising a respective rasterized RGB image of the subject having a respective estimated 3D pose of the body of the subject or the part thereof; compare the or each preliminary output with a corresponding RGB image of the received multiple RGB images having the respective estimated 3D pose of the body of the subject or the part thereof to determine a respective image similarity measurement for the or each preliminary output; and update the set of candidate parameters in dependence on the image similarity measurement(s) to form a set of updated parameters.
[0043] By using 3D Gaussians in the candidate model, this can preserve the local appearance details when the model us deformed with a novel body pose. Comparing the preliminary outputs with RGB images and updating the model in dependence on a similarity measure can result in a more precise pose estimation.
[0044] BRIEF DESCRIPTION OF THE FIGURES
[0045] Figure 1 schematically illustrates the concept of markerless motion capture.
[0046] Figure 2 schematically illustrates the estimation of a motion sequence performed by a subject and the formation of an animatable avatar in known techniques.
[0047] Figure 3 schematically illustrates a pipeline for the known approach of Human Gaussian Splatting (HuGS).
[0048] Figures 4(a), 4(b) and 4(c) schematically illustrate an overview of known 3D human body models and how they are used for 3D pose estimation.
[0049] Figure 5 schematically illustrates the simultaneous estimation of a motion sequence performed by a subject and the formation of an animatable avatar in embodiments of the present invention.
[0050] Figure 6 schematically illustrates a 3D Gaussians layer defining primitives locally in each triangle of a mesh.
[0051] Figure 7 schematically illustrates the optimization of motion sequence and body shape estimation through RGB alignment.
[0052] Figure 8 schematically illustrates a summary and exemplary implementation for embodiments of the present invention. Figure 9(a) schematically illustrates an implementation example of video processing.
[0053] Figure 9(b) schematically illustrates a further implementation example of video processing.
[0054] Figure 9(c) schematically illustrates a further implementation example of video processing.
[0055] Figure 10 schematically illustrates an example of a method for training a model to estimate a 3D motion sequence of a subject.
[0056] Figure 11 schematically illustrates an example of a device for implementing a model trained as described herein.
[0057] Figure 12 schematically illustrates an example of a device for training a model to estimate a 3D motion sequence of a subject.
[0058] DETAILED DESCRIPTION
[0059] Embodiments of the present invention may simultaneously solve two technical problems: markerless motion capture and automated animatable avatar creation.
[0060] An overview of the approach described herein is illustrated in Figure 5. Based on RGB videos 501 (which are preferably multiview videos), a model 502 can perform markerless 3D motion capture using a Gaussian-based avatar reconstruction. The output of the model is both the estimated 3D motion sequence performed by the subject in the input multi-view RGB videos and a photorealistic and animatable avatar, which can then be displayed to perform the estimated motion sequence, or a motion sequence determined from another capture process.
[0061] Unlike prior approaches, such as those exemplified in Figure 2, the present approach can perform the 3D motion estimation and form the animatable avatar simultaneously.
[0062] In the approach described herein, motion capture is performed by aligning a 3D Gaussian-based avatar with RGB image observations. A hybrid mesh-Gaussians animatable avatar is used as the human body model. The human body model is able to precisely fit the body shape and appearance of the observed subject.
[0063] In one implementation, this model contains three layers that represent different aspects of the human body.
[0064] A first layer is a skeleton representation layer. This layer contains body joints (3D points that represent the articulations of the body) connected to each other in a kinematic tree that defines the connection between joints. This layer is similar to the skeleton model presented in Figure 4(a). This layer is used to animate the body model. Given a 3D body pose (i.e. a 3D rotation for each body joint), the skeleton deformation can be computed by traversing the kinematic tree, resulting in a 3D transformation of each body joint.
[0065] A second layer is a mesh layer. The mesh layer defines the surface of the human body (which may be clothed). The mesh is initialized with a parametric mesh template (see Figure 4(b)) and then iteratively optimized by modifying the position of the vertices of triangles of the mesh to fit the subject’s body shape. This mesh can be animated by Linear Blend Skinning (or any other skinning algorithm). A skinning weight vector is attached to each vertex and defines the relationship between body joint transformations from the skeleton and the transformations of vertices of the triangles. These skinning weights can be imported from a template and eventually optimized. A third layer is a 3D Gaussians layer that represents the appearance of the subject. This layer comprises a set of 3D primitives shaped as one or more 3D Gaussians with optimizable parameters (such as position, scaling, orientation, opacity and viewdependent color). These primitives can be used for photorealistic image rendering via tile-based rasterization (see for example “3D Gaussian Splatting for Real-Time Radiance Field Rendering", Kerbl et al., SIGGRAPH 2023). Each primitive is defined locally in a parent triangle from the mesh layer, such that the 3D Gaussians are rigged to the mesh (i.e. the Gaussians are deformed rigidly with their parent triangle). The one or more 3D Gaussians may be allowed to shift along the direction normal to the triangle by a learned offset.
[0066] Each triangle of the mesh contains a first 3D Gaussian. The first 3D Gaussian represents the background on the triangle in a rendered image. The shape of the first 3D Gaussian for a respective triangle can be computed from the vertices of a respective triangle of the mesh as the Steiner ellipse of the respective triangle (see, for example https: / / en.wikipedia.org / wiki / Steiner_ellipse), which may also be referred to as the Steiner inellipse. The first 3D Gaussian is defined on the surface of the mesh. By defining the 3D Gaussians on the surface of a mesh, this can help to preserve the local appearance of details when the model is deformed with a novel body pose.
[0067] One or more triangles of the mesh may contain one or more further 3D Gaussians on the surface of the mesh. The one or more further 3D Gaussians may represent texture in the respective triangle in a rendered image. The one or more further 3D Gaussians may each have a learned position represented as barycentric coordinates of the triangle.
[0068] Each 3D Gaussian primitive has its own set of parameters. These may include a 3D center position, a 3D orientation, scaling parameters and view-dependent colour, which may be stored as spherical harmonics. The one or more further 3D Gaussians representing texture may have one or more of a learned scale, a learned orientation, a learned opacity and a learned viewdependent colour. The position(s) of the one or more further 3D Gaussians may be computed by barycentric interpolation of the coordinates of the vertices of the respective triangle.
[0069] In other words, each triangle contains one background Gaussian (referred to herein as a first Gaussian). The background Gaussian can be shaped as the Steiner ellipse of the triangle. The shape of the background Gaussian is computed explicitly from the vertices as the Steiner ellipse. The opacities of each of the background Gaussians can be set to 1. The view-dependent color for the background Gaussians can be learned with spherical harmonics.
[0070] One or more triangles of the mesh may also contain one or more further 3D Gaussians representing texture in the triangle (also referred to herein as texture Gaussians) with learned position represented as triangle barycentric coordinates, as well as learned scale, orientations, opacity and view-dependent colors. The number of further 3D Gaussians per triangle may be pre-determined or iteratively increased during optimization of the model parameters. The optional texture Gaussians represent the local appearance on the triangle surface. The inclusion of one or more texture Gaussians in each triangle, as well as a background Gaussian, can notably increase the image rendering of the Gaussian model, resulting in a higher quality avatar.
[0071] The 3D Gaussians layer is depicted in Figure 6. The complete mesh for the body of the subject is illustrated at 601. The mesh approximates the body shape. For each triangle 602 of the mesh, the 3D Gaussians layer defines primitives 603-606 locally in each triangle. The one or more 3D Gaussians 603-606 are defined locally on the triangle. The one or more 3D Gaussians 603- 606 can move rigidly with their parent triangle 602. This arrangement can be used to render a high-fidelity image 607.
[0072] The use of 3D Gaussians for motion capture can enable differentiable photorealistic rendering of the subject, and to use RGB reconstruction of the captured image as one optimization objective (loss function) of the motion capture process. In contrast, previously used body models are generally not expressive enough to be able to reconstruct the images, limiting the optimization process to keypoints of the subject’s body and segmentation alignment.
[0073] Both background and texture Gaussians are located on the surface of the triangle at initialization, but are allowed to shift marginally on the triangle normal direction by a learned offset, enabling to precisely fit the local geometry. By defining the Gaussians on the surface of a mesh, the local appearance of details can be preserved when the model is deformed with a novel body pose. Without the static rigging of Gaussians to triangles, images can be aligned by changing the positions of the Gaussians instead of the body pose, resulting in poor motion capture. In the present approach, Gaussians are rigged to the triangles to overcome such issues.
[0074] In some implementations, an optional mesh densification mechanism may also be used. Parametric mesh templates used at initialization typically have a relatively low resolution, limiting their capacity to fit precisely the body shape of a person. The resolution of the mesh can be incrementally increased during the optimization to obtain a more precise mesh.
[0075] To densify the mesh, a chosen triangle can be split into three novel triangles by introducing a new vertex initially located at the triangle center. The background 3D Gaussian from the original triangle is removed and replaced by three novel background 3D Gaussians for each new triangles, whereas textures 3D Gaussians are each attributed to the closest new triangle.
[0076] Using a high-resolution mesh can result in a highly detailed avatar with improved photorealism that is suitable for RGB alignment.
[0077] The body model presented above is used to perform motion capture in an optimization pipeline. An example will now be described with reference Figure 7, in which optimization of the motion sequence and body shape of the subject is performed through RGB alignment.
[0078] Shown at 701 is the canonical mesh, in which the body is in a ‘T-pose’. The canonical mesh is deformed using the current estimate of the 3D body pose 702 that corresponds to a selected real RGB image 703, which can be extracted from an input video (not shown). In this example, the canonical mesh is deformed with Linear Blend Skinning 704. The obtained initial deformed mesh shown at 705 can be referred to as the ‘posed mesh'.
[0079] The posed mesh 705 is translated and rotated with the global transformation that corresponds to the estimated 3D pose in the selected RGB image 703 to give an updated deformed mesh, shown at 706.
[0080] 3D Gaussians are computed from the deformed mesh 706 and the learned 3D Gaussians local parameters, as shown at 707.
[0081] The 3D Gaussians are rasterized into a rendered RGB image 708 using the corresponding camera pose.
[0082] The rendered RGB image 708 for a given estimated 3D pose of the subject is then compared with the corresponding real RGB image 703 from which the given 3D pose was estimated. A similarity measurement is determined from each rendered RGB image.
[0083] In this example, the similarity measurement is performed by forming a set of RGB colour differences comprising, for each pixel, a difference between an RGB value of that pixel in the rendered RGB image 708 and a corresponding pixel (i.e. at the same position in the image) of the real RGB image 703. This similarity measurement is used as an RGB loss function. The 3D Gaussian positions, vertices positions, global transformation and 3D body pose are then updated. This may be carried out using gradient descent methods.
[0084] The 3D body pose can be stored for each timestep of an input video to give an output motion sequence.
[0085] Therefore, motion capture is performed by aligning a 3D body model comprising a 3D Gaussian-based avatar with 2D RGB image observations. Using a more expressive avatar enables RGB alignment and thus improved precision of the motion capture. Using RGB alignment enables to align the 3D body model at a pixel-level, resulting in a very precise pose estimation, with both body pose and shape optimization. The motion sequence can be optimized such that the rendered images from the avatar align with observations in the RGB space. Improving motion capture can lead to an improved photorealism quality of the animatable avatar.
[0086] Although the examples described herein use RGB alignment between the rasterized RGB image and the corresponding real RGB image to determine the similarity measure, other methods may be used, such as the structural similarity index measure (see, for example, https: / / en.wikipedia.org / wiki / Structural_similarity_index_measure). Other options also exist to compare RGB images and determine a similarity measurement may be used.
[0087] An implementation example of a system that can perform pose and shape optimization with RGB alignment with a hybrid mesh-Gaussians animatable avatar is shown in Figure 8.
[0088] A data capture step is shown generally at 801.This step comprises collecting videos of a subject performing movements. This can be done with a single camera or, more preferably, multiple cameras.
[0089] Videos can then be optionally pre-processed in a data preparation stage, shown at 802, to compute camera poses (if not already known) and / or segmentation masks which can be used to perform background removal.
[0090] A captured RGB image and reference mask are shown at 803 and 804 respectively. Multiple captured RGB image frames are formed from the captured video(s).
[0091] The learnable parameters of the candidate model are initialized, as shown generally at 805. The candidate parameters for the global transformation and body pose are illustrated at 806 and 807 respectively.
[0092] In an exemplary implementation, global translations are initialized with zeros, global rotation with the identity rotation, body poses with 'T-pose' (or canonical pose), the canonical mesh with a parametric template mesh, and non-rigid motion and Gaussians parameters with zeros. Alternatively, body poses can be initialized with the predictions of a human pose regressor model.
[0093] Once the learnable model parameters have been initialised (i.e. to form an initial set of parameters), the optimization process begins. This optimization process aims to estimate the correct learnable model parameters that enable the model to accurately reconstruct the captured images. This is an iterative process where, at each iteration, one or several RGB images, each having a given 3D pose, are selected and the following steps are performed.
[0094] The non-rigid motion offset from the non-rigid motion 808 can be added to the canonical mesh 809. In one implementation, non-rigid motion is a 3D translation offset learned for each vertex and for each frame. The non-rigid motion may be initialised with zeros. The non-rigid motion is useful to model the motion that cannot be explained by the body pose, such as garment wrinkles or hair movement. This optional non-rigid motion 808 could be modelled with different approaches, such as a neural network.
[0095] In this example, the canonical mesh is deformed with Linear Blend Skinning using the current estimate of the 3D body pose that corresponds to the selected RGB image 803. This initial deformed mesh can be referred to as the ‘posed mesh' 810.
[0096] The posed mesh 810 is translated and rotated with the global transformation 806 that corresponds to the selected RGB image 803 to give an updated deformed mesh 811.
[0097] 3D Gaussians are computed from the updated deformed mesh 811 and the learned local parameters of the 3D Gaussians 812 (Gaussian positions 813, Gaussian scaling 314 and Gaussian orientations 815). Gaussian positions 813 are computed by barycentric interpolation of vertices coordinates. Gaussian scaling 814 and orientations 815 are computed using the Steiner ellipses for background Gaussians and learnable parameters for texture Gaussians.
[0098] 3D Gaussians 816 are rasterized into a rendered RGB image 817 using the corresponding camera pose. A mask 818 is also rendered by overriding the Gaussian colors with a uniform color.
[0099] Rendered images 817 and masks 818 are compared to the captured RGB image 803 and the reference segmentation mask 804 computed as described above to compute loss functions, shown at the RGB loss 819 and mask loss 820. Optionally, a keypoints loss could also be computed at this step.
[0100] Additional regularization loss functions can be computed to guide the learnable parameters to a physically plausible solution. In this implementation, the following regularization losses are used: a temporal smoothness loss 821, ensuring that nearby frames have similar motion sequence parameters, a mesh normals loss 822 that forces nearby faces of the canonical mesh to have similar normal directions, and a Gaussian offset loss 823 that enforces Gaussians to remain close to the mesh surface. Regularization losses 821, 822, 823 are summed with the RGB loss 819 and mask loss 820.
[0101] In this example, the gradients of the learnable parameters with respect to the loss function are computed, enabling to update all the learnable parameters with the gradient descent principle.
[0102] After the optimization process, the model with the learned parameters enables to deploy the animatable avatar and the motion sequence. The animatable avatar and the motion sequence can be used for tasks such as motion transfer, novel view synthesis or novel pose synthesis, as indicated at 824.
[0103] As described above, optionally, the mesh may be iteratively densified by selecting new faces in each triangle at 825 based on the Gaussians offset and splitting the initial triangles of the mesh. The densified mesh 826 can then be used in the canonical mesh 809. The resolution of the mesh can be increased based on an empirical measure of the local quality of geometry. A triangle may be densified when the 3D Gaussians attached to it have a large offset with the triangle surface, which may indicate insufficient details in the current mesh geometry.
[0104] The entire training pipeline is differentiable, allowing to back-propagate the image reconstruction error to update both body shape and motion sequence parameters. It should be noted that this pipeline can also be combined with keypoints and segmentation alignment used in previous methods. Performing motion capture on long video sequences can in some cases be challenging for methods based on optimization. Different strategies can be used to learn the motion sequence of a given video. Figures 9(a)-9(c) show three possible embodiments that are compatible with the above-described pipeline.
[0105] The implementation examples shown in Figures 9(a)-9(c) include two main processing steps: avatar optimization and motion tracking. The computation of both of these steps is the optimization process described in Figure 8 (excluding data capture, preparation and deployment). The only difference between avatar optimization and motion tracking is the set of learnable parameters which are being optimized. During avatar optimization, both the motion sequence (global transformation, body pose and non-rigid motion) and the avatar (canonical mesh, 3D gaussians) are being optimized at the same time. During motion tracking, the avatar is already available and only the motion sequences parameters are being optimized to align renderings of the avatar with images.
[0106] Figure 9(a) shows an incremental implementation. In this example, the process starts with the first frame which is used to create the avatar. Then motion tracking is done frame per frame along the time dimension. This approach is effective because the motion between consecutive frames is small. However, the avatar initialized on a single frame may be less accurate.
[0107] Figure 9(b) shows a joint training implementation. In this example, all of the parameters are optimized at the same time using all of the data available. Each optimization step is done for a random timestep for which the avatar is deformed with the estimated motion at this timestep and compared to the corresponding image. This implementation is the easiest to implement, but a good initialization of the motion sequence before starting is desirable. It is most suitable for short video sequences.
[0108] Figure 9(c) illustrates avatar creation from keyframes and sequences tracking. In this implementation, the avatar is first optimized with a set of keyframes, either sampled along the sequence or manually chosen. This step is similar to the joint training implementation of Figure (9) except that keyframes only are used, leading to a more detailed avatar and compatibility with long sequences. In a second step, the remaining frames are tracked sequence by sequence. The sequences are defined by the set of frames between two keyframes. The motion parameters of the entire sequence are trained jointly, enabling to enforce temporal smoothness.
[0109] A model trained as described above can be implanted to output an animatable avatar performing an estimated 3D motion sequence for the subject, based on input RGB images of the subject in movement (which may be derived from one or more videos of the subject).
[0110] A video of the animatable avatar performing the estimated 3D motion sequence may be generated. The model may therefore be used for motion transfer. Alternatively, a video of the animatable avatar performing a 3D motion sequence imported from another motion capture operation may be generated. The model may therefore be used for novel pose synthesis.
[0111] The model may also be used to render a video or image of the animatable avatar performing a movement of the 3D motion sequence observed from any camera pose. This may allow the model to be used for novel view synthesis.
[0112] The approach may also be applied to parts of the body of the subject, such as the head. For example, the method may be applied to face data to simultaneously track the head 3D pose and facial expressions using the model. This can be performed where it is desirable to create a Gaussian-based avatar driveable with novel expressions.
[0113] Figure 10 shows an example of a method for training a model to estimate a 3D motion sequence of a subject. At step 1001, the method comprises receiving multiple RGB images depicting a subject in movement. At step 1002, the method comprises forming a candidate model representing a body of the subject or a part thereof, the candidate model comprising a set of primitives each shaped as 3D Gaussians and having a set of candidate parameters. At step 1003, the method comprises forming one or more preliminary outputs using the candidate model, the or each preliminary output comprising a respective rasterized RGB image of the subject having a respective estimated 3D pose of the body of the subject or the part thereof. At step 1004, the method comprises comparing the or each preliminary output with a corresponding RGB image of the received multiple RGB images having the respective estimated 3D pose of the body of the subject or the part thereof to determine a respective image similarity measurement for the or each preliminary output. At step 1005, the method comprises updating the set of candidate parameters in dependence on the image similarity measurements) to form a set of updated parameters.
[0114] Figure 11 shows an example of a device 1100 configured to implement a model trained using the methods described herein. The device 1100 comprises a processor 1101 and a memory 1102. The memory 1102 stores in a non-transient way code that is executable by the processor 1101 to implement the respective entity in the manner described herein. The device 1100 may be implemented by hardware or may be service-based computing device, for example it may be implemented as a cloud-based computing device.
[0115] Figure 12 shows an example of a device 1200 for training a model to estimate a 3D motion sequence of a subject. The device 1200 may be configured to implement the methods described herein. The device 1200 comprises a processor 1201 and a memory 1202. The memory 1202 stores in a non-transient way code that is executable by the processor 1201 to implement the respective entity in the manner described herein. The device 1200 may be implemented by hardware or may be service-based computing device, for example it may be implemented as a cloud-based computing device. Therefore, the method may be deployed in multiple ways, for example in the cloud, on the device, or in dedicated hardware.
[0116] The devices 1100, 1200 may in some implementations also comprise a transceiver that is capable of communicating over a network with other entities. For example, the device may receive videos and / or images from other entities. Those entities may be physically remote from the device 1100, 1200. The network may be a publicly accessible network such as the internet. The entities may in some cases be based in the cloud. These entities may be logical entities. In practice they may each be provided by one or more physical devices such as servers and data stores, and the functions of two or more of the entities may be provided by a single physical device. Each physical device implementing an entity comprises a processor and a memory.
[0117] By using a hybrid mesh-Gaussians animatable avatar, defining the 3D Gaussians on the surface of a mesh, the local appearance of details can be preserved when the model is deformed with a novel body pose. The approach described herein, where the Gaussians can be deformed rigidly with their parent mesh triangle, can allow for improved motion capture. Without the static rigging of Gaussians to triangles, images can be aligned by changing the Gaussians positions instead of the body pose, which may result in failure of the motion capture.
[0118] Prior methods that estimate the pose with keypoints alignment are sensitive to occlusions and to failures of the keypoint detector and neglect most of the information present in the image. Using a 3G Gaussians body model with a similarity measurement, such as RGB alignment, for body pose and body shape optimization enables to align the 3D body model at a pixel level, resulting in a very precise pose estimation. In contrast, RGB alignment with a standard mesh model may fail because of the lack of photorealism.
[0119] Furthermore, by using a mesh densification mechanism to iteratively increase the mesh resolution during the optimization process, the avatar geometry can be represented with a high-resolution mesh, enabling further improvement when fitting fine details, such as garment wrinkles, with improved accuracy. The applicant hereby discloses in isolation each individual feature described herein and any combination of two or more such features, to the extent that such features or combinations are capable of being carried out based on the present specification as a whole in the light of the common general knowledge of a person skilled in the art, irrespective of whether such features or combinations of features solve any problems disclosed herein, and without limitation to the scope of the claims. The applicant indicates that aspects of the present invention may consist of any such individual feature or combination of features. In view of the foregoing description it will be evident to a person skilled in the art that various modifications may be made within the scope of the invention.
Claims
CLAIMS1. A method (1000) for training a model to estimate a 3D motion sequence of a subject, the method comprising: receiving (1001) multiple RGB images (803) depicting a subject in movement; forming (1002) a candidate model representing a body of the subject or a part thereof, the candidate model comprising a set of primitives each shaped as 3D Gaussians and having a set of candidate parameters; forming (1003) one or more preliminary outputs using the candidate model, the or each preliminary output comprising a respective rasterized RGB image (817) of the subject having a respective estimated 3D pose of the body of the subject or the part thereof; comparing (1004) the or each preliminary output with a corresponding RGB image (803) of the received multiple RGB images having the respective estimated 3D pose of the body of the subject or the part thereof to determine a respective image similarity measurement for the or each preliminary output; and updating (1005) the set of candidate parameters in dependence on the image similarity measurements) to form a set of updated parameters.
2. The method as claimed in claim 1, wherein the candidate model is a model of the subject as a 3D Gaussian-based animatable avatar.
3. The method as claimed in claim 1 or claim 2, wherein determining the respective similarity measurement for the or each preliminary output comprises forming a respective set of RGB colour differences comprising, for each pixel of a respective preliminary output having a given 3D pose, a difference between an RGB value of that pixel and a corresponding pixel of the RGB image (803) having the given estimated 3D pose.
4. The method as claimed in any preceding claim, wherein the method comprises forming the candidate model by defining one or more 3D Gaussians (603, 604, 605, 606) on the surface of a mesh (601) representing the body of the subject or the part thereof, the mesh comprising multiple triangles (602).
5. The method as claimed in claim 4, wherein the method further comprises forming a respective rasterized RGB image (817) by deforming a canonical mesh using an estimated 3D pose corresponding to a selected RGB image of the multiple RGB images to form a posed mesh and rotating the posed mesh with a global transformation and / or a 3D translation offset that corresponds to the selected RGB image.
6. The method as claimed in claim 4 or claim 5, wherein each triangle (602) of the mesh (601) contains a first 3D Gaussian (603), wherein the shape of the first 3D Gaussian for a respective triangle is computed from the vertices of a respective triangle as the Steiner ellipse of the respective triangle.
7. The method as claimed in claim 6, wherein one or more triangles (602) of the mesh (601) contain one or more further 3D Gaussians (604, 605, 606) on the surface of the mesh representing texture in the respective triangle.
8. The method as claimed in claim 7, wherein the one or more further 3D Gaussians each have a learned position represented as barycentric coordinates of the triangle.
9. The method as claimed in claim 8, wherein the positions) of the one or more further 3D Gaussians are computed by barycentric interpolation of the coordinates of the vertices of the respective triangle.
10. The method as claimed in any of claims 7 to 9, wherein the one or more further 3D Gaussians have one or more of a learned scale, a learned orientation, a learned opacity and a learned view-dependent colour.
11. The method as claimed in any of claims 7 to 10, wherein the number of further 3D Gaussians per triangle is pre-determined or iteratively increased during optimization.
12. The method as claimed in claim any of claims 4 to 11, wherein the one or more 3D Gaussians are allowed to shift along the direction normal to the triangle by a learned offset.
13. The method as claimed in any preceding claim, wherein the method comprises performing the steps of any preceding claim until a predetermined level of convergence is reached.
14. The method as claimed in any of claims 4 to 13, wherein the method comprises iteratively increasing the resolution of the mesh.
15. The method as claimed in claim 14 as dependent on any of claim 6 to 11 or on any of claims 12 to 14 as dependent on any of claims 6 to 11 , the method comprising increasing the resolution of the mesh by: splitting an original triangle into three new triangles by introducing a new vertex initially located at the centre of the original triangle; and replacing the first 3D Gaussian from the original triangle with three new 3D Gaussians for each new triangle.
16. The method as claimed in claim 15 as dependent on any of claims 7 to 11 or on any of claims 12 to 15 as dependent on any of claims 7 to 11, wherein each of the one or more further 3D Gaussians is attributed to the closest new triangle to the respective further 3D Gaussian.
17. The method as claimed in any preceding claim, wherein the candidate model comprises a skeleton layer for the subject.
18. The method as claimed in any preceding claim, wherein the set of candidate parameters and the set of updated parameters each comprise parameters for both body shape and the estimated motion sequence of the subject.
19. The method as claimed in any preceding claim, wherein the method further comprises deploying the model having the set of updated parameters as a model for performing motion capture and generating an animatable avatar of the subject.
20. A device (1100) comprising one or more processors (1101), the one or more processors being configured to implement a model trained according to the method (1000) of any preceding claim, the system being configured to receive multiple RGB images of a subject in movement and output an animatable avatar performing an estimated 3D motion sequence for the subject.
21. The device (1100) as claimed in claim 20, wherein the one or more processors (1101) are configured to generate a video of the animatable avatar performing the estimated 3D motion sequence.
22. The device (1100) as claimed in claim 20, wherein the one or more processors (1101) are configured to generate a video of the animatable avatar performing a 3D motion sequence imported from another motion capture operation.
23. The device (1100) as claimed in any of claims 20 to 22, wherein the one or more processors (1101) are configured to render a video or image of the animatable avatar performing a movement of the 3D motion sequence observed from any camera pose.
24. A device (1200) for training a model to estimate a 3D motion sequence of a subject, the device comprising one or more processors (1201) configured to : receive (1001) multiple RGB images (803) depicting a subject in movement; form (1002) a candidate model representing a body of the subject or a part thereof, the candidate model comprising a set of primitives each shaped as 3D Gaussians and having a set of candidate parameters; form (1003) one or more preliminary outputs using the candidate model, the or each preliminary output comprising a respective rasterized RGB image (817) of the subject having a respective estimated 3D pose of the body of the subject or the part thereof; compare (1004) the or each preliminary output with a corresponding RGB image (803) of the received multiple RGB images having the respective estimated 3D pose of the body of the subject or the part thereof to determine a respective image similarity measurement for the or each preliminary output; and update (1005) the set of candidate parameters in dependence on the image similarity measurement(s) to form a set of updated parameters.