Method and device for creating controllable avatar

The method addresses the challenge of creating photorealistic and controllable avatars from sparse monocular recordings by initializing avatars with object image priors, enhancing realism and generalization through 3D Gaussian splattering and neural radiance fields.

JP2025188044APending Publication Date: 2025-12-25TOYOTA JIDOSHA KK +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025098423
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-14
Filing Date
2025-06-12
Publication Date
2025-12-25

AI Technical Summary

Technical Problem

Existing methods struggle to create photorealistic and controllable avatars from sparse monocular recordings, particularly for human heads with limited head rotations and facial expressions, due to the under-constrained nature of the reconstruction problem, leading to difficulties in animating unseen poses and facial expressions.

Method used

A computer-implemented method that initializes a controllable avatar, obtains object image priors from a pre-trained image generation model, and learns parameters based on both the image set and object image priors, using techniques like 3D Gaussian splattering and neural radiance fields, to enhance realism and generalization.

Benefits of technology

The method produces more realistic controllable avatars with improved generalization to unseen poses and facial expressions, even with sparse input, by leveraging object image priors to regularize the distribution of novel views and expressions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025188044000001_ABST
    Figure 2025188044000001_ABST
Patent Text Reader

Abstract

To provide a computer-implemented method of creating a controllable avatar of an animated subject from at least one image set of the animated subject.SOLUTION: A method disclosed herein comprises the steps of initializing a controllable avatar, obtaining at least one subject image prior from a pre-trained image generation model (20), and learning parameters of the controllable avatar based on at least one image set (10) and the at least one subject image prior.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of animation and image reconstruction, and more particularly to devices and computer-implemented methods for creating controllable avatars of animated objects. [Background technology]

[0002] Creating controllable avatars of animated subjects such as articulated objects, body parts, and human or animal heads has been a long-standing challenge in computer vision and graphics. In particular, the ability to render photorealistic dynamic avatars from arbitrary viewpoints enables numerous applications in games, filmmaking, immersive telepresence, augmented reality, and virtual reality. For such applications, it is important to be able to control the avatar, and therefore to generalize well to novel poses and facial expressions.

[0003] Reconstructing a 3D representation that can capture the appearance, geometry, and dynamics of an object, such as a human head, poses a significant challenge for high-fidelity avatar generation. The underconstrained nature of this reconstruction problem significantly complicates the task of achieving a representation that combines the realism of novel-view rendering with pose and facial controllability. Furthermore, in the case of a human head, extreme facial expressions and facial details, such as wrinkles, mouth cavity, and hair, are difficult to capture and can result in visual artifacts that are easily noticeable to humans. Similar challenges arise, mutatis mutandis, for other types of objects.

[0004] Recently, a method was proposed to achieve photorealistic 4D (i.e., three spatial dimensions and time) reconstruction and realistic animation of avatars from multi-view videos captured in a professional studio. However, this capture setup is very demanding, and such rich input is often unavailable. In particular, this method faces challenges when reconstructing 4D avatars from monocular recordings from commodity cameras, such as portrait videos captured with a smartphone. This is in fact an under-constrained problem, as there are few or no observations that can serve as constraints for monitoring novel views. As a result, such methods struggle to estimate holdout views far from the input view when using only monocular RGB video as input. With regard to human heads, this problem becomes even more pronounced when applied to observations of short videos with limited captured head rotations and facial expressions, making it difficult to animate unseen poses and facial expressions. Therefore, new devices and methods are needed to create controllable avatars of animated subjects. The following non-patent literature discloses various methods related to the field of reconstruction from images, animations or human body modeling: [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] Shenhan Qian, Tobias Kirschstein, Liam Schoneveld, Davide Davoli, Simon Giebenhain, Matthias Nieszner. Gaussian Avatars: Photorealistic head avatars with rigged 3D gaussians. arXiv preprint arXiv:2312.02069 (2023). [Non-patent document 2] Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuhler, George Drettakis. 3D Gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics (ToG), 42(4):1–14, 2023. [Non-patent document 3] Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, Ren Ng. NeRF: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021. [Non-patent document 4] Tianye Li, Timo Bolkart, Michael J. Black, Hao Li, Javier Romero. Learning a model of facial shape and expression from 4D scans. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 36(6):194:1–194:17, 2017. [Non-patent document 5] Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, Matthias Nieszner. Face2face: Real-time face capture and reenactment of rgb videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2387–2395, 2016. [Non-patent document 6] Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, Kfir Aberman. 2023. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. pp. 22500–22510. [Non-Patent Document 7] Lvmin Zhang, Anyi Rao, Maneesh Agrawala. 2023. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE / CVF International Conference on Computer Vision. 3836-3847. [Non-patent document 8] Ben Poole, Ajay Jain, Jonathan T Barron, Ben Mildenhall. 2022. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988(2022). [Non-Patent Document 9] Jiaming Song, Chenlin Meng, Stefano Ermon. 2020a. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502 (2020). [Non-Patent Document 10] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer. 2021. High-Resolution Image Synthesis with Latent Diffusion Models. arXiv:2112.10752[cs.CV] [Non-Patent Document 11] Wojciech Zielonka, Timo Bolkart, Justus Thies. 2022. Towards metrical reconstruction of human faces. European Conference on Computer Vision. Springer, 250–269. [Non-Patent Document 12] Diederik P Kingma, Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. [Non-Patent Document 13] Tobias Kirschstein, Shenhan Qian, Simon Giebenhain, Tim Walter, Matthias Nieszner. 2023. Nersemble: Multi-view radiance field reconstruction of human heads. ACM Transactions on Graphics (TOG) 42, 4 (2023), 1–14. Summary of the Invention [Problem to be solved by the invention]

[0006] In this regard, the present disclosure relates to a computer-implemented method for creating a controllable avatar of an animated object from at least one set of images of the animated object, the method comprising: initializing a controllable avatar; Obtaining at least one object image prior from a pre-trained image generation model; and learning parameters of the controllable avatar based on the at least one image set and the at least one object image prior.

[0007] For the sake of brevity, this method will be referred to below as the creation method.

[0008] For brevity herein, and unless otherwise indicated by context, "a," "an," and "the" (e.g., an image set) refer to "at least one" or "each" (e.g., an image set) and are intended to include the plural. Conversely, the general use of the plural may include a singular element.

[0009] An avatar may be a digital representation of a subject's appearance. An avatar is controllable, meaning that the avatar's parameters can be changed to alter the avatar's appearance, as described in more detail below. An avatar may be a three-dimensional avatar. The controllability of an avatar over time may be considered a fourth dimension of the avatar.

[0010] The image set may be a single-view image set or a multi-view image set. A multi-view image set may be an image set in which each element or time step corresponds to multiple images of an object taken simultaneously from multiple viewpoints. Conversely, a single-view image set, also called a monocular image set, may be an image set in which each element or time step corresponds to a single image, with successive images taken from the same viewpoint if the camera is fixed, or from different respective viewpoints if the camera is moving.

[0011] The image set may be composed of multiple consecutive frames of a video, i.e., the image set may be a video. Alternatively, the image set may be multiple non-consecutive frames of a video, for example, selected randomly or regularly with a given sampling frequency (e.g., every third frame of the video). Alternatively, the image set may be one or more still images, such as photographs. Using frames from a video makes it possible to ensure visual consistency within the image set, both with respect to subject matter (hairstyle, accessories, etc.) and context (lighting, environment, etc.).

[0012] The image set may be acquired via an image acquisition module, for example a camera, or may be retrieved as an already acquired image set from a database, for example a local or remote server.

[0013] The step of initializing the controllable avatar may be performed in any manner to provide a controllable avatar for optimization. For example, the initialization step may be performed by associating geometric primitives with a parametrically morphable model of the animated object, e.g., according to [Publication ID: 1]. A geometric primitive is a representation of the local visual information of the avatar. The geometric primitive contains the information necessary to display the avatar, possibly after processing such as image rendering. A geometric primitive may be defined by parameters such as position, orientation, shape, size or scale, color, and opacity. Examples of geometric primitives include point clouds, 3D Gaussians [Publication ID: 2], other three-dimensional probability distributions, or more expressive geometric primitives that can be bent or stretched, for example. Expressiveness refers to the ability of a geometric primitive to convey the appearance of an object.

[0014] Furthermore, the model of the animated object is parametric and morphable. The model is defined by parameters, and the avatar can be controlled by changing the model's parameters, such as pose and facial expression, as well as the shape. Hereinafter, unless otherwise indicated in the context, "model" refers to the parametric and morphable model of the animated object. The model may also represent the overall shape of the animated object.

[0015] For example, the model may be a three-dimensional surface model. In these embodiments, each base element of the model (e.g., a cell of a mesh) is two-dimensional and has a 3D position and orientation relative to other elements of the model.

[0016] When associating multiple geometric primitives with a parametrically morphable model, the geometric primitives are rigged to the model. Thus, the geometric primitives add visual information, including shape information, to the model. In doing so, the geometric primitives together form a dynamic representation. Furthermore, this model allows for efficient and consistent animation of geometric primitives, which would otherwise be very difficult to control.

[0017] As mentioned above, learning or reconstructing the parameters of a controllable avatar based solely on an image set can be an under-constrained problem if the image set is sparse. To overcome this problem, the present construction method provides that the parameters of the controllable avatar are learned not only based on the image set, but also based on at least one object image prior obtained from a pre-trained image generative model.

[0018] An image generation model is a model, e.g., a machine learning model, that can generate an image based on a desired input, such as a text prompt and / or another image. The image generation model is pre-trained, i.e., trained before use in the generation method. Any training method suitable for the type of image generation model may be considered.

[0019] In particular, the image generation model is configured to generate images used as object image priors, i.e., priors related to the object. As commonly used in the field of machine learning, "prior" means that images are generated based on some knowledge or assumption about what the object should look like, and that the images are used to supplement the information in the image set to train a controllable avatar, where such knowledge or assumption is included in the image generation model. The image generation model is not necessarily configured to generate images of an object that are identical to the image of the object for which the avatar is desired. However, the image generation model is configured to generate images of objects of the same type or category (e.g., class) as the desired object, such as a human head, if the controllable avatar is intended to replicate a human head.

[0020] The learning step updates the parameter values ​​to increase the similarity between the controllable avatar and the animated object. The learning step may include multiple iterations.

[0021] Object image priors enable us to regularize the distribution of novel view and facial expression synthesis from a controllable avatar. These priors explicitly guide the controllable avatar (e.g., the images rendered from the controllable avatar) toward a latent manifold of realistic images, thereby enhancing its fidelity and realism. Thus, by learning the parameters of the controllable avatar based on both the image set and the object image priors, we can obtain more realistic controllable avatars with satisfactory generalization ability to unseen poses and facial expressions, even when the input image set of the animated object is sparse.

[0022] Optionally, the method further includes fine-tuning the image generation model to personalize it for the animated object before obtaining the at least one object image prior. The fine-tuning may include training the image generation model further than pre-training. This fine-tuning may provide greater consistency between the image set and the object image priors, improving the learning of the controllable avatar.

[0023] Optionally, the fine-tuning step includes fine-tuning the image generation model on at least one image set, i.e., the image set is used as a basis for further training the image generation model. Therefore, no further input or information about the animated object is required. Therefore, the fine-tuning step can be performed efficiently.

[0024] Optionally, the fine-tuning step includes generating augmented data by the image generation model and fine-tuning the image generation model on the augmented data. Given that the image set may be sparse, it is attractive to obtain additional data for fine-tuning the image generation model. In this regard, the image generation model itself may generate additional data, also known as augmented data, and the image generation model may then be fine-tuned on the additional data. As described in more detail below, the generation of augmented data may take into account additional inputs to complement those that would be provided by the image generation model alone, thereby enabling the image generation model to be fine-tuned based on richer information.

[0025] Fine-tuning the image generation model on the image set and on the augmented data may be combined. Fine-tuning the image generation model on the augmented data allows to avoid the viewpoint bias problem that can arise when the image generation model is fine-tuned only on the image set, i.e., the image generation model tends to generate images primarily from the viewpoints of the image set, which may correspond to a limited subset of all possible viewing angles.

[0026] In some cases, the step of fine-tuning the image generation model may be performed first on the image set and then on the augmented data, so that the image generation model is already somewhat fine-tuned when used to generate the augmented data, although this two-step fine-tuning procedure may also be performed in the reverse order.

[0027] Optionally, the augmented data is generated by conditioning an image generation model based on three-dimensional information of the animated object. The three-dimensional (3D) information may include at least one of a depth map (i.e., a map indicating the depth of points on the object relative to a reference point or plane), a normal map (e.g., a map indicating the direction of vectors normal to the local surface of the object), landmarks (e.g., in the case of a head, a map indicating the 2D / 3D positions of notable features such as the nose, mouth, ears, eyes, etc., or parts thereof, whose placement follows a known 3D pattern), etc. Conditioning an image generation model means constraining the image generation model with additional inputs called conditions. Specifically, conditioning the image generation model based on 3D information causes the image generation model to take the 3D information into account when generating the augmented data. This allows the image generation model to generate more consistent data that matches the 3D information of the object.

[0028] Optionally, the three-dimensional information is obtained from at least one image set, so that no further input or information about the animated object is required beyond that already available to the creation method. The three-dimensional information may be obtained directly or indirectly from the at least one image set, for example, based on the controllable avatar after initialization and / or based on any information derived from the at least one image set.

[0029] Optionally, the augmented data includes at least one set of images, each image set showing reconstructions of the animated object from multiple viewpoints. In other words, the augmented data may include multi-view images as defined above. In such an embodiment, the augmented data improves the controllable avatar's generalization ability to unseen poses of the animated object.

[0030] Optionally, at least one object image prior is used as a pseudo ground truth for the learning step. For example, the learning step may include rendering an image from a controllable avatar and updating parameters of the controllable avatar based on a comparison of the rendered image to the pseudo ground truth image formed by the object image prior.

[0031] Optionally, obtaining at least one object image prior includes rendering at least one new image from the controllable avatar and enhancing the at least one new image with an image generative model. A new image is an image that does not belong to the image set. In these embodiments, the image generative model is configured to take an image as input and provide an enhanced version of the image through its pre-training and, optionally, fine-tuning. In this way, an object image prior that may correspond to the enhanced new image allows for more accurate information about the object to be infused for training the controllable avatar.

[0032] Optionally, the at least one new image shows a new view of the animated object with the same facial expression as an image in the at least one image set. In other words, the new image may show a new perspective of a known facial expression of the animated object, which helps the controllable avatar generalize to new poses. The facial expression refers to the posture of the animated object, i.e., the relative position of each animated part of the object with respect to one another. In the case of a human head, the facial expression may correspond to a facial expression. In the case of a robotic arm, the facial expression may correspond to the pose of the arm segments with respect to one another.

[0033] Alternatively or additionally, the at least one new image exhibits a new facial expression of the animated subject relative to the at least one set of images. The new facial expression is characterized by parameters of the controllable avatar taking values ​​that are different from the values ​​those parameters implicitly have in the set of images. The new facial expression can be viewed from any viewpoint (e.g., a known viewpoint or a new viewpoint in the set of images), or even from multiple viewpoints. Taking into account new images exhibiting new facial expressions increases the animatability of the controllable avatar across multiple poses and facial expressions.

[0034] The at least one new image may include one or more images showing at least one new view and / or one or more images showing at least one new facial expression.

[0035] Optionally, the at least one new image is generated by conditioning the image generation model on three-dimensional information of the animated object, optionally obtained from at least one image set. With regard to the step of fine-tuning the image generation model, the same comments as above apply mutatis mutandis.

[0036] Optionally, at least one object image prior is updated during the training step. By updated, we mean that the object image prior may be partially updated or replaced with a newer version. Because the controllable avatar is refined from one training iteration to a subsequent training iteration, the new version (partial or full) is generally more accurate.

[0037] Optionally, the image generation model includes a diffusion model and / or a text-to-image model. A diffusion model is a machine learning model known per se in the art, which generally takes an input and generates a modified version of the input based on some conditioning. This may be performed by adding some noise to the input and iteratively denoising the noisy input based on the conditioning. Furthermore, a text-to-image model is a machine learning model known per se in the art, which generally takes a text prompt as input and generates an image based on the text prompt. The image generation model may combine both approaches or may be a text-to-image diffusion model. However, other types of image generation models, such as generative adversarial networks (GANs), are also within the scope.

[0038] Optionally, at least one image set is a monocular image set, optionally sampled from a monocular video.

[0039] The present disclosure further relates to a device for creating a controllable avatar from at least one set of images of an animated subject, the device comprising: Initialize the controllable avatar, obtaining at least one object image prior from a pre-trained image generation model; The system is configured to learn parameters of the controllable avatar based on the at least one image set and the at least one object image prior.

[0040] This device, hereinafter referred to as the creation device, may be configured to carry out the above creation method and may have some or all of the above features. This device may have the hardware structure of a computer. More generally, this device may comprise a processor configured to carry out the above steps. Memory may also be provided, if necessary.

[0041] The present disclosure is further directed to a system comprising the above production device provided with a video or image capture module for capturing at least one image set, which may be a camera, for example a monocular camera.

[0042] The present disclosure also relates to a computer program set comprising instructions for performing the steps of the above-described method when the program set is executed by at least one computer, which program set may use any programming language and may be in the form of source code, object code, or intermediate code between source and object code, such as partially compiled, or any other desired form.

[0043] The present disclosure further relates to a storage medium readable by at least one computer and having recorded thereon at least one computer program comprising instructions for executing the steps of the above-described method. The storage medium may be any entity or device capable of storing a program. For example, the storage medium may include a mass storage device such as a hard drive. In general, mass storage devices suitable for physically embodying computer program instructions and data include all forms of non-volatile memory, such as, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks and removable disks, and magneto-optical disks.

[0044] Alternatively, the recording medium may be an integrated circuit in which the program is embedded, the circuit being adapted to perform, or for use in the performance of, the method in question.

[0045] The features, advantages, technical and industrial significance of exemplary embodiments of the present invention will now be described with reference to the accompanying drawings, in which like elements are designated by like reference numerals. [Brief explanation of the drawings]

[0046] [Figure 1] FIG. 1 is a block diagram illustrating a method of creation according to one embodiment. [Figure 2] FIG. 1 is a block diagram illustrating steps for fine-tuning an image generation model according to one embodiment. [Figure 3] 1A-1C illustrate example outputs of the present creation method and comparison method for a cross-reproduction task, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0047] One embodiment of the creation method is described with reference to Figure 1. As described above, the creation method is a computer-implemented method for creating a controllable avatar of an animated object from at least one image set of the animated object, and includes the steps of initializing the controllable avatar, obtaining at least one object image prior from a pre-trained image generation model, and learning parameters of the controllable avatar based on the at least one image set and the at least one object image prior.

[0048] Although FIG. 1 is described in terms of a method, each illustrated block can be considered a corresponding function of a production device.

[0049] The embodiment shown in Figure 1 details the animated object being a human head, but the principles described herein also apply to other animated objects, such as articulated (including flexible) objects, bodies, animals, plants, etc.

[0050] Furthermore, the embodiment shown in FIG. 1 can be considered as a set of monocular images I={I i}, e.g., sampled from a monocular video. While video can be useful for effectively tracking an object and obtaining its model while ensuring visual consistency between successive images, the principles of the construction method described herein do not require video; the image set may be images of an object that are discontinuous in time and / or viewpoint. Furthermore, the proposed construction method is particularly advantageous when the input image set is sparse, e.g., monocular (or more generally, single-view) images, although the image set may also include multi-view images if available.

[0051] As shown in Figure 1, the input to the method is a set of images 10, e.g., video, e.g., a monocular video recording of a human head. The method takes as input a sequence of images I = {I i} (e.g., an RGB image, e.g., a monocular image) as output, we generate a dynamic head avatar (a controllable avatar O={O i The aim is to reconstruct the avatar, which can be rendered from various viewpoints to produce a photorealistic image. To achieve high rendering quality with real-time performance, the head avatar may be represented using 3D Gaussian splattering (Non-Patent Document 2). However, other techniques such as Neural Radiance Fields (Non-Patent Document 3) are also within the scope.

[0052] The controllable avatar may be initialized according to Non-Patent Document 1, as summarized below, although the controllable avatar may be initialized in any other suitable manner. To enable the reconstructed avatar to be animated with various head poses and facial expressions, the controllable avatar may be based on a parametrically morphable head model such as FLAME (Non-Patent Document 4). While the present embodiment details a controllable avatar that relies on a parametrically morphable model that includes a mesh with cells, particularly triangular cells (also referred to herein as triangles), the principles described herein also apply to cells of other shapes (e.g., quadrilaterals, etc.), and more generally to other types of models.

[0053] FLAME mesh M={M i To obtain}, head tracking may be performed, such as monocular head tracking, as known per se in the art. The FLAME mesh may then be associated with a 3D Gaussian to create a controllable avatar, which can be driven to create novel head animations by modifying the pose and facial parameters of FLAME. The controllable avatar O then maps the input image I i and O i The rendered view from

number

[0054] However, reconstructing a believable head avatar that can be rendered into photorealistic novel view and expression images from a small number of observations of a human head with limited viewpoints and expressions in the input is a highly under-constrained problem.

[0055] To this end, the method includes the steps of obtaining at least one object image prior from a pre-trained image generation model 20 and learning parameters of the controllable avatar based on the object image prior. In other words, image priors are extracted from a pre-trained image generation model, as will be described in more detail below, and used to optimize the Gaussian splat. i Rendering a new view from

number

number

number

number

number

[0056] As mentioned above, the image generation model may include a diffusion model and / or a text-to-image model. In one embodiment, the image generation model includes a text-to-image diffusion model that takes an image as input, adds noise to the input, and denoises the noisy image given a text prompt to output a modified image.

[0057] As an example, animatable Gaussian splattering can be constructed according to [1] and [2], as briefly summarized below by way of example.

[0058] In [2], a scene is parameterized using a set of discrete geometric primitives known as 3D Gaussian splats. Each splat is characterized by a covariance matrix Σ centered at a position μ. The covariance matrix needs to be semi-positive definite for physical interpretation. [1] utilizes a parametric ellipsoid definition to obtain the covariance matrix Σ = RSS with a scaling matrix S and a rotation matrix R. T R T These matrices are independently optimized, and the scaling vector

number

number

number

[0059] During rendering, we use a tile-based rasterizer to alpha-blend all 3D Gaussians that overlap a pixel in a tile. To respect visibility order and avoid the expense of per-pixel sorting, we sort splats based on their depth value within each tile before blending.

[0060] To make Gaussian splats animatable for a head avatar, we rig 3D Gaussian splats using a FLAME mesh, following the principles of [1]. First, we connect each triangle of the FLAME identity mesh with a 3D Gaussian, and then transform the 3D Gaussian based on the triangle's deformation over different time steps. While the splats are stationary in the local space of the attached triangle, they can dynamically evolve in the global metric space as the attached triangle is rotated, translated, and scaled.

[0061] For each triangle, the method calculates the average position T of its three vertex coordinates as the origin of local space. The method defines a rotation matrix R that describes the triangle's orientation in global space. This rotation matrix consists of three column vectors derived from the direction vector of one edge, the normal vector of the triangle's face, and their cross product. In addition, the triangle scaling k is determined by calculating the average length of one edge of the triangle and its normal. The 3D Gaussian splat is parameterized by the position μ, rotation r, and anisotropic scaling s in the local space of the parent triangle. During the initialization phase, the position μ may be set to zero, the rotation r may be set as an identity matrix, and the scaling s may be set as a unit vector. During rendering, these properties are converted from local space to global space as follows: r' = Rr, (1) μ'=kRμ+T, (2) s'=ks (3) is converted by

[0062] To capture fine detail, a sufficient number of Gaussians are required for rendering. Therefore, an adaptive density control procedure may be employed that dynamically adds splats based on screen-space position gradients. This densification step occurs in local space, with newly added splats inheriting connectivity from the original splats. This process allows the creation of a new Gaussian for each triangle, improving the model's ability to represent local regions.

[0063] Optionally, to prevent identity drift in the object image priors, the method may include fine-tuning the image generation model 20 to personalize it for the animated object before obtaining the object image priors. In other words, the pre-trained priors are customized to the reconstructed identities, for example, via an innovative fine-tuning procedure that uses only the provided input images (i.e., image set 10).

[0064] An example of a two-stage personalization procedure is shown in FIG. 2, although each stage may be performed independently. In the first step, as shown on the left side of FIG. 2, fine-tuning involves fine-tuning an image generation model 20 on an image set 10. Specifically, a pre-trained text-to-image diffusion model may be fine-tuned using the input image set 10 based on a fine-tuning method such as DreamBooth (Non-Patent Document 6). As shown in FIG. 2, the fine-tuning may use a text prompt to condition the noise removal performed by the image generation model 20. The text prompt may be manually entered to describe the identity of the object to be learned (e.g., "portrait photo of a young Caucasian male") or may be automatically generated via, for example, a recognition algorithm or other computer vision technique run on the image set 10.

[0065] In a second step, which may be performed after the first step, the fine-tuning step includes generating augmented data by the image generation model and fine-tuning the image generation model 20 on the augmented data. As shown in the central part of Figure 2, the augmented data may include at least one set of images 22, for example multi-view images, i.e. images showing reconstructions of the animated object from multiple viewpoints.

[0066] Such augmented data may be generated by conditioning the image generation model 20 based on three-dimensional information of the animated object. The three-dimensional information may include a depth map 24, although other variations are also within the scope, as discussed above. The depth map 24 may be obtained from a controllable avatar and thus indirectly from the input image set 10. However, in a variation, the 3D information may be obtained directly from the image set 10. The 3D information may be provided to a conditioning model 26, such as ControlNet (Non-Patent Document 7), configured to provide additional conditioning to the image generation model 20 based on the 3D information. Note that when generating the augmented data, the parameters of the image generation model 20 may be temporarily fixed.

[0067] The generated augmented data may be used as is and / or combined with the input image set 10 to create a view-augmented dataset, as shown in Figure 2. The pre-trained image generation model 20 may then be further fine-tuned using the augmented dataset, as shown in the right part of Figure 2. This process obtains a personalized object image prior that effectively models the distribution of the reconstructed identity multi-view images, thereby mitigating the view bias problem and improving the personalized image generation model's ability to generate multi-view images.

[0068] Referring again to Fig. 1, the object image prior, whether personalized or not, can be used for animatable Gaussian reconstruction, for example, using the Score Distillation Sampling (SDS) loss [8]. If x denotes a rendered image (e.g., a novel image) from a controllable avatar, we add noise ∈ to the input image x at time step t to generate a noisy image x t and feed it into a denoising network (e.g., an image generation model 20) to remove noise.

number

number

number

[0069] For each training iteration, a random frame i and a random viewpoint φ from the image set 10 may be sampled, and then a random new image i from the controllable avatar

number

number

number

[0070] To accelerate the production rate, the DDIM sampling procedure (Non-Patent Document 9) was adopted.

number

[0071] To facilitate matching of the facial expressions of the generated facial images with the input views, the new images may be generated by conditioning the image generation model 20 on three-dimensional information 30 of the animated object, optionally obtained from at least one image set. Specifically, for example, the FLAME tracking mesh M i Depth map D rendered from i =R(M i , φ) may be used as a conditional guide for potential diffusion.

[0072] Then, z0 is set as the pseudo ground truth

number

[0073] New Image

number

number

[0074] Alternatively or additionally, to increase the degree of animation across poses and expressions, the new images or enhanced versions thereof may show new expressions of the animated subject relative to the image set 10. For example, such new images may be generated by randomly sampling FLAME pose and expression parameters from a physically meaningful space for animating Gaussian splats to generate one or more new expression renderings I. expr Similarly, the corresponding pseudo ground truth

number

number

number

number

number

number

number

number

[0075] As mentioned above, the object image priors are optionally updated during the training step. Since the 2D diffusion model does not have 3D recognition capabilities, the initially generated image object priors

number

number

number

number

number

number

[0076] In subsequent iterations, the updated controllable avatar O provides the multi-view consistent images as the initial input to perform image denoising,

number

[0077] The step of learning the parameters of the controllable avatar itself may be supervised by a loss function that combines one or more differences in input views, novel views, novel expressions, scale regularization, and translation regularization. For example, the following loss function may be used:

number

[0078] position regularization term L pos ensures that the Gaussians stay close to the triangles they are attached to during optimization. The position regularization term regularizes the local position of each Gaussian by

number

[0079] The scale regularization term mitigates the formation of large Gaussians that can lead to jitter problems due to small rotations of the triangle. We regularize the local scale of each Gaussian by:

number

[0080] The scale of the Gaussian is, for example, ∈ scale If σ is less than 0.6, this term is invalid.

[0081] In equation (4), the aforementioned L rec L corresponding to img (I rec ,I) is the splat position L pos and the scale regularization L scale Similarly to the section, we ensure the learning of a controllable avatar based on an image set 10 according to Non-Patent Document 1.

number

number

[0082] The following describes implementation details according to an example. The image generation model 20 can include Stable Diffusion (Non-Patent Document 10). As mentioned above, ControlNet (Non-Patent Document 7) can be used to inject 3D information for facial expression and view control. The guide and conditioning scales can be 7.5 and 1.0, respectively.

[0083] The FLAME mesh can initially be obtained from monocular video using the MICA tracker (Non-Patent Document 11). Animatable Gaussians can be optimized (i.e., controllable avatar parameters can be learned) for, e.g., 30,000 iterations using Adam (Non-Patent Document 12), with learning rates of 5e-5, 1.7e-2, 1e-3, 2.5e-3, and 5e-2 for splat position, scaling coefficient, rotation quaternion, color, and opacity, respectively. If the position gradient is greater than 0.0002, adaptive densification is performed every 1,000 iterations until 20,000 iterations are reached. Gaussians with opacity less than 0.005 are also removed. During training, FLAME parameters may be fine-tuned using learning rates of 1e-6, 1e-5, and 1e-3 for translation, joint rotation, and expression coefficients, respectively.

[0084] To achieve 360-degree capture of the head avatar, novel viewpoints are randomly sampled. The azimuth angle range is [-180°, 180°], the elevation angle range is [-45°, 45°], and the distance between the camera and the head avatar position is 1.0 to 1.5 units. To make the sampled novel expressions physically meaningful, the range of pose and facial expression parameters can be obtained from the training dataset and then sampled within the calculated range.

[0085] The new image and / or image object priors can be rendered at a resolution of 768x768. In equation (4), λ view , λ expr , λ pos , λ scale can be set to 1, 0.2, 0.01, and 1, respectively.

[0086] We conducted experiments on video recordings from the NerSemble dataset (NPL 13), using single-view videos (i.e., image set 10) as input. Figure 3 shows qualitative results for the cross-identity reconstruction task, in which a reconstructed avatar with unseen head poses and facial expressions is driven from motion sequences of different identities. Specifically, column (a) shows the driving frames that determine the facial expressions and viewpoints to be reproduced, column (b) shows the source identity that reconstructs the controllable avatar, column (c) shows the output results of the Gaussian avatar (NPL 1) for the task, and column (d) shows the output of our proposed construction method. As can be seen in the two examples in rows (1) and (2), our proposed method (d) outperforms the Gaussian avatar (c), exhibiting more lifelike facial expressions and more realistic rendering, especially in areas not visible in the source identity images.

[0087] Furthermore, ablation studies conducted by the present inventors have revealed the following. Although personalization helps to prevent the avatar from straying from its original identity, our proposed creation method outperforms prior art methods even without personalization of the image generative model. Our proposed two-stage personalization outperforms one-stage personalization based only on input images without augmentation data. The explicit use of object image priors to learn the parameters of the controllable avatar, especially when further augmented with, for example, a diffusion loss as described above, effectively mitigates the oversaturation problem commonly found in standard score distillation sampling losses. By conditioning the personalization and / or object image prior generation with 3D information, e.g., depth information, it is possible to avoid floating artifacts. Compared to landmark-guided methods, depth maps produce more realistic images around the mouth and back of the head.

[0088] Furthermore, our proposed method proves to be highly robust: we tested varying the number of input images and found that our creation method did not experience any significant performance degradation in quantitative results, even when provided with only eight frames. Our creation method overcomes the limitations of avatar fidelity, even when generated from commodity devices.

[0089] Although the present disclosure refers to certain exemplary embodiments, modifications can be made to these examples without departing from the general scope of the invention as defined by the claims. In particular, individual features of the different embodiments shown / described can be combined into additional embodiments. Accordingly, the description and drawings should be considered in an illustrative rather than a restrictive sense.

Claims

1. 1. A computer-implemented method for creating a controllable avatar of an animated object from at least one image set (10) of said animated object, comprising: initializing the controllable avatar; At least one object image prior from a pre-trained image generation model (20). [Equation 1] obtaining a said at least one image set (10) and said at least one object image prior [Equation 2] and learning parameters of the controllable avatar based on the

2. 2. The method of claim 1, further comprising, prior to obtaining the at least one object image prior, fine-tuning the image generation model (20) to personalize the image generation model (20) for the animated object.

3. The method of claim 2 , wherein the fine-tuning step comprises fine-tuning the image generation model (20) on the at least one image set (10).

4. The fine-tuning step includes: generating augmented data by said image generation model (20); and fine-tuning the image generation model (20) on the augmented data.

5. the augmented data is generated by conditioning the image generation model (20) on three-dimensional information of the animated object; 5. The method of claim 4, wherein optionally, said three-dimensional information is obtained from said at least one image set (10).

6. 5. The method of claim 4, wherein the augmented data includes at least one set of images (22), each image set showing reconstructions of the animated object from multiple viewpoints.

7. the at least one object image prior [Equation 3] The method of any one of claims 1 to 3, wherein is used as a pseudo ground truth for the learning step.

8. At least one object image prior [Equation 4] The method of any one of claims 1 to 3, wherein the step of obtaining comprises the steps of rendering at least one new image from the controllable avatar and augmenting the at least one new image with the image generation model (20).

9. The at least one new image is a new view (I) of the animated object having the same facial expression as an image of the at least one image set (10). view 9. The method of claim 8, wherein the at least one image set (10) is displayed in a manner that indicates at least one of: a new facial expression of the animated object compared to the at least one image set (10);

10. the at least one new image is generated by conditioning the image generation model (20) on three-dimensional information of the animated object; 9. The method of claim 8, wherein optionally, said three-dimensional information is obtained from said at least one image set (10).

11. During the learning step, the at least one object image prior [Equation 5] The method according to any one of claims 1 to 3, wherein:

12. The method according to any one of claims 1 to 3, wherein the image generation model (20) comprises a diffusion model and / or a text-to-image model.

13. The method of any one of claims 1 to 3, wherein said at least one image set (10) is a monocular image set, optionally sampled from a monocular video.

14. A device for creating a controllable avatar of an animated object from at least one set of images (10) of said animated object, comprising: initializing the controllable avatar; At least one object image prior from a pre-trained image generation model (20). [Equation 6] Get said at least one image set (10) and said at least one object image prior [Equation 7] a device configured to learn parameters of the controllable avatar based on the

15. A computer program set comprising instructions for performing the steps of the method of any one of claims 1 to 3 when said computer program set is executed by at least one computer.

16. A recording medium readable by at least one computer and having recorded thereon at least one computer program comprising instructions for carrying out the steps of the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Photorealistic content generation from animated content by neural radiance field diffusion guided by vision-language models

    WO2024164030A2