Apparatus and method for reconstructing a three-dimensional conversational model for a subject

The 3D conversational head model integrates audio and visual features using 3D Gaussians to create a lifelike, interactive avatar with synchronized facial and body gestures, addressing the limitations of existing technologies in dynamic 3D avatar reconstruction.

WO2026082267A1PCT designated stage Publication Date: 2026-04-23HUAWEI TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2024-10-14
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Current avatar reconstruction technologies fail to create lifelike 3D avatars that can dynamically change angles, move in 3D space, and synchronize facial expressions with body language, lacking full control over body and head movements, and do not support text and audio language capabilities.

Method used

A method and device for reconstructing a 3D conversational head model using 3D Gaussians, integrating audio and visual features to modify opacity and color, and integrating with a body model for seamless facial and body gesture synchronization, driven by a Large Language Model for real-time interaction.

Benefits of technology

Enables a photorealistic, interactive 3D avatar capable of real-time facial expression and body movement synchronization, supporting conversational capabilities and immersive user interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024078913_23042026_PF_FP_ABST
    Figure EP2024078913_23042026_PF_FP_ABST
Patent Text Reader

Abstract

A method (1300) for reconstructing a model of a 3D conversational head of a subject, comprising: receiving (1301) images depicting a head and audio data corresponding to speech; extracting (1302) visual features and audio features; fusing (1303) the features to form multiple fused audio features; inputting (1304) the fused features to a candidate model defining a set of primitives shaped as 3D Gaussians; for each 3D Gaussian: in dependence on the fused features, modifying (1305) the opacity and colour of the 3D Gaussian to form an audio-dependent opacity and colour; using the audio-dependent opacity and colour of the respective 3D Gaussian, forming (1306) preliminary outputs for comparing (1307) with a corresponding ground-truth image to determine a respective similarity measurement; and updating (1308) parameters of the candidate model. This may result in a photorealistic 3D conversational head model that can respond to user inputs to animate expressions of a 3D avatar.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] APPARATUS AND METHOD FOR RECONSTRUCTING A THREE-DIMENSIONAL CONVERSATIONAL

[0002] MODEL FOR A SUBJECT

[0003] FIELD OF THE INVENTION

[0004] This invention relates to three-dimensional models, for example that may be used to create animatable avatars, which may be used in computer vision applications.

[0005] BACKGROUND

[0006] Capturing, reconstructing and animating virtual humans that look and behave realistically is central for a number of applications, such as videogames, simulation robotics, telepresence, augmented reality (AR), virtual reality (VR) and content generation. As virtual interactions become increasingly prevalent, there is a growing need for more lifelike and engaging human-computer interaction.

[0007] There are several approaches in the state-of-the-art that aim to solve a technical problem of an animatable avatar reconstruction. However, there is no current solution that can address the problem holistically by reconstructing the whole 3D human body, enabling text and audio language capabilities and enabling full control of body and head, together with free camera viewpoint navigation.

[0008] Currently, avatars generally exist in the following forms: 2D avatars with, for example, diffusion or generative adversarial network (GAN) backbones, 3D audio-driven head models and 3D full body avatars.

[0009] Existing 2D avatars provide only 2D reconstruction (diffusion based or with GAN backbones). Since these avatars are fundamentally 2D, they cannot fully capture the depth, spatial orientation, or perspective of a real 3D human. This limits their realism in dynamic scenes where the avatar needs to change angles or move in 3D space (for example, walking around, turning their head). The illusion of depth is either artificially created or omitted entirely, making these avatars less immersive in scenarios that require 3D interaction, such as VR, AR or video games with 3D environments.

[0010] 2D avatars may also have static or limited movement. Although diffusion models can generate high-quality images, they often struggle with producing smooth, coherent motion over time. Most 2D avatar systems, including Synthesia and VASA- 1 , rely on pre-recorded or heavily constrained animations, which can result in unnatural movement or overly simplistic gestures. The lack of fluid body and facial movements reduces the overall believability of the avatar.

[0011] 3D audio-driven head models provide parts-based reconstruction or control interface (e.g. face only, body-only). However, they are not always realistic in generating facial expressions based on speech.

[0012] Existing 3D full-body avatars do not generally provide language capabilities (i.e. speech re-enactment comes from motion capture (MoCap) transfers or real audio). Most existing 3D full-body avatars do not have any speech capabilities. They have no mechanism for the avatar to generate appropriate speech-related animations or gestures in response to the user's inputs. Avatars based on pre-recorded MoCap transfers or real audio (rather than spontaneously generated audio) are limited in their ability to be customized for different users or scenarios. The body movements and speech re-enactment are tied to specific motion capture recordings, making it difficult to personalize the avatar’s gestures and facial expressions for individual users. They also generally have limited facial and body gesture coordination. While some 3D full-body avatars can animate facial expressions and body gestures independently, synchronizing them to create a cohesive and natural appearance is challenging. For example, full-body avatars may struggle to align facial expressions with body language in a way that looks natural, such as nodding while smiling or turning the head during speech.

[0013] It is desirable to develop an approach that may overcome at least some of the above issues.

[0014] SUMMARY OF THE INVENTION

[0015] According to a first aspect, there is provided a method for reconstructing a model of a 3D conversational head of a subject, the method comprising: receiving one or more images depicting a head of a subject and audio data corresponding to speech by the subject; extracting visual features and audio features from the one or more images and the audio data respectively; fusing the visual features and the audio features to form multiple fused audio features; inputting the multiple fused audio features to a candidate model representing a 3D head of the subject, the candidate model defining a set of primitives shaped as 3D Gaussians, wherein each 3D Gaussian has an opacity and a colour; for each 3D Gaussian: in dependence on the multiple fused audio features, modifying the opacity and colour of the respective 3D Gaussian to form an audio-dependent opacity and colour for the respective 3D Gaussian; using the audio-dependent opacity and colour of the respective 3D Gaussian, forming one or more preliminary outputs; comparing the or each preliminary output with a corresponding ground-truth RGB image to determine a respective image similarity measurement for the or each preliminary output; and updating parameters of the candidate model in dependence on the image similarity measurement(s).

[0016] This may result in a photorealistic 3D conversational head model for the subject that can respond to user input prompts to animate facial expressions of a 3D avatar of the subject.

[0017] The method may comprise, for each 3D Gaussian: projecting the multiple fused audio features onto a learned audio- visual basis space by multiplying the multiple fused audio features by a projection vector to form a set of fused audio features in the audiovisual basis space; from the set of fused audio features in the audio-visual basis space, forming a frame- specific feature; inputting the frame- specific feature to an opacity modifier to modify the opacity and colour of the respective 3D Gaussian to form the audio-dependent opacity and colour for the respective 3D Gaussian. This may allow the opacity and colour of the 3D Gaussian to be modified appropriately for a specific image frame.

[0018] The received audio data corresponding to speech by the subject may correspond to the received one or more images depicting the head of the subject. For example, the one or more images may be image frames of a video and the audio data may be corresponding audio data of the video. Each received image frame may have a corresponding part of the received audio data. Visual features may be extracted from a respective image frame and audio features may be extracted from the audio data (for example, the part of the audio data) corresponding to the respective image frame. The visual features extracted from the respective image frame and the audio features extracted from the corresponding audio data may be fused to form the fused audio features. The fused audio features corresponding to a respective image frame may be projected onto the learned audiovisual basis space. The frame- specific feature may correspond to the respective image frame. The ground-truth image compared with a respective preliminary output formed in dependence on the fused audio features for a respective image frame may be the respective image frame from which the visual features were extracted. The above method may be performed for each received image and its corresponding audio data. The respective 3D Gaussian may have an audio- visual feature basis that is blended via the multiple fused audio features to obtain the frame-specific feature. This may condition the model to perform fast animations via modifications to the Gaussians’ opacities and colours.

[0019] The frame- specific feature may be input to the opacity modifier alongside the position of the respective 3D Gaussian to obtain the audio-dependent opacity and colour for the respective 3D Gaussian. This may allow the respective 3D Gaussian to be rendered accurately in dependence on received audio.

[0020] The method may comprise receiving a set of camera poses. The one or more images may comprise one or more image frames per camera pose. This may allow multiple images to be processed to provide input for reconstructing the model for a particular camera pose.

[0021] The visual features may comprise one or more eye features extracted from the input image(s) and / or face landmarks (which may be provided as input, and the method may comprise receiving the face landmarks) for the subject. This may allow facial features of the subject to be reconstructed.

[0022] The face landmarks may comprise 2D face keypoints of one or more of the eyes, nose and chin. This may allow for more accurate reconstruction of facial features.

[0023] The method may further comprise cropping the one or more images depicting the head of the subject and / or removing background from the one or more images. This may allow irrelevant areas to be removed from the images.

[0024] The one or more preliminary outputs may be formed using a Gaussian rasterizer. The audio-dependent colour and opacity of the respective 3D Gaussian may be input to the Gaussian rasterizer to render the one or more preliminary outputs. This may allow a photorealistic image to be appropriately rendered in dependence on audio features.

[0025] The candidate model may be a neural network. This may be a convenient implementation for reconstructing the head model and updating the parameters of the candidate model.

[0026] The method may comprise integrating a 3D conversational head model having the updated set of parameters with a body model, and / or with a scene (for example, containing an image of a body). The head model and the body model and / or scene may be integrated to create a full-bodied avatar. This may allow for a fully integrated head and body model of a subject that can be used as a 3D avatar.

[0027] The 3D conversational head model and the body model may be separate models. This may allow the models to be reconstructed separately and then integrated together to be used during inference. This may be a convenient implementation.

[0028] The body model may comprise one or more body 3D Gaussians. The method may comprise: extracting one or more positions of the head of the subject (which may indicate a border of the head); and merging one or more 3D Gaussians of the 3D conversational head model and one or more body 3D Gaussians of the body model in dependence on the extracted head positions). This may allow the head and body models to be merged in a realistic manner.

[0029] The method may comprise performing border-specific Gaussian densification for a border between the 3D conversational head model and the body model. This may allow for a more photorealistic 3D avatar. The method may comprise performing joint harmonization for 3D Gaussians of the 3D conversational head model and the body model to homogenize colours of the head and body of the subject. This may allow for a more photorealistic 3D avatar.

[0030] The method may comprise purging (i.e. removing) non-desired 3D Gaussians located at the head of the subject using a spherical pruning strategy. This may all the interface between the head and the body models to be optimized.

[0031] The candidate parameters may be updated to form the updated parameters using the gradient descent principle. This may allow the model to be iteratively updated before deployment.

[0032] The method may comprise outputting a three-dimensional conversational head model having the updated set of parameters. The output model may be used for inference, for example as part of a 3D conversational avatar for the subject.

[0033] According to a second aspect, there is provided a device comprising one or more processors, the one or more processors being configured to output and / or implement a model reconstructed according to the method of any preceding claim, the device being configured to: receive one or more images depicting a head of a subject; receive audio data corresponding to speech by the subject; and output a three-dimensional conversational head model for the subject. The device may form the three dimensional head model of the subject using the reconstruction method having any of the features described above.

[0034] This may result in a photorealistic 3D conversational head model for the subject that can respond to user input prompts to animate facial expressions of a 3D avatar of the subject.

[0035] The one or more processors may be configured to render a video of the 3D conversational head performing movement corresponding to an input speech segment or input text prompt. This may allow the conversational head to be used as a 3D avatar.

[0036] According to a third aspect, there is provided a device for reconstructing a model of a 3D conversational head of a subject, the device comprising one or more processors configured to: receive one or more images depicting a head of a subject and audio data corresponding to speech by the subject; extract visual features and audio features from the one or more images and the audio data respectively; fuse the visual features and the audio features to form multiple fused audio features; input the multiple fused audio features to a candidate model representing a 3D head of the subject, the candidate model defining a set of primitives shaped as 3D Gaussians, wherein each 3D Gaussian has an opacity and a colour; for each 3D Gaussian: in dependence on the multiple fused audio features, modify the opacity and colour of the respective 3D Gaussian to form an audio-dependent opacity and colour for the respective 3D Gaussian; using the audio-dependent opacity and colour of the respective 3D Gaussian, form one or more preliminary outputs; comparing the or each preliminary output with a corresponding ground-truth RGB image to determine a respective image similarity measurement for the or each preliminary output; and update parameters of the candidate model in dependence on the image similarity measurement(s).

[0037] This may result in a photorealistic 3D conversational head model for the subject that can respond to user input prompts to animate facial expressions of a 3D avatar of the subject.

[0038] By using 3D Gaussians in the candidate model, this can preserve the local appearance details when the model is deformed. Comparing the preliminary outputs with RGB images and updating the model in dependence on a similarity measure can result in a more precise head model. The received images may be observations in the RGB space. The received images may be real images. The received images may be 2D RGB images. The images may be from one or more videos captured using one or more cameras. This may allow the rendered images produced by the candidate model to be compared with the real RGB images to iteratively update the model.

[0039] Determining the respective similarity measurement for the or each preliminary output may comprise forming a respective set of RGB colour differences comprising, for each pixel of a respective preliminary output (for example, having a given 3D pose), a difference between an RGB value of that pixel and a corresponding pixel of the ground truth image. Using RGB alignment can enable to align the model at a pixel level. This may also allow for optimization of the estimated motion sequence such that the rendered images from a 3D Gaussian-based avatar align with observations in the RGB space.

[0040] The body model may comprise a skeleton layer for the subject. Together with the 3D Gaussians, this may form the model of the body of the subject or part thereof. This may allow the body shape of the subject to be more realistically represented.

[0041] According to a further aspect, there is provided one or more computer programs for instructing a computer comprising one or more processors to implement the method above.

[0042] According to a further aspect there is provided a data carrier storing in non-transitory form the one or more computer programs above.

[0043] BRIEF DESCRIPTION OF THE FIGURES

[0044] Figure 1 schematically illustrates an exemplary pipeline of an animatable 3D head model reconstruction process.

[0045] Figure 2 schematically illustrates a general block diagram showing an exemplary pipeline of a dynamic 3D body reconstruction process.

[0046] Figure 3 schematically illustrates an exemplary implementation of a pipeline for the dynamic 3D body reconstruction process

[0047] Figure 4 schematically illustrates an exemplary pipeline of the head-body integration process.

[0048] Figure 5 schematically illustrates an example of the head-body integration process.

[0049] Figure 6 schematically illustrates a general overview of an example of a body pose reanimation process that may be used in the present approach where a skeleton model is used to drive the movements of the human.

[0050] Figure 7 schematically illustrates an exemplary pipeline of the body pose reanimation process.

[0051] Figure 8 schematically illustrates an exemplary pipeline of a large language model-driven speech re-enactment process.

[0052] Figure 9 shows an example of a real-time interactive viewer for the animatable avatar.

[0053] Figure 10 shows further examples of a real-time interactive viewer. Figures 11(a), 11(b) and 11(c) show examples of a real-time interactive viewer illustrating head-body integration. Figure 11(a) shows the dynamic 3D full-body model (with face Gaussians pruned using a head-body integration method). Figure 11(b) shows an animatable 3D head model with audio-driven facial expression. Figure 11(c) show a fully integrated model combining 3D head and body reconstructions, animated and responding to user's inputs in real-time.

[0054] Figure 12 schematically illustrates an overview of exemplary components of a real-time interactive conversational avatar viewer system.

[0055] Figure 13 schematically illustrates an example of a method for reconstructing a model of a 3D conversational head of a subject.

[0056] Figure 14 schematically illustrates an example of a device for implementing a model reconstructed as described herein.

[0057] Figure 15 schematically illustrates an example of a device for reconstructing a model of a 3D conversational head of a subject.

[0058] DETAILED DESCRIPTION

[0059] Embodiments of the present invention may allow for the reconstruction and rendering of 3D photorealistic avatar which is interactive, easy to control and has conversational capabilities. This may address problems such as face expression control via audio signals (speech), automatic expressions driven by a Large Language Model (LLM), photorealistic and real-time capabilities (i.e. in current consumer-grade graphics processing units (GPUs)), full-body dynamic 3D human reconstruction and body-pose control via skeleton-guided deformations.

[0060] As will be described in more detail below, the present approach uses an animatable 3D human head model that is reconstructed to capture facial expressions and speech animation. Once reconstructed, it can be driven by text and audio input signals in realtime.

[0061] Dynamic 3D human body reconstruction allows for the creation of a highly detailed and lifelike 3D human avatar. Real-time updates to the avatar's movements can be supported, ensuring a consistent and realistic representation.

[0062] The above aspects further allow for head-body integration. In such implementations, the avatar can seamlessly integrate facial expressions and head movements with full-body gestures, creating a cohesive and natural interaction experience.

[0063] Automatic expressions of the head model can be driven by a LLM to understand and process user input. The 3D head model can re-enact speech by a user, driven by the LLM. This can enable interactive conversation. The avatar can engage in meaningful conversations, understand context, and generate appropriate responses. The interactive conversational 3D human avatar can respond to the user's input prompts via the LLM which in turn animates the facial expressions of the avatar. The corresponding body pose animation is generated accordingly.

[0064] Body pose re-animation can be used to allow full-body movement of the avatar. The avatar's movements and expressions can be rendered in real-time, allowing for smooth and immediate responses to user interactions. Optimized rendering pipelines can be used to minimize latency and ensure a high-quality visual experience.

[0065] Users can engage with the avatar through various modes of interaction, including text, voice, and other inputs (for example, AR / VR interfaces). The reconstruction and optional further optimization of the animatable 3D head model reconstruction will now be described.

[0066] These reconstruction and optimization steps may be performed offline, prior to the deployment of the model for inference.

[0067] These steps may comprise data pre-processing, model training and optimization stages.

[0068] The reconstruction method uses audio- visual features (extracted from audio data, 2D images and optionally face landmarks) that are fused and used to modify the colours and opacity of 3D Gaussians of a head model. In a preferred implementation, the fused features can be projected onto a learnt audio- visual basis space, which in turn can condition a model, such as a neural network, to perform fast animations via modifications to the Gaussian’s opacities and colours. The reconstruction method can optimize the audio-visual feature basis, a set of neural networks weights and a pointcloud of Gaussian splats.

[0069] Figure 1 schematically illustrates a pipeline of the 3D animatable head model reconstruction (or training) process. In this example, the input is audio data 101, in this example comprising an audio feature vector corresponding to a section of speech, multiple RGB image frames 102 including the subject’s head corresponding to the section of speech, and 2D face landmarks 103.

[0070] The input may also comprise a set of P camera poses (not shown in Figure 1 ), with F image frames per camera pose.

[0071] To reconstruct the audio-driven animatable 3D human head model, face parsing may be performed to extract the head region from the input image frames and 2D landmarks (for example, 2D keypoints of one or more of the eyes, nose, chin etc.) and video matting to focus in on the area of the head and remove background from the input images.

[0072] A head tracking pipeline can be run to export camera poses (which may be initially head poses that have been converted to camera poses, to map all observations to the canonical head pose) and eye expressions (for example, blinking). The corresponding audio features for each image frame can be extracted by passing the user's input speech through an audio encoder.

[0073] The extraction of the eye features from an image frame and the extraction of the corresponding audio features from the audio data are indicated at 104 and 105 respectively in Figure 1. The audio features and corresponding visual features may be extracted using audio and visual feature extraction models dedicated for this aim (such as a tracker).

[0074] The two sets of features (eye features extracted at 104 and audio features extracted at 105) are then combined to create multiple fused audio features in a feature fusion step at 106. The feature fusion may be performed using, for example, a concatenation operation or a small neural network.

[0075] Camera poses, the fused audio features 107, and the input images (which as discussed above may be cropped and / or have background removed) can then be used to train the animatable 3D conversational head model.

[0076] In the examples described herein, the model is a Gaussian Splatting (3DGS) model. The model defines a set of primitives shaped as 3D Gaussians. Each 3D Gaussian has parameters such as opacity, colour, rotation, scale and position. The candidate head model is indicated at 108 in Figure 1.

[0077] In the model, every 3D Gaussian may contain an audio-feature basis. The audio- visual basis space can be learnt during the model reconstruction process. The audio-visual basis can be blended via the fused audio features to obtain a frame-specific feature 110. The multiple fused audio features 107 can be projected onto the learned audio-visual basis space by multiplying the multiple fused audio features by a projection vector to form a set of fused audio features in the audio-visual basis space, shown at 109. As mentioned above, the fused audio features are projected using the audio- visual basis. This may be performed using a vector to matrix multiplication (see, for example, https: / / mathinsight.org / matrix_vector_multiplication), which includes a summation operation. After the projection, the frame-specific or expression-specific feature 110 is produced.

[0078] The frame-specific feature can be input to an opacity modifying neural network (such as an MLP) alongside each 3D Gaussian's position to obtain audio-dependent colour and opacity for each 3D Gaussian, as indicated at 111.

[0079] The audio-dependent colour and opacity for each respective 3D Gaussian are then fed to a Gaussian Splatting rasterizer 112 alongside other parameters of the respective 3D Gaussian, which may include rotation R, scale S and position it. to render the learned Gaussians and produce a preliminary output in the form of rendered image 113, Ir.

[0080] The model is optimized by comparing the rendered image 113, Ir, against the corresponding ground truth image 114, Igt. A respective similarity measurement (such as a structural similarity index measure (SSIM)) may be determined for the or each preliminary output. This process may comprise forming a respective set of RGB colour differences comprising, for each pixel of a respective preliminary output having a given 3D pose, a difference between an RGB value of that pixel and a corresponding pixel of the ground truth image. Using RGB alignment can enable to align the model at a pixel level, such that the rendered images from the 3D Gaussian-based head model align with observations in the RGB space.

[0081] The parameters of the candidate model 108 are then updated in dependence on the similarity measurements, as illustrated at 115, for example using a gradient descent approach.

[0082] Therefore, features extracted from the input images, together with the corresponding audio features, are fused and fed through a neural network to learn changes to 3D Gaussian's attributes (specifically, colour and opacity). The result is an animatable 3D head model represented by 3D Gaussians which can be animated by input audio signals and rendered in real-time.

[0083] Figure 2 is a general block diagram showing an exemplary pipeline 200 for a dynamic 3D body reconstruction process.

[0084] In this example, the input is a set of P camera poses and F image frames per camera, as indicated at 201. As shown at 202, the F frames are partitioned into N sliding windows of frames. As shown at 203, a dynamic 3D Gaussian Splatting model (see, for example “3D Gaussian Splatting for Real-Time Radiance Field Rendering” Kerbl et al., ACM Transactions on Graphics, volume 42(4), July 2023) is trained for each sliding window. As shown at 204, each model is fine-tuned sequentially to enforce temporal consistency. The output is a 3D Gaussian model for every frame F in the sequence of input images, as indicated at 205.

[0085] Figure 3 shows a pipeline for a specific exemplary implementation for the dynamic 3D body reconstruction process. To build the dynamic 3D human body model, the following steps may be performed.

[0086] A first data processing phase is illustrated at 301. The input is a set of P camera poses and F image frames per camera. The input video frames are pre-processed to produce 2D human masks of the human that are extracted from every image. A conventional optical flow estimator can also be run to predict the optical flow between frames for each image.

[0087] In a second phase, as illustrated at 302, the frames are partitioned into N sliding windows based on the optical flow. This partitions the input sequence into manageable chunks of sliding windows. Each sliding window overlaps with the next sliding window so that temporal consistency can be enforced on the overlapping frames. Once the accumulated optical flow exceeds a threshold, a new sliding window is created. The predicted optical flow is used to adaptively partition the sequence into sliding windows, i.e. when the optical flow is high, a greater number of sliding windows are created, and when it is low, fewer sliding windows are created.

[0088] In a third phase, as illustrated at 303, a dynamic 3D Gaussian Splatting model is trained for each sliding window to represent the 3D body and its motion. 3D points (x, y, z) corresponding to the positions of 3D Gaussians and time frame t are input to a spatial-temporal encoder to extract voxel features f (which may be learned). The voxel features are then fed through a neural network (such as a multi-layer perceptron (MLP)) to predict offsets for each Gaussian's position x, scale s and rotation r. The voxel features can be passed through the neural network to learn displacements to the Gaussian's attributes (position, scale, and rotation). The final image at frame time t is created by rendering the displaced 3D Gaussians using a 3D Gaussian Splatting rasterizer.

[0089] In a fourth phase, as illustrated at 304, each model is fine-tuned sequentially to enforce temporal consistency by using a temporal consistency loss applied to the overlapping frames of each sliding window. The output is a 3D Gaussian model for every frame in the sequence.

[0090] The entire model can be optimized by comparing the rendered image to the ground truth and learning the neural network weights accordingly.

[0091] The reconstructed 3D conversational head model can be integrated with the 3D body model if desired. Integration of the 3D Gaussian conversational head to a full body, and / or to a (dynamic) scene can be performed such that the resulting avatar is full- bodied. As will be described in more detail below, the integration pipeline can perform border-specific Gaussian densification, followed by joint harmonization steps where colours can be homogenized and non-desired Gaussians located in the head can be purged with a spherical pruning strategy.

[0092] One way to combine the reconstructed 3D conversational head model with a 3D body model to create a photorealistic human will now be described with reference to the head-body integration pipeline shown in Figure 4.

[0093] The exemplary implementation of Figure 4 involves six steps, as described below. The animatable head model is combined with a dynamic 3D full-body model. The body model may be any suitable model. Here, the body model is a Gaussian splatting model, for example as described in "Human Gaussian splatting: Real-time rendering of animatable avatars ' . Moreau et al., Conference on Computer Vision and Pattern Recognition 2024. The body model may have other forms. The body model can be trained separately to the head model.

[0094] Once trained, the head and body models are integrated. This can be done by aligning and merging the 3D Gaussians of the head model and the body model which may be referred to as Gaussian Splat Merge, as indicate at 401 in Figure 4.

[0095] For a seamless integration, the head position (for example, the face position) can be determined, as indicated at 402. The determined head position may indicate a border of the head. One or more additional 3D Gaussians can be added at the border region of the head. The additional 3D Gaussians may be additional 3D Gaussians of the body model. This can "blend" the models together, as indicated at 403.

[0096] Any of the optional steps shown at 404 may be implemented for 4-dimensional (4D) and / or body optimization.

[0097] As indicated at 405, undesired 3D Gaussians can be removed from the body model by "pruning". Gaussians can be pruned (removed) from the body model that he within a spherical region centred on the head position (for example, the face position). The dynamic 3D body Gaussian's atributes can be optimized, as indicated at 406. 4D optimization refers to 3D space plus time. Optimization may be performed by comparing rendered images to the ground truth images (for example, using a similarity measure, as described previously) and updating neural network weights of the body model.

[0098] As indicated at 407, the head model may be optimized to adjust one or more colours of the head (for example, colours of the face) to more closely match one or more colours of the body model (for example, skin colour of the face and hands). This may be performed by optimizing a last layer of neural network (e.g. MLP) weights of the head model. This may be used to adjust the colours of the face for more harmonious colours and a more seamless unification of the head and body models.

[0099] The result is a fully-integrated head and body model represented by 3D Gaussians which can be rendered using the 3D Gaussian Splating technique (see, for example “3D Gaussian Splatting for Real-Time Radiance Field Rendering” Kerbl et al., ACM Transactions on Graphics, volume 42(4), July 2023).

[0100] A real- world example of the head-body integration process for a particular avatar is shown in Figure 5. In the image indicated at 501, images rendered using the separate head and body models are shown. At 502, Gaussians of the head and body are merged and rendered image 503 has been formed using the integrated models. At 504, face colour optimization is performed to more seamlessly blend the skin colours of the chest, arms and face in the rendered image 505.

[0101] The combined head and body models can be used for body pose re-animation. The dynamic 3D body reconstruction provides a dynamic human body model that can be rendered in motion.

[0102] During interactive rendering, to generate realistic body motions, one option is to replay the body motion observed during a capture sequence while novel user-generated speech is pronounced by the animatable head model. Shorter or longer body motion sequences can be generated from the captured body movement observations by selecting keyframes from the sequence and joining sections of body movement together. Either by looping sections of the video or playing the sequence backwards, a continuous body animation sequence can be created at the desired length, i.e. the body movements should match the length of the audio spoken by the head model to make the resulting animation look more natural.

[0103] Optionally, the body model can be also fully animatable, i.e. can be rendered with novel body poses not observed in the capture (described below with respect to Figure 6). In this case, the system is able to adapt the body motion with the content of the speech and the interaction in a more realistic way.

[0104] Figure 6 shows a general overview of the body pose reanimation process. The input is a set of P camera poses and F image frames per camera. In this example, a skeleton model is used to drive the movements of the human.

[0105] In this example, the pipeline uses a Human Gaussian Splating (HuGS) approach (as described in Moreau et al., “Human Gaussian Splatting: Real-time rendering of Animatable Avatars”, Moreau et al., Conference on Computer Vision and Pattern Recognition 2024). Given a 3D body pose as input, the canonical representation (or “T” pose), shown at 601, is deformed and 3D Gaussians are rendered into RGB images for the different poses, as shown at 602. The deformation model can be trained using RGB images, as shown at 603.

[0106] Figure 7 schematically illustrates a pipeline of this exemplary body pose reanimation process. Given a target body pose that is input, the canonical human (in “T” pose) is deformed to generate the 3D Gaussians in the observation space. Images can then be rendered from any camera view. In this example, the body pose re-animation process is performed by rigging 3D Gaussians to a skeleton model. As shown in the overview of Figure 7, canonical positions and orientations are first deformed with Linear Blend Skinning (LBS) using learned skinning weights.

[0107] The 3D Gaussian conversational head model can be driven from audio- visual LLM.

[0108] Figure 8 shows a pipeline of the LLM-driven speech re-enactment process. Audio input is provided by the user. In this example, a user can provide an input text prompt or audio input. This can be provided as audio input in the form of speech 801, which may be transcribed by, for example, an Automatic Speech Recognition (ASR) module 802. Alternatively, a text prompt 803 may be provided (in this example “Recite a good line from a movie”).

[0109] To enable the user to converse with the 3D virtual human avatar in an interactive way in real-time, a state-of-the-art LLM can be used for converting text prompts to text responses. The input text prompt is fed to the LLM module 804 (if the input is audio, it is transcribed by the ASR module 802). The LLM module 804 can give a text response to the prompt.

[0110] The text response output by the LLM can be converted to spoken audio by a TTS module 805, which outputs an audio sequence. Audio features are extracted from the audio sequence from the TTS module 805 by the audio feature extraction module 806.

[0111] These audio features are used then to animate the 3D Gaussian head model. The eye features extracted during reconstruction at 104 are reused during inference (i.e. the eye part is not reanimated, but just replayed from the training samples). The eye features are fused with the audio features extracted from the novel speech or text input.

[0112] The fused audio features are fed to the animatable 3D Gaussian head model having multiple 3D Gaussians 807 to obtain audiodependent colour and opacity values. The fused audio features are converted to the audiovisual feature basis, as shown at 808, and then input to an opacity modifier to perform opacity and colour adjustment, as shown at 809. The head model can be integrated with the body model at 810.

[0113] Images can be rasterized by a Gaussian rasterizer 811 using the obtained audio-dependent colour and opacity values for the Gaussians and the user’s input speech or the text response from the LLM can be used as audio for the animated avatar. The Gaussian rasterizer outputs rendered images.

[0114] After reconstruction of the models, real-time interactive rendering can be implemented within an interactive viewer. This can enable a user to freely navigate the 3D environment from any camera viewpoint and converse with the 3D human avatar through text prompts or audio inputs.

[0115] Figure 9 shows an example of a real-time interactive viewer. The integrated head and body models can be rendered and animated in real-time inside the interactive viewer. This allows the user to fully explore the virtual environment with a roaming virtual camera. A text-question input by the user is fed to the LLM module (or, if the input is audio, it is first transcribed by the ASR module). The LLM's text-answer is converted to spoken audio by the TTS module. Audio features extracted by an audio feature extraction module are then used then to animate the 3D Gaussian head model while the body pose animates accordingly.

[0116] Figure 10 shows further examples of the real-time interactive viewer. A range of different camera poses and facial expressions of the avatar are shown. The user can freely explore the scene and interact with the conversational avatar in real-time. The avatar may be integrated within a 3D virtual background for greater realism.

[0117] Figures 1 l(a)-l 1(c) illustrate real-time interactive viewer head-body Integration. In Figure 11(a) shows the dynamic 3D fullbody model (with face Gaussians pruned using the head-body integration method described herein). Figure 11(b) shows the animatable 3D head model with audio-driven facial expression. Figure 11(c) shows the fully integrated model combining 3D head and body reconstructions, animated and responsive to user's inputs in real-time.

[0118] Figure 12 shows an overview of exemplary components of the real-time interactive conversational avatar viewer system 1200. The system comprises three primary modules: database 1201, viewer 1202 and server 1203. The database module 1201 may house background assets 1204, head / face-body base models 1205, and precomputed transformation components 1206. In Figure 12, the database module 1201 is linked to the viewer module 1202 via a dotted line, indicating the flow of static data necessary for rendering. The viewer module 1202 incorporates components for talking head generation 1207, body gesture generation 1208, and rendering 1209. It is responsible for synthesizing the visual representation of the conversational virtual human avatar, combining the base models and transformation data from the database module 1201 with dynamic features received from the server 1203. The viewer 1202 is the main platform that brings together the head and body models with speech and language features. The viewer can seamlessly combine these parts to create a smooth and immersive 3D human avatar that a user can have a conversation with.

[0119] The viewer 1202 may be based on a conventional solution such as 3D Gaussian Splatting (for example “3D Gaussian Splatting for Real-Time Radiance Field Rendering” Kerbl et al., ACM Transactions on Graphics, volume 42(4), July 2023), which is excellent at rendering high-quality static images. However, conventional 3GDS does not support real-time conversations needed for a lifelike avatar. To overcome this, the 3DGS viewer may be enhanced by connecting it to one or more remote servers 1203. The viewer 1202 may send text and / or audio to the server 1203. The server 1203 may be used to process one or more of the following: 1 ) User input, such as text input and / or audio input (for example, captured by a microphone). This can be sent to the server 1203 as shown at 1210. 2) ASR: the server may convert audio data (e.g. representing spoken words) into text. This may be performed by ASR module 1211. 3) LLM processing: generating contextually relevant responses. This may be performed by LLM module 1212. 4) TTS audio generation: converting text responses into speech. This may be performed by TTS module 1213. 5) Audio feature generation. This may comprise extracting parameters needed for facial expressions. This may be performed by talking face feature generation module 1214. 6) Animatable head and body models. For example, for driving the avatar's movements and expressions. Such models may be reconstructed at the server 1203.

[0120] Figure 13 shows an example of a method for reconstructing a model of a 3D conversational head of a subject. At step 1301, the method comprises receiving one or more images depicting a head of a subject and audio data corresponding to speech by the subject. At step 1302, the method comprises extracting visual features and audio features from the one or more images and the audio data respectively. At step 1303, the method comprises fusing the visual features and the audio features to form multiple fused audio features. At step 1304, the method comprises inputting the multiple fused audio features to a candidate model representing a 3D head of the subject, the candidate model defining a set of primitives shaped as 3D Gaussians, wherein each 3D Gaussian has an opacity and a colour. At step 1305, the method comprises, for each 3D Gaussian, in dependence on the multiple fused audio features, forming an audio-dependent opacity and colour for the respective 3D Gaussian. At step 1306, the method comprises using the audio-dependent opacity and colour of the respective 3D Gaussian, forming one or more preliminary outputs. At step 1307, the method comprises comparing the or each preliminary output with a corresponding ground-truth RGB image to determine a respective image similarity measurement for the or each preliminary output. At step 1308, the method comprises updating parameters of the candidate model in dependence on the image similarity measurements).

[0121] Figure 14 shows an example of a device 1400 configured to output and / or implement a model reconstructed using the methods described herein. The device 1400 comprises a processor 1401 and a memory 1402. The memory 1402 stores in a non- transient way code that is executable by the processor 1401 to implement the respective entity in the manner described herein. The device 1400 may be implemented by hardware or may be service-based computing device, for example it may be implemented as a cloud-based computing device.

[0122] Figure 15 shows an example of a device 1500 for reconstructing a 3D conversational head model. The device 1500 may be configured to implement the methods described herein. The device 1500 comprises a processor 1501 and a memory 1502. The memory 1502 stores in a non-transient way code that is executable by the processor 1501 to implement the respective entity in the manner described herein. The device 1500 may be implemented by hardware or may be service-based computing device, for example it may be implemented as a cloud-based computing device. Therefore, the method may be deployed in multiple ways, for example in the cloud, on the device, or in dedicated hardware.

[0123] The devices 1400, 1500 may in some implementations also comprise a transceiver that is capable of communicating over a network with other entities. For example, the device may receive videos and / or images from other entities. Those entities may be physically remote from the device 1400, 1500. The network may be a publicly accessible network such as the internet. The entities may in some cases be based in the cloud. These entities may be logical entities. In practice they may each be provided by one or more physical devices such as servers and data stores, and the functions of two or more of the entities may be provided by a single physical device. Each physical device implementing an entity comprises a processor and a memory.

[0124] Traditional 2D avatars lack the depth and realism needed to provide an immersive experience. Embodiments of the present invention may address this need by combining cutting-edge 3D reconstruction techniques with advanced conversational artificial intelligence, enabling users to interact with a virtual humanthat behaves and responds in a natural and lifelike manner.

[0125] The viewer may operate hand-in-hand with separate modules running on remote servers. Each module, such as speech recognition or language processing, works independently but communicates efficiently with the viewer. By moving heavy processing tasks to the servers, the viewer on the local client device runs faster and focuses on rendering and essential communication. Specifically, the viewer can handle: 1 ) Sending: user inputs (audio from the microphone or text), 2) Receiving: selected processed outputs (transcribed text from ASR, responses from the LLM, synthesized audio from TTS, and audio parameters from the feature extraction module), 3) Animating the avatar: using the received data to generate the avatar's face and body move naturally, 4) Rendering: displaying the final image on the screen, combining background, body, and face elements. This setup minimizes communication between the viewer and the server to essential data exchanges, allowing the server to handle complex computations.

[0126] Embodiments of the present invention may provide a complete system that allows users to interact with a highly realistic and dynamic 3D human avatar through natural language conversations. The avatar is designed to respond in real-time, with humanlike gestures, facial expressions, and body movements, creating an immersive and interactive experience. The system leverages advanced technologies in 3D digital human modelling, natural language processing, and real-time rendering to provide a seamless and engaging user experience. Unlike prior approaches, this solution addresses the problem holistically, by reconstructing the whole 3D human body, enabling text and audio language capabilities, and allowing full control of body and head, together with free camera viewpoint navigation.

[0127] An enhancement of the present system is the ability to synchronize facial expressions, lip movements, and body gestures with spoken words in real time. When the LLM generates a text response, the TTS module converts it into audio, which then drives the avatar's facial animations to reflect appropriate emotions. To ensure the avatar appears natural even when idle, a human movement module may be incorporated within the viewer. This module can generate subtle movements and eye blinks by adaptively compositing and looping through a sequence of pre-recorded motions. This can create a lifelike idle state.

[0128] For smooth transitions between idle and speaking states, audio playback and the application of facial animations can be started as soon as the necessary input data is ready. Rendering may occur at a steady rate of, for example, 30 frames per second. While mouth movements can be synchronized with speech, eye blinks and other subtle expressions may be reused from recorded data to maintain naturalness. To align body gestures with speech, the avatar's movements may be adjusted based on the length of the audio response. This may ensure that all aspects of the avatar's behaviour are harmoniously synchronized.

[0129] Users can engage in real-time conversations with the avatar using their voice or by typing. The viewer can send these inputs to the remote server, where the LLM can processes them and generate relevant responses. The viewer can then animate the avatar to speak and move naturally. Depending on how a user interacts, for example whether by speaking or typing, the system may adjust its communication with the server. This may ensure a smooth experience.

[0130] The approach also allows for optimized rendering performance. While delivering high-quality, photorealistic images, the viewer is optimized for real-time performance. The present approach can implement efficient rendering techniques and utilize hardware acceleration (GPU processing) to ensure that visual quality does not compromise responsiveness. Recognizing that adding extra modules directly into the viewer could slow it down, only the essential facial animation module, which predicts facial attributes, may be used within the viewer. Other processing-heavy modules may be implemented at the server, as discussed earlier, allowing the viewer to run smoothly at, for example, 100 frames per second while being capable of generating the avatar's responses during conversations.

[0131] This present approach represents a significant advancement over existing technologies by integrating several sophisticated components into a single, interactive platform. The combination of speech-driven animatable facial expressions, dynamic 3D body reconstruction, body-pose control, seamless head-body integration, advanced natural language processing capabilities, and real-time rendering can result in a truly lifelike and engaging conversational avatar. The approach has the ability to deliver a high-fidelity interactive experience with natural expressions, gestures and responses, which is not currently achievable with existing systems.

[0132] The approach has applications in the following non-limiting exemplary areas: virtual customer service agents or personal assistants, virtual customer service for e-commerce platforms, virtual instructors or guides in educational software, interactive characters in gaming or entertainment industries, personalized avatars in social VR platforms or meta verse platforms and virtual healthcare providers in telemedicine, medical consultations and therapy sessions.

[0133] The applicant hereby discloses in isolation each individual feature described herein and any combination of two or more such features, to the extent that such features or combinations are capable of being carried out based on the present specification as a whole in the light of the common general knowledge of a person skilled in the art, irrespective of whether such features or combinations of features solve any problems disclosed herein, and without limitation to the scope of the claims. The applicant indicates that aspects of the present invention may consist of any such individual feature or combination of features. In view of the foregoing description, it will be evident to a person skilled in the art that various modifications may be made within the scope of the invention.

Claims

CLAIMS1. A method (1300) for reconstructing a model of a 3D conversational head of a subject, the method comprising: receiving (1301) one or more images (102) depicting a head of a subject and audio data (101 ) corresponding to speech by the subject; extracting (1302) visual features and audio features from the one or more images and the audio data respectively; fusing (1303) the visual features and the audio features to form multiple fused audio features (107); inputting (1304) the multiple fused audio features to a candidate model (108) representing a 3D head of the subject, the candidate model defining a set of primitives shaped as 3D Gaussians, wherein each 3D Gaussian has an opacity and a colour; for each 3D Gaussian: in dependence on the multiple fused audio features (107), modifying (1305) the opacity and colour of the respective 3D Gaussian to form an audio-dependent opacity and colour for the respective 3D Gaussian; using the audio-dependent opacity and colour of the respective 3D Gaussian, forming (1306) one or more preliminary outputs (113); comparing (1307) the or each preliminary output (113) with a corresponding ground-truth RGB image (114) to determine a respective image similarity measurement for the or each preliminary output; and updating (1308) parameters of the candidate model in dependence on the image similarity measurement(s).

2. The method as claimed in claim 1 , wherein the method further comprises, for each 3D Gaussian: projecting the multiple fused audio features onto a learned audio-visual basis space by multiplying the multiple fused audio features by a projection vector to form a set of fused audio features in the audio- visual basis space (109); from the set of fused audio features in the audio-visual basis space, forming a frame- specific feature (110); and inputting the frame- specific feature to an opacity modifier to modify the opacity and colour of the respective 3D Gaussian to form the audio-dependent opacity and colour for the respective 3D Gaussian.

3. The method as claimed in claim 2, wherein the respective 3D Gaussian has an audio-visual feature basis that is blended via the multiple fused audio features (107) to obtain the frame-specific feature (110).

4. The method as claimed in any preceding claim 2 or claim 3, wherein the frame-specific feature (110) is input to the opacity modifier alongside the position of the respective 3D Gaussian to obtain the audio-dependent opacity and colour for the respective 3D Gaussian.

5. The method as claimed in any preceding, wherein the method comprises receiving a set of camera poses, wherein the one or more images (102) comprise one or more image frames per camera pose.

6. The method as claimed in any preceding claim, wherein the visual features comprise one or more eye features extracted from the input image(s) and / or face landmarks for the subject.

7. The method as claimed in claim 6, wherein the face landmarks comprise 2D face keypoints of one or more of the eyes, nose and chin.

8. The method as claimed in any preceding claim, wherein the method further comprises cropping the one or more images depicting the head of the subject and / or removing background from the one or more images.

9. The method as claimed in any preceding claim, where the one or more preliminary outputs are formed using a Gaussian rasterizer (112), wherein the audio-dependent colour and opacity of the respective 3D Gaussian are input to the Gaussian rasterizer to render the one or more preliminary outputs.

10. The method as claimed in any preceding claim, wherein the candidate model is a neural network.

11. The method as claimed in any preceding claim, wherein the method comprises integrating a 3D conversational head model having the updated set of parameters with a body model, and / or with a scene.

12. The method as claimed in claim 11, wherein the 3D conversational head model and the body model are separate models.

13. The method as claimed in claim 11 or claim 12, wherein the body model comprises one or more body 3D Gaussians and wherein the method comprises: extracting a position of the head of the subject; and merging one or more 3D Gaussians of the 3D conversational head model and one or more body 3D Gaussians of the body model in dependence on the extracted head position.

14. The method as claimed in claim 13, wherein the method comprises performing border-specific Gaussian densification for a border between the 3D conversational head model and the body model.

15. The method as claimed in claim 13 or claim 14, wherein the method comprises performing j oint harmonization for the 3D Gaussians of the 3D conversational head model and the body 3G Gaussians of the body model to homogenize colours of the head and body of the subject.

16. The method as claimed in any of claims 13 to 15, wherein the method comprises purging non-desired 3D Gaussians located at the head of the subject using a spherical pruning strategy.

17. The method as claimed in any preceding claim, wherein the method comprises performing the steps of any preceding claim until a predetermined level of convergence is reached.

18. A device (1400) comprising one or more processors (1401), the one or more processors being configured to output and / or implement a model reconstructed according to the method of any preceding claim, the device being configured to: receive one or more images (102) depicting a head of a subject; receive audio data (101) corresponding to speech by the subject; and output a three-dimensional conversational head model for the subject.

19. The device as claimed in any of claim 18, wherein the one or more processors are configured to render a video of the 3D conversational head performing movement corresponding to an input speech segment or input text prompt.

20. A device (1500) for reconstructing a model of a 3D conversational head of a subject, the device comprising one or more processors (1501) configured to:receive (1301) one or more images (102) depicting a head of a subject and audio data (101) corresponding to speech by the subject; extract ( 1302) visual features and audio features from the one or more images and the audio data respectively; fuse (1303) the visual features and the audio features to form multiple fused audio features (107); input (1304) the multiple fused audio features (107) to a candidate model (108) representing a 3D head of the subject, the candidate model defining a set of primitives shaped as 3D Gaussians, wherein each 3D Gaussian has an opacity and a colour; for each 3D Gaussian: in dependence on the multiple fused audio features, modify (1305) the opacity and colour of the respective 3D Gaussian to form an audio-dependent opacity and colour for the respective 3D Gaussian; using the audio-dependent opacity and colour of the respective 3D Gaussian, form (1306) one or more preliminary outputs (113); compare (1307) the or each preliminary output (113) with a corresponding ground-truth RGB image (114) to determine a respective image similarity measurement for the or each preliminary output; and update (1308) parameters of the candidate model in dependence on the image similarity measurement(s).

21. One or more computer programs for instructing a computer comprising one or more processors to implement the method (1300) as claimed in any of claims 1 to 17.

22. A data carrier (1502) storing in non-transitory form the one or more computer programs as claimed in claim 21.18