Audio-driven body motion synthesis using a diffusion probabilistic model

EP4616368A1Pending Publication Date: 2025-09-17MOTORICA AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023805906
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-16
Filing Date
2023-11-08
Publication Date
2025-09-17

AI Technical Summary

Technical Problem

Current methods for audio-driven body motion synthesis, particularly for gestures and dance, face challenges in coherence and variability, with deterministic approaches failing to capture the full range of plausible motion and data-driven techniques resulting in unnatural 'average' gestures.

Method used

A diffusion probabilistic model is employed for audio-driven body motion synthesis, allowing for the generation of natural-looking body-pose sequences that can be controlled by style parameters, enabling speech-driven gesture synthesis and music-driven dance synthesis with minimal manual labeling and offering unlimited motion variation.

Benefits of technology

The model effectively generates realistic and varied body motion that matches desired styles and moods, improving the coherence and naturalness of synthesized gestures and dance, enhancing applications such as animation and virtual agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 1.1
    Figure 1.1
Patent Text Reader

Abstract

There is disclosed a method (100) for providing a model suitable for audio-driven body motion synthesis, the method comprising: obtaining (110) combined audio training data and body-motion training data; forming (120) conditioning data, wherein the conditioning data includes the audio training data; providing (122) a model of body pose sequences using a trainable diffusion probabilistic model; and training (124) the model on the basis of the body-motion training data and the conditioning data. To perform audio-driven body motion synthesis using this trained model, a further method disclosed herein comprises: obtaining an audio signal; and generating a sequence of body poses from the model conditioned upon at least the audio signal. This synthesis may optionally use classifier-free guidance for all or parts of the conditioning.
Need to check novelty before this filing date? Find Prior Art

Description

AUDIO-DRIVEN BODY MOTION SYNTHESIS USING A DIFFUSION PROBABILISTIC MODEL TECHNICAL FIELD

[0001] The present disclosure relates to audio-driven body motion synthesis, especially synthesis of gestures based on speech and synthesis of dance based on music. More precisely, it proposes methods and devices for generating a natural- looking body-pose sequence to accompany an audio signal. BACKGROUND

[0002] The ability to automatically synthesize body motion is a key endeavor to provide compelling and relatable characters for many applications including animation, crowd simulation, virtual agents and social robots. In the particular case of gesture synthesis, this has however proved to be a particularly difficult problem. A major challenge is the lack of coherence in gesture production – the same speech utterance is usually accompanied by different gestures from speaker to speaker and time to time. Previous rule-based or deterministic methods fail to model this massive variation. Data-driven regression techniques minimizing a mean square error instead lead to “average” gestures that are unlikely to be seen in real life. In order to model realistic motion, there is a need to move from deterministic to generative models that are capable of modelling the full space of plausible motion.

[0003] For any motion synthesis it is desirable to control or modify the style of the output motion. In gesture synthesis, use cases include artistic control over gesturing style to match a desired personality or mood, or automatic control over, e.g., gesture or gaze direction. Research has found that a) different styles of gesticulation are perceived as associated with the mental state [AONB17], or the personality of the speaker [SN17], and b) that this effect correlates with motion statistics like average gesture velocity, spatial extent and height [SN17, CN19, KG10].

[0004] In WO2021234151A1, the applicant has proposed a method for providing a model suitable for speech-driven gesture synthesis, as well as a method for speech- driven gesture synthesis using such a model. The model disclosed in WO2021234151A1 is a model of speaker pose sequences controllable by conditioning data, such as past and future speech data and gesture-control data, and it is constructed from normalizing flows [KD18, PNR19, HAB19].SUMMARY

[0005] One objective of the present disclosure is to present a probabilistic generative model for audio-driven body motion synthesis. A further objective is to present a model that allows the style of the output body motion to be controlled or modified by a user. A particular objective is to enable speech-driven gesture synthesis relating to upper-body gestures or full-body gestures, preferably with control over character location and direction. A further particular objective is to enable music- driven synthesis of dance-like body motion, including rhythmically consistent body poses with optional root motion. The body motion synthesis should furthermore require little or no mandatory manual labelling of input data, and it should be non- deterministic in the sense that it provides unlimited body-motion variation.

[0006] At least some of these objectives are achieved by the invention as defined by the independent claims. The dependent claims relate to advantageous embodiments of the invention.

[0007] In a first aspect of the present disclosure, there is provided a method for providing a model suitable for audio-driven body motion synthesis. The method comprises: obtaining combined audio training data and body-motion training data; forming conditioning data, wherein the conditioning data includes the audio training data; providing a model of body pose sequences using a diffusion probabilistic model; and training the model on the basis of the body-motion training data and the conditioning data.

[0008] The inventors have realized that the hitherto unknown use of diffusion probabilistic models with these characteristics in audio-driven body motion synthesis is a promising prospect that may bring several advantages, and they propose a framework for this new use in the present disclosure. The proposed body motion synthesis framework has been experimentally validated for the representative cases speech-to-gestures and music-to-dance.

[0009] In some embodiments of the invention, the body-motion training data is accompanied by a time series of one or more style-control parameters. The style- control parameters form part of the conditioning data. The time series of style- control parameters may be extracted from the body-motion training data, or it may be obtained as manual or semi-automatic annotations associated with the gesturetraining data. The term “time series” does not presuppose a variation over time, but a style-control parameter may have a constant value throughout the body-motion training data, i.e., conceptually the time series consists of repetitions of this value. The style-control parameters may refer to measurable aspects of style (e.g., hand height, hand speed, gesticulation radius, root motion, degree of correlation of right- and left-hand movements etc.) and to moods, mental states, expressions, personality (e.g. the Big 5 personality traits) and visible physical conditions (e.g., happiness, sadness, frustration, anger, calmness, status, certainty, engagement, formality, age, sickness, gender). When a sequence of body poses is generated using the trained body-motion model, the conditioning on style-control parameters allows the style of the body-pose sequence to be controlled or modified in accordance with the desires of an end user of the body-motion model. This holds independently of the measurability of the style aspect in question; indeed, even a style parameter with a meaning merely in the sphere of human cognition (e.g., happiness), which an annotator person may have added manually, is useful for style control as long as it is understandable to the end user.

[0010] In a second aspect of the present disclosure, there is provided a method for audio-driven body motion synthesis using a model obtainable by the method of any of the preceding claims. This method comprises: obtaining an audio signal, and generating a sequence of body poses from the model conditioned upon the audio signal, and optionally conditioned upon style-control parameters.

[0011] The second aspect shares many of the effects and advantages of the first aspect, and it can be implemented with a corresponding degree of technical variation. Further, the sequence of body poses can be applied to a digital or physical character, to enable an ampler experience of the audio. For example, the intelligibility of speech can be improved. In one use case, the audio signal is an audio signal with a recorded safety-oriented message directed to members of the general public (e.g., passengers in a vehicle or vessel), which become more receptive or more attentive if the speech is accompanied by a rendering of natural-looking gestures. The communication could become richer or more efficient. Similarly, the playback of a musical tune can be accompanied by a simultaneous rendering of a digital or physical character who performs synthesized dance movements.

[0012] In some embodiments of the second aspect of the invention, the synthesis of the body poses is conditioned upon style-control parameters, as described above. The generation may also be preceded by preprocessing of the audio signal corresponding to such preprocessing that was previously applied to the audio training data in the training phase.

[0013] Independent protection for devices suitable for performing the above methods is claimed. The invention further relates to a computer program containing instructions for causing a computer to carry out the above methods. The computer program may be stored or distributed on a data carrier. As used herein, a “data carrier” may be a transitory data carrier, such as modulated electromagnetic or optical waves, or a non-transitory data carrier. Non-transitory data carriers include volatile and non-volatile memories, such as permanent and non-permanent storage media of magnetic, optical or solid-state type. Still within the scope of “data carrier”, such memories may be fixedly mounted or portable.

[0014] Generally, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field, unless explicitly defined otherwise herein. All references to “a / an / the element, apparatus, component, means, step, etc.” are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, step, etc., unless explicitly stated otherwise. The steps of any method disclosed herein do not have to be performed in the exact order described, unless explicitly stated. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Aspects and embodiments of the invention are now described, by way of example, with reference to the accompanying drawings, on which: figures 1 and 2 are flowcharts of methods according to embodiments of the present invention; figure 3 illustrates a denoising process in an audio-driven diffusion model for body- motion generation, wherein ^^:^represents conditioning information,represents the output body poses, boxes represent vectors and / or sequences, rounded boxes represent learned components, and circles and ovals represent fixed operations. The process uses ^ steps;figure 4 is a combination of superimposed semitransparent (‘ghosted’) images of a rendered character in a sequence of body poses generated by a diffusion probabilistic model having been trained in accordance with embodiments herein; figure 5 contains four further combinations of superimposed semitransparent images of a rendered character, as in figure 4, wherein the body poses have been generated based on an audio signal representing speech and under the influence of style-control parameters representing happiness, sadness, disagreement and old age, respectively; figure 6 is a snapshot of two body poses generated for two different values (left: 0.5, right: 2.0) of the style-control parameter representing old age introduced in figure 5, to illustrate effects of such embodiments where a quantitative variability in style is possible; figure 7 contains two further combinations of superimposed semitransparent images of a rendered character, as in figure 4, wherein the body poses have been generated based on an audio signal representing music and under the influence of style-control parameters representing one jazz and one casual dancing style; figure 8 contains snapshots of a rendered character wherein the body poses have been generated by combining two different diffusion-model predictions in a product- of-experts combination, both based on the same audio signal representing speech but under the influence of different style-control parameters representing old age (style 1, second from left) and anger (style 2, second from right), respectively, with each snapshot using a different weighted combination between the two (^^= 1.25, 1, 0.75, … , 0, −0.25 and ^^= 1 − ^^), so as to demonstrate the effect that interpolation as well as extrapolation based on the styles has on the overall character shape and posture; and figure 9 contains overlaid snapshots a rendered character performing 10 seconds of motion wherein the body poses have been generated by combining two different diffusion-model predictions in a product-of-experts combination, both based on the same audio signal representing speech but under the influence of different style- control parameters representing no motion (style 1, second from left) and public speaking (style 2, second from right), respectively, with each position using a different weighted combination between the two (^^= 1.25, 1, 0.75, … , 0, −0.25 and^^= 1 − ^^), so as to demonstrate the effect that interpolation as well as extrapolation based on the styles has on the generated character motion. DETAILED DESCRIPTION

[0016] The aspects of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, on which certain embodiments of the invention are shown. These aspects may, however, be embodied in many different forms and should not be construed as limiting; rather, these embodiments are provided by way of example so that this disclosure will be thorough and complete, and to fully convey the scope of all aspects of the invention to those skilled in the art. Like numbers refer to like elements throughout the description. Related work

[0017] Gestures are essential to human non-verbal communication. McNeill [McN92] categorizes co-speech gestures into iconics, metaphorics, beats, deictics and emblems.

[0018] Synthesis of body motion and, in particular, gestures has recently shifted from rule-based systems – comprehensively reviewed in [WMK14] – towards data- driven approaches. Below, only data-driven methods will be discussed, since this line of research is continued.

[0019] Data-driven human body-motion generation. Several recent works have used neural networks to generate body-motion aspects such as locomotion [HHS17, HKS17, HAB19], lip movements [SSKS17] and head motion [GLM17, SB18]. A challenge in these domains is the large variation in the output given the same control. Different approaches have been employed to overcome this issue. For locomotion synthesis, studies have leveraged constraints from foot contacts to simplify the problem [HSK16, HKS17, HHS17]. Unfortunately, this is not applicable to speech-driven gestures. Closer to the domain of the present disclosure is speech- driven head-motion synthesis, where Greenwood et al. [GLM17] apply a conditional variational autoencoder (CVAE) while Sadoughi & Busso [SB18] use conditional generative adversarial networks, but these methods have not been evaluated for gesture synthesis.

[0020] Deterministic and probabilistic gesture generation. Like body motion in general, data-driven methods are on the rise in gesture generation. Levine et al. [LKTK10] used an intermediate state between speech and gestures and a hidden Markov model to learn the mapping. They selected motions from a fixed library, which limits the range of gestures their approach can generate. The present model, in contrast, is capable of generating unseen gestures.

[0021] Recently, Hasegawa et al. [HKS18] designed a speech-driven neural network capable of producing 3D motion sequences. Kucherenko et al. [KHH19] extended this work to incorporate representation learning for the motion, achieving smoother gestures as a result. Yoon et al. [YKJ19] meanwhile used neural-network sequence- to-sequence models on TED-talk data to map text transcriptions to 2D gestures. Some recent works used adversarial loss terms in their training to avoid mean- collapse, while still remaining deterministic [FNM19, GBK19]. In another recent work, Ahuja et al. [AMMS19] conditioned pose prediction not only on the audio of the agent, but also on the audio and pose of the interlocutor. All these methods produce the same gesticulation every time for a given input, while the method presented in this disclosure is probabilistic and can produce different gestures for the same input through random sampling.

[0022] Several researchers have applied probabilistic methods to gesture generation. For example, Bergmann & Kopp [BK09] applied a Bayesian decision network to learn a model for generating iconic gestures. Their approach is a hybrid between data-driven and rule-based methods because they have rules, but they learn them from data. Chiu & Marsella [CM11] took a regression approach: a network based on restricted Boltzmann machines (RBMs) was used to learn representations of arm gesture motion, and these representations were subsequently predicted based on prosodic speech-feature inputs by another network also based on RBMs. Later, Chiu et al. [CMM15] proposed a method to predict co-verbal gestures using a machine learning model which is a combination of a feed-forward neural network and Conditional Random Fields (CRFs). They limited themselves to a set of 12 discrete, pre-defined gestures. Sadoughi & Busso [SB19] used a probabilistic graphical model for mapping speech to gestures, but only experimented on three hand gestures and two head motions. The inventors believe that methods that learn and predictarbitrary movements, like the one proposed herein, represent a more flexible and scalable approach than the use of discrete and pre-defined gestures.

[0023] Style control. Control over animated motion can be exerted at different levels of abstraction. while animators and actors have explicit control over motion, it is often of interest to control higher-level properties that relate to how they are perceived. The relation between low-level motion and these properties has been extensively studied. Studies have uncovered a significant correlation between statistical properties of the motion (such as gesticulation height, velocity and spatial extent) and the perception of personality along the Big Five personality traits [Lip98, KG10, SN17] and emotion [NLK13, CN19]. In particular, Smith & Neff [SN17] modify statistical properties of existing gestures and demonstrate that these modifications create distinctly perceived personalities. Normoyle et al. [NLK13] used motion editing to identify links between motion statistics and the emotions and emotion intensities recognized by human observers.

[0024] Style control can also be exerted from motion data labeled with style descriptions. Aberman et al. [AWL20] and Ghorbani et al. [GFHTC22] present neural networks containing sub-networks to extract style-embeddings, i.e., latent vectors describing the motion style. At synthesis time, the style of the generated motion can be controlled by either providing input motion sequences containing the desired style, or sampling a style from the latent style-space.

[0025] Another line of research considers how to use machine learning to modify motion expression, based not on emotional categories or low-level statistics but on transferring stylistic properties from other recordings onto the target motion [HPP05, XWCH15, HHKK17, SCNW19]. This is known as style transfer. Style can also be controlled in some underlying parameter space. Aristidou et al. [AZS17] present a system to modify emotional expression (valence and arousal) of a given dance motion, while Brand & Hertzmann [BH00] jointly synthesize both style and choreography without motion as an explicit input.

[0026] In the present disclosure, the inventors similarly pursue the synthesis of novel motion with continuous and instantaneous control of expression. The approach of the present disclosure is agnostic to the level of abstraction of the desired control space, and will be referred to broadly as style control, regardless of the style being described with labeled descriptions or statistical properties of the motion.

[0027] Probabilistic generative sequence models. This sub-section reviews probabilistic models of complex sequence data, especially multimedia, to connect the preferred method of this disclosure – an adapted version of the causal approach presented in MoGlow [HAB19] – to related methodologies and applied work.

[0028] Early works on probabilistic human locomotion modelling investigated Gaussian process dynamical models [WFH08], along with their predecessors GP- LVMs [GMHP04, LWH12], as approaches that combined autoregressive aspects with a continuous-valued hidden state.

[0029] To escape inflexible distributional assumptions, variational autoencoders (VAEs) [RMW14, KW14] can generate samples from more complex distributions by incorporating an unobservable (latent) variable. Lately, generative adversarial networks (GANs) [GPAM14, Goo16] – another deep-learning method using a latent variable – have been the state-of-the-art in, e.g., natural image generation [BDS19]. Especially notable for this disclosure are applications of GANs to synthesizing speech-driven head motion [SB18] and video of talking faces [VPP19, PAM18, PWP18]. While GANs have been found to be capable of producing highly convincing random samples, they are notoriously difficult to train [LKM18].

[0030] Normalizing flows. In the prior disclosure WO2021234151A1, the applicant used normalizing flows [KD18, PNR19] for speech-driven gesture generation. Flows have gained interest since they have the same advantage as GANs of generating output by non-linearly transforming a latent noise variable, but by using a reversible neural network to do this it becomes possible to compute and maximize the likelihood of the training data, just like in classical probabilistic models like GMMs. Recent work has shown that normalizing flows successfully can generate complex data such as natural images [KD18, CBDJ19], audio waveforms [PVC19] and motion data [HAB19] with impressive quality. WO2021234151A1 built on the latter work by adapting it to gesture generation. Flows have subsequently been combined with VAEs for gesture generation [TWGM21], or to generate dance motion from audio [VPHB21].

[0031] Diffusion models, also known as diffusion probabilistic models, are a class of latent variable models. These models describe a process of generating a distribution via a Markov chain and can be trained using variational inference. Diffusion models are reported to have been applied to a variety of tasks, includingimage denoising, inpainting, super-resolution, and image generation. For example, an image generation model would start with a random noise image and then, after having been trained to reverse the diffusion process on natural images, the model would be able to generate new natural images. [TRG22] introduces Motion Diffusion Model (MDM), a classifier-free diffusion-based generative model for the human motion domain. [KKC22] and [ZCP22] also disclose diffusion model-based text- driven motion generation frameworks. In all cases, motion can be generated from natural language text descriptions of the properties of the desired motion (for example “A person walks fast 4 steps”). [HZP22] discloses a zero-shot text-driven framework for 3D avatar generation and animation, where custom avatars with basic motion can be generated based on similar text descriptions text, although their method does not use a diffusion model to accomplish this. None of these demonstrate the possibility of using diffusion models to generate co-speech motion from speech, whether that speech is represented as audio waveforms, acoustic features, or in any other form. Method

[0032] This section introduces diffusion probabilistic models and how they can be used to model audio-driven body motion. Bold type signifies vectors, and non-bold type scalars, including vector elements. Limits of summation are written in upper case, with lower case denoting indexing operations and colons delimit ranges of indexing into sequences.

[0033] Diffusion probabilistic models. In this disclosure, body-motion models are used to describe the distribution of body-pose sequences ^^:^= ,  … , ^^]of a speaking character or a character listening and moving to music using diffusion probabilistic models (or just “diffusion models”) [SDWMG15, SE19, HJA20]. The diffusion probabilistic models are a highly general technique for estimating continuous-valued probability distributions ^(^) in a manner that allows efficient training without restricting the generality of the distributions that can be described (unlike, e.g., the requirement for invertibility with normalizing flows). It is assumed that ^ has been normalized as is standard in deep learning, such that each element has zero mean and unit variance.

[0034] Diffusion probabilistic models are rooted in non-equilibrium thermodynamics. The idea is to define a simple Markov chain of steps ^ = 0, … ,  ^(the forward-time diffusion process) that add noise to, and possibly also shrink, the value from the previous step, thus gradually turning an observation ^ = ^^into noise ^^, and then use deep learning to learn to invert these steps and gradually shape an initial vector of noise back into a randomly drawn sequence from the learned probability distribution (the reverse-time, denoising process).

[0035] The diffusion process operating on the initial observation ^^is typically defined as a Markov chain of successive Gaussian distributions ,where ^(^; ^, ^)denotes the multivariate Gaussian probability density function with mean ^ and covariance matrix ^ evaluated at ^. The set of ^^^- and ^^^-values completely determine the diffusion process. In practice, ^^^is defined in terms of ^^^, with ^^^=^1 − ^^^([SDWMG15, HJA20]) or ^^^= 1 ([SSDK21]) being common choices; we will use the former throughout. The diffusion process is then completely ^ specified by the values. [KPH21] uses linearly spaced ^^^-values in the interval[10^^, 0.05].

[0036] Because a sum of Gaussian random variables also is Gaussian, we can write the distribution of the noisy sequence at any ^ given the initial sequence as^ where ^^and ^^are straightforward to compute analytically from ^. For leading diffusion models [KSPH21] the signal-to-noise ratio (SNR) ^^⁄^^starts high at ^ = 0 but decreases with every step, getting close to zero at ^ = ^. A common setup is to choose ^^[SDWMG15, HJA20], which ensures that ^ decreases as ^ increases.

[0037] To invert the diffusion process, as needed for generating plausible observations from noise ^^~^(^, ^), we define another Markov chain ^(^^^^|^^)that, so to speak, runs ^ in reverse. From this we can obtain the distribution ^^(^^)of ^^through marginalization .The subscript ^ indicates that the distribution ^ is parameterized by a parameter ^ ∈ Θ, typically the weights of a deep neural network. It turns out [SDWMG15] that ^(^^^^|^^)can be well approximated by a Gaussian distributionif the amount of noise added in step ^ is small relative to the magnitude of ^^^^. The covariance ^^is often ignored and set equal to a scaled version of the identity matrix [HJA20], in which case the predicted mean value ^^(^^, ^)of the previous variable in the Markov chain is the central component defining the learned distribution. Diffusion models generally perform this prediction using a neural network.

[0038] Learning in a diffusion model is the task of finding ^ such that ^^(^^)is similar to the true data distribution ^(^). Many diffusion models are based on a technique called score-matching, which leads [KSPH21] to a loss function of the form,where ^^is drawn from the training dataset ^ ={^^}, ^~^(^, ^)is white Gaussian noise, ^ is uniformly random on{1, … ,  ^}, ^^are constants that can be computed from{^^}, and ^^(^, ^)is a neural network that defines the learned aspects of the denoising process. This loss amounts to looking at a noisy example ^^= ^^^^+ ^^^ and predicting the noise ^ that has been added to it. This is mathematically equivalent to predicting the denoised observation from the noisy one, although one or the other formulation may work better in practice for machine learning [KSPH21].

[0039] It can be shown that minimizing a loss of the form ^(^|^)maximizes a variational lower bound on the training-data likelihood. The bound is generally made tighter, and models made better, by increasing the number of diffusion steps ^ [KSPH21]. In the limit ^ → ∞ a stochastic differential equation (SDE) is obtained [SSDK21]. However, models that optimize a slightly modified training loss where ^^= 1 have been found to generate perceptually more convincing results [HJA20,KSPH21]. Training optimizes the parameters ^ to minimize the loss ^(^|^), possibly in combination with other loss terms like in in [TRG22].

[0040] Conventionally, drawing random samples from a diffusion model consists of sampling a noise variable ^^~^(^, ^), and then iteratively sampling ^^^^given ^^. Specifically, each step ^^^^is computed as a weighted linear combination of ^^, ^^(^, ^), and random noise, where the weights depend on ^^^^, ^^and ^^[SME21]. However, [SSDK21, SME21] showed that the iterative sampling can be replaced by a deterministic procedure that still describes the same distribution and uses the same trained function ^^. (Then only ^^is random, similar to a GAN or a normalizing flow.) In both cases, the process can be approximated in a faster procedure by only sampling a subset of the ^ diffusion steps used during training, as described by [ND21].

[0041] Conditional diffusion models. A conditional diffusion model describes a conditional distribution ^(^^|^)with conditioning ^ by learning a Markov chain of conditional distributions ^^(^^^^|^^, ^)from a dataset of pairs of observations and their conditioning variables ^ ={(^^, ^)}. In practice, this is done by letting the function ^^(the neural network) depend on ^, as in ^^(^, ^, ^). However, it is also possible to train a classifier on noisy examples ^^and use it to guide the diffusion process so as to draw samples with more distinctive features, which is called classifier-guided diffusion [DN21]. The effect of such guidance can efficiently be approximated by combining a conditional and an unconditional diffusion model, ^^(^, ^, ^)and ^^(^, ^),respectively, to guide the diffusion process, which is called classifier-free guidance [HS21]. In this approach, samples generated by combining conditional and unconditional noise predictions as ^^(^, ^, ^) = ^^(^, ^) + γ^^^(^, ^, ^) − ^^(^, ^)^. This is straightforward to implement and also allows exaggerating the effect of ^ by setting ^ (guidance factor) greater than unity, which often is preferred in applications.

[0042] In practice the unconditional model is usually trained at the same time as the conditional model by occasionally dropping out all or parts of the conditioning. The dropped parts can be set to zero or replaced by a mask token ∅ representing the absence of information, depending on for which parts of ^ one wants to applyclassifier-free guidance. This allows training one single network that can describe both conditional and unconditional distributions.

[0043] A model of audio-driven body pose sequences using diffusion. There have been several extensions of diffusion probabilistic models to describe sequence data such as audio waveforms [KPH21, CZZ21] and motion (pose sequences ^^:^= [^^,) [ZCP22, TRG22]. The body-motion synthesis framework disclosed herein is validated using DiffWave [KPH21], which is a diffusion model for conditional audio waveform generation that can generate sequences of arbitrary length. We extend the model to vector-valued observations ^^that represent the character pose at time ^. We also replace the dilated convolutions in DiffWave by multi-head self-attention [CDL16, VSP17], which gave better loss-function values and subjectively better-looking motion.

[0044] The overall architecture of the generating of the body-pose sequence is illustrated in figure 3A. A denoising network 300 that defines the Markov chain ^^is used. The denoising network 300 first uses sinusoidal embeddings to encode the diffusion step ^ as a higher-dimensional vector ^(^), which is then passed through a feedforward network 320 (figure 3D) to obtain a modified embedding ^′(^). The noisy input is similarly passed through a projection layer (which essentially is equivalent to one-dimensional convolution with kernel size 1) and subjected to a ReLU nonlinearity. The network architecture then comprises a number of ^ blocks that function as residual layers, each with a residual block 310 (figure 3B). Each block ^ takes as input theof the previous step ^ − 1, the modified embedding ^′(^)representing ^, and the conditioning information ^. One of the two outputs of the block is added to the input towhich is sent to the next block, making the construction a residual network, while the other output is sent via a skip connection to the final output-generating subnetwork. The output-generating subnetwork adds all the activation vectors from the ^ skip connections and passes the result through a fully connected network that produces the final output value ^^(^, ^, ^).

[0045] Internally, each residual block 310 contains several operations, shown in figure 3C. The embedding of ^ is passed through an affine projection layer and added to the residual input at each time ^, obtaining a vector ^^,^,^(^^,^for short). A stack ofone or more Transformers [VSP17] or Conformers [GCQ20], carrying the reference 330, both with multi-head self-attention [CDL16, VSP17] and GELU activations [HG16], is applied across the entire timeof such vectors, in a setup where the resulting output vectors have twice the dimensionality of the input ones. Each residual block 310 has a Conformer whose convolution uses kernels of length 3 and a fixed dilation, but that dilation value cycles through the values{0, 2, 4}across the different blocks, where 0 denotes using a regular Transformer instead of a Conformer. To the output of the Transformer / Conformer 330 stack is added an affine projection of (an embedding of) the conditioning information at each ^. (This and all affine projection of time-dependent vectors in the architecture are implemented using 1D one-by-one convolutions.) The result is then split in two subvectors of the same size as the original ^^, one of which is fed into a hyperbolic tangent nonlinearity (the activation) and the other through a sigmoid nonlinearity (the gate). These two vectors are then multiplied together, creating a generalization of FiLM conditioning [PSDV18]. The final outputs for the block are obtained by subjecting the vector of gated activations to two affine projections: one projection yielding the contribution to the residual output, and another, parallel, projection instead yielding the activations for the skip connection.

[0046] The detailed Transformer / Conformer 330 architecture used for the demonstration system in this disclosure is shown in figure 3E. Multi-head neural (self-)attention [CDL16, VSP17] that doubles the dimensionality is followed by a gating, where half of the output vector is passed through a sigmoid nonlinearity and then multiplied element-wise with the other half. This is followed by dropout and the result is added to the input to the self-attention (i.e., a residual connection). Layer normalization [BKH16] is applied to the result. As motion is expected to be invariant to translation in time and space, a translation-invariant [WH21, RSR20] scheme was used to describe the impact of sequence position (time) ^ on the self-attention, specifically TISA [WH21]. This ensures that invariance to temporal translation does not need to be learned. After the attention block, there is a second network with a residual connection. This network uses a 1D convolution with kernel size 1 (for Transformer layers) or 3 (for Conformers) and GELU activation [HG16]. Conformers use a dilation that is determined by the current dilation value for the current residual block. Like the self-attention block, this is followed by dropout, a residual connection,and finally layer normalization. This feeds immediately into the next Transformer / Conformer in the stack, unless this is the end of the stack.

[0047] As an alternative to the architecture illustrated in figure 3, it is noted that the denoising process within the scope of this disclosure may include two or more neural networks.

[0048] Relating to the special case of speech-driven gesture generation, it is noted that human gestures that co-occur with speech can be considered divided into a preparation, stroke, and a retraction phase. In order to synchronize gestures to accompany speech (e.g., perform beat-gestures concurrently with prosodic emphasis in the acoustic features), the gestures must be prepared in advance. For this reason, the model may optionally take into account not only the current control inputs ^^at time instance ^, but also surrounding control inputs, either all of them or a fixed window. A window of speech features ^^^^:^^^with lookahead ^ is useful when gestures are to be generated incrementally, for example in online applications. The lookahead ^ in these windows is set in such manner that a sufficient amount of future information can be taken into account. Subjectively, the inventors found 1 s (one second) to be sufficient for generating beat gestures in synchrony with speech, while 0.5 s is too short. In embodiments devoted to other use cases, it may be sufficient to condition the model only on the current control input.

[0049] In addition to letting body poses depend on audio, one may wish to exert further control over the style or other properties of the gesticulation. It is proposed according to the present invention to add such style-control input values ^^alongside the audio-related inputs ^^, in order to train a style-controllable body-motion generation system (model). By appending control vectors to each time frame, control inputs are allowed to change over time with the same granularity as the output motion. In the next section, a few control schemes that modify meaningful properties of the body motion will be explored.

[0050] Relating again to the use case of speech-driven gesticulation, it is noted that the gesticulation may be co-occurring with locomotion (root motion), for example when people are engaged in conversation while walking. The locomotion component may be generated by the model, or it may be imposed by way of style-control parameters; in other words, an end user wishing to generate body motion pre- specifies the desired root motion path while gestures and other body-pose aspects aregenerated by the model. To model such scenarios and exert control over the path the character takes, one can additionally condition the model on the velocity (along the ground) and the angular velocity (around the up-axis) of the character. In contrast to WO2021234151A1, the inventors here suggest a data-augmentation scheme, where motion data of locomotion without speech audio is augmented with silent audio. This helps the model discern the lower-body motion correlated with locomotion from the upper-body gesticulation.

[0051] Product-of-expert diffusion models. Mathematically, classifier-free guidance corresponds to a weighted combination of the denoising predictions of two diffusion models [HS21]. Some embodiments rely on a generalization of classifier- free guidance, using a weighted combination of two or more diffusion model terms representing two or more conditional diffusion models and optionally one or more unconditional diffusion model, as follows: .This constitutes interpolation between several models (a convex combination) if all weights= 1, and extrapolation otherwise. As mentioned, ^^(^, ^)represents a conditional diffusion probabilistic model for at least two values of the index ^. This proposed procedure thus allows generating styles that are a mixture of the styles of different models. It is possible for the constituent models to be different models, or to be the same model evaluated at different conditional inputs, ^^(^, ^)= ^^(^, ^^, ^); we may call this type of interpolation guided interpolation. One or more of the terms can also be unconditional diffusion models.

[0052] Score matching operates on the gradients of the logarithm of the probability density [KSPH21, HS21], so sums of denoising processes amount to products of densities. Applying this to the denoising process densities, one obtains .This type of model is generally known as a product-of-experts model and can be described as a type of ensemble model or ensemble. If the sum of weights is unity, this is a barycentric combination, whereas the use of sums greater than unity amounts to sampling from reverse process steps whose distribution have a reducedGibbs temperature. This approach can contribute to more consistent (and often preferred) sampled output from probabilistic models. As a special case, the product of experts can be used to interpolate between different styles in a style-conditional model ^^, .

[0053] The recipe proposed above is different from conventional style interpolation in many typical motion-generation models, which rely on creating network inputs ^′ =∑^ ^^^^^^that are weighted combinations of different style representations (whether the styles are one-hot vectors, or latent-space style representations from an encoder like in [GFHTC22], or something else). These conventional approaches may not generalize well, since the averaged input value ^′ might not have been encountered at training time. The approach according to this embodiment is furthermore not the same as blending ^′ ^ ^,^= ∑^^^^^^,^between distinct gesture realizations (averaging in the output space), which for the case of images would give a trivial output resembling a double exposure. Instead, each step of the reverse diffusion process seeks outputs that are consistent with the predictions of several denoising processes, leading to a compromise that all constituent models can agree on (which is a characteristic of products of experts). This is appealing for synthesis tasks – like here – since it avoids outputs that any one “expert” considers to be low probability, i.e., unnatural. This is the right kind of inductive bias for generative tasks [TvdOB16].

[0054] While mixture of expert models have been considered with diffusion models before [BNH22, ZBLZ22], those proposals have either not been product of expert diffusion models, but binary trees of diffusion experts [BNH22] (where only one expert is used at any step of the denoising process), of been a product of experts where only one model has been a diffusion model and the other models have been other energy-based models derived by removing the last layer from conventional deep neural classifiers [ZBLZ22]. System setup and training

[0055] Training-data processing. For the experiments, the system was trained and tested on several datasets. For gesture synthesis, a configuration of the methodwithout style was evaluated with the Trinity Speech Gesture I (TSG) dataset (available at https: / / trinityspeechgesture.scss.tcd.ie / data / Trinity%20Speech- Gesture%20I / GENEA_Challenge_2020_data_release / ) and a style-conditioned configuration with the ZeroEGGS (ZEG) dataset (https: / / github.com / ubisoft / ubisoft-laforge-ZeroEGGS). The TSG dataset is a large database of synchronized speech and gestures collected by Ferstl et al. [FM18]. The data consists of 244 minutes of motion capture (bvh) and audio (wav) of one male actor speaking spontaneously on different topics. The ZEG dataset, collected by Ghorbani et al. [GFHTC22], contains 67 recordings of monologue speech and gesture of a single female actress. In this dataset, the actress was instructed to speak and gesture in 19 different styles. In total the data contains 134 minutes of motion capture (bvh format) and audio (wav format) with 3-12 minutes of data for each style category. In both datasets, the subject is standing and free to move around, shifting stance or to take a few steps back and forth.

[0056] To demonstrate the general capabilities of the method for audio-driven motion synthesis, the inventors also trained the system to generate dance motion from music audio. Here, the data from [VPHB21] was used. This dataset is pooled from AIST Dance DB [TFHG19] together with data recoded by the authors, and has a mix of different dance styles, including street dance, ballet, Greek folkdance and casual dancing. To this dataset, containing 1240 minutes of music and motion, an addition of 3 jazz styles (43 minutes) were added recorded by the inventors. This dataset is henceforth referred to as DanceDB.

[0057] For gesture synthesis, the audio signal was transformed to 20-channel mel- frequency cepstral coefficients (MFCCs) for the TSG data and 16-channel MFCCs for ZEG. For dance synthesis, 14 audio features were derived from the music files from DanceDB: MFCC (5), spectral flux (1), chroma (6), the activation of the RNNDownBeat-processor model (1) and beat tracking (1). These features were obtained using the Madmom toolbox (https: / / github.com / CPJKU / madmom).

[0058] To process the motion data, the inventors decomposed the global motion into features representing the global path of the character along the floor and local features describing the motion relative to this. Following [HSK16, HAB19], three features for the root translation and rotation were extracted, namely the frame-wise delta ^- and ^-translations together with the delta ^-rotation of the floor-projected,smoothed hip pose. The smoothing was set to 0.5 s (window length) for translation and 0.5 s for rotation.

[0059] All skeletal joint angles were then expressed relative to a reference T-pose of the character and transformed to an exponential map representation. Other angle representations were also explored, such as using the coordinates of two of the axes of the rotation matrix of each joint. As this resulted in similar results, only the exponential map representation is reported here.

[0060] The inventors provided an augmented version of the data with mirrored joint angles together with the unaltered speech, thus increasing the available amount of gesture training data. This augmentation was only done for TSG and DanceDB, as ZEG already had provided mirrored data.

[0061] The inventors further down-sampled the data in TSG and ZEG to 30 fps (using frames ^ = 0,2,4, …, and ^ = 1,3,5, thus obtaining twice the material) and sliced it into 256 frame-long (8.5 s) sequences starting at each consecutive frame (^ = 0,1,2, … ). For DanceDB, a frame rate of 20 fps was used.

[0062] Unlike the method in WO2021234151A1, some embodiments presented herein consider style as an internal property of the multimodal communication expressed in both speech and motion rather than statistical properties of the hands’ motion alone. To demonstrate this, the inventors constructed style-features for the ZEG and DanceDB datasets. The style of each frame of data was represented with a one-hot vector ^^= [^^, … ^^, … , ^^] where ^^= 1 if the frame was expressed with style ^ in the training data, and 0 otherwise, S denoting the number of styles represented in the data.

[0063] Specifically, since hand motion is central to speech-driven gestures, the inventors studied control over various aspects of the motion of the wrist joints (whose positions were computed, in hip-centric coordinates, using forward kinematics). This joint position data was then used to calculate the hand height (right hand only), the hand speed (sum of left and right hands) and the gesticulation radius (the sum of the hand distances to the up-axis through the root node). Each of these three quantities were then averaged using a four-second sliding window and the resulting, smoothed time-series used as an additional input ^^to train style-controllable model. In addition, the inventors also computed the correlation between right and left handmovements (mirrored along the ^-axis) across 4 s sliding windows, to enable learning of control over the symmetry of generated gestures.

[0064] The inventors noted that the style-control approach of the present disclosure is highly general: If it is possible to associate each frame in the data with a feature or style vector (which may vary for each time ^ or be constant per character, recording, etc.), this can be used to train a system with style input to the synthesis; the four lower-level style attributes discussed here are only intended as examples.

[0065] Audio and optional style features were used as input features to the system and joint angles and root motion as output. In a separate system, allowing user- controlled locomotion, the root motion was employed as input instead of output.

[0066] Network tuning and training. The inventors tuned the model hyperparameters using grid-search. The final model was composed of ^ = 10 blocks of residual layers, each having a stack of 4 self-attention layers (8 heads, 256 attention channels, 1024 channels in the feedforward network) and a convolution cycle of length 3. The noise levels were set to 100 steps linearly distributed between 1.0 ⋅ 10^^and 5.0 ⋅ 10^^. For optimization, the Adam optimizer [KB15] was used together with warmup and stepwise learning rate decay. The final learning rate values were lr^^^= 1.0 ⋅ 10^^, warmup for 10k steps, and decay with a factor of 0.5 ⋅ 10^^every 10 steps. Additionally, the style-controlled models had a dropout rate of 0.2 applied to the style features intended for guidance of the synthesized motion. The proposed models were trained for 150,000 optimization steps for TSG, 100,000 steps for ZEG and 200,000 steps for DanceDB.

[0067] Proposed systems and baselines. Following parameter tuning, the inventors trained two systems for gesture synthesis and one for dance synthesis. The gesture synthesis systems are denoted DG and DGS, where DG was trained on the TSG dataset and conditioned only on speech (as audio training data), and DGS was trained on the ZEG data and conditioned on both speech and style. The dance synthesis system, denoted DGD, was trained on DanceDB and conditioned on music (as audio training data) and style. The two systems used the same hyperparameters identified in the preceding subsection. (In the terminology of the claims, DG constitutes a “second sub-model”, while DGD and DGS can be used as a “first sub- model”.)

[0068] Synthesis. After training the systems, the inventors synthesized new gesture motion using held-out data from the datasets. For DG, the audio clips from the test data of TSG were used, cutting each of the clips into 10 s long snippets and processing them into the input features to the model. The resulting synthesized motions were then transferred to a skinned 3D character and rendered together with the audio into videos. The videos can be found at the following URLs:

[0069] Figure 4 shows the 3D character gesticulating overlaid with 4 ghosted prior time steps. The prior time steps are equidistant, so that the variability in speed along the motion path can be perceived. The overlaid time steps will also illustrate the approximate duration for which different regions of space are occupied by the character’s arms.

[0070] For the DGS system, one of the takes was held out during training and used for testing. In this case, the test audio was repeatedly used to generate different styles of motion, by only changing the one-hot style vector. These videos can also be found at the above-identified URLs.

[0071] Figure 5 shows the gesticulating character, overlaid with 4 ghosted prior time steps like in figure 4, in four different styles: happy, sad, disagreeing and old. It is noted how the character posture changes according to style.

[0072] All the previous tests were carried out using no guidance during synthesis. To show the effect of guidance, i.e. that a quantitative variability in style is achievable, the inventors additionally synthesized styles using guidance factors 0.5 (left) and 2 (right) in combination with the style parameter ‘old’. Examples in the form of body pose snapshots are shown in figure 6 and in videos at the URLs.

[0073] For the DGD model, held-off data was used to synthesize dances in several styles. The results can be seen as still images in figure 7 and as video clips at the URLs. More precisely, figure 7 contains two further combinations of superimposed semitransparent images of a rendered character, as in figure 4, wherein the body poses have been generated based on an audio signal representing music and under the influence of style-control parameters representing one dancing style identified as ‘street’ (upper image) and one dancing style identified as ‘casual’ (lower image).None of the previous examples show the effect of product-of-expert combinations of style- conditional models. To demonstrate the ability to interpolate and extrapolate between two styles, the inventors synthesized outputs from several different product- of-experts configurations of models with several different mixing coefficients for each.

[0074] Figure 8 demonstrates the effect that such interpolation has on body posture in a number of snapshots interpolating between the style parameter ‘old’ (a hunched figure) on the left and the style parameter ‘angry’ (upright) on the right. A weightedcombination of the outputs of two sub-models has been used in the denoising steps, one output from a sub-model trained on conditioning data including the ‘old’ style parameter, and one output from a sub-model trained on conditioning data including the ‘angry’ style parameter. The two sub-models may be two distinct models (e.g., which have been trained in separate processes and / or on separate training data), or the two sub-models may refer to a common model evaluated at two distinct values of the style parameter (e.g.,^^, ^), ^^(^, ^^, ^), where ^^, ^^differ with regard to at least one style parameter).

[0075] Figure 9 demonstrates the effect that interpolation has on motion dynamics by showing poses from 10 seconds of animation overlaid on top of each other, where the style parameter goes from ‘still’ (hardly any motion) on the left to the style parameter ‘oration’ (a highly animated style of public speaking) on the right. In both figures, a barycentric combination of two style-conditional reverse diffusion processes are used, where the coefficient ^^of the style on the left in each figure goes from 1.25 on the far left to −0.25 on the far right. The characters on the extreme ends thus represent extrapolation, the poses next to them are pure styles, and the ones in- between illustrate three points on a continuum interpolating between the two pure styles.

[0076] The still images and videos confirm that the inventors have successfully achieved the objective of enabling probabilistic audio-driven body motion synthesis that permits optional style control. It is foreseen that the body motion synthesis framework can be validated by a subjective study using human raters, comparing the systems to an upper baseline of ground truth examples from the held-out data and baseline systems such as WO2021234151A1. Particular embodiments of the invention

[0077] Figure 1 is a flowchart of a method 100 for providing a model suitable for audio-driven body motion synthesis. The method 100 may be implemented by a device equipped with memory and processing circuitry including one or more processor cores. The device may, for example, be a general-purpose computer with input / output capabilities allowing it to obtain training data (e.g., as a data file) and make the trained model available for use in body motion synthesis. Such general- purpose computer may optionally be equipped with a hardware accelerator suitable for machine-learning computations. Alternatively, the device is a networked (orcloud) processing resource. The model may either remain on the device and be put to use there for body motion synthesis, or it may be exported in a transferable format for use on a different device or a different execution platform.

[0078] The method 100 begins with a step of obtaining 110 combined audio and body-motion training data. The audio training data may be an audio signal, and the audio signal may contain sound identifiable as speech or music. Further, the audio training data may be a synthesized speech signal generated from text. Still further, the audio training data may be a structured representation of speech, such as a sequence of onset times for words and syllables, optionally annotated with intonation and other semantic or stylistic information. Similarly, the audio training data could be a structured representation of music, such as a file in Musical Instrument Digital Interface (MIDI) or a similar format.

[0079] The body motion training data may be based on a sequence of video frames captured by one or more video cameras. The body motion training data may be structured data including motion-capture data, or joint-angle data in particular. The audio and body motion training data are “combined” in the sense they contain timing indicators or other metadata allowing them to be aligned in time. The audio and body motion training data may be downsampled, for example, to a consistent rate of 60 frames per second.

[0080] In an optional step 112, the audio training data is preprocessed to obtain a spectrogram. Such preprocessing eliminates phase information, yet provides a robust basis for the model training to follow. For example, the spectrogram may be a power spectrogram, a magnitude spectrogram, a mel-frequency power spectrogram, a log- power spectrogram, or it may be composed of mel-frequency cepstrum coefficients (MFCC).

[0081] In a further optional step 114, the body-motion training data is preprocessed. The preprocessing 114 may convert the body-motion training data into an exponential-map representation. Alternatively or additionally, the preprocessing 114 may include time synchronization, downsampling, and / or coordinate conversion such as a conversion to T-pose coordinates. Further alternatively, the body-motion training data may be preprocessed 114 into a representation using two axes out of three from 3×3 rotation matrices. (This is a compact coding / representation, knowing that an orthogonal matrix is completely defined by the elements of two of its rows ortwo of its columns.) It is noted that the exponential-map representation can be used for some joints and the rotation-matrix representation can be used for other joints of a same body.

[0082] In a further optional step 116, additional training data may be provided by combining the audio training data with processed body-motion training data. The body-motion training data may be processed by right–left mirroring, i.e., by discarding data relating to the right or left half of the body and replacing it with a mirror image. The additional training data provided in step 116 may be described as a mirrored version of the originally obtained body-motion data (joint angles) together with the unaltered audio.

[0083] In a still further optional step 118, a time series of one or more style-control parameters is obtained from the body-motion training data. The style-control parameters may refer to measurable aspects of style (e.g., hand height, hand speed, gesticulation radius, root motion (or locomotion), degree of correlation of right- and left-hand movements etc.) and to moods, mental states, expressions, personality (e.g. the Big 5 personality traits) and visible physical conditions or other aspects in the sphere of human cognition (e.g., happiness, sadness, frustration, anger, calmness, status, certainty, engagement, formality, age, sickness, gender). The time series of style-control parameters may be extracted from the body-motion training data, e.g., by tracking visual features of the body and computing said hand height, hand speed etc. Alternatively, the style-control data may be obtained as manual or semi- automatic (elicited) annotations associated with the gesture training data. This allows the style of the output motion to be controlled or modified according to an end user’s wishes. In other words, the style-control parameters obtained in the training phase constitute (manual or semi-automatic) annotations of the gesture training data, reflecting what a viewer may perceive as various degrees of hand height, hand speed etc. In the body-motion synthesis phase, then, the values of the style-control parameters constitute inputs to the body-motion synthesis process and they control certain quantitative aspects of the style of the synthesized body poses, such as hand height, hand speed, etc. (It is noted that quantitative style control can alternatively be achieved by classifier guidance, classifier-free guidance and interpolation between multiple sub-models, as detailed below.)

[0084] In a next step 120 of the method 100, the (optionally preprocessed) audio training data (sequence of acoustic features ^^:^) is used as conditioning data (sequence ^^:^). In some embodiments of the invention, the style-control parameters (sequence ^^:^) are also included in the conditioning data. As illustrated by figure 3, the conditioning data used when generating ^^can, in some embodiments, include not only past data but also future data. Near the beginning of a data set, dummy past data (e.g., silence) may be used, and similarly dummy future data may be used at the end of a data set. In particular, the conditioning data may include future audio data.

[0085] In one embodiment, conditioning on future speech data is used in step 120 so as to support speech-driven gesture generation. More precisely, human gestures that co-occur with speech can be described as segmented into a preparation, a stroke and a retraction phase. In order to synchronize gestures with speech (e.g., perform beat gestures concurrently with prosodic emphasis in the acoustic features), the invention allows the model to prepare gestures in advance. The gestures can then be executed in synchrony with the speech. In particular, conditioning input information used at a given time instance may contain not only the current speech features but also all surrounding speech features or a window thereof, including future speech features.

[0086] (It is noted that the body-motion synthesis framework could alternatively be implemented as an autoregressive system, where the body poses are generated incrementally. In this case, a lookahead ^^:^^^for the audio features, with ^ a number of frames representing a duration of between 0.5 and 1.0 s, has been found to be suitable for body motion generation, especially for speech-to-gesture synthesis. This does not preclude also using the preceding acoustic features.)

[0087] In a next step 122, after the conditioning data has been formed, a model of body pose sequences is provided using a diffusion probabilistic model. The model may be a Markov chain with one or more steps, and / or it may be a differential equation representing a Markov chain with an infinite number of steps. These steps may themselves be stochastic or deterministic. The behavior of the steps in the Markov chain is determined by one or more (trainable) neural networks. The neural network may include convolution operations, a convolutional neural network (CNN), a transformer, a conformer, a recurrent neural network (RNN), a long short-term memory unit (LSTM) unit, and / or a gated recurrent unit (GRU). The neural networkmay include at least one neural attention mechanism, at least one instance of self- attention, and / or at least one convolutional layer. The neural attention mechanism may include attention weights which are dependent on sequence position 1 through ^. Optionally, the dependence of the attention weights on sequence position is encoded using a scheme which considers position difference but is invariant with respect to translation in sequence position ^, e.g., as in the models in [WH21]. Specifically, the neural attention mechanism can be implemented – possibly as part of a transformer and / or conformer component – as a neural self-attention mechanism or a neural dot-product self-attention mechanism.

[0088] The method 100 concludes with a training step 124, in which the model is trained on the basis of the body-motion training data and the conditioning data. It is recalled that the conditioning data may include the audio data but also, in some embodiments, values of the style-control parameters. A nonzero dropout rate may be applied to the neural network, or its inputs, or both. Dropped inputs to the neural network can be replaced by special values representing absence of information. Further optionally, a Markov chain that adds noise to the data may be used to train the model.

[0089] After a sufficient quantity of training has been completed, the audio-driven body motion synthesis model is ready for use in body-motion synthesis, in accordance with a user’s requests. The model may be made available as part of a body-motion generation system.

[0090] Figure 2 is a flowchart of a method 200 for audio-driven body-motion synthesis. The method 200 may use a body motion synthesis model which has been obtained by an execution of the method 100 illustrated in figure 1. Alternatively, the method 200 may use a probabilistic model with equivalent properties that was obtained by a different process. The properties include that the model shall be based on a diffusion probabilistic model, shall relate to body-pose sequences and be conditioned upon at least audio data. An optional property is that the body motion synthesis is conditioned upon past and future conditioning data, such as future audio data or future speech data. The method 200 may be executed online (e.g., using a live audio signal) or offline. Similar to the model providing method 100, the synthesis method 200 may be implemented by a device equipped with memory and processing circuitry including one or more processor cores, as discussed above.

[0091] In an initial step 210 of the method 200, an audio signal is obtained. Within the scope of the present disclosure, the audio signal may for example be a speech signal, such as a signal that encodes recorded speech, or a signal containing speech synthesized from text. Still within the scope of the present disclosure, the audio signal may contain recorded or synthesized music. As explained, at least the following audio signals could be used, provided the body-motion model has been trained properly: recorded speech or music; a synthesized speech signal generated from text; structured representation of speech, such as a sequence of onset times for words and syllables, optionally annotated with intonation and other semantic or stylistic information; a structured representation of music, such as a MIDI file.

[0092] In an optional next step 212, the audio signal is preprocessed. The preprocessing may correspond to any preprocessing that was applied to the audio training data when the body-motion model was trained. As explained in connection with step 112 above, the preprocessing 212 may return a power spectrogram or magnitude spectrogram.

[0093] In a further optional step 214, values of one or more style-control parameters are obtained. More precisely, in the case where the trained body-motion model has been conditioned upon a set of style-control parameters (in the sense explained above) in addition to audio, step 214 includes obtaining values of these corresponding to a request from a user. For example, the values of the style-control parameters may be entered by a human user via a user interface, or they may be contained in a call or request message from an executing process form which the sequence of body poses is to be generated.

[0094] In a further optional step 216, as deemed necessary, time-aligned segmentations of the audio signal and the style-control parameter values are provided. This allows execution in discrete time, as schematically illustrated in figure 3. It is noted that commonly practiced techniques for audio-signal segmentation, in which sequentially overlapping windowing functions are utilized, are understood to produce segmentations in the sense of the present disclosure.

[0095] The execution of the method 200 then goes on to a step 218 of generating a sequence of body poses from the body-motion model conditioned upon the conditioning data, i.e., at least conditioned upon the audio signal. The generating 218 of a body pose may be characterized as a sampling from a probability distributiondescribed by the trained diffusion probabilistic model. The perceived quality of the output can sometimes be improved by tuning (making small adjustments to) the standard deviation of the underlying probability distribution on which the normalizing flows act and / or by adjusting the stochasticity in the steps of the Markov chain used during sampling. Likewise, the speed of the process may be tuned by using a different number of Markov chain steps than used during training. This may mean that the sampling proceeds under slightly different conditions than the training.

[0096] In step 218, the impact of the conditioning information on the body-pose generation can in some embodiments be controlled, in a substep 218.1, by means of guided diffusion. The diffusion process may be classifier-guided. A further option is to apply classifier-free guidance to the diffusion. This allows the style of the sequence of body poses to be amplified or reduced. Accordingly, the quantitative magnitude of measurable aspects of style can be varied, as can the degree of various moods, mental states, expressions or physical conditions (e.g., happy / happier, sad / less sad, younger / older). In implementations where guided diffusion is used to amplify or reduce the style, above-described step 214 may include obtaining a quasi-binary value of the style-control parameters to be used. Indeed, although the body-motion model has been conditioned upon a style-control parameter in the form of a continuous variable, it can be assigned a unit or default value, such as 1.0, since the quantitative variability in style is anyhow taken care of by the guidance.

[0097] An alternative or additional way of achieving a quantitative variability in style is to use a body-motion model that includes a first sub-model having been trained on the basis of conditioning data including the style-control parameters, and a second sub-model having been trained on the basis of conditioning data not including the style-control parameters. In this connection, the first and second sub- models may in fact refer to the same basic model, wherein the first sub-model is the basic model evaluated for a non-zero (or active) value of the conditional input and the second sub-model is the basic model evaluated for a zero (or inactive, or neutral) value of the conditional input. As explained above, to generate a body-pose sequence using a diffusion probabilistic model, a plurality of so-called denoising steps 218.2 are performed. With access to the first and second sub-models of the diffusion model, a weighted combination of the output of the first sub-model and the output of the second sub-model is applied as guidance. The relative weight given to the first sub-model in this combination (cf. guidance factor γ above) can be chosen in accordance with the magnitude of the style desired by the end user, i.e., the magnitude at which the style is to be expressed. In particular, the guidance applied in each denoising step 218.2 may comprise a barycentric combination of the first and second sub-models’ denoising predictions. Either of these options is useful for guiding the reverse diffusion process so as to make the effect of the conditioning more or less pronounced. Again, when two sub-models are used to achieve the quantitative variability, it is sufficient to obtain a unit or default value (quasi-binary value) of the desired active style-control parameter.

[0098] A further alternative way of achieving a quantitative variability in style is to use a body-motion model that has been trained using training data with multiple values of each style-control parameter. The total time and effort spent on training could however be somewhat higher.

[0099] In step 218, further, a combination of two styles can be achieved by applying as guidance a weighted combination of the output of a first sub-model and a second sub-model, wherein the first sub-model has been trained on the basis of conditioning data including a first set of one or more style-control parameters, and the second sub- model has been trained on the basis of conditioning data including a second set of one or more style-control parameters, which is distinct (or in particular disjoint) from said first set. This enables an interpolation between the first set and second set, wherein the weighting corresponds to a ratio of the magnitudes (relative magnitudes) at which the first and second sets of style-control parameters are to be expressed. Here, the first and second sub-models may refer to the same basic model, wherein the first sub-model is the basic model evaluated for a first (non-zero) value of the conditional input and the second sub-model is the basic model evaluated for a second (non-zero) value of the conditional input.

[0100] Optionally, the weighted combination to be used as guidance in the denoising step may be a combination of three or more sub-models. Further optionally, at least one of the sub-models in the combination may be a sub-model having been trained on the basis of conditioning data not including any style-control parameter. This allows the magnitude of the interpolated style expression to be varied. In embodiments where the first and second sub-models refer to the same basic model, the sub-model trained on conditioning data not including any style-control parameter can be replaced with the basic model evaluated for a zero (or inactive, or neutral) value of the conditional input.

[0101] In one example, the basic model is trained on the basis of conditioning data with a first and a second scalar style-control parameters ^^, ^^representing ‘old’ and ‘angry’, wherein the first sub-model may correspond to the style-control parameter assignment(^^, ^^)= (1,0) to the basic model, and the second sub-model may correspond to the style-control parameter assignment(^^, ^^)= (0,1). It is understood that the training has been performed in such manner that the value 1 corresponds to activity / presence and the value 0 to inactivity / absence. To obtain a style expression with equal contributions of the ‘old’ the ‘angry’ style, a weighted combination of the output of the first and second sub-models should be used as guidance. To weaken this style expression ‘old and angry’, the weighted combination can be extended with a desired amount of a third sub-model corresponding to the neutral style-control parameter assignment(^^, ^^)= (0,0) to the basic model.

[0102] The sequence of body poses thus obtained may subsequently be applied 220 to a digital or physical character. It may be rendered as a video sequence to be played back on a computer. Alternatively, the sequence of body poses may be transformed into control signals to be applied to the actuators of a robot representing the moving body. A “digital” (or virtual) character in this sense may be a representation of a human, humanoid, animal or imaginary character in the form of visual elements of a graphical user interface. A “physical” character may be a robot, a puppet or a similar artificial representation of a human, humanoid or animal character.

[0103] The aspects of the present disclosure have mainly been described above with reference to a few embodiments. However, as is readily appreciated by a person skilled in the art, other embodiments than the ones disclosed above are equally possible within the scope of the invention, as defined by the appended patent claims. REFERENCES [AMMS19] AHUJA C., MA S., MORENCY L.-P., SHEIKH Y.: To react or not to react: End-to-end visual pose forecasting for personalized avatar during dyadic conversations. In Proc. ICMI (2019), pp.74–84.[AONB17] ALEXANDERSON, S., O’SULLIVAN, C., NEFF, M., & BESKOW, J.: Mimebot—investigating the expressibility of non-verbal communication across agent embodiments. ACM T. Applied Perception.14, 4 (2017), 1–13. [AWL20] ABERMAN K., WENG Y., LISCHINSKI D., COHEN-OR D., CHEN B.: Unpaired motion style transfer from video to animation. ACM Trans. Graph.39, 4 (2020), 64:1–64:12. [AZS17] ARISTIDOU A., ZENG Q., STAVRAKIS E., YIN K., COHEN-OR D., CHRYSANTHOU Y., CHEN B.: Emotion control of unstructured dance movements. In Proc. SCA (2017), p.9. [BDS19] BROCK A., DONAHUE J., SIMONYAN K.: Large scale GAN training for high fidelity natural image synthesis. In Proc. ICLR (2019). [BH00] BRAND M., HERTZMANN A.: Style machines. In Proc. SIGGRAPH (2000), pp.183–192. [BK09] BERGMANN K., KOPP S.: GnetIc–using Bayesian decision networks for iconic gesture generation. In Proc. IVA (2009), pp.76–89. [BKH16] BA J. L., KIROS J. R., HINTON G. E.: Layer normalization. In Proc. NIPS Deep Learning Symposium (2016). [BNH22] BALAJI Y., NAH S., HUANG X., VAHDAT A., SONG J.,KREIS K., AITTALA M., AILA T., LAINE S., CATANZARO B., ET AL.: eDiff-I: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324 (2022). [CBDJ19] CHEN R. T. Q., BEHRMANN J., DUVENAUD D., JACOBSEN J.-H.: Residual flows for invertible generative modeling. In Proc. NeurIPS (2019), pp.9913– 9923. [CDL16] CHENG J., DONG L., LAPATA M.: Long short-term memory-networks for machine reading. In Proc. EMNLP (2016), pp.551–561. [CM11] CHIU C.-C., MARSELLA S.: How to train your avatar: A data driven approach to gesture generation. In Proc. IVA (2011), pp.127–140. [CMM15] CHIU C.-C., MORENCY L.-P., MARSELLA S.: Predicting co-verbal gestures: A deep and temporal modeling approach. In Proc. IVA (2015).[CN19] CASTILLO G., NEFF M.: What do we express without knowing?: Emotion in gesture. In Proc. AAMAS (2019), pp.702–710. [CZZ21] CHEN N., ZHANG Y., ZEN H., WEISS R. J., NOROUZI M., CHAN W.: WaveGrad: Estimating gradients for waveform generation. In Proc. ICLR (2021). [DN21] DHARIWAL P., NICHOL A.: Diffusion models beat GANs on image synthesis. In Proc. NeurIPS (2021), pp.8780–8794. [FM18] FERSTL Y., MCDONNELL R.: Investigating the use of recurrent motion modelling for speech gesture generation. In Proc. IVA (2018), pp.93–98. [FNM19] FERSTL Y., NEFF M., MCDONNELL R.: Multi-objective adversarial gesture generation. In Proc. MIG (2019), pp.3:1–3:10. [GBK19] GINOSAR S., BAR A., KOHAVI G., CHAN C., OWENS A., MALIK J.: Learning individual styles of conversational gesture. In Proc. CVPR (2019), pp.3497– 3506. [GFHTC22] GHORBANI, S., FERSTL, Y., HOLDEN, D., TROJE, N. F., & CARBONNEAU, M. A.: ZeroEGGS: Zero-shot example-based gesture generation from speech. arXiv preprint arXiv:2209.07556. [GH00] GHAHRAMANI Z., HINTON G. E.: Variational learning for switching state- space models. Neural Comput.12, 4 (2000), 831–864. [GLM17] GREENWOOD D., LAYCOCK S., MATTHEWS I.: Predicting head pose from speech with a conditional variational autoencoder. In Proc. Interspeech (2017), pp. 3991–3995. [GMHP04] GROCHOW K., MARTIN S. L., HERTZMANN A., POPOVIĆ Z.: Style- based inverse kinematics. ACM T. Graphic.23, 3 (2004), 522–531. [Goo16] GOODFELLOW I.: NIPS 2016 tutorial: Generative adversarial networks. arXiv preprint (2016). arXiv:1701.00160. [GPAM14] GOODFELLOW I., POUGET-ABADIE J., MIRZA M., XU B., WARDE- FARLEY D., OZAIR S., COURVILLE A., BENGIO Y.: Generative adversarial nets. In Proc. NIPS (2014), pp.2672–2680.[GQC20] GULATI A., QIN J., CHIU C.-C., PARMAR N., ZHANG Y., YU J., HAN W., WANG S., ZHANG Z., WU Y., PANG R.: Conformer: Convolution-augmented transformer for speech recognition. In Proc. Interspeech (2020), pp.5036–5040. [HAB19] HENTER G. E., ALEXANDERSON S., BESKOW J.: MoGlow: Probabilistic and controllable motion synthesis using 35ormalizing flows. arXiv preprint (2019). arXiv:1905.06598. [HG16] HENDRYCKS D., GIMPEL K.: Gaussian error linear units (GELUs). arXiv preprint arXiv:1606.08415 (2016). [HHKK17] HOLDEN D., HABIBIE I., KUSAJIMA I., KOMURA T.: Fast neural style transfer for motion data. IEEE Comput. Graph.37, 4 (2017), 42–49. [HHS17] HABIBIE I., HOLDEN D., SCHWARZ J., YEARSLEY J., KOMURA T.: A recurrent variational autoencoder for human motion synthesis. In Proc. BMVC (2017). [HJA20] HO J., JAIN A., ABBEEL P.: Denoising diffusion probabilistic models. In Proc. NeurIPS (2020), pp.6840–6851. [HKS17] HOLDEN D., KOMURA T., SAITO J.: Phase-functioned neural networks for character control. ACM T. Graphic.36, 4 (2017), 42:1–42:13. [HKS18] HASEGAWA D., KANEKO N., SHIRAKAWA S., SAKUTA H., SUMI K.: Evaluation of speech-to-gesture generation using bidirectional LSTM network. In Proc. IVA (2018), pp.79–86. [HPP05] HSU E., PULLI K., POPOVIĆ J.: Style translation for human motion. In ACM T. Graphic. (2005), vol.24, pp.1082–1089. [HS97] HOCHREITER S., SCHMIDHUBER J.: Long short-term memory. Neural Comput.9, 8 (1997), 1735–1780. [HSK16] HOLDEN D., SAITO J., KOMURA T.: A deep learning framework for character motion synthesis and editing. ACM T. Graphic.35, 4 (2016), 138:1–138:11. [HS21] HO J., SALIMANS T.: Classifier-free diffusion guidance. In Proc. NeurIPS Workshop on DGMs and Applications (2021).[HZP22] HONG F., ZHANG M., PAN L., CAI Z., YANG L., LIU Z.: AvatarCLIP: Zero- shot text-driven generation and animation of 3D avatars. ACM T. Graphic.41, 4 (2022), 1–19. [JKEB19] JONELL P., KUCHERENKO T., EKSTEDT E., BESKOW J.: Learning non- verbal behavior for a social robot from YouTube videos. In Proc. ICDL-EPIROB Workshop Nat. Non-Verbal Affect. Hum.-Robot Interact. (2019). [KB15] KINGMA D. P., BA J.: Adam: A method for stochastic optimization. In Proc. ICLR (2015). [KD18] KINGMA D. P., DHARIWAL P.: Glow: Generative flow with invertible 1×1 convolutions. In Proc. NeurIPS (2018), pp.10236–10245. [KG10] KOPPENSTEINER M., GRAMMER K.: Motion patterns in political speech and their influence on personality ratings. J. Res. Pers.44, 3 (2010), 374–379. [KHH19] KUCHERENKO T., HASEGAWA D., HENTER G. E., KANEKO N., KJELLSTRÖM H.: Analyzing input and output representations for speech-driven gesture generation. In Proc. IVA (2019), pp.97– 104. [KJvW20] KUCHERENKO T., JONELL P., VAN WAVEREN S., HENTER G. E., ALEXANDERSON S., LEITE I., KJELLSTRÖM H.: Gesticulator: A framework for semantically-aware speech-driven gesture generation. arXiv preprint (2020). arXiv:2001.09326. [KKC22] KIM J., KIM J., CHOI S.: FLAME: Free-form language-based motion synthesis & editing. arXiv preprint arXiv:2209.00349 (2022). [KPH21] KONG Z., PING W., HUANG J., ZHAO K., CATANZARO B.: DiffWave: A versatile diffusion model for audio synthesis. In Proc. ICLR (2021). [KSPH21] KINGMA D., SALIMANS T., POOLE B., HO J.: Variational diffusion models. In Proc. NeurIPS (2021), pp.21696–21707. [KW14] KINGMA D. P., WELLING M.: Auto-encoding variational Bayes. In Proc. ICLR (2014). [Lip98] LIPPA R.: The nonverbal display and judgment of extraversion, masculinity, femininity, and gender diagnosticity: A lens model analysis. J. Res. Pers.32, 1 (1998), 80–107.[LKM18] LUCIC M., KURACH K., MICHALSKI M., GELLY S., BOUSQUET O.: Are GANs created equal? A large-scale study. In Proc. NeurIPS (2018), pp.698–707. [LKTK10] LEVINE S., KRÄHENBÜHL P., THRUN S., KOLTUN V.: Gesture controllers. ACM T. Graphic.29, 4 (2010), 124. [LWH12] LEVINE S., WANG J. M., HARAUX A., POPOVIĆ Z., KOLTUN V.: Continuous character control with low-dimensional embeddings. ACM T. Graphic. 31, 4 (2012), 28. [McN92] MCNEILL D.: Hand and Mind: What Gestures Reveal about Thought. University of Chicago Press, 1992. [ND21] NICHOL A. Q., DHARIWAL P.: Improved denoising diffusion probabilistic models. In Proc. ICML (2021), pp.8162–8171. [NLK13] NORMOYLE A., LIU F., KAPADIA M., BADLER N. I., JÖRG S.: The effect of posture and dynamics on the perception of emotion. In Proc. SAP (2013), pp.91–98. [PAM18] PUMAROLA A., AGUDO A., MARTINEZ A. M., SANFELIU A., MORENO- NOGUER F.: GANimation: Anatomically-aware facial animation from a single image. In Proc. ECCV (2018), pp.818–833. [PNR19] PAPAMAKARIOS G., NALISNICK E., REZENDE D. J., MOHAMED S., LAKSHMINARAYANAN B.: Normalizing flows for probabilistic modeling and inference. arXiv preprint (2019). arXiv: 1912.02762. [PSDV18] PEREZ E., STRUB F., DE VRIES H., DUMOULIN V., COURVILLE A.: FiLM: Visual reasoning with a general conditioning layer. In Proc. AAAI (2018), vol. 32. [PVC19] PRENGER R., VALLE R., CATANZARO B.: WaveGlow: A flow-based generative network for speech synthesis. In Proc. ICASSP (2019), pp.3617–3621. [PWP18] PHAM H. X., WANG Y., PAVLOVIC V.: Generative adversarial talking head: Bringing portraits to life with a weakly supervised neural network. arXiv preprint (2018). arXiv:1803.07716. [RMW14] REZENDE D. J., MOHAMED S., WIERSTRA D.: Stochastic backpropagation and approximate inference in deep generative models. In Proc. ICML (2014), pp.1278–1286.[RSR20] RAFFEL C., SHAZEER N., ROBERTS A., LEE K., NARANG S., MATENA M., ZHOU Y., LI W., LIU P. J., ET AL.: Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res.21, 140 (2020), 1–67. [SB18] SADOUGHI N., BUSSO C.: Novel realizations of speech-driven head movements with generative adversarial networks. In Proc. ICASSP (2018), pp.6169– 6173. [SB19] SADOUGHI N., BUSSO C.: Speech-driven animation with meaningful behaviors. Speech Commun.110 (2019), 90–100. [SCNW19] SMITH H. J., CAO C., NEFF M., WANG Y.: Efficient neural networks for real-time motion style transfer. ACM T. Graphic.2, 2 (2019), 13. [SDWMG15] SOHL-DICKSTEIN J., WEISS E., MAHESWARANATHAN N., GANGULI S.: Deep unsupervised learning using nonequilibrium thermodynamics. In Proc. ICML (2015), pp.2256–2265. [SE19] SONG Y., ERMON S.: Generative modeling by estimating gradients of the data distribution. In Proc. NeurIPS (2019). [SME21] SONG J., MENG C., ERMON S.: Denoising diffusion implicit models. In Proc. ICLR (2021). [SN17] SMITH H. J., NEFF M.: Understanding the impact of animated gesture performance on personality perceptions. ACM T. Graphic.36, 4 (2017), 49. [SSDK21] SONG Y., SOHL-DICKSTEIN J., KINGMA D. P., KUMAR A., ERMON S., POOLE B.: Score-based generative modeling through stochastic differential equations. In Proc. ICLR (2021). [SSKS17] SUWAJANAKORN S., SEITZ S. M., KEMELMACHER-SHLIZERMAN I.: Synthesizing Obama: learning lip sync from audio. ACM T. Graphic.36, 4 (2017), 95. [TFHG19] TSUCHIDA, S., FUKAYAMA, S., HAMASAKI, M., & GOTO, M.: AIST Dance Video Database: Multi-Genre, Multi-Dancer, and Multi-Camera Database for Dance Information Processing. In ISMIR 1, 5 (2019). [TRG22] TEVET G., RAAB S., GORDON B., SHAFIR Y., COHEN-OR D., BERMANO A. H.: Human motion diffusion model. arXiv preprint arXiv:2209.14916 (2022).[TvdOB16] THEIS L., VAN DEN OORD A., BETHGE M.: A note on the evaluation of generative models. Proc. ICLR (2016). [TWGM21] TAYLOR S., WINDLE J., GREENWOOD D., MATTHEWS I.: Speech- driven conversational agents using conditional Flow-VAEs. In Proc. CVMP (2021), pp.1–9. [VPHB21] VALLE-PÉREZ G., HENTER G. E., BESKOW J., HOLZAPFEL A., OUDEYER P.-Y., ALEXANDERSON S.: Transflower: Probabilistic autoregressive dance generation with multimodal attention. ACM T. Graphic.40, 6 (2021), 1:1–1:13. [VPP19] VOUGIOUKAS K., PETRIDIS S., PANTIC M.: Realistic speech-driven facial animation with GANs. Int. J. Comput. Vision (2019), 1–16. [VSP17] VASWANI A., SHAZEER N., PARMAR N., USZKOREIT J., JONES L., GOMEZ A. N., KAISER Ł., POLOSUKHIN I.: Attention is all you need. In Proc. NIPS (2017), vol.30. [WFH08] WANG J. M., FLEET D. J., HERTZMANN A.: Gaussian process dynamical models for human motion. IEEE T. Pattern Anal.30, 2 (2008), 283–298. [WH21] WENNBERG U., HENTER G. E.: The case for translation-invariant self- attention in transformer-based language models. In Proc. ACL-IJCNLP (2021), pp. 130–140. [WMK14] WAGNER P., MALISZ Z., KOPP S.: Gesture and speech in interaction: An overview. Speech Commun.57 (2014), 209–232. [XWCH15] XIA S., WANG C., CHAI J., HODGINS J.: Realtime style transfer for unlabeled heterogeneous human motion. ACM T. Graphic.34, 4 (2015), 119. [YKJ19] YOON Y., KO W.-R., JANG M., LEE J., KIM J., LEE G.: Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots. In Proc. ICRA (2019). [ZBLZ22] ZHAO M., BAO F., LI C., ZHU J.: EGSDE: Unpaired image-to-image translation via energy-guided stochastic differential equations. arXiv preprint arXiv:2207.06635 (2022). [ZCP22] ZHANG M., CAI Z., PAN L., HONG F., GUO X., YANG L., LIU Z.: MotionDiffuse: Text-driven human motion generation with diffusion model. arXiv preprint arXiv:2208.15001 (2022).

Claims

CLAIMS 1. A method (100) for providing a model suitable for audio-driven body motion synthesis, the method comprising: obtaining (110) combined audio training data and body-motion training data; forming (120) conditioning data, wherein the conditioning data includes the audio training data; providing (122) a model of body pose sequences using a trainable diffusion probabilistic model; and training (124) the model on the basis of the body-motion training data and the conditioning data.

2. The method of claim 1, wherein the audio training data is speech training data, and the body-motion training data is gesture training data.

3. The method of claim 1, wherein the audio training data is music training data, and the body-motion training data is dance training data.

4. The method of any of the preceding claims, wherein the conditioning data includes, for each time instance, past and future data.

5. The method of any of the preceding claims, wherein the conditioning data includes future audio data.

6. The method of any of the preceding claims, further comprising: obtaining (118) a time series of one or more style-control parameters from the motion training data, wherein the conditioning data further comprises the style-control parameters.

7. The method of claim 6, wherein the style-control parameters include style information computed or extracted from the body-motion training data, such as one or more of: hand height, hand speed, gesticulation radius, correlation of right- and left-hand movements, root motion.

8. The method of claim 6 or 7, wherein the style-control parameters include manual or semi-automatic annotations associated with the body-motion training data.

9. The method of any of the preceding claims, further comprising: providing (116) additional training data by combining the audio training data with processed motion training data, such as right-left mirrored motion data.

10. The method of any of the preceding claims, further comprising: preprocessing (112) speech training data constituting the audio training data to obtain a spectrogram, such as a mel-frequency power spectrogram, log-power spectrogram, or mel-frequency cepstrum coefficients (MFCCs), or preprocessing (112) music training data constituting the audio training data to derive MFCC, spectral flux, chroma, activation of a beat model, and / or beat tracking.

11. The method of any of the preceding claims, wherein the body-motion training data comprises motion-capture data, such as joint-angle data.

12. The method of any of the preceding claims, further comprising: preprocessing (114) the body-motion training data into an exponential-map representation.

13. The method of any of the preceding claims, further comprising: preprocessing (114) the body-motion training data into a representation using two axes out of three from 3×3 rotation matrices.

14. The method of any of the preceding claims, further comprising: preprocessing (114) the body-motion training data by one or more of the following: time synchronization, downsampling, coordinate conversion, conversion to coordinates relative to a T-pose.

15. The method of any of the preceding claims, wherein the diffusion probabilistic model includes a neural network.

16. The method of claim 15, wherein the neural network includes one or more of: a neural attention mechanism; convolution operations; a convolutional neural network, CNN; a transformer; a conformer; a recurrent neural network, RNN; a long short-term memory unit, LSTM unit; a gated recurrent unit, GRU.

17. The method of claim 16, wherein the neural network includes a neural attention mechanism having attention weights dependent on sequence position,wherein optionally the dependence of the attention weights on sequence position is encoded using a translation-invariant scheme.

18. The method of claims 15 to 17, wherein a nonzero dropout rate is applied to the neural network and / or its inputs.

19. The method of claim 18, wherein dropped inputs to the neural network are replaced by special values representing absence of information.

20. A method (200) for audio-driven body motion synthesis using a model obtainable by the method of any of the preceding claims, the method comprising: obtaining (210) an audio signal; and generating (218) a sequence of body poses from the model conditioned upon at least the audio signal.

21. The method of claim 20, further comprising: obtaining (214) values of one or more style-control parameters, wherein the model, from which the sequence of body poses is generated, is further conditioned upon the obtained style-control parameters.

22. The method of claim 21, further comprising: providing (216) a segmentation of the audio signal and a segmentation of the style- control parameter values, wherein the segmentations comprise respective sequences of time segments which are pairwise aligned in time between the segmentations.

23. The method of claim 21 or 22, wherein generating (218) the sequence of body poses includes: controlling (218.1) a magnitude of the conditioning upon the style-control parameters by means of classifier-free guidance, for thereby amplifying or reducing a style of the sequence of body poses.

24. The method of claim 21 or 22, wherein the model includes - a first sub-model having been trained on the basis of conditioning data including the style-control parameters, and - a second sub-model having been trained on the basis of conditioning data not including the style-control parameters,wherein said generating (218) the sequence of body poses comprises a plurality of denoising steps (218.2), in each of which a weighted combination of the output of the first sub-model and the second sub-model is applied as guidance, wherein the weighting corresponds to magnitude at which the style-control parameters are to be expressed.

25. The method of claim 21 or 22, wherein the model includes - a first sub-model constituting a basic model which has been trained on the basis of conditioning data including the style-control parameters and which is evaluated for a first value of a conditioning input, and - a second sub-model constituting the basic model evaluated for a second value of the conditioning input, wherein said generating (218) the sequence of body poses comprises a plurality of denoising steps (218.2), in each of which a weighted combination of the output of the first sub-model and the second sub-model is applied as guidance, wherein the weighting corresponds to magnitude at which the style-control parameters are to be expressed.

26. The method of claim 25, wherein the first value or the second value of the conditioning input corresponds to an absence of style expression.

27. The method of claim 25, wherein the first value and the second value of the conditioning input correspond to a presence of style expression.

28. The method of any of claims 24 to 27, wherein the guidance applied in each denoising step (218.2) comprises a barycentric combination of the first and second sub-models’ denoising predictions.

29. The method of claim 21 or 22, wherein generating (218) the sequence of body poses contains a weighted combination of several denoising-model predictions in a product-of-experts configuration.

30. The method of claim 29, where the different denoising model predictions are given by the same model but with different conditional inputs.

31. The method of claim 21 or 22, wherein the model includes - a first sub-model having been trained on the basis of conditioning data including a first set of one or more style-control parameters, and- a second sub-model having been trained on the basis of conditioning data including a second set of one or more style-control parameters, which is distinct from said first set, wherein said generating (218) the sequence of body poses comprises a plurality of denoising steps (218.2), in each of which a weighted combination of the output of the first sub-model and the second sub-model is applied as guidance, wherein the weighting corresponds to a ratio of the magnitudes at which the first and second sets of style-control parameters are to be expressed.

32. The method of claim 21, wherein the first and second sets of style-control parameters are disjoint.

33. The method of claim 21, 22, 25 or 27, wherein the model further includes - a third sub-model having been trained on the basis of conditioning data not including any style-control parameter, and, in each denoising step (218.2), a weighted combination of the outputs of the first, second and third sub-models is applied as guidance.

34. The method of claim 20 to 33, wherein the audio signal represents speech.

35. The method of claim 20 to 33, wherein the audio signal represents music.

36. The method of any of claims 20 to 35, further comprising: applying (220) the sequence of body poses to a digital or physical character.

37. The method of claims 20 to 36, further comprising: preprocessing (212) the audio signal as specified in claim 10.

38. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any of the preceding claims.

39. A device comprising memory and processing circuitry configured to carry out the method of any of claims 1 to 37.