Method, device and computer program

JP2025505340A5Pending Publication Date: 2026-01-09MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024536188
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-01-31
Filing Date
2023-01-06
Publication Date
2026-01-09

AI Technical Summary

Benefits of technology

【0010】 本明細書において開示される第3の態様によると、第1の態様の方法を実行するためのプロセッサにより実行可能な命令を含むコンピュータ可読ストレージデバイスが提供される。この「発明の概要」は、以下の「発明を実施するための形態」において更に説明される概念のうちの選択されたものを単純化形式で導入するために提供される。この「発明の概要」は、請求される主題の重要特徴又は必須特徴を特定することを意図していないし、請求される主題の範囲を制限するために使用されることも意図していない。請求される主題もまた、本明細書に記載の任意の又はすべての欠点を解決する実装形態に制限されない。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

1. A computer-implemented method comprising: receiving video data from a user from a user device; training a first machine learning model based on the video data to provide a second machine learning model personalized to the user, the second machine learning model being trained to predict the user's movements based on the audio data; receiving another audio data from the user; determining predicted movements of the user based on the another audio data and the second machine learning model; and using the user's predicted movements to generate an animation of an avatar for the user.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Technical Field The present disclosure relates to methods, apparatus, and computer programs for animating a user's avatar. In particular, some examples relate to generating a signature movement of the user in response to voice input from the user, which in some examples may be used in a communication session with other users. [Background technology]

[0002] background Users may use avatars to represent themselves while communicating with other users, and such avatars may be customized by users to change the appearance of the avatar.

[0003] The avatar may be based on the user's appearance to provide a likeness to the user, allowing the user to express their personality to other users.

[0004] When communicating with other users using only voice and avatars generated according to known methods, the user experience is degraded compared to video calling: the characteristic movements of the user (e.g. facial movements, facial expressions, head pose) are lost when using a generic avatar to communicate as opposed to using a video or user-specific avatar to communicate.

[0005] Not only the visual appearance but also the visual movement of the avatar is part of expressing the user's personality. The avatar movement can be enabled by animation derived from video or audio signals. For a given communication service, this animation can be generic for all users or can be customized to be unique for each user.

[0006] In prior art systems, when communicating with others by using an avatar that is animated by audio signals only, the person's visual features face, head and / or body movements cannot be visually tracked to enable animation. The lack of visual movement by the avatar can degrade a user's ability to express their own personality or to perceive the personality of another user.

[0007] In conventional systems, avatars animated by audio signals alone may have visual-motor animations ranging from very simple, such as mouth flapping with no other motion, to full motion. The mapping of audio to motions may be artificial (i.e., unrelated to audio except for the presence or absence of audio), or based on a general model (i.e., by using some known audio to animate some known motions, such as rolling the mouth while saying "O"), or some combination thereof. Current audio-driven animations focus on the mouth or on the mouth and facial expressions. Other head or body movements tend to be absent or tend to be quite artificial and general to all users of the service. Summary of the Invention [Means for solving the problem]

[0008] overview According to a first aspect disclosed herein, a computer-implemented method is provided. The method includes receiving video data from a user device from a user. The method further includes training a first machine learning model based on the video data to generate a second machine learning model personalized to the user, the second machine learning model being trained to predict the user's movements based on the audio data. The method still further includes receiving audio data from the user and determining the user's predicted movements based on the separate audio data and the second machine learning model. The method also includes using the user's predicted movements to generate an animation of the user's avatar.

[0009] According to a second aspect disclosed herein, there is provided an apparatus configured to perform the method of the first aspect.

[0010] According to a third aspect disclosed herein, a computer-readable storage device is provided that includes instructions executable by a processor for performing the method of the first aspect. This Summary is provided to introduce in a simplified form selected concepts that are further described in the Detailed Description below. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. The claimed subject matter is also not limited to implementations that solve any or all of the disadvantages described herein.

[0011] BRIEF DESCRIPTION OF THE DRAWINGS To assist in understanding the present disclosure and to show how embodiments may be carried into effect, reference is made to the accompanying drawings, by way of example only, in which: [Brief description of the drawings]

[0012] [Figure 1] FIG. 1 is a schematic block diagram of an exemplary computing system for performing the methods disclosed herein. [Figure 2A] FIG. 1 is a schematic diagram illustrating an exemplary training aspect of the model disclosed herein. [Figure 2B] FIG. 1 is a schematic diagram illustrating an example personalization aspect of the model disclosed herein. [Diagram 3] FIG. 2 is a schematic diagram of an exemplary testing aspect of the model disclosed herein. [Figure 4] 1 is a diagram illustrating an example method flow according to some embodiments disclosed herein. [Diagram 5] 1 is an exemplary user interface according to certain embodiments disclosed herein. [Figure 6]1 is an exemplary neural network that may be used in the examples disclosed herein. [Figure 7] 1 illustrates an exemplary user device. [Figure 8] 1 illustrates an exemplary computing device. [Figure 9] An exemplary method flow is shown. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Detailed Description The present disclosure relates to systems, methods, and computer-readable storage devices for providing a user avatar that provides an authentic representation of a user by exhibiting characteristic movements of the user. Such characteristic movements may include the user's facial movements, facial expressions, and head poses. In some examples, the user avatar may provide a identifiable similarity to the user by not only having a similar appearance to the user but also exhibiting the user's characteristic movements.

[0014] An exemplary application is for digitally or virtually engaging one or more other users, for example, in a call or other form of communication session. The system may receive voice input from a user represented by an avatar. Based on the voice input, the system may generate characteristic movements of the user that are used to drive the animation of the avatar that is displayed as the voice is output. When used during a phone call, this allows one or more other users on the call to not only identify that the avatar is being used by an expected person based on their characteristic movements, but also to get along with and fully engage with the avatar. This may reduce cases of identity theft associated with the use of avatars during communication, thus increasing security.

[0015] There are several reasons why a user may choose to be represented by an avatar during a communication (e.g., a phone call) with one or more other users. For example, a user may not be in an appropriate environment to be recorded during a video call and therefore may not want their actual appearance to be used during the communication. Such reasons may include one or more of the following: the user may not have appropriate lighting; the user may not be appropriately dressed; the user may not want to reveal their actual location that may be inferred from the background during the phone call due to privacy / security issues.

[0016] Another situation in which the use of an avatar to communicate with other users may help a user is a situation in which the user has limited bandwidth and / or limited processing resources. Using an avatar instead of a video representation of the user may reduce the required bandwidth and / or processing resources to communicate with other users compared to a situation in which the user is represented by a video. Although it may not be possible to record and transmit the user's streamed video to other users, the use of an avatar may allow the option for the user's representation to be displayed with a smaller bandwidth or with fewer required processing resources. For example, a scenario may be considered in which a user is about to board a train. The user may realize or be informed during the train trip that he or she will not have enough bandwidth to be represented by a streamed video during the train trip. The user may then choose to be represented by an avatar rather than being represented by a streamed video.

[0017] Another example is envisaged where avatars are used within a mixed reality environment. Some embodiments of the present invention improve the representation of a user by an avatar by representing characteristic movements of the user through the avatar. Such characteristic movements may include, for example, facial movements, facial expressions and head pose. For example, during a mixed reality experience, when a user is represented by an avatar and speech from the user is output, a prediction of how the user will move during the speech output may be made and represented by the avatar. For example, if a user will nod their head when articulating some phrases or speaking at a certain volume or intonation, this may be represented by the avatar. Facial movements may be considered to include movements of facial features (e.g., eyebrow movements, forehead movements, mouth movements, nose movements, eye movements, etc.). Head pose may be considered to include: roll angle; pitch angle; yaw angle, which provides the orientation of the head. Thus, there are three degrees of freedom of movement of the head pose. According to some examples, each of the roll angle, pitch angle and yaw angle may have a respective limit to represent the limit of neck flexibility of the avatar.

[0018] A more comprehensive avatar experience may be provided by representing the characteristic appearance and movements of the user during communication with other users. The characteristic movements of the user, particularly the user's face and head, may be captured and represented by the avatar. This provides an experience that represents the user according to the user's personality and the user's culture, and allows any physical condition (e.g., a medical condition such as facial spasm) to be shown and represented. Such representations may be customized by the user. This customization may allow the user to select how they are represented by their avatar.

[0019] By using an avatar to represent the characteristic movements of a user, an impression of the user's presence may be more accurately provided to other users during a communication session (e.g., during a mixed reality experience, video call, etc.). This provides a more authentic experience to the user during the communication session. This also provides more secure communications since users who know the person represented by the avatar and / or have already spoken to the person by using the avatar may recognize the characteristic movements of the person represented by the avatar, ensuring that they are communicating with that person rather than someone else. Additionally, as discussed above, bandwidth and processing resources may be conserved by using avatars during a communication session.

[0020] According to some examples, an avatar may be considered to include a visual digital representation of a user. An avatar may change pose during communication with other users. The pose change may be made dependent on audio data received from the user represented by the avatar. The pose change may include a change in the avatar's head position. In some examples, an avatar may be based on the user's appearance. An avatar may be generated based on image data corresponding to a user. In some examples, the image data may be used to generate a mesh corresponding to the user's appearance, and the image data may then be used to overlay additional details (e.g., skin, hair, makeup) onto the mesh to generate an avatar corresponding to the user's appearance. Thus, an avatar may represent a high likeness of the user. However, in some examples, a user may choose to customize or modify the appearance of the avatar such that the avatar does not represent a high likeness of the user. An avatar may be considered to include a virtual representation of a user. In some examples, an avatar may include a representation of the user from the shoulders up or from the neck up. As an example, mesh 212 may be a unique geometric shape designed by a designer (e.g., an artist) and informed by a three-dimensional scan of a human head. However, any mesh representation of the head geometry can be used.

[0021] The appearance of the avatar may be based on the mesh 212 using any known method. In some examples, the mesh 212 may be based on a different initial mesh that is fitted to correspond to the user's appearance. A texture model may also be overlaid on top of the mask to provide an avatar that corresponds to the user's appearance. Exemplary methods of fitting 3D models are discussed, for example, in U.S. Patent No. 11127225B1, entitled "Fitting 3D models of composite objects." In some examples, the appearance of the user's avatar may be determined separately from the method of FIG. 2B. In some examples, the appearance of the avatar may be determined during the process of FIG. 2B based at least in part on image data received at audio and image data 208. In some examples, the avatar provides an approximate representation of the user such that a similarity between the user appearance and the avatar appearance is provided. As discussed further below, in some examples, the avatar appearance determined at this stage may be later customized by the user.

[0022] Figure 1 illustrates an example system 100 that may be used to implement some examples of the present invention. It will be appreciated that system 100 is not limited to only the devices illustrated in Figure 1. The avatar modeling system may be accessed by each of the devices in system 100 and may be hosted across one or more devices in system 100.

[0023] The user device 102a may include any suitable user device for using an avatar modeling application. For example, the user device 102 may include a mobile phone, a smartphone, a head mounted device (HMD), smart glasses, smart wearable technology, a laptop, a tablet, a HoloLens, etc.

[0024] User device 102a may communicate with one or more other user devices 102b-102n over network 104. In some examples, one or more other user devices may communicate directly (e.g., by using Bluetooth technology) without using network 104. Network 104 may include any suitable network (e.g., an Internet network, an Internet of things (IoT) network, a local area network (LAN)), a 5G network, a 4G network, a 3G network, etc. A user of user device 102a may be represented by an avatar when communicating with one or more of user devices 102b-102n.

[0025] The user device 102a may be used to receive a user's voice data and corresponding image data. In some examples, the image data and voice data may include video. The video data may include voice data synchronized with the image data. In some examples, the user's voice data and video data may be captured by directly using the user device 102a's voice receiver (e.g., microphone) and image receiver (e.g., camera). In some examples, the user's video data may be downloaded or uploaded to the user device 102a. The user's video data may be used by a machine learning (ML) model to learn how the user moves when communicating. For example, the user may discuss some subjects, say some words, make some noises (e.g., laugh, cry), use some intonations, perform some facial movements or facial poses while increasing the volume of their speech or decreasing the volume of their speech. The user device 102a may also be configured to receive voice data without the corresponding image data.

[0026] The one or more ML models may be implemented in the user device 102a or in one or more of the computing devices 106a and 106n. Each of the computing devices 106a and 106n may include one or more computing devices for processing information. Each computing device may include at least one processor and at least one memory as well as other components. In some examples, one or more of the computing devices 106a, 106n may be hosted in the cloud such that one or more of the computing devices 106a-106n include cloud computing devices.

[0027] In some examples, one or more ML models may be implemented on a combination of one or more of the user device 102a, computing device 106a-106n. In some examples, training of the ML model may occur on the user device 102a. In some examples, training of the ML model may occur on the computing device 106a. In some examples, training of the ML model may occur partially on the user device 102a and partially on 106a. In some examples, if the user device 102a is in a low power state or has limited resources, it may be useful for the training of the ML model to occur at least partially on the computing device 106a. One or more ML models described herein may be implemented in the cloud or in the user device for reasons of scale: for example, in the case of a broadcast, lecture, or game show (such as a 1v100 / quiz show) where many participants are watching a few animated avatars of lecturers, actors, game show participants, etc., it may be useful to generate the avatar animations in a cloud device. Where bandwidth is limited or where there is greater symmetry of participant interaction and / or hardware (such as in a conference), animation can occur at the user device or in some variation of the two ways. There are also scenarios where different participants in the same conference may choose local or cloud computing based on their preferences, and indeed where part of the animation occurs in the cloud and the rest occurs on the device.

[0028] Once an ML model is trained that can predict a user's movements during a communication based on the voice data, the ML model can be used to predict the user movements and to animate an avatar accordingly during the communication. In some examples, the ML model can be used to generate avatar movements at the user device 102a. In some examples, the ML model can be used to generate avatar movements at the computing device 106a. In some examples, the ML model can be used to generate avatar movements partially at the user device 102a and partially at the computing device 106a. In some examples, it can be useful for the generation of avatar movements to occur at least partially at the computing device 106a when the user device 102a is in a low power state or has limited resources.

[0029] FIG. 2A illustrates an exemplary method 200 for training a base model. The exemplary method trains a “base” model f that is considered “personalizable” for a particular user. θ The basic model f θ is trained (or meta-trained) on an existing training dataset. The existing dataset will have many users, with each user having one or more corresponding videos of the user speaking. In some examples, the one or more corresponding videos of each user will include a small number of videos of each user. In some examples, the small number of videos may range from 1 to 5. Each of the videos in the training dataset will have labeled head vertices. Each of the videos may include audio data synchronized with the image data.

[0030] In general, the basic model θ During training (or meta-training) to provide a large dataset of users, where each user has one or more videos of the user, each of which has labeled vertices of the user's face, a generic (non-individualized) model g that takes in audio and images (video) and outputs the corresponding head pose and motion.θ 209. By doing this across a large dataset, the model f θ knows how to individuate: once deployed, the base model f θ can be given a very small number of images of a completely new user and knows how to process those images to become personalized to that user. Thus, the basic model f θ can be considered as a personalizable model. Next, the basic model f θ can be used to provide personalized avatar motion to any new user with very few images of them. In some examples, the base model f θ may be considered to include the “first” ML model. In some examples, model g θ 209 can be considered to include a "third" ML model.

[0031] The above paragraph is the basic model θ We will explain the general method for determining the basic model f θ A specific method for determining f is now considered with respect to FIG. 2A. FIG. 2A shows a “base” model f that is considered “personalizable” at the end of training. θ FIG. 2A illustrates a specific example of a method 200 for determining a personalizable model f θ2A shows an example of how the user's image 209 may be determined, but it will be understood that other suitable methods may be used. The first user's dataset 203 may be randomly sampled from the training dataset. The dataset 203 may include a first video 205a, a second video 205b, a third video 205c, and a fourth video 205d. Each of the first video 205a, the second video 205b, the third video 205c, and the fourth video 205d may have a labeled user's head vertex. Each of the first video 205a, the second video 205b, the third video 205c, and the fourth video 205d may correspond to the same user. The dataset 203 may be considered to include a "context dataset" of the first user. In other examples, the dataset 203 may include more or less than the four videos shown in FIG. 2A. Data set 203 is of the same user as data set 221, which includes audio track 1) 207a, audio track 2) 207b, and audio track 3) 207c. Data set 221 may also include corresponding video for each of audio track 1) 207a, audio track 2) 207b, and audio track 3) 207c. Data set 221 may be considered to include a "target dataset" for the same user as context dataset 203.

[0032] Model G θ 209 includes a neural network used to predict user head movements. θ 209 may, in some examples, be initially set based on predictions of multiple users' head movements while speaking. θ The weights of 209 can be initialized with default values. θ The weightings of 209 may be initialized by user-defined values.

[0033] The context data set 203 is used to generate an updated model g′ that is adapted to a particular user by using the context data set 203. θ Model G to provide 211 θIn some examples, the context data set 203 is used to update the weights of the model g such that the model EQ 209 can generate head movements for the avatar based on the audio track of the video. θ 209. The difference between the head vertex motion and the video corresponding to the audio track user can then be minimized to provide the difference between the head vertices of the model avatar during the motion. This results in an updated model g′ that is adapted to the particular user represented in the context dataset 203. θ This may be performed for each of the images 205a, 205b, 205c, and 205d in the context dataset 203 to provide the model g′ θ can be considered to include a "fourth" ML model.

[0034] Model G θ 209 to adapt the updated model g′ to the particular user represented in the context data set 203. θ The target data set 211, which includes audio track 1) 207a, audio track 2) 207b, and audio track 3) 207c, is then provided as an updated model g′ θ 211. It should be noted that the target dataset 211 may contain more or less than the three audio tracks shown in the example of FIG. 2A. Next, the updated model g′ θ 211 uses input audio track 1) 207a to provide predicted head movement for input audio track 1) 213a, input audio track 2) 207b to provide predicted head movement for input audio track 2) 213b, and input audio track 3) 207c to provide predicted head movement for input audio track 3) 213c.

[0035] In 215, the predicted head movements 213a, 213b, and 213c are compared to the video corresponding to each of the audio tracks 1), 2), and 3). The comparison may be based on the true vertices of the labeled user's head in the video and the vertices of the predicted head movements in 213a, 213b, and 213c. An error is calculated based on the predicted head movement of audio track 1) 213a and the video corresponding to audio track 1) 207a. An error / loss is calculated similarly for the predicted head movement of audio track 2) 213b and the predicted head movement of audio track 3) 213c.

[0036] In 217, the error calculated in 215 is backpropagated to the original model g θ 209. This may be done, for example, by using a gradient descent algorithm so that the error is minimized in 217 for the users of data set 203 and data set 221.

[0037] In 219, the process of 203 to 217 is a personalizable basic model f θ In some examples, the process of 203-217 is repeated for randomly sampled users until the error converges from one user to the next in the sample.

[0038] FIG. 2B shows a personalized model f' for a particular user. θ 2 illustrates an example personalization aspect using system 201 to provide 210, which may predict characteristic movements of a user during a communication (e.g., during a conversation). The personalized ML model may in some instances be considered to include a "second" ML model.

[0039] The system 201 may be implemented in a user device, in a computing device (e.g., a cloud computing device) connected to a network of user devices, or across a combination of user devices and computing devices. It may be useful to provide the personalization aspects of the system 201 at least partially in a computing device such as device 106a when a user device (e.g., user device 102a) is in a low power state or has limited resources. The system 201 includes a personalizable base model f, which may be generated by using a method similar to that described above with respect to FIG. 2A. θ In some examples, the basic model f θ may be stored on a user device (eg, user device 102a).

[0040] The user audio and image data 208 may include audio data of the user synchronized with image data showing one or more facial expressions of the user during the communication. The user audio and image data 208 may be considered to include video data. For example, the image data may include one or more images showing head poses and facial expressions corresponding to certain audio segments of the audio data.

[0041] In some examples, audio and image data 208 may include one or more videos of a user speaking, each of which may include an audio track and synchronized video data.

[0042] The user audio and image data 208 may be captured by an image receiver and an audio receiver of a user device (e.g., the user device 102a). The user audio and image data may also be downloaded or uploaded to the user device 102a.

[0043] In some examples, to use an application on a user device, such as user device 102a, a user may be required to provide a video of themselves speaking into the application. For example, the application may require a 30-second clip of the user speaking into the application. This video may be used to provide user audio and image data 208.

[0044] User voice and image data 208 and the personalizable base model f θ Based on this, in operation 221, a personalized model f' θ 210 is prepared for the user. Unlike the video used in the training phase (e.g., in the method of FIG. 2A), the audio and image data may not have labeled vertices. According to some examples, a personalized model f′ θ 210 is optimized by using a few-shot learning technique. Exemplary few-shot learning techniques may include: ●Finetuning methods (e.g., How transferable are features in neural networks? Yosinski et al, 2014 CoRR abs / 1411.1792 https: / / arxiv.org / abs / 1411.1792) ● Gradient-based meta-learning methods (e.g., Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks, Finn et al., Proceedings of the 34th International Conference on Machine Learning, PMLR 70:1126-1135, 2017); Metric-based meta-learning methods (e.g., ProtoNets such as those described in Prototypical Networks for Few-shot Learning, Snell et al., 19 June 2017); and ●Model-based meta-learning methods (e.g., CNAPs such as those described in Fast and Flexible Multi-Task Classification Using Conditional Neural Adaptive Processes, Advances in Neural Information Processing Systems 32 (2019)7957-7968, Requiema et al., 2019).

[0045] Based on the voice data and image data 208, in 221, the system 201 generates a personalized model f' of the user. θ 210 (which can predict how the user's facial expressions, head pose, and lip movements will change based on the received voice input from the user). Next, we optimize the personalization model f' θ 210 may be used to animate an avatar while a user is providing voice input, as discussed below with respect to FIGURE 3. This may be used to communicate with other users in a communication session (e.g., a mixed reality experience).

[0046] Operation 221 is to generate a personalized model f' of the user. θ To provide 210 basic models F θ Then, as discussed below with respect to FIG. θ 210 can be used to predict changes in head pose, facial expression and lip movement for future speech input into the system. θ is trained to learn how to personalize the model according to the input video data by using the method discussed above with respect to FIG. 2A, so that the video data (user voice and image data 208) is θ When input to the personalizable base model f θ is the personalized model f' θThe operation 221 may be personalized to provide 210. Operation 221 may include a few-shot learning technique as disclosed herein.

[0047] Personalized model f' θ 210 may also include a mesh 212 that corresponds to the user's face and may be used as an avatar. The mesh may be generated based on video that includes an image of the user's face. In some examples, the mesh 212 may be generated during operation 221. In other examples, the mesh may be generated in a separate step based on an image of the user. In other examples, the mesh 212 may correspond to a pre-set mesh selected by the user. The mesh 212 may be two-dimensional or three-dimensional. The mesh 212 may be used to generate an avatar representing the user. Movements such as changes in head pose, changes in facial expression, and lip movement may be used to move the mesh 212.

[0048] Figure 3 shows how the ML model f' is used to predict characteristic movements based on the user's input voice. θ 210 shows an example of how the method of FIG. 3 may be used. The method of FIG. 3 may be considered to include an inference phase. In some examples, it may be useful for the inference phase to occur at least partially in computing device 106a when user device 102a is in a low power state or has limited resources. The inference phase of system 300 may be implemented in a user device and in a computing device (e.g., a cloud computing device) connected to a network of user devices or implemented across a combination of user devices and computing devices. It may be useful to provide the inference phase of system 300 at least partially in a computing device such as device 106a when a user device (e.g., user device 102a) is in a low power state or has limited resources.

[0049] The speech data 314 is used to generate a personalized model f' θThe voice data 314 and the user voice and image data 208 may be from the same user. The voice data 314 may be different data relative to the voice data of the voice and image data 208. The personalization model f' θ 210 may be pre-trained as described herein (e.g., with respect to FIG. 2A and then with respect to FIG. 2B). The speech data is then used to generate a personalization model f′ to output predictions 318 of the user's lip movements, facial expressions, and head pose depending on the speech data. θ Sent to 210.

[0050] The lip movement, facial expression, and head pose predictions 318 may be used at 320 to generate avatar animation at 322. In some examples, the operations 320 may include one or more further actions based on the user's predicted lip movement, facial expression, and head pose. For example, it may be determined not to provide one or more lip movements, facial expression changes, and head pose changes in animating the avatar. This determination may be based on user input or may be set by the system based on processing constraints. Such user input may be provided by a user customizing the avatar animation. For example, a user may determine not to provide some facial expression changes if the user has a condition such as Tourette syndrome. Whether or not to animate the avatar may also be determined at 320, and furthermore, shape parameters may be changed at 320 by the user to change the shape of the mesh underlying the avatar. In this manner, users may customize their avatar movements to represent their personalities according to their preferences.

[0051] At 322 , avatars and their respective animations are provided based on the predictions of 318 and the additional information received at 320 .

[0052] At 324, the avatars and their respective animations may be output to one or more of: a user device of the user represented by the avatar (e.g., user device 102a); a computing device (e.g., computing device 106a); or another user device belonging to another user (e.g., user devices 102b, 102n). The avatars and their respective animations may be output in parallel with the audio data 314.

[0053] FIG. 4 illustrates an example method flow between a user device 102a, a computing device 106a (which may include a cloud computing device), and a user device 102n.

[0054] In 430, a personalizable base model f θ is accessed by at least one of the user device 102a and the computing device 106a. θ can be generated as described above with respect to FIG. 2A. θ may be stored on at least one of the user device 102a and the computing device 106a. In some examples, the computing device 106a may store the personalizable base model f θ to the user device 102a. In some examples, the user device 102a may transmit the personalizable base model f θ to the computing device 106a.

[0055] 432 and 434 correspond to Fig. 2A. At 432, the user device 102a receives video data. This may be captured by the user device 102a or uploaded or downloaded to the user device 102a. The video data may include one or more videos. The videos may include audio data and synchronized image data. The videos may include videos of a user operating the user device 102a.

[0056] In 434, the personalizable basic model fθ is the personalized model f' θ This personalized model f' is personalized by using the user's video data to provide θ As discussed above, the personalizable basic model f θ may be personalized at the computing device 106a, at the user device 102a, or across both the user device 102a and the computing device 106a. 434 represents the personalization model f′ θ This may include using a few-shot learning technique to provide

[0057] At 436, the user optionally selects the personalized model f′ trained at 432. θ This is discussed further below with respect to FIG. 5. Such customization can be achieved by customizing the personalized ML model f′ θ Some of the customization operations may be hosted at the computing device 106a in some examples. Selection of the customization options is performed at the user device 102a.

[0058] 436, 438, 440 and 442 correspond to Figure 3. In some examples, the user may choose to use an avatar at 435 during the communication session prior to step 436. At 436, voice data is received by the user device 102a from the user.

[0059] At 438, the speech data received at 436 is processed to generate a personalized model f′ to predict changes in lip movement, facial expression, and head pose corresponding to the speech data. θThe avatar is passed into and used. As discussed above, this may be done at the user device 102a, at the computing device 106a, or across both the user device 102a and the computing device 106a. The avatar may be further modified by the user by specifying other constraints at 438. In some examples, the shape parameters of the avatar may be modified at 438.

[0060] In some examples, at 440, avatars and animations corresponding to the audio data received at 436 may be selected for use by a user.

[0061] At 442, the animated avatar can be used along with the voice data received at 436 to communicate with the user 102n in a communication session. This could be in a video call (where the user of the user device 102a is replaced by an avatar), in a mixed reality environment, or in any other suitable communication session that uses a representation of the user of the user device 102a in parallel with the user's voice.

[0062] 5 illustrates an exemplary user interface 500. It should be understood that the user interface 500 is merely an example and that fewer or more user selections may be displayed to the user. It should also be understood that different layouts of the user interface 500 may be provided. The user interface may be provided on a user device, such as user device 102a.

[0063] At 550, the user may view their avatar's lip movements, facial expression changes, and head pose changes as speech input is provided. The speech input may have been previously provided to an application running user interface 500. The user may choose to provide an alternative speech input at 554 so that they may view the avatar's lip movements, facial expression changes, and head pose changes as other speech inputs are provided. Selection of option 554 may open another menu allowing the user to select or provide another speech input.

[0064] In options 552, users may be able to further customize their avatar animation. Users may be able to increase or decrease some features or characteristic movements of the animation. Thus, the user may change the degree to which at least one of the predicted movements is made. This may be useful if the user has certain characteristic movements that they would like to emphasize or not show. As shown in the example of FIG. 5, sliders may be used to increase or decrease these features or characteristic movements of the animation, although other options are envisioned (e.g., providing a numerical value from some threshold that the user may change for each feature). Although three features are shown in FIG. 5, more or fewer features may be provided to the user. As an example, feature X may correspond to eyebrow movement, feature Y may correspond to head movement, and feature Z may correspond to the overall expressiveness of the avatar. These features and characteristic movements may be increased by at least one of the following: increasing the avatar's movement vector when the user increases the feature; or increasing the frequency with which the avatar performs the feature. Increasing the movement vector may include increasing the amount that the avatar's feature moves. For example, increasing the motion vector of eyebrow movement may include increasing the distance the eyebrows move during an avatar animation. In another example, increasing the overall expressiveness of the avatar may include increasing the frequency with which the avatar moves during an avatar animation. Option 552 may be used in some examples to dampen the user's twitching, if desired.

[0065] In some examples, a user of a user device receiving an animation of an avatar of another user (e.g., a user using user device 102 n) is provided with options similar to those described with respect to 552 to scale up or down the characteristic movements displayed to the receiving user.

[0066] In some examples, the changes in options 552 may be reflected in a corresponding change to the underlying ML model (e.g., by changing the weights of the neural network of the ML model). In some examples, the changes in options 552 may be reflected by making changes to the avatar animation predicted by the underlying ML model, without changing the ML model itself (i.e., without post-processing). This could include, for example, changing the motion vectors of the avatar motion predictions made by the ML model.

[0067] A user may have one or more different avatar personas. The personas may be based on different sets of corresponding image and voice data. In some examples, the personas are further edited by using personalization options such as the options shown at 552. In some examples, the personas may be based on the same corresponding image and voice data after different personalization by the user. Users may want to have more than one persona for their avatar in different social settings. For example, a user may want an avatar used with friends to be more expressive and an avatar used in a work setting to be less expressive. By selecting option 555, users may select a different persona previously set up and stored for their avatar, or may create a new persona. These personas may be stored for selection by the user as needed.

[0068] At 556, users may personalize the appearance of their avatar. This may include adding accessories to the avatar, changing clothing, changing hairstyles, etc.

[0069] At 560, the user is presented with the effect of training data on the output avatar animation. Different portions of the input image data and audio data may have different levels of influence on the output avatar animation. In the example of figure 560, a pie chart is used, but any other suitable illustration (e.g., a list of the percentages of each video) may be provided to the user. In the example of FIG. 5, three videos were used (although it will be understood that a different number of videos may be used in other examples). Thus, the portion of audio data and the corresponding portion of image data may be considered to include three videos. Video 2 had a greater influence on the output avatar animation than videos 1 and 3. This indicates that video 2 is dominant in training the underlying ML model. This could be, for example, because video 2 is longer than videos 1 and 3 or because the user was more expressive during videos 1 and 3. If the user is not satisfied with the way their avatar is animated, the user may choose to use option 558 to remove video 2. The user may also add another video by using this option. This option may be particularly useful if the user has a medical condition that affects their communication and varies in intensity over different times.

[0070] Figure 6 shows the neural network g θ , g' θ , f θ , f' θ6 illustrates an example neural network 600 that may be used in the examples of any of the above. It should be noted that the neural network is shown as an example only, and thus neural networks having different structures may be used in the methods and systems described herein. It should also be noted that neural networks having more layers and / or nodes may be used. Each node (e.g., node 662a, node 662b, node 662c, node 664) of the neural network 600 represents a floating point value. Each edge (e.g., edge 666a, edge 666b, edge 666c, etc.) represents a floating point value and is known as a "weight" or "parameter" of the model. These weights are updated as the model is trained according to the process described above.

[0071] In this example of neural network 600, the input is a 3-dimensional vector and the edges are 24-dimensional vectors (12 edges connect the input to layer 1, and 12 edges connect layer 1 to the output).

[0072] In one example, to calculate the value of a node in a given layer: take each node connected to it from the previous layer, multiply each node by its associated edge value / weighting, and then add these together. Then a non-linear function (e.g. sigmoid or tanh) is typically applied to this value. This is repeated to get the value of each node in a given layer, and so on for each layer until the output layer is reached.

[0073] In the example of neural network 600, the output is a three-dimensional vector that may represent the probability that the model predicts that the input belongs to three possible object classes. In the example of neural network 600, the neural network has one hidden layer of nodes. For deep neural networks, which may in some examples be used as neural networks in the methods disclosed herein, there are usually two or more hidden layers.

[0074] The value of node 662a is 2=n 11 and the value of node 662b is 6=n 12 and the value of node 662c is 1=n 13 Consider an example where weighting 666a has a value of 1=w1, weighting 666b has a value of 4=w2, and weighting 666c has a value of 2=w3, then the value of node 664 is defined as follows: n 21 =σ(n 11 w1+n 12 w2+n 13 w3) = σ(2 * 1+6 * 4+1 * 2)=σ(2+24+2)=σ(28).

[0075] In the example where σ is the tanh function, n 21 would be equal to 1. Similar formulas could be used to determine other values ​​for neural networks.

[0076] In the example of Figure 6, a three-dimensional input is input to the neural network 600 to provide a three-dimensional output. According to this example, the input may represent three dimensions of the speech input (e.g., any three of amplitude, frequency, pitch, intonation, metadata representing sentence structure, etc.) and the output may represent a three-dimensional spatial vector representing the user's head pose. The dimensionality of the output may be increased if different facial component poses (lip, eye, etc. poses) are used. More or fewer components of the speech input may also be taken into account by increasing or decreasing the dimensionality of the input.

[0077] It may be emphasized that, as disclosed herein, any neural network may be used as the ML model, and Fig. 6 is shown merely as an example. For example, any of the following may be used: a variational autoencoder (VAE); a convolutional neural network; a generative model; a discriminative model; etc.

[0078] An exemplary wireless communication device will now be described in more detail with reference to FIG. 7, which shows a schematic diagram of a communication device 1000. Such devices may include, for example, one or more of the following: user device 102a; user device 102b; user device 102n. A suitable communication device may be provided by any device capable of transmitting and receiving wireless signals. Non-limiting examples include mobile stations (MS) or mobile devices, such as mobile phones or what are known as "smartphones", computers equipped with wireless interface cards or other wireless interface facilities (e.g., USB dongles), personal data assistants (PDAs) or tablets equipped with wireless communication capabilities, or any combination thereof, and so on. Mobile communication devices may provide communication of data, for example to carry voice, electronic mail (email), text messages, multimedia, and other communications. A user may thus be offered and provided with a myriad of services via the communication device. Non-limiting examples of these services include two-way or multi-way telephony, data communication or multimedia services, or simply access to a data communication network system such as the Internet. A user may also be provided with broadcast or multicast data. Non-limiting examples of content include downloads, television and radio programs, videos, advertisements, various alerts and other information.

[0079] A wireless communication device may, for example, be a mobile device (i.e., a device that is not fixed to a particular location) or a stationary device. A wireless device may or may not require human interaction for communication.

[0080] The wireless device 1000 may receive signals over the air or wireless interface 1007 via suitable arrangements for reception and may transmit signals via suitable arrangements for transmitting wireless signals. In Fig. 7 the transceiver arrangement is generally designated by block 1006. The transceiver arrangement 1006 may for example be provided by a radio part and an associated antenna arrangement. The antenna arrangement may be arranged internally or externally to the wireless device.

[0081] A wireless device typically comprises at least one data processing entity (e.g., processor) 1001, at least one memory 1002, and other possible components 1003 for use in software- and hardware-assisted execution of the tasks it is designed to perform, including controlling access to and communication with the access system and other communication devices. The data processing, storage, and other related control devices may be provided on a suitable circuit board and / or in a chipset. This feature is represented by reference 1004. A user may control the operation of the wireless device by a suitable user interface (keypad 1005, voice commands, touch-sensitive screen or pad, combinations of these, etc.). A display 1008, a speaker, and a microphone may also be provided. Furthermore, a wireless communication device may include suitable connectors (either wired or wireless) to other devices and / or for connecting external accessories (e.g., hands-free devices) to it. The communication devices 1002, 1004, 1005 may access the communication system based on various access technologies. The user device 1000 may be connectable to one or more networks, such as the network 104, to receive and / or transmit information. Such networks may include, for example, the Internet network. The communication device 1000 may include equipment for recording video and audio.

[0082] FIG. 9 illustrates an example of a computing device 1100. In some examples, the computing device 1100 may have a similar structure to the computing device 106a, the computing device 106n, etc. The computing device 1100 may include at least one memory 1101, at least one data processing unit 1102, 1103, and an input / output interface 1104. Through the interface, the computing device may be coupled to a network, such as the network 104. The receiver and / or the transmitter may be implemented, for example, as a radio front end or a remote radio head. For example, the computing device 1100 or the data processing units 1101, 1102, in combination with the memory 1101, may be configured to execute appropriate software code to provide the functionality performed by the computing device as disclosed herein. The computing device 1100 may be connectable to one or more networks, such as the network 104, to receive and / or transmit information. Such a network may include, for example, the Internet network. In some examples, the computing device 1100 may include a cloud-implemented device.

[0083] 9 illustrates an example method flow 900. The method flow 900 may be performed by a user device, such as user device 102a; by a computing device, such as computing device 106a; or by a combination of a user device (e.g., user device 102a) and a computing device (e.g., computing device 106a).

[0084] At 901, the method includes receiving video data from a user device from a user. In some examples, this may include capturing the video data from the user device.

[0085] At 902, the method includes training a first machine learning model based on the video data to provide a second machine learning model. The second machine learning model may be personalized to the user and may be configured to predict the user's movements based on the audio data. The first machine learning model may be a personalizable base model f θ The second machine learning model may include a personalization model f' θ (e.g., model 210).

[0086] At 903, the method 900 includes receiving other audio data from the user. The other audio data may be different from the audio data used to train the first machine learning model at 902. The other audio data may also be different from the audio data of the video data received at 901.

[0087] At 904, the method 900 includes determining predicted movements of the user based on the further audio data and the second machine learning model.

[0088] At 905, the method 900 includes using the predicted movements of the user to generate an animation of an avatar of the user.

[0089] One or more elements of the above-described systems may be controlled by a processor and associated memory containing computer-readable instructions for controlling the system. The processor may control one or more devices for implementing the methods disclosed herein. Circuitry or a processing system may also be provided for controlling one or more systems.

[0090] It is understood that the processors or processing systems or circuitry referred to herein may in fact be provided by a single chip or integrated circuit, or by multiple chips or integrated circuits (optionally provided as a chipset), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a graphic processing unit (GPU), etc. A chip or chips may include circuitry (and possibly firmware) for embodying at least one or more of a data processor or processors, a digital signal processor or processors, baseband circuitry, and radio frequency circuitry that may be configured to operate according to exemplary embodiments. In this regard, exemplary embodiments may be implemented at least in part by computer software stored in a (non-transitory) memory and executable by a processor, or by hardware, or by a combination of tangibly stored software and hardware (as well as tangibly stored firmware).

[0091] Reference is made herein to data storage for storing data. This may be provided by a single device or a number of devices. Suitable devices include, for example, hard disks and non-volatile semiconductor memories (e.g., solid state drives or SSDs).

[0092] At least some aspects of some embodiments described herein with reference to the accompanying drawings include computer processes carried out in a processing system or processor, but the invention also extends to computer programs (particularly computer programs on or in a carrier) adapted to realize the invention. The programs may be in the form of non-transitory source code, object code, code intermediate source and object code (such as in specially compiled form) or any other non-transitory form suitable for use in the implementation of the process according to the invention. The carrier may be any entity or device capable of carrying a program. For example, the carrier may include a storage medium such as a solid state drive (SSD) or other semiconductor-based RAM; ROM (e.g. CD ROM, semiconductor ROM); magnetic recording medium (e.g. floppy disk or hard disk); optical memory devices in general; and so on.

[0093] According to a first aspect, a computer-implemented method is provided that includes receiving video data from a user device; training a first machine learning model based on the video data to provide a second machine learning model personalized to the user, the second machine learning model being configured to predict movements of the user based on audio data; receiving another audio data from the user; determining predicted movements of the user based on the another audio data and the second machine learning model; and using the user's predicted movements to generate an animation of an avatar for the user.

[0094] According to some examples, the user's movements include: the user's lip movements; changes in the user's head pose; and changes in the user's facial expressions.

[0095] According to some examples, the method includes communicating with another user device, and an animation of the avatar is used during the communication.

[0096] According to some examples, training the first machine learning model based on the video data to provide the second machine learning model is performed at least in part by a cloud computing device.

[0097] According to some examples, training the first machine learning model based on the video data to provide the second machine learning model is performed at least in part by the user device.

[0098] According to some examples, using the user's predicted movements to generate an animation of the user's avatar includes: receiving at least one input from a user that varies a degree of at least one of the user's predicted movements to provide a user customized movement; using the user's customized movement to generate an animation of the user's avatar.

[0099] According to some examples, the video data includes two or more images of a user.

[0100] According to some examples, the method includes: generating a first external personality of the user based on predicted motion and based on a first customization of the avatar animation from the user; predictive motion based on video data from the user and a second customization of the avatar animation from the user; different predicted motion of the user by using different video data; different predicted motion of the user by using different video data from the user and a third customization of the avatar animation from the user, and generating a second external personality of the user based on at least one of the following: storing the user's second external personality; and providing the user with an option to select either the user's first external personality or the user's second external personality to provide the avatar animation.

[0101] According to some examples, the method includes: receiving information from a user to edit an appearance of the avatar; and updating the avatar based on the received information.

[0102] According to some examples, the method includes: determining a representation that has a similar appearance to the user; and basing an appearance of the avatar on the representation.

[0103] According to some examples, the training dataset includes two or more videos of each user of a plurality of users, each video having at least one labeled vertex of the respective user's head; the method includes: i) training a third machine learning model based on one or more videos of one of the plurality of users to provide a fourth machine learning model; ii) predicting head movements of the one of the plurality of users for one or more portions of the audio data; iii) calculating an error in predicted head movements of the one of the plurality of users for the one or more portions of the audio data by using the at least one labeled vertex of the head; iv) back-propagating the error in predicted head movements of the one of the plurality of users to update parameters of the third machine learning model, the method further including: repeating steps i)-iv) on a random sample of the plurality of users until the error converges from one user to the next user within the sample; and subsequently using the third machine learning model as the first machine learning model.

[0104] According to some examples, at least one of the first machine learning model, the second machine learning model, the third machine learning model, and the fourth machine learning model includes a convolutional and / or sequential neural network configured to operate on the audio data to predict changes in head pose, facial expressions, and lip movement.

[0105] According to some examples, training the first machine learning model based on the video data to provide the second machine learning model includes using one or more few-shot learning techniques.

[0106] According to a second aspect, an apparatus is provided that includes at least one processor; and at least one memory including computer program code; the at least one memory and the computer program code configured to cause the apparatus, by the at least one processor, to do the following: receive video data from a user device from a user; train a first machine learning model based on the video data to provide a second machine learning model, where the second machine learning model may be personalized to the user, and where the second machine learning model is trained to predict the user's movements based on the audio data; receive further audio data from the user; determine predicted movements of the user based on the further audio data and the second machine learning model; and use the user's predicted movements to generate an animation of an avatar of the user.

[0107] At least one memory and computer program code configured to cause, by at least one processor, the apparatus of the second aspect to perform any of the steps of the example method of the first aspect.

[0108] According to a third aspect, a computer-readable storage device is provided that includes instructions executable by a processor for: receiving video data from a user device from a user; training a first machine learning model based on the video data to provide a second machine learning model, where the second machine learning model can be personalized to the user, and where the second machine learning model is trained to be capable of predicting the user's movements based on the audio data; receiving another audio data from the user; determining predicted movements of the user based on the another audio data and the second machine learning model; and using the user's predicted movements to generate an animation of an avatar for the user.

[0109] The processor-executable instructions of the third aspect may be for performing any of the steps of the example method of the first aspect.

[0110] According to a fourth aspect, there is provided a computing device comprising: a memory comprising one or more memory units; and a processing device comprising one or more processing units; the memory storing code arranged to execute on the processing device, the code configuring on the processing device to perform a method of the first aspect or a method of any of the examples of the first aspect.

[0111] According to a fifth aspect, a computer-implemented method is provided that includes: receiving video data of a user from a user device, the video data including audio data and image data corresponding to the audio data; training a first machine learning model based on the video data, thereby resulting in a second trained machine learning model, the second trained machine learning model being personalized to the user, the second trained machine learning model being configured to predict movements of the user; receiving another audio data from the user device; inputting the another audio data into the second trained machine learning model, thereby resulting in predicted movements of the user; receiving at least one input from the user device that changes an extent of at least one of the user's predicted movements, thereby resulting in customized movements of the user; and generating an animation of the user's avatar by using the user's customized movements.

[0112] According to some examples, the predicted movements of the user include at least one of: lip movements of the user; changes in the user's head pose; and changes in the user's facial expressions.

[0113] According to some examples, the method includes communicating with another user device, and an animation of the avatar is used during communication with the other user device.

[0114] According to some examples, training the first machine learning model based on the video data is performed at least in part by a cloud computing device.

[0115] According to some examples, training the first machine learning model based on the video data is performed at least in part by a user device.

[0116] According to some examples, the video data includes two or more images of a user.

[0117] According to some examples, the method includes: generating a first external personality of the user based on predicted motion and based on a first customization of the avatar animation from the user device; generating a second external personality of the user based on at least one of: predicted motion based on the user's video data and a second customization of the avatar animation from the user device; different predicted motion of the user using second video data different from the video data; different predicted motion of the user using third video data of the user different from the video data and a third customization of the avatar animation from the user device, the method further including: storing the user's second external personality; and providing the user with an option to select either the user's first external personality or the user's second external personality to provide the avatar animation.

[0118] According to some examples, the method includes receiving information from a user device that edits an appearance of the avatar; and updating the avatar based on the received information from the user device that edits an appearance of the avatar.

[0119] According to some examples, the method includes: determining a representation that has a similar appearance to the user; and basing an appearance of the avatar on the representation.

[0120] According to some examples, the training dataset includes two or more videos of each user of the plurality of users, each video having at least one labeled vertex of the respective user's head; the method includes: i) training a third machine learning model based on at least one video of one user of the plurality of users to provide a fourth machine learning model; ii) predicting head movements of the user of the plurality of users based on at least a portion of the audio data and the fourth machine learning model, each portion of the at least a portion of the audio data having a corresponding video; iii) calculating a predicted head movement error of the at least a portion of the audio data by using the predicted head movement and the at least one labeled vertex of the user's head in the corresponding video of each portion of the at least a portion of the audio data; iv) updating parameters of the third machine learning model by back-propagating the predicted head movement error of the one user of the plurality of users; the method further includes: repeating steps i)-iv) on a random sample of the plurality of users until the error converges from one user to the next user within the sample; and subsequently using the third machine learning model as the first machine learning model.

[0121] According to some examples, at least one of the first machine learning model, the second trained machine learning model, the third machine learning model, and the fourth machine learning model include at least one of the following: a convolutional neural network configured to predict changes in a user's head pose, changes in a user's facial expression, and changes in a user's lip movement; a successive neural network configured to predict changes in a user's head pose, changes in a user's facial expression, and changes in a user's lip movement.

[0122] According to some examples, the first machine learning model includes at least one of: a convolutional neural network configured to operate on the speech data to predict changes in a user's head pose, changes in a user's facial expression, and changes in the user's lip movement; and a successive neural network configured to operate on the speech data to predict changes in a user's head pose, changes in a user's facial expression, and changes in the user's lip movement.

[0123] According to some examples, training the first machine learning model based on the video data includes using at least one few-shot learning technique.

[0124] According to a sixth aspect, there is provided an apparatus including at least one processor; and at least one memory including computer program code; the at least one memory and the computer program code configured to cause the apparatus by the at least one processor to: receive video data from a user device from a user, the video data including audio data and image data corresponding to the audio data; train a first machine learning model based on the video data, thereby resulting in a second trained machine learning model personalized to the user, the second trained machine learning model being trained to predict movements of the user; receive another audio data from the user device; input the another audio data into the second trained machine learning model, thereby resulting in predicted movements of the user; receive at least one input from the user device that changes an extent of at least one of the predicted movements of the user, thereby resulting in customized movements of the user; and generate an animation of the user's avatar by using the customized movements of the user.

[0125] At least one memory and computer program code configured to cause, by at least one processor, the apparatus of the sixth aspect to perform any of the steps of the example method of the fifth aspect.

[0126] According to a seventh aspect, a computer-readable storage device is provided that includes instructions executable by a processor for: receiving video data of a user from a user device, the video data including audio data and image data corresponding to the audio data; training a first machine learning model based on the video data, thereby resulting in a second trained machine learning model, the second trained machine learning model being personalized to the user, the second trained machine learning model being capable of predicting movements of the user; receiving another audio data from the user device; inputting the another audio data into the second trained machine learning model, thereby resulting in predicted movements of the user; receiving at least one input from the user device that changes an extent of at least one of the user's predicted movements, thereby resulting in customized movements of the user; and generating an animation of the user's avatar by using the user's customized movements.

[0127] The processor-executable instructions of the seventh aspect may be for performing any of the steps of the example method of the fifth aspect.

[0128] According to an eighth aspect, there is provided a computing device comprising: a memory comprising one or more memory units; and a processing device comprising one or more processing units; the memory storing code arranged to execute on the processing device, the code being configured on the processing device to perform the method of the fifth aspect or the method of any of the examples of the fifth aspect.

[0129] The examples described herein should be understood as illustrative examples of embodiments of the present invention. Further embodiments and examples are envisioned. Any feature described with respect to any one example or embodiment may be used alone or in combination with other features. In addition, any feature described with respect to any one example or embodiment may also be used in combination with one or more features of any other of the examples or embodiments, or in any combination of any other of the examples or embodiments. Furthermore, equivalents and modifications not described herein may also be employed within the scope of the present invention, as defined in the claims.

Claims

1. receiving user video data from a user device (102a), the video data including audio data and image data corresponding to the audio data; training a first machine learning model based on the video data, thereby generating a second trained machine learning model (210) personalized to the user, the second trained machine learning model (210) configured to predict movements of the user; receiving another voice data (314) from the user device (102a); inputting the additional audio data (314) into the second trained machine learning model (210), thereby generating predicted movements (318) of the user; receiving at least one input from the user device (102a) that changes a degree of at least one of the predicted movements of the user, thereby resulting in a customized movement of the user; generating an animation of the user's avatar using the customized movements of the user; and communicating with a further user device (102n), wherein the animation of the avatar is used during communication with the further user device (102n); 11. A computer-implemented method comprising:

2. The predicted movement (318) of the user is: lip movements of the user; a change in the user's head posture; Changes in the user's facial expression The method of claim 1 , comprising at least one of:

3. The method described in claim 1, wherein at least one degree of the predicted movement of the user is changed by changing the weighting of the second trained machine learning model (210).

4. The method of claim 1 , wherein training the first machine learning model based on the video data is performed at least in part by a cloud computing device (106 a).

5. The method of claim 1 , wherein training the first machine learning model based on the video data is performed at least in part by the user device (102 a).

6. The method of claim 1 , wherein the video data includes two or more videos of the user.

7. generating a first external persona for the user based on the predicted movement (318) and based on a first customization of the animation of the avatar from the user device (102a); generating a second external persona for the user based on at least one of the predicted movement based on the video data of the user and a second customization of the animation of the avatar from the user device (318), a different predicted movement of the user using the second video data different from the video data, and a different predicted movement of the user using third video data different from the video data and a third customization of the animation of the avatar from the user device (102a); Including, storing the second external persona of the user; providing the user with an option to select either the first external persona of the user or the second external persona of the user to provide animation of the avatar; The method of claim 1 further comprising:

8. receiving information from the user device (102a) to edit the appearance of the avatar; updating the avatar based on the received information from the user device (102a) editing the appearance of the avatar; The method of claim 1 , comprising:

9. determining a representation having a similar appearance to the user; basing the appearance of the avatar on the representation; The method of claim 1 , comprising:

10. 2. The method of claim 1, wherein the training dataset (203) includes two or more images (205a, 205b) of each user of a plurality of users, each of the images having at least one labeled vertex of the head of the respective user, the method comprising: i) training a third machine learning model (209) based on at least one video of a user of the plurality of users to provide a fourth machine learning model (211); ii) predicting head movements of the users of the plurality of users based on at least a portion of the audio data and the fourth machine learning model (211), wherein each portion of the at least a portion of the audio data has a corresponding video; iii) calculating an error in the predicted head movement of the at least one portion of audio data by using the predicted head movement and at least one labeled vertex of the user's head in a corresponding image of each portion of the at least one portion of audio data; iv) updating parameters of the third machine learning model (209) by backpropagating the error in the predicted head movement of the user of the plurality of users; and the method further comprises: repeating steps i) to iv) for the random sample of users until the error converges from one user to the next within the sample; and then using the third machine learning model (209) as the first machine learning model; A method comprising:

11. 11. The method of claim 10, wherein at least one of the first machine learning model, the second trained machine learning model (210), the third machine learning model (209), and the fourth machine learning model (211) comprises at least one of a convolutional neural network configured to predict changes in head pose of the user, changes in facial expression of the user, and lip movements of the user, and a successive neural network configured to predict changes in head pose of the user, changes in facial expression of the user, and lip movements of the user.

12. 2. The method of claim 1, wherein the first machine learning model comprises at least one of a convolutional neural network configured to operate on speech data to predict changes in the user's head pose, changes in the user's facial expression, and lip movements of the user, and a successive neural network configured to operate on speech data to predict changes in the user's head pose, changes in the user's facial expression, and lip movements of the user.

13. The method of claim 1 , wherein training the first machine learning model based on the video data includes using at least one few-shot learning technique.

14. At least one processor (1001), and At least one memory (1002) containing computer program code An apparatus (1000) comprising: The at least one memory (1002) and the computer program code are configured to cause the at least one processor (1001) to receiving video data from a user device (102a) from a user, the video data including audio data and image data corresponding to the audio data; training a first machine learning model based on the video data, thereby generating a second trained machine learning model (210), the second trained machine learning model (210) being personalized to the user, the second trained machine learning model (210) being trained to predict the user's movements; receiving another voice data (314) from the user device (102a); inputting the additional audio data (314) into the second trained machine learning model (210), thereby generating predicted movements (318) of the user; receiving at least one input from the user device (102a) that changes a degree of at least one of the predicted movements of the user, thereby resulting in a customized movement of the user; generating an animation of the user's avatar by using the customized movements of the user; and communicating with a further user device (102n), wherein the animation of the avatar is used during communication with the further user device (102n); An apparatus (1000) configured to perform the above.

15. receiving user video data from a user device (102a), the video data including audio data and image data corresponding to the audio data; training a first machine learning model based on the video data, thereby generating a second trained machine learning model (210) personalized to the user, the second trained machine learning model (210) capable of predicting the user's movements; receiving another voice data (314) from the user device (102a); inputting the additional audio data (314) into the second trained machine learning model (210), thereby generating predicted movements (318) of the user; receiving at least one input from the user device (102a) that changes a degree of at least one of the predicted movements of the user, thereby resulting in a customized movement of the user; generating an animation of the user's avatar by using the customized movements of the user; and communicating with a further user device (102n), wherein the animation of the avatar is used during communication with the further user device (102n); 1. A computer-readable storage device containing processor-executable instructions for: