Personalized representation and animation of humanoid characters

The method of creating a personalized data library using machine learning to enhance 3D model animations addresses the limitations of existing file formats, enabling accurate and personalized animations in VR and AR environments.

JP2026510388APending Publication Date: 2026-04-02PICTORYTALE AS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-14
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing 3D file formats for representing humans in VR and AR environments lack the ability to capture nuanced anatomical and individual characteristics, resulting in generic and inaccurate animations that fail to convey personalized facial expressions and gestures.

Method used

A method and apparatus for creating a library of personalized data by analyzing video streams to detect and encode unique gestures and expressions, using machine learning to enhance 3D model animations with personalized data layers.

Benefits of technology

Enables richer and more personalized representations of humanoid characters in VR and AR environments by capturing and applying unique anatomical and individual characteristics, improving animation accuracy and personalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026510388000001_ABST
    Figure 2026510388000001_ABST
Patent Text Reader

Abstract

A computer device maintains a library of personalized data for use in animating 3D models representing human or humanoid characters in a shared VR or AR environment. The personalized data includes at least descriptions of animations, textures, or colors that represent an individual's gestures or habits, for example, when performing gestures or facial expressions. The personalized data is stored as an indexed file or library that is made available to other participants in the shared environment. When animation description data is generated, an index referencing the detected behavior is sent to the shared environment along with the animation description data. Other participants apply the referenced behavior from the library along with the animation data when rendering the 3D model. The embodiment includes a method and apparatus for creating personalized data that includes a personalized data index having animation data, and for applying the personalized data when rendering animations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the representation of humans or humanoid characters in 3D animations such as virtual reality (VR) and augmented reality (AR) environments. In particular, the present invention relates to the personalization of such representations, particularly, but not limited to, movement, gesture, and expression.

Background Art

[0002] Humans are typically represented in AR and VR environments as 3D characters or solid characters. 3D characters are generally represented according to common 3D file formats used in the film industry or for AR and VR purposes. Such file formats include, for example, FBX, OBJ, GLB, and USD. 3D files typically store one or more representations of four types of data: model shape, model surface texture, scene details, and model animation. File formats do not necessarily store the same type of data for representation. For example, OBJ files do not store animation data, while FBX files store all animation data.

[0003] All of these file formats are essentially general-purpose. This means that they are intended to represent not only 3D representations of humans or humanoid characters but also any 3D model. In other words, these 3D file formats do not utilize features and limitations specific to the anatomical structure of humans and do not include features specifically adapted to represent the mannerisms unique to an individual, but are within the limitations indicated by the anatomical structure or meaning. For example, gestures and expressions are limited by the anatomical structure of humans, and their significance is limited by the extent to which an observer associates a given gesture or expression with meaning. (Of course, motion capture and skeletal representations are limited by the anatomical structure of humans but are not limited in the sense described herein)

[0004] Traditional file formats cannot capture the nuances of human anatomy and individual characteristics, resulting in several limitations and drawbacks, particularly in live interactions between participants represented as 3D characters in shared AR or VR environments. Such shared environments are sometimes referred to as metaverses. This term should not be interpreted as referring to only one specific implementation of a shared environment, nor should it be interpreted as referring to any particular technique regarding the representation of an environment, whether it is represented on one or more computers, whether it is a single shared environment, or a collection of several interconnected shared environments.

[0005] To illustrate the shortcomings of conventional file formats, consider a metaverse where User A is represented as a 3D character in an AR environment, and this representation of User A can be observed by User B, who is in a different location. In this scenario, User A stands in front of a camera and interacts with User B, who sees User A as a 3D character in AR. User A's representation performs gestures (body and face) synchronized with gestures actually performed by User A. This is achieved by using a motion capture module that generates the representation of movement as animation data, which is transmitted over the network and applied to User A's 3D representation displayed on User B's device. These animations are typically generic in the sense that all movements are represented in the same way, regardless of what moves or how. For example, in the case of joint-based animation or skeleton animation, vertices are captured, transmitted over the network, and then applied to User A's 3D character.

[0006] Two drawbacks associated with this method become immediately apparent. Firstly, the accuracy of the animation is good in terms of vertex capture and transfer, as long as the circumstances allow in each real-world situation. Typically, a limited range of vertex movement can be captured with only a single camera, or sometimes even just a handheld mobile phone. In production settings, advanced motion capture tools use either multiple cameras surrounding the subject or a single camera with post-processing by software tools. While this provides rich and accurate vertex movement information, it is not suitable for live transfers in situations such as meetings or games.

[0007] Secondly, the animation information is not personalized. Because the 3D character file format is generic and the animation lacks precision, captured facial expressions and gestures cannot convey personalized details. For example, one person may create a smile with equal lip movement on both sides of their cheeks, while another person's natural smile may have different lip movements, perhaps due to asymmetry. While capturing such details is possible in a professional studio environment, it is not easily achieved with standard 3D model animations created by users.

[0008] With this in mind, it is desirable to introduce new methods and systems that can easily make the animation of 3D characters representing users interacting in VR or AR environments richer and more personalized. [Overview of the project]

[0009] In consideration of the aforementioned shortcomings, the present invention provides methods and apparatus aimed at providing richer and more personalized representations of humanoid characters by providing methods for capturing and presenting actions, including personalized behaviors in the form of gestures, facial expressions, etc. The present invention includes three interrelated products or embodiments, namely, methods and apparatus for creating personalized data; methods and apparatus for using or including personalized data when creating animations; and methods and apparatus for applying personalized data when rendering an animated 3D model to present a personalized animation of the model. In embodiments of the present invention, two or more of these embodiments may be combined, but typically, the creation of personalized data is performed in conjunction with the creation of the 3D model (i.e., creating a “vocabulary” of personalized actions), independently of the actual session in which the animation data is exchanged. The use of personalized data is performed when animation data is created to animate the 3D model (i.e., this is the use of the personalized “vocabulary” on the sender or part of the session). Finally, the application of personalized data to the rendered animation of the 3D model is performed on the receiver for presentation (i.e., the receiving and presentation of actions from the “vocabulary”).

[0010] According to a first aspect, a method is provided for creating a library of personalized data for use in animating a 3D model representing a human or humanoid character in a shared environment in a computer system. The method includes the steps of: acquiring a video stream of a user while the user is performing a series of actions; providing frames from the video stream as input to a computerized process that analyzes the frames to detect one or more actions performed by the user; identifying one or more detected actions; and for one or more detected actions, extracting action description data from the video frames and storing the action description data in a library of personalized data. The action description data includes an encoded form of one or more detected actions, and each action description is associated with an index.

[0011] In some embodiments, frames from a video stream are provided as input to a computerized photogrammetry process to generate 3D information about the user. The 3D information can then be used to generate a 3D representation of the user, which can be stored as a 3D model that can be represented and animated in a shared environment.

[0012] Embodiments of the present invention may utilize a machine learning subsystem trained to perform at least one of detecting and identifying actions within a stream of video frames in order to analyze frames, detect one or more actions, and identify one or more detected actions. Such a machine learning subsystem may include at least one artificial neural network.

[0013] Some embodiments of the present invention may be configured to infer an estimated action description for at least one undetected action based on action description data for at least one detected action. The estimated action descriptions may then be indexed and stored in a library of personalized data along with the action descriptions based on the detected actions.

[0014] The behavior description may include at least one of the following: animation data, texture data, and color data.

[0015] A computer device according to a first embodiment is also provided for creating a library of personalized data for use in animating 3D models representing humans or humanoid characters in a shared environment. Such a computer device includes at least one video camera, a personalized data creation module including a submodule configured to receive frames from at least one video camera, analyze the received video frames, and detect actions performed by a user depicted in the video frames, and a submodule configured to identify the detected actions, extract action description data including encoded forms of one or more detected actions, and associate each action description with an index. The device also includes a storage unit configured to receive and store the personalized data received from the data creation module.

[0016] A computer device according to this aspect of the present invention may also include a 3D modeling module configured to receive frames from at least one video camera and generate a 3D model based on images of people represented in the received video frames using photogrammetry processing of the video frames. At least one of the submodules configured to detect motion and the submodules configured to identify motion may include an artificial neural network.

[0017] A second aspect of the present invention provides a method for providing personalized data along with animation data for animation of a 3D model representing a human or humanoid character in a shared environment in a computer system. The method includes the steps of: connecting to a shared environment; sending a library of personalized data to the shared environment, which holds motion description data including encoded embodiments of one or more motions, each motion description associated with an index; acquiring a video stream of a user while the user is performing an action; providing frames from the video stream as input to a computerized motion capture process to generate animation data from the user's movements represented in the video stream; providing frames from the video stream as input to a computerized process that analyzes the frames to detect and identify at least one action corresponding to an action represented in the library of personalized data; and sending the generated animation data and index for the detected and identified at least one action to the shared environment.

[0018] A method according to this embodiment can perform at least one of detecting one or more actions and identifying one or more detected actions, using a machine learning subsystem trained to perform at least one of detecting and identifying actions in a stream of video frames. Such a machine learning subsystem may include at least one artificial neural network. The action description includes at least one of animation data, texture data, and color data.

[0019] A computer device according to a second embodiment includes: a storage unit that stores a library of personalized data holding motion description data, each motion description associated with an index; a video camera; an animation module configured to receive frames from at least one video camera, and a submodule configured to perform motion capture processing on the received video frames to generate animation data from movements performed by a person represented in the video stream; a submodule configured to analyze the video frames and detect movements performed by a person depicted in the video frames; a submodule configured to identify the detected movements and retrieve an index associated with the identified movements from the storage unit; and a communication interface configured to transmit the library of personalized data, the animation data, and the motion index to a shared environment.

[0020] A third aspect of the present invention provides a method for applying personalized data to animation data when animating a 3D model representing a human or humanoid character in a shared environment in a computer system. The method includes the steps of: connecting to a shared environment; receiving a library of personalized data holding motion description data including encoded modes of one or more motions, each motion description being associated with an index; receiving animation data and at least one index referencing motions represented in the library of personalized data; using at least one index to retrieve motion description data from the library of personalized data; applying the retrieved motion description data to animation data to generate personalized animation data; and rendering and animating a 3D model according to the personalized animation data.

[0021] In some embodiments, such a method includes the step of receiving a 3D model along with a library of personalized data. The library of personalized data may be received from a repository connected to a computer network, and animation data and at least one index referencing the behavior represented in the library of personalized data may be received from a device participating in a shared environment. In some embodiments, the personalized data is generated independently of the generation of the 3D model.

[0022] A computer device configured to operate according to this aspect of the present invention includes: a library of personalized data holding action description data including encoded modes of one or more actions, each action description associated with an index; a communication interface for receiving animation data and an index referencing actions represented in the library of personalized data from a shared environment; a rendering module (206) configured to include a 3D model in a local representation (209) of the shared environment, retrieve action description data referenced by an index received from the library of personalized data, apply the retrieved action description data to animation data to generate personalized animation data, and render and animate the 3D model according to the personalized animation data; and a display unit configured to visualize at least a portion of the shared environment including the rendered and animated 3D model.

[0023] The present invention will now be described in more detail with reference to the drawings. [Brief explanation of the drawing]

[0024] [Figure 1] This shows simplified representations of three faces with different personalized expressions. [Figure 2]A block diagram showing an embodiment of a client device 200 configured to operate in accordance with the present invention. [Figure 3] A flowchart showing an exemplary embodiment of a method for creating personalized data. [Figure 4] A flowchart showing an exemplary embodiment of a method for creating and transmitting animation data and a personalized data index. [Figure 5] A flowchart showing an exemplary embodiment of a method for receiving animation data and a personalized data index when animating and rendering a 3D model, and enhancing the animation data using the referenced personalized data. [Figure 6] A block diagram showing the flow of data between two devices connected to the same shared environment.

Embodiments for Carrying Out the Invention

[0025] The present invention generally relates to 3D animation. More specifically, the present disclosure describes methods and systems for animating 3D representations of humans or humanoid characters, particularly 3D representations of a user in a VR or AR environment. More specifically, embodiments of the present invention can be configured to provide personalization of such 3D representation animations based on captured gestures, expressions, and other quirks that are unique to an individual user or otherwise related to a particular character or representation.

[0026] The following descriptions of various embodiments refer to the drawings, where similar reference numbers indicate the same or corresponding elements. While many features are necessary for a computer-based system to function, it will be understood that some features are well known in the art and present in most systems. This disclosure does not unnecessarily describe such details. Instead, features whose description would facilitate understanding of the invention are prioritized, while less important details may be given a somewhat simplified or schematic presentation. For some features, it is assumed that those skilled in the art can provide the necessary contextual information from general knowledge of the art. Therefore, certain prior art elements may be excluded for the purpose of illustrating the principles of the invention rather than cluttering the drawings and disclosure with details that do not contribute to understanding these principles.

[0027] Unless otherwise specified, it should be noted that different features or elements can be combined with each other, whether or not they are described together as part of the same embodiment below. The combinations of features or elements in the exemplary embodiments are made to facilitate understanding of the invention, rather than limiting the scope of the invention to a limited set of embodiments, and are intended to be interchangeable insofar as alternative elements having substantially the same function are shown in each embodiment; however, for the sake of brevity, no attempt has been made to disclose a complete description of all possible substitutions of features.

[0028] Furthermore, those skilled in the art will understand that the present invention can be implemented without many of the details included in this detailed description. Conversely, some well-known structures or functions may not be illustrated or described in detail to avoid unnecessarily obscuring the relevant descriptions of various embodiments. The terms used in the following description are intended to be interpreted in their broadest and most reasonable way, even when used in conjunction with the detailed descriptions of specific embodiments of the present invention.

[0029] The terms personalized or personalized are used repeatedly in this disclosure. These terms should be understood in a technical sense. In many, or most, cases, personalization refers to something specific to an individual, reality, or artificial object, and several points should be kept in mind. First, personalization does not require uniqueness. Personalized data is data that describes a manner of representation and animation that can be experienced as unique, but in principle, personalized data can be identical for representations of different individuals. Second, personalization does not need to be derived from the represented individual. Instead, personalized data created by other means can be applied to an individual's representation. In other words, personalized data may be derived from the represented individual and may be specific to that individual, but does not need to be. For consistency, data is referred to as personalized data regardless of how the data was created and whether the same data is applied to multiple representations of multiple individuals. Similarly, the term individual is used to refer to what is represented, whether it is a real person or an imaginary character, and whether the individual is pre-recorded or interacting with a VR or AR environment in real time.

[0030] The following terms are primarily used, but deviations from them should be interpreted from the context. A user is a person who uses or interacts with a system that implements one or more features of the present invention. An individual is a representation in the system of a human or humanoid character, which may be real (user) or fictional (character). A model, or 3D model, is a data structure that realizes or is a representation of a representation. Therefore, a 3D representation is a 3D model that represents a user or character in a shared environment.

[0031] Therefore, in the context of this disclosure, a 3D representation is a 3D model of a human or humanoid character. According to the present invention, this representation includes, is supplied with, or is associated with personalized data. A 3D representation typically includes four types of data: model geometry, surface texture, conventional animation data, and personalized data. Model geometry can typically be represented as meshes, skeletons, geometric primitives, and combinations thereof, connected to each other in defined relationships that influence how the movement of one point or vertex in a geometric model may affect the movement of other parts of the model. Personalized data can interact with or influence at least one of the texture data and animation data during rendering. For example, as schematically shown in Figure 1, a typical smile may be a symmetrical upward movement of the rim of the mouth, represented by a first face 101. A person may blush when smiling, as represented by a second face 102. This blushing of the cheeks 104 can be represented by a change in the texture data of the relevant part of the 3D model. Here, another person represented by a third face 103 may have a distorted or biased smile 105, perhaps raising one eyebrow 106 while smiling. Humans pay close attention to such small changes in facial expressions, but these are not easily captured by systems designed to capture not only humans but all kinds of movements and shapes, and in particular, relatively subtle changes that are meaningful to the human observer are not given priority.

[0032] Similar reasoning can be applied to gestures. For example, relatively small movements of the arms, hands, and fingers can represent significant meaning specific to an individual, or even individual habits, and thus, when observed, make the 3D representation of that individual more personally recognizable as that particular individual.

[0033] To enable the representation of these small variations in gestures and expressions without improving the overall capabilities of the recording device or significantly increasing the amount of data, the present invention provides a solution in which an additional data type or data layer is added. For the purposes of this disclosure, we assume that the starting point is a 3D modeling format or 3D representation including the 3D model itself, animation, and textures. However, the present invention is not limited to such a format and may be used in a context in which, for example, one of the animation and texture data is not included in the model itself but is treated separately. This may depend, for example, on the file format being used.

[0034] Herein, we refer to Figure 2, a block diagram showing one embodiment of a client device 200 configured to operate according to the present invention. It should be noted that the functions provided by the present invention include three main activities: capturing and storing personalized data, capturing and transmitting animation data, and rendering. These activities will be described in more detail below. The client device shown in Figure 2 is configured to have features and capabilities relevant to all of these activities. However, providing a device configured for only one or two of these activities is consistent with the principles of the present invention. For example, a device configured to perform the capture and storage of personalized data may include more sophisticated hardware, but it may be too large to use during animation (e.g., gameplay or AR conferencing). Therefore, personalized data may be captured by a device specifically configured for that purpose, and the capture and rendering of animation may be performed by a dedicated device. Another possibility is a broadcast setting in which personalization and animation are captured in one type of device (e.g., a studio setup) while rendering is being performed by a client device configured to act as a receiver.

[0035] In consideration of this, it will be understood that the descriptions of different modules or features in this disclosure are not repeated for all possible distributions of features across the types of devices. Where two or more modules are described herein as being interactable, they may be so by being provided together in a single device, or by communicating with each other via a communication link, such as a computer network. Furthermore, they may be provided together in the same device, whether or not that device includes further modules described as part of one embodiment herein. In other words, a device consistent with the principles of the present invention may be provided by combining the modules or features described herein, even if the combination is not explicitly described or illustrated. Instead, the embodiments described are selected not because they constitute a complete catalog of possible embodiments, but because they facilitate the understanding of the present invention.

[0036] The first module in the embodiment shown in Figure 2 is a personalized data creation module 201. This module includes or is connected to one or more cameras 202. The cameras 202 are configured to capture video of a user standing in front of them while performing one or more frequently occurring actions such as turning, standing up, sitting down, various hand gestures, and facial expressions such as smiling, looking angry, or looking disappointed. This process is described in more detail below with reference to Figure 3.

[0037] A 3D model of the user is created from a video stream delivered by the camera to the personalized data creation module 201. The 3D model can be created using photogrammetry, which in embodiments having multiple cameras 202 may include stereophotogrammetry or epipolar geometry. The personalized data creation module 201 also scans the video data stream to detect different movements, gestures, and expressions, collectively referred to herein as motions. The detected motions are categorized, indexed, and data describing the motions is stored. The data describing the motions may be at least one of animation data and texture data. The categorized and indexed data is referred to as personalized data. In the figure, the personalized data creation module 201 is shown as including three submodules. The first submodule 221 is the 3D model creation module. This module can create a 3D model using photogrammetry. Some embodiments of the present invention do not include this module and instead can utilize a 3D model created (or created) outside of the device 200. Furthermore, the 3D model creation module 201 may not be part of the personalized data creation module 201, but may be a separate module. Whether or not the apparatus 200 according to the present invention includes a 3D model creation module does not largely depend on whether or not it includes other optional features, or how those other features are configured.

[0038] A second submodule 222 is a module that may be configured to analyze received video frames and detect motion. A third submodule 223 may be configured to identify detected motion and extract motion description data from the video frames, such as motion, texture, or color information that represents the nature of the motion. For detection and classification / identification, these modules may include, for example, an artificial neural network. Motion capture technology can be used to extract motion description data.

[0039] The personalized data creation module 201 includes, or can be connected to, a storage unit 203 capable of storing personalized data. This storage unit 203 may be part of the personalized data creation module 201, a separate local device, or a cloud-based service. 3D model data may be stored in the storage unit 201 together with the personalized data, separately or as part of the same data file, or stored in a separate device. In some embodiments, the 3D model may be created separately from the personalized data, either beforehand (in which case the 3D model may or may not be available when the personalized data is created) or afterward. These options and some of their implications are described in more detail below.

[0040] The next module is the animation module 204. The animation module 204 is connected to camera 202. In embodiments such as those shown in Figure 2, the personalized data creation module 201 and the animation module 204 may utilize the same camera 202, or they may be connected to separate cameras, for example, if it is determined during the design process of a particular embodiment that personalized data creation requires high-resolution video to capture different behavioral details, while animation requires lower resolution to limit the bandwidth required for data transfer. Other optical or image processing capabilities may also differ, which must be determined as part of determining the specifications required for a particular embodiment of the invention. If personalized data creation and animation are to be performed by completely separate devices, possibly separated in space and time, they will require their own cameras.

[0041] The animation process performed by animation module 204 is based on video capture of an individual performing some activity related to, for example, a game or an online meeting (VR or AR meeting room). To animate the individual's 3D representation so that it can be presented to other users in a shared environment, animation module 204 may include a submodule 231 configured to perform motion capture, where the movement of various vertices is described as animation data. Motion capture, whether conventional marker-based motion capture, newer markerless technology, joint-based motion capture, or facial motion capture, is well known and understood in the art and will not be described in further detail here.

[0042] In addition to performing motion capture, the animation module 204 processes the video stream in much the same way as the personalized data creation module 201. However, for the animation module 204, it is sufficient to simply detect and identify the motion. Data describing the detected motion is not captured unless additional parameters such as start time and duration are required. Instead, the animation module obtains an index of the detected motion and includes that index in the animation data. For this purpose, the animation module may include a submodule 232 configured to detect motion and a submodule 233 configured to identify the detected motion and obtain the index associated with the identified motion from the storage unit 203.

[0043] When a user connects to the shared environment, the animation module 204 retrieves a 3D model (3D representation) and associated personalized data representing the individual from the memory unit 203 and can send the personalized data to the shared environment using the communication interface 205. The personalized data is then made available to all similar client devices connected to the shared environment. While the interaction session is in progress, the animation module 204 sends animation data and personalized data indices to the shared environment. The operation of the animation module 204 is described in more detail below with reference to Figure 4.

[0044] The last module shown in Figure 2 is the rendering module 206. The rendering module 206 is connected to a communication interface 205 and a display unit 207, such as a VR or AR headset. In some embodiments, the device 200 may be embedded in such a headset. Also shown in Figure 2 is a communication network 208 to which the communication interface 205 is connected, allowing the device 200 to communicate with a server 209 that operates the shared environment and other devices 210 that also participate in the shared environment 209. Although the drawing shows the server 209 as a single computer, it should be noted that the environment may be implemented on many computers, and each computer may contain one or more processors. All these servers may be located in the same place or distributed across many locations. Thus, in the figure, the server 209 represents one or any combination of many computers, any combination or distribution of remote services associated with the shared environment, and the shared environment itself. Where this disclosure refers to one server or more servers, this is intended to cover both possibilities in each example, as well as embodiments in which server 209 is part of one of the participating devices 210 or is distributed across several participating devices, for example, in a peer-to-peer solution.

[0045] The shared environment, whether it is a virtual reality environment, an augmented reality environment, a metaverse, or any other such term, is referred to as the shared environment in this specification and has the same reference number as server 209. Other participants in the shared environment 209 are represented in the diagram as a single display unit (headset) 210, but in principle there is no limit to the number of participants, as long as the hardware and software used to manage the shared environment 209 have sufficient resources to handle them all.

[0046] At the start of a session, the rendering module 206 can receive the 3D representations and associated personalization data of the participating individuals. During the session, the rendering module 206 receives animation data and personalization index information from the shared environment 209 via the communication interface 205. The rendering module can then render (display) any 3D representation visible to the active user and animate that rendering based on the received animation data. Using the received personalization index data, the rendering module can obtain appropriate personalization data behavior from the personalization data received at the start of the session, and then use this personalization data to modify the rendering and / or animation of the 3D representation in a manner further described below with reference to Figure 5.

[0047] For the sake of simplicity, the memory unit 203 is shown as receiving data only from the personalized data creation module 201 and delivering data only to the animation module 204, and then sending this information to the shared environment 209, or to some online repository accessible to the participants of the shared environment 209 (or directly to each participant 210). It will be understood that the memory unit 203 may be a storage device that is accessible to read and write by all modules. For example, when the rendering module 206 receives a participant's 3D representation and associated personalized data, Figure 2 assumes that this information is stored in the working memory which is part of the rendering module 206. However, regardless of what other features or capabilities the embodiment may include, the device may have any combination of storage devices and memory devices known in the art, and the memory may be shared among all modules and accessible by all modules, or some modules may control memory or memory space that is not accessible by other modules.

[0048] Referring here to Figure 3, a more detailed description of an embodiment of the personalized data creation module 201 is given, along with a description of the method for creating personalized data. Unless otherwise specified, all optional features, steps, or configurations described with reference to this drawing can be freely combined with any embodiment of modules outside the personalized data creation module 201.

[0049] The method described with reference to Figure 3 assumes that the creation of a 3D model representing an individual and the creation of personalized data describing the individual's gestures or habits are part of the same process. As mentioned above, these processes may be separated so that the personalized data is created independently of the creation of the 3D model and the two are combined later. This has specific implications, which will be described in more detail below, including the fact that generic personalized data (also called default or placeholder personalized data) may be generated and that personalized data may be transferred between individuals. The emphasis here is on the creation of personalized data, and embodiments that do not create a 3D model may not perform any steps or actions that are solely related to this.

[0050] As described above, this process involves a user (individual) standing in front of a camera 202 while performing actions that are captured and used by a personalized data creation module 201 to create a 3D model (in some embodiments) and personalized data. In the first step 301, the user is video captured while performing a specific standardized task. These tasks may include a standard repertoire of movements, gestures, and facial expressions, such as turning to one side, turning to the other side, sitting down, standing up, and specific gestures using the arms or hands. In addition, the user may express, for example, a smile, a smirk, disappointment, excitement, anger, etc. In some embodiments, the user may also perform unscripted actions or their own expressions.

[0051] In the next step 302, the video is processed by the personalized data creation module 201 to generate a 3D model of the user. This can be done using methods well known in the art. For example, the video frames can be processed and a 3D model can be generated using photogrammetry. In embodiments where multiple cameras 202 are available during the creation of the 3D model and personalized data, this can include techniques such as stereophotogrammetry and epipolar geometry. If only one camera is available, the 3D information can still be inferred, for example, based on how different parts or points of the user's body move relative to each other when the user turns. Pose estimation techniques may also be used.

[0052] After the 3D model is generated, the process proceeds to the generation of personalized data. This step may be further divided into step 303, which detects specific actions; step 304, which identifies actions; step 305, which encodes actions; step 306, which indexes actions; and step 308, which stores actions, and this may be repeated for several actions until the entire input video has been analyzed. Although these steps are shown in the diagram as being performed sequentially, performing actions in parallel, for example by detecting additional actions while already detected actions are still being encoded, is consistent with the principles of the present invention. Also, step 308, which stores encoded actions, is shown as being performed after all actions have been detected, encoded, and indexed, but of course, it may be stored immediately after encoding and indexing.

[0053] Action detection 303 can be aided by an approximate knowledge of when actions occur, to the extent that the video is scripted (i.e., the user is given instructions on which actions to perform in which sequence, or the user is prompted for each action). However, to improve action detection and better distinguish the start and end of each detected action, actions can be detected using artificial intelligence (AI), particularly machine learning (ML) methods based on deep neural networks, such as convolutional neural networks (CNNs). These methods can be combined with other methods, such as conventional feature detection. For example, feature detection can be used to detect actions and determine their start and end, while neural networks can be used to identify actions. In some embodiments, AI detection of actions may also be able to detect a wider range of actions than those included in the script. Thus, the user can perform action selection based on their preferences and select and remember all actions that the AI ​​can detect as recognizable actions. Thus, the user can generate personalized data to some extent, not only in the sense that actions are described in a way that is based on the user's own gestures, but also in the sense that the personalized action selection can be unique to each user. Again, this does not mean that each user has a unique choice of behavior, and two users do not need to have the same choice of behavior represented in their personalized data.

[0054] If an action is detected, the action must be classified 304. This simply means that after it is determined in step 303 that a series of frames contain an action, the contents of those frames are analyzed to determine what action the user was performing, for example, whether it was a smile, a wink, a yawn, a specific hand gesture, etc. In some embodiments, detection 303 and classification 304 may be performed by individual neural networks. In other embodiments, a single neural network performs detection and classification, in which case detection and classification may be performed as a single step. These methods can also be combined with other methods such as conventional feature detection, pattern recognition, and motion detection. For example, motion detection can be used to detect an action and determine the start and end of the action, feature detection can be used to classify the action as a gesture or facial expression, and a neural network can be used to identify the action.

[0055] The detected and classified behaviors may be encoded as animations in the form of blend shapes, or blend shapes with displacement, pure frame-by-frame vertex changes, or information within a mesh.305 These personalized changes to a 3D model mesh over time, regardless of the technical solution chosen, are referred to herein as animation descriptions. In embodiments where the personalized data can include texture information, including color, this is referred to as animation and / or texture descriptions.

[0056] Each detected and encoded action is indexed in step 306. The index serves to identify a particular action, and when that particular action is detected during user interaction with the shared environment 209, a corresponding description of the animation and / or texture can be obtained, as will be described in more detail below. It will be understood that in different embodiments, indexing may follow a different scheme. A library containing a set of actions for a particular individual will be associated with that individual, so the indexes for the various actions may only need to be locally unique. However, it is consistent that we operate with a globally unique index, i.e., an index that identifies not only the action but also the individual to which it is associated. Furthermore, some embodiments may operate with a static index in the sense that, for example, a smile will always have the same index for all individuals, while other embodiments may generate an index randomly as actions are detected and encoded.

[0057] The process of determining whether all actions have been processed is shown as the next step, 307. As long as there are remaining frames that have not been analyzed or actions that have not been classified, coded, or indexed, the process returns to step 303. (It will be understood that two or more loops may run simultaneously. Thus, action detection can continue until all frames have been processed, classification can be repeated until all detected actions have been classified, coding can be repeated until all classified actions have been coded, and indexing can be repeated until all coded actions have been indexed.)

[0058] In step 307, after it is determined that all detected actions have been processed, they are stored in the storage unit 203 as a library of personalized data in step 308.

[0059] In some embodiments of the present invention, the detection algorithm 303 may also be configured to perform predictions. In such embodiments, the algorithm infers undetected actions, i.e., actions that the user did not perform in the video, based on the detected related actions. For example, if the user provides a smile as one of the detected actions, the algorithm can predict how the user will show excitement and generate a description of the corresponding animation. In this way, it is possible to generate more personalized data than what is actually displayed by the user, providing a rich repertoire of personalized gestures and expressions. It will be understood that the number of actions included in the personalized data will increase the file size. Therefore, in some embodiments, actions can be prioritized based on available memory space or bandwidth, or this can be a user-configurable parameter.

[0060] After personalized data is generated by the personalized data creation module 201 and stored in the storage unit 203, the personalized data can be used as part of the user's interaction with the shared environment 209. Referring to Figure 4, Figure 4 shows a flowchart illustrating the principle of how a client device can establish a connection to the shared environment 209 and begin interacting with it.

[0061] In the first step 401, the device 200 connects to the shared environment 209. This process may include authentication and authorization based on user credentials (e.g., passwords) and other handshake procedures to set the parameters of the underlying communication protocol. This is well known in the art and will not be described further herein.

[0062] In the next step 402, the user's 3D representation is sent to the shared environment 209 along with personalized data. This allows the server operating the shared environment 209 to distribute this representation and personalization to other participating devices 210. In some embodiments, for example, in a peer-to-peer solution or one-to-one interaction, this information is sent directly to the corresponding device rather than to the server computer.

[0063] It should be noted that the initial steps or processes described above initialize the session and enable the device 200 to interact with the shared environment. In most cases, these steps are performed only once when the session starts. However, the following steps in Figure 4 are ongoing during the session and should be understood as a process pipeline rather than individual steps performed sequentially. These steps are primarily handled by the animation module 204.

[0064] After initialization, the process proceeds to step 403, in which the user is captured in video using camera 202 connected to animation module 204. As already mentioned, this may be the same camera or a different camera used by personalized data creation module 201, depending on the design and configuration choices made for a particular embodiment. Step 403 is an ongoing process that can continue with or without interruption as long as the user interacts with the shared environment 209.

[0065] The video captured in step 403 is delivered as a stream of video frames to two processes that may run in parallel. In one process 404, animation data is generated from the video captured by motion capture. This may be performed by a motion capture submodule, which is part of the animation module 204, using techniques well known in the art. The animation data may be in the form of a description of the movement of joints of a skeletal representation or vertices of a mesh. In a process running in parallel with the motion capture performed in process 404, process 405 is performed to detect motion and find the corresponding motion index. This process 405 may be based on artificial intelligence and corresponds to what is performed by the personalized data creation module 201 to detect and identify motion, as illustrated with reference to Figure 3. The motion performed by the user (e.g., a gesture or facial expression) is identified from the received video frames, and an index referring to the description of that motion in the personalized data library is provided as an output, along with any necessary parameters such as start time and duration. Next, in process 406, the animation data from process 404 is sent to the shared environment 209 along with the operation index and parameters from process 405.

[0066] The process described above may include details regarding color, texture, additional objects in the scene (i.e., non-user objects), light, and sound. These aspects will not be described in further detail, but it will be understood that they can be integrated with the information (or parts thereof) already described. In particular, animation data can be generalized to include texture and color in all cases. Furthermore, the detection and identification of motion may be supported by information other than that present in the video frames. For example, if laughter or a specific word is detected in the audio signal, this may be used to identify the corresponding motion, and the index of that motion may be included in the data transmitted in process 407.

[0067] Next, with reference to Figure 5, the processing performed by the rendering module 206 will be described. As with the example above, this is done in the form of an exemplary embodiment, and its variations are within the scope of the present invention.

[0068] In the first step 501, the device 200 connects to the shared environment 209. This corresponds to step 401 in Figure 4. In fact, for a device that interacts with the shared environment by receiving data from and sending data to the shared environment, this step may be performed only once to connect both the animation module 204 and the rendering module 206 to the shared environment. Thus, steps 401 and 501 may be the same step.

[0069] In step 502, the 3D representations of one or more other participants in the shared environment 209 are received along with their respective personalized data. This information may be stored in working memory accessible by the rendering module 206, as described above.

[0070] Similar to Figure 4, the first two steps in Figure 5 initialize the session, enabling device 200 to interact with the shared environment. The subsequent steps in Figure 5 focus on receiving remote information from the shared environment or individual remote devices, and processing this information, which is performed by the rendering module 206. Again, these steps can be understood as processes within a pipeline, rather than as sequentially executed steps.

[0071] In step 503, the animation data, along with the personalized index, is received from at least one remote device, either directly from the remote device or from a server operating the shared environment. Whether the data is received directly from the remote device or from a server depends on the underlying platform embodiment, and it will be understood that the choice may be based on the need to reduce latency that can be achieved by direct communication between devices and the need to coordinate the position and movement of many objects, including characters, that can be achieved by utilizing a server. Any alternative may be chosen for the purposes of this disclosure, and therefore this embodiment will not be described in further detail.

[0072] In the next step, or process 504, the received animation data is extracted and prepared for application to the local representation in the shared environment. In particular, the animation of the 3D representation of the remote user (or other character) is prepared. This is mainly done according to conventional animation, but may include preparation for the next step.

[0073] In parallel with the extraction of animation data, step or process 505 extracts a personalized index and any associated parameters from the received data stream. The personalized index refers to a specific behavior in the personalized data file or library received in step 502, and associated data describing the referenced behavior may be obtained here 506. In a subsequent process 507, the animation data from step 504 may be modified or enhanced based on the obtained personalized data and any received parameters. The output from this process may then be rendered 508 as a representation or animation of the shared environment 209, or as one or more 3D representations that are part of that environment.

[0074] The result that is visible to the user is that 3D representations are not animated based solely on motion capture. Instead, the animation is enhanced by personalized behaviors in other ways that are highly noticeable to the user, such as gestures, facial expressions, changes in color or texture, or behaviors that may be too subtle to register by motion capture but are related to actions that humans pay particular attention to. Furthermore, by providing a library of such behavior descriptions during session initialization, or even before session initialization, it is only necessary to send indices that reference these behaviors during the session. This can significantly reduce the bandwidth required and thereby reduce latency.

[0075] In embodiments of the apparatus 200 including a personalized data creation module 201, an animation module 204, and a rendering module 206, it will be understood that the apparatus 200 may be configured to perform all of the methods described above. However, providing only one of these modules and only the corresponding method within the apparatus is consistent with the principles of the present invention. An apparatus providing only the personalized data creation module 201 may be intended for use, for example, in a studio environment, to create 3D models and personalized data that can later be used by other apparatuses. An apparatus providing only the animation module 204 may be intended for use, for example, to broadcast animations of a performance by an actor or musician on stage. Correspondingly, an apparatus having only the rendering module 206 may be intended for receiving such broadcasts. Furthermore, an apparatus having the personalized data creation module 201 and the animation module but not the rendering module 206 may be intended for studio production and broadcasting, while an apparatus having the animation module 204 and the rendering module 206 but not the personalized data creation module 201 may be intended for interaction in a shared environment based on 3D representations created in advance using different apparatuses.

[0076] In principle, it is also possible to provide a device that includes a personalized data creation module 201 and a rendering module 206 but does not include an animation module 204; however, the useful use cases for such a configuration may be limited.

[0077] In embodiments having two or more of the modules described and configured to correspond to two or more of the methods described above, it will be understood that the embodiments of each module and method are not strictly dependent on one another. Accordingly, embodiments of each module can be combined with any embodiment of the other modules described herein.

[0078] As already mentioned, 3D models and personalized data can, in principle, be created independently. However, this may require standardization with respect to representation and animation. This gives rise to certain possibilities. One such possibility is that an existing user 3D model can later be enhanced with personalized data. Another possibility is the establishment of default personalized data. In this case, personalized data should not be understood as personal in the sense that it is created by or for a specific person or character, but rather as a substitute for such data to enable gestures and expressions that may make the 3D representation appear more personal through the user's human gestures and expressions, even if they are general in this case. This may be useful when a 3D representation is available but personalized data is not. Such general personalized data can be made available, for example, from an online repository associated with the server running the shared environment 209.

[0079] A further development of this embodiment involves applying personalized data for one user to a 3D representation of a different user (or synthetic character). This means that if user A's personalized data is applied to user B's 3D representation, user B will still look the same, but the 3D representation will begin to display mannerisms reminiscent of user A. For example, if user A has a distinctive laugh that involves raising one eyebrow, user B will also begin to laugh in the same way. This principle can be applied to synthetic characters and / or synthetic personalized data (i.e., characters or personalized data that are artificially designed rather than based on real people). The result could be, for example, a 3D representation of a real user beginning to perform actions reminiscent of an anime character, or an anime character beginning to move and make stern faces like a famous musician on stage.

[0080] The default scenario described above involves motion detection by the transmitting animation module 204, but the present invention can provide additional flexibility or extensions. If a participant in the shared environment 209 is using a device 200 with animation capabilities but no personalized data is available, the animation information is sent to the shared environment without the personalized data. Then, some embodiments of the present invention can include motion detection capability in the rendering module 206. In this case, it will be understood that motion detection cannot be based on video frames and must be based on animation data and other information, such as audio. Some examples may be that the detection of clapping hands in animation data may trigger a smile from personalized data, or the word "bravo" in an audio stream may trigger a clapping gesture. In this case, it will be understood that the device providing the animation data may not have a library or file of personalized data available (although this information may be available, for example, from an online repository, even if the device currently used by the user does not have motion detection). If personalized data is unavailable, the rendering module may use default or placeholder personalized data as described above.

[0081] Similarly, the rendering module 206 can be configured to enhance the behavior of the 3D representation by adding actions not identified by any received personalized index, based on the identified actions. For example, the rendering module 206 can be configured to add clapping hands when an index referencing an excited expression is received.

[0082] Referring here to Figure 6, which illustrates the data flow between two devices 200 and 210 during an interaction session in a shared environment. The devices communicate via a communication network 208, which in this case may include any additional participants, as well as any or more servers used to manage the shared environment. The same numbers are used to refer to elements of both devices 200 and 210. With respect to data present on both devices, the reference numbers are prefixed with A and B, respectively, to indicate the source of the data. Users A and B are observed by their respective cameras 202 and thus represented in the stream of video frames provided as input.

[0083] Several protocols are known in the art and may be used in embodiments of the present invention, but in this example, Internet Protocol (IP) 601 is used to carry User Datagram Protocol (UDP) 602. Additional protocol layers may exist on top of these, but they may be application-specific and can be thought of as containers for application data, shown as containing two parts: animation data 603 and motion index 604. The animation data 603 is obtained as output from a motion capture process 605, and the motion index, along with any associated parameters, is provided from a motion detection process 606. These processes are illustrated with reference to Figure 4. Both of these processes receive video frames from a video camera 202 and can run in parallel.

[0084] The UDP and IP layers are used to transmit this data from the local device to the shared environment and from the shared environment to the local device. With respect to device 200 in the figure, it is shown as the generated and transmitted animation data 603A and operation index 604A, as well as the animation data 603B and operation index 604B received from another device, in this case device 210. With respect to device 210, it will be noted that the transmitted data includes animation data 603B and operation index 604B, and the received data includes animation data 603A and operation index 604A generated and transmitted by device 200.

[0085] The received animation data 603 is enhanced with personalized data identified by the received motion index 604 in the process described above, referring to Figure 5. The output from this process 607 is applied to the 3D representation of the remote user and can then be rendered by the display device 207.

[0086] The various modules, features, and configurations described herein can be implemented using combinations of hardware and software components, as will be readily understood by those skilled in the art. General-purpose components such as processors, communication buses and interfaces, user interfaces, power supplies, memory circuits and devices are well known to those skilled in the art and are therefore not described in detail.

Claims

1. A method in a computer system for creating a library of personalized data for use in the animation of 3D models representing humans or humanoid characters in a shared environment, The steps include (301) acquiring the user's video stream while the user is performing a series of actions, The steps include providing frames from the video stream as input to a computerized process that analyzes the frames in order to detect one or more actions performed by the user (303), (304) A step of identifying one or more detected actions, (308) A step of extracting action description data from the video frame for the detected one or more actions and storing the action description data in a library of personalized data, wherein the action description data includes an encoded form of the detected one or more actions. A method comprising the step of associating each stored action description with an index (306).

2. The steps include providing frames from the video stream as input to a computerized photogrammetry process to generate 3D information about the user, A step of generating a 3D representation of the user using the aforementioned 3D information, The further step includes storing the user's 3D representation as a 3D model that can be represented and animated in the shared environment. The method according to claim 1.

3. The method according to claim 1 or 2, wherein at least one of the steps of analyzing the frames to detect one or more actions and identifying the one or more detected actions is performed by a machine learning subsystem (222, 223) trained to perform at least one of the steps of detecting and identifying actions in a stream of video frames.

4. The method according to claim 3, wherein the machine learning subsystem includes at least one artificial neural network.

5. A step of inferring an estimated action description of at least one undetected action based on the action description data of at least one detected action, The steps include storing and indexing the estimated behavioral descriptions in the personalized data library, The method according to any one of claims 1 to 4.

6. The method according to any one of claims 1 to 5, wherein the operation description includes at least one of animation data, texture data, and color data.

7. A method in a computer system that provides personalized data along with animation data for animating a 3D model representing a human or humanoid character in a shared environment, The steps include connecting to the shared environment (401), Step (402) of sending a library of personalized data to the shared environment, which holds action description data including encoded forms of one or more actions, wherein each action description is associated with an index, The steps include (403) acquiring the user's video stream while the user is performing an action, The steps include providing frames from the video stream as input to a computerized motion capture process in order to generate animation data from the user's movements represented in the video stream (404), The steps include providing frames from the video stream as input to a computerized process that analyzes the frames in order to detect (405) and identify at least one action corresponding to an action represented in the library of personalized data, The steps include (406) transmitting the generated animation data and the index of the detected and identified at least one action to the shared environment, method.

8. The method according to claim 7, wherein at least one of the steps of detecting one or more actions and identifying the one or more detected actions is performed by a machine learning subsystem trained to perform at least one of the steps of detecting and identifying actions in a stream of video frames.

9. The method according to claim 8, wherein the machine learning subsystem includes at least one artificial neural network.

10. The method according to any one of claims 7 to 9, wherein the operation description includes at least one of animation data, texture data, and color data.

11. A method in a computer system for applying personalized data to animation data when animating a 3D model representing a human or humanoid character in a shared environment, The steps include connecting to the aforementioned shared environment, A step of receiving a library of personalized data that holds action description data including encoded embodiments of one or more actions, wherein each action description is associated with an index, The steps include receiving animation data and at least one index that references an action represented in the library of personalized data, The steps include: obtaining behavioral description data from the personalized data library using at least one of the aforementioned indexes; The steps include: applying the acquired motion description data to the animation data to generate personalized animation data; A method comprising the steps of rendering and animating the 3D model according to the personalized animation data.

12. The method according to claim 11, further comprising the step of receiving the 3D model together with the library of personalized data.

13. The method according to claim 11, wherein the library of personalized data is received from a repository connected to a computer network, and the animation data and the at least one index referencing the actions represented in the library of personalized data are received from a device participating in the shared environment.

14. The method according to claim 13, wherein the personalized data is generated independently of the generation of the 3D model.

15. A computer device for creating a library of personalized data for use in the animation of 3D models representing humans or humanoid characters in a shared environment, At least one video camera (202), A personalized data creation module (201) includes a submodule (222) configured to receive frames from at least one video camera (202), analyze the received video frames, and detect actions performed by the user depicted in the video frames; and a submodule (223) configured to identify the detected actions, extract action description data including encoded forms of one or more detected actions, and associate each action description with an index (306); A computer device comprising: a storage unit (203) configured to receive and store personalized data received from the data creation module (201);

16. The computer device according to claim 15, further comprising a 3D model creation module (221) configured to receive frames from at least one video camera (202) and to generate a 3D model based on an image of a person represented in the received video frame using photogrammetry processing of the video frame.

17. The computer device according to claim 15 or 16, wherein at least one of the submodules (222) configured to detect motion and the submodules (223) configured to identify motion includes an artificial neural network.

18. A computer device for providing personalized data along with animation data for animation of a 3D model representing a human or humanoid character in a shared environment (209), A storage unit (203) that stores a library of personalized data holding action description data including encoded forms of one or more actions, wherein each action description is associated with an index, Video camera (202), An animation module (204) includes a submodule (231) configured to receive frames from at least one video camera (202), perform motion capture processing on the received video frames to generate animation data from movements performed by a person represented in the video stream, an animation module (204) configured to analyze video frames and detect actions performed by a person depicted in the video frames, and an animation module (204) configured to identify detected actions and obtain an index associated with the identified actions from the storage unit (203), A computer device including a communication interface (205) configured to transmit the library of personalized data, animation data, and motion index to the shared environment (209).

19. The computer device according to claim 18, wherein at least one of the submodules (222) configured to detect motion and the submodules (223) configured to identify motion includes an artificial neural network.

20. A computer device for applying personalized data to animation data when animating a 3D model representing a human or humanoid character in a shared environment, A communication interface (205) for receiving a library of personalized data holding action description data including encoded forms of one or more actions, wherein each action description is associated with an index, animation data, and an index that references an action represented in the library of personalized data from the shared environment (209), A rendering module (206) is configured to include the 3D model within the local representation of the shared environment (209), acquire motion description data referenced by an index received from the personalized data library, apply the acquired motion description data to animation data to generate personalized animation data, and render and animate the 3D model according to the personalized animation data. A computer device comprising: a display unit (207) configured to visualize at least a portion of the shared environment (209) including rendered and animated 3D models.