CUSTOMIZED REPRESENTATION AND ANIMATION OF HUMANOID CHARACTERS
By creating a library of personalization data through video analysis and machine learning, the method addresses the limitations of generic 3D file formats, achieving improved animation accuracy and personalization of humanoid characters in VR and AR environments.
Patent Information
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Current Assignee / Owner
- PICTORYTALE AS
- Filing Date
- 2024-03-14
- Publication Date
- 2026-07-14
AI Technical Summary
Existing 3D file formats for representing humans or humanoid characters in VR and AR environments are generic and fail to exploit the specificities of human anatomy and individual mannerisms, leading to limited animation accuracy and lack of personalization in gestures and facial expressions, especially in live interactions.
A method and device for creating a library of personalization data by analyzing video streams to capture and store personalized gestures and facial expressions, using machine learning to detect and index actions, and applying this data to animate 3D models in shared environments.
Enhances animation accuracy and personalization of 3D representations, allowing for richer and more recognizable interactions by incorporating individual-specific movements and expressions.
Smart Images

Figure 00000033_0000 
Figure 00000033_0001 
Figure 00000033_0002
Abstract
Description
1 / 27 CUSTOMIZED REPRESENTATION AND ANIMATION OF HUMANOID CHARACTERS TECHNICAL FIELD
[0001] The present invention relates to representations of humans or humanoid characters in 3D animation, such as in virtual reality (VR) and augmented reality (AR) environments. In particular, the invention relates to the customization of such representations, particularly, but not exclusively, with respect to movements, gestures, and expressions. BACKGROUND
[0002] Humans are typically represented in AR and VR environments as 3D characters or volumetric characters. 3D characters are generally represented according to popular 3D file formats used in the film industry or for AR and VR purposes. Examples of such file formats include FBX, OBJ, GLB, and USD. A 3D file typically stores representations of one or more of four data types: model geometry, model surface texture, scene details, and model animation. File formats do not necessarily store the same data types for representation. For example, an OBJ file does not store any animation data, while an FBX file stores all animation data.
[0003] All of these file formats are generic in nature. This means they are intended to represent any 3D model, not just 3D representations of human or humanoid characters. In other words, these 3D file formats do not take advantage of specific features and limitations of human anatomy, nor do they include features specifically adapted to represent individual idiosyncrasies, but rather operate within limitations dictated by anatomy or meaning. For example, gestures and facial expressions are limited by human anatomy, and their significance is limited by the extent to which an observer associates a particular gesture or expression with a meaning. (It is true that motion capture and skeletal representations are limited by human anatomy, but not in the sense discussed here).
[0004] The inability of traditional file formats to exploit the specificities of human anatomy and individual mannerisms results in several limitations and disadvantages, especially in live interaction between participants. Petition 870250080960, dated 09 / 09 / 2025, page 11 / 60 2 / 27 represented as 3D characters in a shared AR or VR environment. Such shared environments may also be referred to as a metaverse or the metaverse. This term should not be interpreted as referring to only one specific implementation of a shared environment, nor should it be interpreted as referring to any specific technology with respect to the representation of the environment, whether it is represented on one or multiple computers, whether it is a single shared environment or a collection of multiple shared environments that are connected to each other, and so on.
[0005] To illustrate the shortcomings of traditional file formats, consider a metaverse where user A is represented as a 3D character in an AR environment, and this representation of user A can be observed by user B, who is in a different location. In this scenario, user A stands in front of a camera and interacts with user B, while user B sees user A as a 3D character in AR. The representation of user A will perform gestures (body and face) that are synchronized with those actually performed by user A. This is achieved through the use of a motion capture module that generates motion representations as animation data, and the animation data is transferred over a network and applied to the 3D representation of user A displayed on user B's device. These animations are generally generic in the sense that all movement is represented in the same way, regardless of what moves and how it moves.For example, in the case of joint-based animations or skeletal animation, the vertices are captured and transferred over the network, which are then applied to the user A's 3D character.
[0006] Two disadvantages associated with this method are readily apparent. First, the animation accuracy is only as good as the real-world situation allows in terms of vertex capture and transfer. Typically, with only a single camera, perhaps just a mobile phone, limited vertex movements can be captured. In production situations, advanced motion capture tools use multiple cameras around the object or a single camera with post-processing by software tools. While this provides rich and accurate information about vertex movement, it is not suitable for live transfer, for example, at a conference or game.
[0007] Second, the animation information is not personalized. Because the file formats of 3D characters are generic and the animations lack precision, the captured expressions and gestures fail to convey detail. Petition 870250080960, dated 09 / 09 / 2025, page 12 / 60 3 / 27 customized. For example, while one person might smile with identical lip movements on both sides of the cheek, another person's natural smile might differ in how the lips move, for example, by being asymmetrical. While it's possible to capture these details in a professional studio environment, it's typically not something that can be easily done in a standard user-created 3D model animation.
[0008] In view of this, it is desirable to introduce new methods and systems that can facilitate richer and more personalized animation of 3D characters representing users engaged in interaction in a VR or AR environment. SUMMARY OF THE INVENTION
[0009] In view of the deficiencies described above, the present invention provides methods and devices aimed at providing a richer and more personalized representation of humanoid characters, providing ways to capture and present actions that include personalized behavior in the form of gestures, facial expressions, and much more. The invention includes three interrelated products or aspects, namely, a method and a device for creating personalization data, a method and a device for using or including personalization data when creating an animation, and a method and a device for applying personalization data when rendering an animated 3D model in order to present a personalized animation of the model.Two or more of these aspects can be combined in embodiments of the invention, but typically the creation of personalization data is performed in conjunction with the creation of the 3D model (i.e., it creates the vocabulary of personalized actions) independently of any actual session where animation data is exchanged. The use of personalization data is performed when the animation data that will be used to animate the 3D model is created (i.e., it is the use of the personalization vocabulary at the transmission end or part of a session). Finally, the application of the personalization data to a rendered animation of a 3D model is performed at the receiving end for presentation (i.e., it is the reception and presentation of the vocabulary actions).
[0010] According to the first aspect, a method is provided for creating, in a computer system, a library of personalization data for use in animating a 3D model representing a human or humanoid character in a shared environment. The method includes, for example, a video stream of a user while he or she performs a sequence of actions, providing frames from the stream of Petition 870250080960, dated 09 / 09 / 2025, page 13 / 60 4 / 27 video as input to a computerized process of frame analysis, in order to detect one or more actions performed by the user, identifying one or more detected actions and, for the detected actions, extracting action description data from the video frames and storing the action description data in a personalization data library. The action description data includes coded aspects of the detected action(s), and the respective action descriptions are associated with an index.
[0011] In some embodiments, frames from the video stream are provided as input to a computerized photogrammetry process to generate 3D information about the user. The 3D information can then be used to generate a 3D representation of the user, and the 3D representation of the user can be stored as a 3D model that can be represented and animated in the shared environment.
[0012] The embodiments of the invention may, in order to analyze the frames and detect one or more actions and identify one or more detected actions, utilize a machine learning subsystem trained to perform at least one of the following actions: detect and identify actions in a video frame stream. Such a machine learning subsystem may include at least one artificial neural network.
[0013] Some embodiments of the invention can be configured to infer, based on the action description data for at least one detected action, an estimated action description for at least one action that was not detected. The estimated action description can then be indexed and stored in the personalization data library, along with the action descriptions based on the detected actions.
[0014] Action descriptions may include at least one of the following: animation data, texture data, and color data.
[0015] A computer device according to the first aspect, for creating a personalization data library for use in animating a 3D model representing a human or humanoid character in a shared environment, is also provided. Such a computer device includes at least one video camera, a data creation personalization module configured to receive frames from at least one video camera and including a submodule configured to analyze received video frames and detect actions performed by a user depicted in the video frames, and a submodule configured to identify detected actions and extract action description data, including encoded aspects of one or more detected actions, and associating each action description with an index. The Petition 870250080960, dated 09 / 09 / 2025, page 14 / 60 The 5 / 27 device also includes a storage unit configured to receive and store customization data received from the data creation module.
[0016] A computer device according to this aspect of the invention may also include a 3D model creation module configured to receive frames from at least one video camera and use photogrammetric processing of the video frames to generate a 3D model based on images of a person represented in the received video frames. At least one of the submodules configured to detect actions and the submodule configured to identify actions may include an artificial neural network.
[0017] According to the second aspect of the invention, a method is provided for, in a computer system, providing personalization data together with animation data for animating a 3D model representing a human or a humanoid character in a shared environment.The method includes connecting to the shared environment, transmitting to the shared environment a personalization data library containing action description data, including encoded aspects of one or more actions, each action description being associated with an index, a video stream of a user while the user is performing actions, providing video stream frames as input to a computerized motion process, capturing to generate animation data from the user's movement represented in the video stream, providing video stream frames as input to a computerized frame analysis process to detect and identify at least one action corresponding to an action represented in the personalization data library, and transmitting the generated animation data and the index for at least one detected and identified action to the shared environment.
[0018] A method according to this aspect may perform at least one of the following actions: detect one or more actions and identify the detected action(s), using a machine learning subsystem trained to perform at least one of the following actions: detect and identify actions in a video frame stream. Such a machine learning subsystem may include at least one artificial neural network. The action description includes at least one of the following: animation data, texture data, and color data.
[0019] A computer device according to the second aspect includes a storage unit that stores a library of personalization data. Petition 870250080960, dated 09 / 09 / 2025, p. 15 / 60 6 / 27 containing action description data, including coded aspects of one or more actions, each action description being associated with an index, a video camera, an animation module configured to receive frames from at least one video camera and including a submodule configured to perform motion capture processing of the received video frames to generate animation data from movements performed by a person depicted in the video stream, a submodule configured to analyze video frames and detect actions performed by a person depicted in the video frames, and a submodule configured to identify detected actions and obtain indices associated with the identified actions from the storage unit, and a communication interface configured to transmit the personalization data library, animation data, and action indices to the shared environment.
[0020] In the third aspect of the invention, a method is provided for applying personalization data to animation data in a computer system when animating a 3D model representing a human or a humanoid character in a shared environment. The method includes connecting to the shared environment, receiving a personalization data library containing action description data, including encoded aspects of one or more actions, each action description being associated with an index, receiving animation data and at least one index referencing an action represented in the personalization data library, using at least one index to retrieve action description data from the personalization data library, applying retrieved action description data to the animation data to generate animation personalization data, and rendering and animating the 3D model according to the animation personalization data.
[0021] In some embodiments, this method includes receiving the 3D model along with the customization data library. The customization data library can be received from a repository connected to a computer network, and the animation data and at least one index referencing an action represented in the customization data library can be received from a device participating in the shared environment. In some embodiments, the customization data was generated independently of the generation of the 3D model.
[0022] A computer device configured to operate in accordance with this aspect of the invention includes a communication interface for receiving a personalization data library containing action description data, including aspects Petition 870250080960, dated 09 / 09 / 2025, page 16 / 60 7 / 27 encoded from one or more actions, each action description being associated with an index, animation data and indexes referencing actions represented in shared environment customization data libraries, a rendering module (206) configured to include the 3D model in a local representation of the shared environment (209), to retrieve action description data referenced by indexes received from the customization data library, to apply retrieved action description data to animation data to generate customization animation data and to render and animate the 3D model according to the customization animation data, and a display unit configured to view at least part of the shared environment including rendered and animated 3D models. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The invention will now be described in more detail with reference to the drawings, where:
[0024] FIG. 1 shows a simplified representation of three faces with different customized facial expressions;
[0025] FIG. 2 illustrates, in a block diagram, one embodiment of a client device (200) configured to operate in accordance with the invention;
[0026] FIG. 3 is a flowchart that illustrates an exemplary embodiment of a method for creating personalization data;
[0027] FIG. 4 is a flowchart that illustrates an exemplary implementation of a method for creating and transmitting animation data and personalization data indexes;
[0028] FIG. 5 is a flowchart illustrating an exemplary embodiment of a method for receiving animation data and personalization data indices and for using referenced personalization data to enhance animation data when animating and rendering a 3D model; and
[0029] FIG. 6 is a block diagram that illustrates the data flow between two devices connected to the same shared environment. DETAILED DESCRIPTION
[0030] The present invention relates generally to 3D animation. More specifically, the present disclosure describes methods and systems for animating 3D representations of humans or humanoid characters, particularly 3D representations of users in VR or AR environments. More specifically, the Petition 870250080960, dated 09 / 09 / 2025, page 17 / 60 8 / 27 embodiments of the present invention can be configured to provide customization of the animation of such 3D representations based on captured gestures, facial expressions, and other mannerisms that are specific to individual users or otherwise related to a specific character or representation.
[0031] In the following description of various embodiments, reference will be made to the drawings, in which similar numerical references indicate the same or corresponding elements. It is possible to see that, although many features are necessary for the operation of a computer system, some are well known in the art and are present in most systems. This disclosure will not dwell unnecessarily on such details. Instead, features whose description will facilitate understanding of the invention will be prioritized, while less important details may be presented in a simplified or schematic way. For some features, it is assumed that a person skilled in the art will be able to provide the necessary contextual information from general knowledge of the field.Thus, certain conventional elements may have been omitted in order to exemplify the principles of the invention, rather than burdening the drawings and disclosure with details that do not contribute to the understanding of these principles.
[0032] It should be noted that, unless otherwise indicated, different features or elements may be combined with each other, regardless of whether they have been described together as part of the same embodiment below. The combination of features or elements in the exemplary embodiments is intended to facilitate understanding of the invention, rather than limiting its scope to a limited set of embodiments. To the extent that alternative elements with substantially identical functionality are shown in the respective embodiments, they should be interchangeable. However, for the sake of brevity, no attempt has been made to disclose a complete description of all possible permutations of features.
[0033] Furthermore, those skilled in the art will understand that the invention can be put into practice without many of the details included in this detailed description. On the other hand, some well-known structures or functions may not be shown or described in detail in order to avoid unnecessarily obscuring the relevant description of the various implementations. The terminology used in the description presented below should be interpreted in its broadest and most reasonable form, even when used in conjunction with a detailed description of certain specific implementations of the invention. Petition 870250080960, dated 09 / 09 / 2025, page 18 / 60 9 / 27
[0034] In this disclosure, the term personalized, or personalization, will be used repeatedly. This term should be understood in a technical sense. In many, or most, cases, personalization will refer to something specific to an individual, real or artificial, and some aspects should be considered. First, personalization does not require exclusivity. Personalization data is data that describes aspects of representation and animation that may be perceived as peculiar, but, in principle, personalization data can be identical for representations of different individuals. Second, personalization does not need to be derived from the individual represented. Instead, personalization data that has been created by other means can be applied to a representation of an individual. In other words, although personalization data can be derived from the individual represented and can be unique to that individual, this does not need to be the case.For consistency purposes, the data will be referred to as personalization data regardless of how it was created and whether the same data is applied to more than one representation of more than one individual. Similarly, the term individual will be used to refer to what is being represented, whether it is a real person or an imaginary character, and whether the individual is pre-recorded or interacts with a VR or AR environment in real time.
[0035] The following terminology will be adopted primarily, but any deviations should be interpreted from the context. A user is a person who uses or interacts with a system that implements one or more features of the present invention. An individual is a representation in a system of a human or humanoid character, which may be real (a user) or fictional (a character). A model, or a 3D model, is a data structure that performs, or is, the representation. Thus, a 3D representation is a 3D model that represents a user or a character in the shared environment.
[0036] Thus, a 3D representation, in the context of the present disclosure, is a 3D model of a human or humanoid character. According to the invention, this representation includes, is provided with, or associated with, personalization data. A 3D representation can typically include four types of data: model geometry, surface texture, traditional animation data, and personalization data. The model geometry can typically be represented as a mesh, a skeleton, geometric primitives, and combinations thereof, connected to each other in defined relationships that influence how a point or vertex moves in the model. Petition 870250080960, dated 09 / 09 / 2025, page 19 / 60 10 / 27 Geometric data can influence movement in other parts of the model. Customization data can interact with or influence at least one of the texture data and animation data during rendering. For example, as schematically illustrated in Figure 1, a generic smile might be a symmetrical upward movement of the edges of the mouth, as represented by the first face (101). A person might blush when smiling, as represented by the second face (102). This reddening of the cheeks (104) could be represented by a change in the texture data of the relevant part of the 3D model. Another person, represented here by a third face (103), might have a crooked or unbalanced smile (105), while perhaps raising an eyebrow (106) when smiling.Humans are very attentive to these small variations in facial expressions, but they are not easily captured by systems designed to capture all types of movements and shapes, not just of humans, and particularly without any priority given to relatively subtle variations that are nevertheless laden with meaning for human observers.
[0037] A similar line of reasoning can be applied to gestures. For example, relatively small movements of arms, hands, and fingers can represent significant meaning or be representative of an individual's specific mannerisms, thus making a 3D representation of that individual more recognizable as that particular individual, more personal, when viewed.
[0038] To represent these small variations in gestures and expressions without enhancing the overall capabilities of the recording equipment and significantly increasing the amount of data, the present invention provides a solution in which an additional data type, or data layer, is added. For the purposes of this disclosure, it is assumed that the starting point is a 3D modeling format, or 3D representation, that includes the 3D model itself, animation, and texture. However, the invention is not limited to such formats and can be used in contexts where, for example, one of the animation and texture data is not included in the model itself but is treated separately. This may depend, for example, on the file format used.
[0039] Reference is now made to FIG. 2, which illustrates in a block diagram, one embodiment of a client device (200) configured to operate in accordance with the invention. It should be noted that the functionality provided by the invention involves three main activities. These are the capture and storage of personalization data, capture and transmission of animation and rendering data. These Petition 870250080960, dated 09 / 09 / 2025, page 20 / 60 11 / 27 activities will be described in more detail below. The client device illustrated in FIG. 2 is configured with features and capabilities related to all of these activities. However, it is consistent with the principles of the invention to provide devices configured for only one or only two of these activities. For example, a device configured to capture and store personalization data may include more sophisticated hardware but be too bulky for use during animation (e.g., a game or an AR conference). Thus, personalization data can be captured by a device specifically configured for that purpose, while animation capture and rendering are performed by a dedicated device.Another possibility is a broadcast setup, where the customization and animation are captured with one type of device (e.g., studio equipment), while the rendering is performed by a client device configured to operate as a receiver.
[0040] In view of this, it should be understood that the description of different modules or functionalities in this disclosure does not apply to all possible distributions of functionality among device types. If two or more modules are described here as being able to interact, they may do so by being provided together in a device or by communicating with each other through a communication link, such as a computer network. Furthermore, they may be provided together in the same device, regardless of whether or not this device includes other modules that have been described as part of an embodiment described herein. In other words, devices that are consistent with the principles of the invention may be provided by combining modules or features that have been described herein, even if this combination has not been explicitly described or shown.Instead, the embodiments described are chosen because they facilitate understanding of the invention, not because they constitute a complete catalog of possible embodiments.
[0041] The first module in the embodiment shown in FIG. 2 is a data creation customization module (201). This module includes or is connected to one or more cameras (202). This camera (202) is configured to capture video of a user standing in front of the camera (202) while performing one or more frequent actions, such as turning around, standing up, sitting down, various hand gestures, as well as facial expressions such as smiling, looking angry, looking disappointed, and so on. This process will be described in more detail below with reference to FIG. 3. Petition 870250080960, dated 09 / 09 / 2025, page 21 / 60 12 / 27
[0042] From the video stream delivered to the data creation customization module (201) by the camera, a 3D model of the user is created. The 3D model can be created using photogrammetry, which, in multi-camera realizations (202), may include stereophotogrammetry or epipolar geometry. The data creation customization module (201) also scans the video data stream to detect different movements, gestures, and expressions, collectively referred to here as actions. The detected actions are classified and indexed, and the data describing the action is stored. The data describing the action can be at least one of the following: animation data and texture data. The classified and indexed data will be referred to as customization data. In the drawing, the data creation customization module (201) is shown including three submodules. A first submodule (221) is a 3D model creation module.This module can use photogrammetry to create the 3D model. Some embodiments of the invention may not include this module and instead utilize a 3D model that was (or will be) created externally to the device (200). The 3D model creation module (201) may also be a separate module, rather than being part of the customization data creation module (201). Whether or not a device (200) according to the invention includes a 3D model creation module is largely independent of whether it includes other optional features or how other features are configured.
[0043] A second submodule (222) is a module that can be configured to analyze received video frames and detect actions. The third submodule (223) can be configured to identify detected actions and extract action description data from video frames, such as motion, texture, or color information representative of aspects of the action. For detection and classification / identification, these modules can, for example, include artificial neural networks. For extracting action description data, motion capture techniques can be used.
[0044] The customization data creation module (201) may include or be connected to a storage unit (203) where the customization data may be stored. This storage unit (203) may be part of the customization data creation module (201), may be a separate local device, or may be a cloud-based service. The 3D model data may be stored together with the customization data on the storage unit (201), separately or as part of the same data file, or may be Petition 870250080960, dated 09 / 09 / 2025, page 22 / 60 13 / 27 stored on a separate device. In some embodiments, the 3D model can be created separately from the personalization data, either beforehand (in which case the 3D model may or may not be available when the personalization data is created) or afterward. These options and some implications will be discussed in more detail below.
[0045] The next module is an animation module (204). The animation module (204) is connected to the camera (202). In embodiments such as that illustrated in Figure 2, the personalization data creation module (201) and the animation module (204) may use the same camera (202) or may be connected to separate cameras, for example, if it is determined during the design process for a specific embodiment that the generation of personalization data requires higher resolution video to capture details of the different actions, while animation requires lower resolution to limit the bandwidth required for data transfer. Other optical or image processing capabilities may also differ and this will have to be determined as part of determining the specifications required for a specific implementation of the invention.If the creation and animation of personalization data are performed by completely separate devices, perhaps separated in space and time, these will require their own cameras.
[0046] The animation process performed by the animation module (204) is based on capturing video of an individual performing some activity, for example, related to games or online conferences (VR or AR conference room). To animate a 3D representation of the individual, as presented to other users in the shared environment, the animation module (204) may include a submodule (231) configured to perform motion capture, and the movement of the various vertices is described as animation data. Motion capture, whether traditional marker-based or more recent markerless techniques, and also joint-based and facial motion capture, is well known and understood in the field and will not be described in further detail here.
[0047] In addition to performing motion capture, the animation module (204) processes the video stream in the same way as the data creation customization module (201). However, for the animation module (204) it is sufficient to detect and identify actions. No data describing the detected actions is captured, except to the extent that additional parameters are needed, such as start time and duration. In Petition 870250080960, dated 09 / 09 / 2025, page 23 / 60 14 / 27 instead, the animation module obtains the index of the detected action and includes this index with the animation data. For this purpose, the animation module may include a submodule (232) configured to detect actions and a submodule (233) configured to identify detected actions and obtain indices associated with the identified actions from the storage unit (203).
[0048] When a user connects to a shared environment, the animation module (204) can obtain the 3D model representing an individual (the 3D representation) and the associated personalization data from the storage unit (203) and use a communication interface (205) to transmit the personalization data to the shared environment. The personalization data will then be available to all similar client devices connected to the shared environment. While an interaction session is in progress, the animation module (204) will transmit animation data and personalization data indexes to the shared environment. The operation of the animation module (204) will be described in more detail below with reference to FIG. 4.
[0049] The last module illustrated in FIG. 2 is a rendering module (206). The rendering module (206) is connected to the communication interface (205) and to a display unit (207), for example a VR or AR headset. The device (200) may, in some embodiments, be incorporated into such a headset. Also illustrated in FIG. 2 is a communication network (208) to which the communication interface (205) is connected and through which the device (200) is able to communicate with a server (209) operating the shared environment and with other devices (210) also participating in the shared environment (209). It should be noted that although the drawing shows the server (209) as a single computer, the environment may be implemented on multiple computers, each of which may include one or more processors. These servers may all be located in the same location or may be distributed across multiple locations.In the drawing, the server (209) therefore represents any combination of one or more computers and any combination or distribution of remote services associated with the shared environment, as well as the shared environment as such. When this disclosure refers to server in the singular or servers in the plural, it is intended, in each case, to encompass both possibilities, as well as realizations in which the server (209) is part of one of the participating devices (210) or distributed among several participating devices, for example, in a peer-to-peer solution. Petition 870250080960, dated 09 / 09 / 2025, page 24 / 60 15 / 27
[0050] The shared environment, whether or not it is called a virtual reality environment, augmented reality environment, metaverse, etc., will be referred to here as the shared environment and will have the same reference number as the server (209). Other participants in the shared environment (209) are represented in the drawing as a single display unit (a headset) (210), but there is, in principle, no limitation on the number of participants, provided that the hardware and software used to manage the shared environment (209) have sufficient resources to handle them all.
[0051] At the beginning of a session, the rendering module (206) can receive 3D representations of the participating individuals and associated personalization data. During the session, the rendering module (206) will receive animation data and personalization index information from the shared environment (209) via the communication interface (205). The rendering module can then render (display) any 3D representations visible to an active user and animate this rendering based on the received animation data. The received personalization index data can be used to retrieve appropriate personalization data actions from the personalization data received at the beginning of the session, and this personalization data can then be used to modify the rendering and / or animation of the 3D representations in a manner that will be described in more detail below, with reference to FIG. 5.
[0052] It should be noted that to keep the design simple, the storage unit (203) is shown receiving data only from the customization data creation module (201) and delivering data only to the animation module (204), which in turn transmits this information to the shared environment (209) or to some online repository where it is accessible to the participants of the shared environment (209) (or directly to each participant (210)). It is understood that the storage unit (203) can be a storage device accessible for reading and writing to all modules. For example, when the rendering module (206) receives 3D representations of participating individuals and associated customization data, FIG. 2 assumes that this information is stored in the working memory that is part of the rendering module (206).However, regardless of what other features or capabilities an embodiment may include, a device may feature any combination of storage and memory known in the art, and the memory may be shared among all modules and accessible to all or some of them. Petition 870250080960, dated 09 / 09 / 2025, page 25 / 60 16 / 27 modules can control memory or memory space, which is not accessible to other modules.
[0053] Returning now to FIG. 3, a more detailed description of the embodiments of the personalization data creation module (201) will be provided together with a description of the methods of creating personalization data, unless otherwise indicated, all features, steps or optional settings described with reference to this drawing may be freely combined with any and all embodiments of modules that are external to the personalization data creation module (201).
[0054] The method described with reference to FIG. 3 assumes the creation of a 3D model representing the individual and the creation of personalization data describing the individual's mannerisms or peculiarities as part of the same process. As mentioned above, these processes can be separated so that personalization data is created independently of the creation of the 3D model, and the two can be combined later. This has certain implications that will be described in more detail below, including the fact that generic personalization data (which may also be called standard or placeholder personalization data) can be generated and that personalization data can be transferred between individuals. The emphasis here is on the creation of personalization data, and realizations that do not create the 3D model may not perform steps or actions related solely to this.
[0055] As mentioned above, this process involves a user (individual) standing in front of a camera (202) while performing actions that are captured and used by the data creation module (201) to create the 3D model (in some embodiments) and the personalization data. In a first step (301), the user is captured on video while performing certain standardized tasks. These tasks may include a standard repertoire of movements, gestures, and facial expressions, for example, turning to one side, turning to the other side, sitting, standing, as well as certain gestures with the arms or hands. In addition, the user may, for example, smile, give an ironic smile, express disappointment, enthusiasm, anger, and so on. In some embodiments, the user may also perform unprogrammed actions or expressions.
[0056] In the next step (302), the video is processed by the personalization data creation module (201) and a 3D model of the user is generated. This can be done by Petition 870250080960, dated 09 / 09 / 2025, page 26 / 60 17 / 27 use of well-known methods in the field. For example, video frames can be processed and the 3D model can be generated using photogrammetry. In realizations where more than one camera (202) is available during the creation of the 3D model and personalization data, this may involve techniques such as stereophotogrammetry and epipolar geometry. With only one camera available, 3D information can still be inferred based on, for example, how different parts or points of the user's body move relative to each other when the user turns around. Pose estimation techniques can also be used.
[0057] After the generation of the 3D model, the process proceeds with the generation of the personalization data. This step can be subdivided into detection (303) of a specific action, identification (304) of the action, encoding (305) of the action, indexing (306) of the action and storage (308) of the action and this can be repeated for several actions until the entire input video has been analyzed. These steps are illustrated in the drawing as being executed sequentially, but it is consistent with the principles of the invention to execute actions, for example, in parallel, for example, detecting additional actions while actions already detected are still being encoded. Furthermore, the storage (308) of the encoded actions is shown as being executed after all actions have been detected, encoded and indexed, but these can, of course, be stored as soon as they are encoded and indexed.
[0058] Action detection (303) can be aided by an approximate knowledge of when the action will occur, insofar as the video is scripted (i.e., the user has been given a description of what actions to perform and in what sequence or the user is prompted to perform each action). However, in order to improve action detection and better delineate the beginning and end of each detected action, actions can be detected using artificial intelligence (AI), in particular a machine learning (ML) method based on a deep neural network, for example, a convolutional neural network (CNN). These methods can be combined with other methods, such as traditional feature detection. For example, feature detection can be used to detect an action and determine its beginning and end, while a neural network is used to identify it.AI action detection can also, in some implementations, be capable of detecting a wider range of actions than those included in a script. Thus, the user can perform a selection of actions based on their own preference, as well as any actions that the AI is able to detect. Petition 870250080960, dated 09 / 09 / 2025, page 27 / 60 18 / 27 as a recognizable action can be selected to be stored. Thus, users can, to some extent, generate personalization data, not only in the sense that actions are described in a way that represents them according to their own peculiarities, but the selection of which actions are personalized can also be specific to each user. Again, this does not mean that each user has a unique selection of actions, but that two users do not need to have the same selection of actions represented in their personalization data.
[0059] When an action is detected, it must be classified (304). This simply means that, after it has been determined in step (303) that a sequence of frames contains an action, the content of these frames is analyzed to determine what action the user was performing, for example, whether it was a smile, a wink, a yawn, a specific hand gesture, and so on. In some embodiments, detection (303) and classification (304) may be performed by individual neural networks. In other embodiments, a single neural network performs both detection and classification, in which case detection and classification may be performed in a single step. These methods can also be combined with other methods, such as traditional feature detection, pattern recognition, and motion detection.For example, motion detection can be used to detect an action and determine the beginning and end of the action, feature detection can be used to classify the action as a gesture or a facial expression, while a neural network can be used to identify the action.
[0060] The detected and classified actions can be coded (305) as animation in the form of blendshapes, or blendshapes with displacements, pure information about frame-by-frame vertex changes in a mesh, etc. These custom changes in the 3D model mesh over time will, regardless of the technical solution chosen, be referred to here as animation description. In realizations where the customization data may include texture information, including color, this will be referred to as animation and / or texture description.
[0061] Each detected and coded action will be indexed in step (306). The index serves to identify a specific action so that, when that specific action is detected during a user's interaction with a shared environment (209), the corresponding description of the animation and / or texture can be retrieved, as will be described in more detail below. It will be noticed that the indexing can follow schemes Petition 870250080960, dated 09 / 09 / 2025, page 28 / 60 19 / 27 different in different realizations. The library comprising a set of actions for a specific individual will be associated with that individual, so the indexes for the various actions only need to be locally unique. However, it is consistent to operate with indexes that are globally unique, i.e., such that the index not only identifies the action but also the individual to which it is associated. Furthermore, some realizations may operate with static indexes in the sense that, for example, a smile always has the same index for all individuals, while other realizations may generate indexes randomly as actions are detected and coded.
[0062] A process to determine if all actions have been processed is illustrated with the next step (307). As long as there are remaining frames that have not been analyzed or detected actions that have not been classified, coded, and indexed, the process will return to step (303). (Note that there may be more than one loop running simultaneously. Thus, action detection may continue until all frames have been processed, classification may be repeated until all detected actions have been classified, coding may be repeated until all classified actions have been coded, and indexing may be repeated until all coded actions have been indexed.)
[0063] After it has been determined in step (307) that all detected actions have been processed, these are stored in step (308) as a personalization data library in storage unit (203).
[0064] In some embodiments of the invention, the detection algorithm (303) can also be configured to perform prediction. In such embodiments, the algorithm infers undetected actions, i.e., actions that the user did not perform in the video, based on related actions that were detected. For example, in data where a user provided a smile as one of the detected actions, the algorithm can predict how the user will demonstrate enthusiasm and generate a description of the corresponding animation. In this way, richer personalization data can be generated than that actually displayed by the user, providing a richer repertoire of personalized gestures and expressions. It will be noticed that the number of actions included in the personalization data will increase the file size. Some embodiments may therefore prioritize between actions based on available storage space or bandwidth, or this may be set as a user-configurable parameter. Petition 870250080960, dated 09 / 09 / 2025, page 29 / 60 20 / 27
[0065] After the personalization data has been generated by the data creation module (201) and stored in the storage unit (203), it becomes available for use as part of the user's interaction with a shared environment (209). Reference is now made to FIG. 4, which illustrates in a flowchart the principles of how a client device can establish a connection with a shared environment (209) and initiate interaction.
[0066] In a first step (401) the device (200) connects to a shared environment (209). This process may involve authentication and authorization based on user credentials (e.g., passwords) and other handshake procedures to configure the parameters of the underlying communication protocols. This is well known in the art and will not be described in more detail in this document.
[0067] In a subsequent step (402) the user's 3D representation is transmitted to the shared environment (209) along with the personalization data. This allows the server or servers operating the shared environment (209) to distribute this representation and personalization to other participating devices (210). In some embodiments, for example, in point-to-point solutions or one-to-one interaction, this information is sent directly to a corresponding device and not to a server computer.
[0068] It should be noted that the initial steps, or processes, described above initialize a session and allow a device (200) to interact with a shared environment. In most cases, these steps are executed only once, when a session is started. The following steps in FIG. 4, however, are continuous during the session and should be understood more as a pipeline of processes than as discrete steps executed sequentially. These steps are primarily the responsibility of the animation module (204).
[0069] After initialization, the process advances to step (403), where the user is captured on video by the camera (202) connected to the animation module (204). As already mentioned, this may be the same camera used by the personalization data creation module (201) or a different camera, depending on the design and configuration choices made for a specific implementation. This step (403) will be a continuous process that may continue with or without interruptions as long as the user interacts with the shared environment (209). Petition 870250080960, dated 09 / 09 / 2025, page 30 / 60 21 / 27
[0070] The video captured from step (403) is delivered as a video frame stream to two processors that may be running in parallel. In one processor (404), animation data is generated from the captured video by means of motion capture. This can be performed by a motion capture submodule that is part of the animation module (204), using techniques well known in the field. The animation data can be in the form of motion description, for example, of joints in a skeletal representation and vertices in a mesh. In a process running in parallel with the motion capture performed in processor (404), a process (405) of action detection and search for a corresponding action index is performed.This process (405) can be based on artificial intelligence and corresponds to what is performed in order to detect and identify actions by the personalization data creation module (201), as described with reference to FIG. 3. The actions (e.g., gestures or expressions) performed by the user are identified from the received video frames and the index referring to the description of this action in the personalization data library is provided as output, along with any necessary parameters, such as start time and duration. Then, in a subsequent processing (406), the animation data from the processing (404) is transmitted to the shared environment (209) along with the action index and process parameters (405).
[0071] The process described above may involve details relating to color, texture, additional objects in the scene (i.e., objects other than the user), light and sound. These aspects will not be described in greater detail, but it will be understood that they may be integrated with (or part of) the information already described. In particular, animation data may, in all cases, be generalized to include texture and color. Furthermore, the detection and identification of actions may be supported by other information besides that present in the video frames. For example, if laughter or certain words are detected in an audio signal, this may be used to identify a corresponding action and the index for that action may be included in the data transmitted in the process (407).
[0072] Returning now to FIG. 5, a description of the processing performed by the rendering module (206) will be presented. As in the examples above, this is carried out in the form of an exemplary embodiment, variations of which are within the scope of the invention. Petition 870250080960, dated 09 / 09 / 2025, page 31 / 60 22 / 27
[0073] In a first step (501) the device (200) connects to a shared environment (209). This corresponds to step (401) in FIG. 4. In fact, for devices that interact with the shared environment receiving data from the shared environment as well as transmitting data to the shared environment, this step can be performed only once to connect both the animation module (204) and the rendering module (206) to the shared environment. Thus, step (401) and step (501) can be one and the same step.
[0074] In a step (502) 3D representations of one or more other participants in the shared environment (209) are received along with their respective personalization data. This information can be stored in the working memory accessible to the magnification module (206), as described above.
[0075] As in FIG. 4, the first two steps in FIG. 5 initialize a session and allow a device (200) to interact with a shared environment. The following steps in FIG. 5 focus on receiving remote information from the shared environment or individual remote devices, and processing this information as performed by the rendering module (206). Again, these are steps that can be understood as processing within a pipeline, rather than steps executed sequentially.
[0076] In step (503) animation data along with personalization indices are received from at least one remote device, either directly from the remote device or from a server operating the shared environment. It is understood that receiving data directly from remote devices or from a server depends on the underlying platform and that choices can be made based on the need to reduce latency – which could be done by direct communication between devices – and the need to coordinate the position and movement of many objects, including characters – which could be done using a server. For the purposes of this disclosure, either alternative can be chosen and therefore this aspect will not be discussed in further detail.
[0077] In a next step, or process, (504) the received animation data is extracted and prepared so that it can be applied to the local representation of the shared environment. In particular, the animation of a 3D representation of a remote user (or other character) is prepared. This is done primarily in accordance with traditional animation, but may include preparation for the next step. Petition 870250080960, dated 09 / 09 / 2025, page 32 / 60 23 / 27
[0078] Parallel to the extraction of animation data, a step, or process, (505) extracts customization indexes and any associated parameters from the received data stream. The customization indexes refer to specific actions in the customization data file, or library, received in step (502), and the relevant data describing a referenced action can now be retrieved (506). In a subsequent process (507), the animation data from step (504) can be modified or enhanced based on the retrieved customization data and any received parameters. The output of this process can then be rendered (508) as a representation or animation of the shared environment (209) or one or more 3D representations that are part of that environment.
[0079] The user-visible result is that 3D representations are animated not only based on motion capture. Instead, the animation is enhanced by custom actions, which can include gestures, facial expressions, changes in color or texture, or other forms that may be too subtle to be recorded by motion capture, but are still very perceptible to the user because they relate to actions to which humans pay special attention. Furthermore, by providing a library with these action descriptions during, or even before, the initialization of a session, it is only necessary to transmit indexes that reference these actions during the session. This significantly reduces the bandwidth required and therefore can also reduce latency.
[0080] It will be noticed that, in embodiments of the device (200) that include a data creation module (201), an animation module (204), and a personalization rendering module (206), the device (200) can be configured to perform all the methods described above. However, it is consistent with the principles of the invention to provide only one of these modules, and only the corresponding method, in a device. A device that provides only the personalization data creation module (201) may be one intended to be used, for example, in a studio environment to create 3D models and personalization data that can be used later by another device. A device that provides only the animation module (204) may be one intended to be used to transmit animation, for example, of a performance by actors or musicians on stage.Similarly, a device with only the rendering module (206) may be intended for receiving such transmissions. In addition, a device with a personalization data creation module (201) and module. Petition 870250080960, dated 09 / 09 / 2025, page 33 / 60 24 / 27 animation, but without the rendering module (206), could be a device intended for studio production and broadcasting, while a device with an animation module (204) and a rendering module (206), but without any personalization data creation module (201), could be one intended for interaction in a shared environment based on 3D representations previously created using a different device.
[0081] In principle, it is also possible to provide devices with a personalization data creation module (201) and a rendering module (206), but without an animation module (204), but the valuable use cases for such a configuration may be limited.
[0082] It will also be noticed that, in realizations with two or more of the modules described and configured correspondingly with two or more of the methods described above, the implementation of the respective modules and methods does not depend on each other in any strict sense. Thus, the realizations of the respective modules can be combined with any realization of the other modules described herein.
[0083] As already mentioned, 3D models and personalization data can, in principle, be created independently. This may require standardization regarding representation and animation. This creates certain possibilities. One of these possibilities is that existing 3D models of a user can be subsequently enhanced with personalization data. Another possibility is the establishment of standard personalization data. In this case, personalization data should not be understood as personal in the sense that it is created by or for a specific person or character, but rather that it replaces this data to allow gestures and expressions that can make a 3D representation seem more personal through human gestures and expressions, even if, in this case, they are generic. This can be valuable in cases where a 3D representation is available but without personalization data.This generic personalization data may be available in an online repository, for example, one that is associated with the server that operates the shared environment (209).
[0084] A further development of this aspect is to apply personalization data for one user to a 3D representation of a different user (or a synthetic character). This would mean that if the personalization data for user A were applied to the 3D representation of user B, user B would still have the same Petition 870250080960, dated 09 / 09 / 2025, page 34 / 60 25 / 27 appearance, but the 3D representation would begin to exhibit mannerisms that would be reminiscent of user A. For example, if user A has a characteristic way of smiling while raising an eyebrow, user B would begin to smile in the same way. This principle can be applied to synthetic characters and / or synthetic personalization data (i.e., characters or personalization data that have been artificially designed rather than based on a real person). The result could, for example, be that a 3D representation of a real user would begin to perform actions reminiscent of a cartoon character, or a cartoon character could begin to move and make faces like a famous musician on stage.
[0085] Although the standard scenario described above involves the detection of actions by the animation module (204) on the transmission side, the invention may provide further flexibility or enhancements.If a participant in a shared environment (209) is using a device (200) with animation capabilities but without any personalization data, the animation information will be transmitted to the shared environment without personalization data. Some embodiments of the invention may then include action detection capabilities in the rendering module (206). It is understood that, in this case, action detection cannot be based on video frames, but rather on animation data and other information, for example, audio. Some examples might be that the detection of clapping in the animation data could trigger a smile from the personalization data or the word bravo in the audio stream could trigger a clapping gesture.It will be noticed that, in this case, the device providing the animation data probably did not provide a library or file of personalization data (although this information may be available, for example, in an online repository, even if the device currently in use by the user does not have action detection). If personalization data is not available, the animation module can use standard or placeholder personalization data, as described above.
[0086] Similarly, the rendering module (206) could be configured to enhance the behavior of a 3D representation by adding actions that were not identified by any received customization index based on actions that present. For example, the rendering module (206) could be configured to add clapping if an index referencing an enthusiastic facial expression is received. Petition 870250080960, dated 09 / 09 / 2025, p. 35 / 60 26 / 27
[0087] Reference is now made to FIG. 6, which illustrates the data flow between two devices (200, 210) during an interaction session in a shared environment. The devices communicate through a communication network (208) which, in this case, can be considered as including any additional participants, as well as any server or servers used to manage the shared environment. The same numerical references will be used to refer to elements of devices (200) and (210). Regarding the data present in both devices, the numerical references are appended with A and B, respectively, to indicate the origin of the data. Users A and B are observed by their respective cameras (202) and are therefore represented in the video frame stream provided as input.
[0088] Although several protocols are known in the art and can be used in embodiments of the invention, this example uses the Internet Protocol (IP) (601) to transport the User Datagram Protocol (UDP) (602). Additional protocol layers may be present on top of these, but may be application-specific and may be considered as containers for the application data, which are illustrated as containing two parts, namely, animation data (603) and action indices (604). The animation data (603) is obtained as output from a motion capture processor (605) and the action indices, along with any relevant parameters, are provided by an action detection processor (606). These processes are described above with reference to Figure 4. Both processors receive video frames from a video camera (202) and can be run in parallel.
[0089] The UDP and IP layers are used to transport this data from the local device to the shared environment, as well as from the shared environment to the local devices. With respect to device (200) in the diagram, this is shown as animation data (603A) and action indices (604A) that are generated and transmitted and animation data (603B) and action indices (604B) that are received from another device, in this case from devices (210). It will be noted that with respect to devices (210) the transmitted data includes animation data (603B) and action indices (604B) while the received data includes the animation data (603A) and action indices (604A) generated and transmitted by device (200).
[0090] The received animation data (603) is enhanced with custom data identified by the received action indices (604) in a process described above with reference to FIG. 5. The output of this processing (607) is applied to Petition 870250080960, dated 09 / 09 / 2025, page 36 / 60 27 / 27 3D representation of the remote user, which can then be rendered by a display device (207).
[0091] The various modules, resources, and configurations described herein can be implemented using a combination of hardware and software components in a manner easily understood by those knowledgeable in the field. Generic components, such as processors, buses and communication interfaces, user interfaces, power supplies, memory circuits and devices, and the like, have not been described in detail, as they are well known to those skilled in the art. Petition 870250080960, dated 09 / 09 / 2025, page 37 / 60
Claims
1 / 5 CLAIMS 1. A method in a computer system for creating a personalization data library for use in animating a 3D model representing a human or humanoid character in a shared environment, characterized in that it comprises: obtaining (301) a video stream from a user while the user is performing a sequence of actions; providing frames from the video stream as input to a computerized process for analyzing the frames in order to detect (303) one or more actions performed by the user; identifying (304) one or more detected actions; for one or more detected actions, extracting action description data from the video frames and storing (308) the action description data in a personalization data library, the action description data including encoded aspects of one or more detected actions; and associating (306) each stored action description with an index.
2. A method according to claim 1, characterized in that it further comprises: providing frames from the video stream as input to a computerized photogrammetry process to generate 3D information about the user; using the 3D information to generate a 3D representation of the user; and storing the 3D representation of the user as a 3D model that can be displayed and animated in the shared environment.
3. Method according to claim 1 or 2, characterized in that at least one of: analysis of the frames and in order to detect one or more actions and identify one or more detected actions is performed by a machine learning subsystem (222, 223) that has been trained to perform at least one of: detecting and identifying actions in a video frame stream.
4. Method according to claim 3, characterized in that the machine learning subsystem includes at least one artificial neural network. Petition 870250080960, dated 09 / 09 / 2025, p. 38 / 60 2 / 5 5. A method according to one of the preceding claims, characterized in that it further comprises: inferring, based on the action description data for at least one detected action, an estimated action description for at least one action that was not detected; and storing and indexing the estimated action description in the personalization data library.
6. A method according to one of the preceding claims, characterized in that the description of the action includes at least one of: animation data, texture data, and color data.
7. A method in a computer system for providing personalization data together with animation data for animating a 3D model representing a human or a humanoid character in a shared environment, characterized in that it comprises: connecting (401) to the shared environment; transmitting (402) to the shared environment a personalization data library containing action description data including encoded aspects of one or more actions, each action description being associated with an index; obtaining (403) a video stream of a user while the user is performing actions; providing frames from the video stream as input to a computerized motion capture process to generate (404) animation data of the user's motion represented in the video stream;provision of video stream frames as input to a computerized frame analysis process in order to detect (405) and identify at least one action corresponding to an action represented in the personalization data library; and transmission (406) of the generated animation data and the index for at least one detected and identified action to the shared environment.; 8. Method according to claim 7, characterized in that at least one of the detected and identified actions is performed by a machine learning subsystem that has been trained to perform at least one of the detection and identification actions in a video frame stream. Petition 870250080960, dated 09 / 09 / 2025, p. 39 / 60 3 / 5 9. Method according to claim 8, characterized in that the machine learning subsystem includes at least one artificial neural network.
10. A method according to one of claims 7 to 9, characterized in that the description of the action includes at least one of: animation data, texture data, and color data.
11. A method in a computer system for applying personalization data to animation data when a 3D model represents a human or a humanoid character in a shared environment, characterized in that it comprises: connection to the shared environment; receiving a personalization data library containing action description data, including coded aspects of one or more actions, each action description being associated with an index; receiving animation data and at least one index referencing an action represented in the personalization data library; using at least one index to retrieve action description data from the personalization data library; applying retrieved action description data to animation data to generate personalized animation data; and rendering and animating the 3D model according to the personalized animation data.
12. Method according to claim 11, characterized in that it further comprises receiving the 3D model together with the customization data library.
13. A method according to claim 11, characterized in that the personalization data library is received from a repository connected to a computer network, and the animation data and at least one index referencing an action represented in the personalization data library are received from a device participating in the shared environment.
14. Method according to claim 13, characterized in that the personalization data was generated independently of the generation of the 3D model. Petition 870250080960, dated 09 / 09 / 2025, pp. 40 / 60 4 / 5 15. Computer device for creating a personalization data library for use in animating a 3D model representing a human or humanoid character in a shared environment, characterized in that it comprises: at least one video camera (202); a personalization data creation module (201) configured to receive frames from at least one video camera (202), and including a submodule (222) configured to analyze received video frames and detect actions performed by a user represented in the video frames, and a submodule (223) configured to identify detected actions and extract action description data, including coded aspects of one or more detected actions, and association (306) of each action description to an index; and a storage unit (203) configured to receive and store personalization data received from the data creation module (201).
16. Computer device according to claim 15, characterized in that it further comprises a 3D model creation module (221) configured to receive frames from at least one video camera (202) and to use photogrammetry processing of the video frames to generate a 3D model based on images of a person represented in the received video frames.
17. Computer device according to claim 15 or 16, characterized in that at least one of the submodules (222) configured to detect actions and the submodule (223) configured to identify actions comprises an artificial neural network.
18. Computer device for providing personalization data together with animation data for animating a 3D model representing a human or humanoid character in a shared environment (209), characterized in that it comprises: a storage unit (203) that stores a personalization data library containing action description data including coded aspects of one or more actions, each action description being associated with an index; a video camera (202); an animation module (204) configured to receive frames from at least one video camera (202), and including a submodule (231) configured to perform motion capture processing of the received video frames to generate Petition 870250080960, dated 09 / 09 / 2025, page.41 / 60 5 / 5 animation data from movements performed by a person depicted in the video stream, a submodule (232) configured to analyze video frames and detect actions performed by a person depicted in the video frames, and a submodule (233) configured to identify detected actions and obtain indexes associated with the identified actions from the storage unit (203); and a communication interface (205) configured to transmit the personalization data library, animation data and action indexes to the shared environment (209).
19. Computer device according to claim 18, characterized in that at least one of the submodules (222) configured to detect actions and the submodule (223) configured to identify actions comprise an artificial neural network.
20. Computer device for applying personalization data to animation data when animating a 3D model representing a human or humanoid character in a shared environment, characterized in that it comprises: a communication interface (205) for receiving a personalization data library containing action description data, including coded aspects of one or more actions, each action description being associated with an index, animation data and indexes that reference actions represented in personalization data libraries of the shared environment (209);a rendering module (206) configured to include the 3D model in a local representation of the shared environment (209), to retrieve action description data referenced by indexes received from the customization data library, to apply retrieved action description data to animation data to generate custom animation data and to render and animate the 3D model according to the custom animation data; and a display unit (207) configured to view at least part of the shared environment (209) including rendered and animated 3D models. Petition 870250080960, dated 09 / 09 / 2025, p. 42 / 60;