Avatar joint semantics representation

By mapping user-defined avatar skeleton joints to a detailed reference model within the MPEG-I Scene Description standard, the solution addresses interoperability issues, enabling accurate feature implementation for user-defined avatars in virtual world platforms.

WO2025149308A1PCT designated stage expired Publication Date: 2025-07-17INTERDIGITAL CE PATENT HOLDINGS SAS

Patent Information

Application Number
PCT/EP2024/086534
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-08
Filing Date
2024-12-16
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing virtual world platforms face challenges in ensuring interoperability and accurate implementation of features like automated action or gesture recognition for user-defined avatars due to variations in skeletal model semantics across different representations.

Method used

The proposed solution involves enriching the MPEG-I Scene Description standard's 'MPEG_node_avatar' extension with new attributes that describe the semantics of the skeletal model used for avatars, mapping user-defined skeleton joints to a highly-detailed reference skeleton model, such as the LOA-4 version of the Web3D humanoid skeleton, ensuring unambiguous joint localization on the human body.

Benefits of technology

This approach enables accurate and consistent implementation of features like foot contact detection and gesture recognition across different avatar representations by providing a standardized reference for joint semantics, allowing platforms to adapt algorithms to user-defined avatars effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024086534_17072025_PF_FP_ABST
    Figure EP2024086534_17072025_PF_FP_ABST
Patent Text Reader

Abstract

In one implementation, we describe the semantics of skeleton joints of the avatar to be animated as a mapping to a highly-detailed skeleton model containing most or all of the joints and bones of human anatomy (the "reference skeleton"). Such a reference skeleton (e.g., the "LOA-4" version of the Web3D humanoid skeleton model) features detailed models of the hand and foot phalanges and joints for every vertebra of the spine. Owing to the level of detail of the reference skeleton, each joint of a user-defined animation skeleton can be mapped to one of the joints of the reference skeleton. The identification and location of each joint in the user-defined skeleton on the human body is thereby accurately and unambiguously described. Such semantics can be used in implementing some features of a virtual world platform, for example, in automated action or gesture recognition.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] AVATAR JOINT SEMANTICS REPRESENTATION This application claims the priority to European Application No. 24305020.0, filed on 8 January 2024, which is incorporated herein by reference in its entirety. TECHNICAL FIELD [1] The present embodiments generally relate to the representation of digital humans and their interaction with 3D virtual environments, more particularly, to the animation of avatars based on a skeleton model. BACKGROUND [2] According to an embodiment, a method is presented, comprising: obtaining a reference skeleton model; obtaining a skeleton model for an avatar; mapping each joint of said skeleton model for said avatar to a corresponding joint in said reference skeleton model; and encoding information indicating said mapping for said avatar in scene description. [3] According to another embodiment, an apparatus is presented, comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to: obtain a reference skeleton model; obtain a skeleton model for an avatar; map each joint of said skeleton model for said avatar to a corresponding joint in said reference skeleton model; and encode information indicating said mapping for said avatar in scene description. [4] According to another embodiment, a method is presented, comprising: obtaining a reference skeleton model; decoding, from scene description, information indicating mapping of each joint of a skeleton model for an avatar to a corresponding joint in said reference skeleton model; and matching each joint of said avatar to said corresponding joint in said reference skeleton model based on said mapping. [5] According to another embodiment, an apparatus is presented, comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to: obtain a reference skeleton model; decode, from scene description, information indicating mapping of each joint of a skeleton model for an avatar to a corresponding joint in said reference skeleton model; and match each joint of said avatar to said corresponding joint in said reference skeleton model based on said mapping. [6] One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform the method according to any of the embodiments described herein. One or more of the present embodiments also provide a computer readable storage medium having stored thereon instructions for processing scene description according to the methods described herein. [7] One or more embodiments also provide a computer readable storage medium having stored thereon a scene description generated according to the methods described above. One or more embodiments also provide a method and apparatus for transmitting or receiving the scene description generated according to the methods described herein. SUMMARY [8] In one implementation, we describe the semantics of skeleton joints of the avatar to be animated as a mapping to a highly-detailed skeleton model containing most or all of the joints and bones of human anatomy (the “reference skeleton”). Such a reference skeleton (e.g., the “LOA-4” version of the Web3D humanoid skeleton model) features detailed models of the hand and foot phalanges and joints for every vertebra of the spine. Owing to the level of detail of the reference skeleton, each joint of a user-defined animation skeleton can be mapped to one of the joints of the reference skeleton. The identification and location of each joint in the user- defined skeleton on the human body is thereby accurately and unambiguously described. Such semantics can be used in implementing some features of a virtual world platform, for example, in automated action or gesture recognition. [9] One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform the method according to any of the embodiments described herein. One or more of the present embodiments also provide a computer readable storage medium having stored thereon instructions for processing scene description according to the methods described herein.

[0010] One or more embodiments also provide a computer readable storage medium having stored thereon scene description generated according to the methods described above. One or more embodiments also provide a method and apparatus for transmitting or receiving the scene description generated according to the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] FIG.1 shows an example architecture of an XR processing engine.

[0012] FIG. 2 shows an example of the syntax of a data stream encoding an extended reality scene description.

[0013] FIG.3 shows an example graph of an extended reality scene description.

[0014] FIG.4 shows an example of an extended reality scene description comprising behavior data.

[0015] FIG.5 illustrates skeleton representation of the reference MPEG-I-SD avatar model.

[0016] FIG.6 illustrates topologies of the reference avatar mesh models for the MPEG-I Scene Description standard, for various levels of detail.

[0017] FIG.7 illustrates the LOA-4 variant of the Web3D HAnim animation skeleton.

[0018] FIG.8A illustrates the reference skeleton and avatar skeleton used in the example scene description, and FIG. 8B illustrates the mapping between joints of the reference skeleton and avatar skeleton.

[0019] FIG. 9 illustrates a block diagram describing the parsing of the new “MPEG_node_avatar” extension attributes, according to an embodiment.

[0020] FIG. 10 illustrates a block diagram of the processing model for an application feature that depends on the avatar skeleton semantics, according to an embodiment. DETAILED DESCRIPTION

[0021] Various XR applications may apply to different context and real or virtual environments. For example, in an industrial XR application, a virtual 3D content item (e.g., a piece A of an engine) is displayed when a reference object (piece B of an engine) is detected in the real environment by a camera rigged on a head mounted display device. The 3D content item is positioned in the real-world with a position and a scale defined relatively to the detected reference object.

[0022] For example, in an XR application for interior design, a 3D model of a furniture is displayed when a given image from the catalog is detected in the input camera view. The 3D content is positioned in the real-world with a position and scale defined relatively to the detected reference image. In another application, some audio file might start playing when the user enters an area close to a church (being real or virtually rendered in the extended real environment). In another example, an ad jingle file may be played when the user sees a can of a given soda in the real environment. In an outdoor gaming application, various virtual characters may appear, depending on the semantics of the scenery which is observed by the user. For example, bird characters are suitable for trees, so if the sensors of the XR device detect real objects described by a semantic label ‘tree’, birds can be added flying around the trees. In a companion application implemented by smart glasses, a car noise may be launched in the user’s headset when a car is detected within the field of view of the user camera, in order to warn him of the potential danger. Furthermore, the sound may be spatialized in order to make it arrive from the direction where the car was detected.

[0023] An XR application may also augment a video content rather than a real environment. The video is displayed on a rendering device and virtual objects described in the node tree are overlaid when timed events are detected in the video. In such a context, the node tree comprises only virtual objects descriptions.

[0024] FIG. 1 shows an example architecture of an XR processing engine 130 which may be configured to implement the methods described herein. A device according to the architecture of FIG.1 is linked with other devices via their bus 131 and / or via I / O interface 136.

[0025] Device 130 comprises the following elements that are linked together by a data and address bus 131: ^ a microprocessor 132 (or CPU), which is, for example, a DSP (or Digital Signal Processor); ^ a ROM (or Read Only Memory) 133; ^ a RAM (or Random Access Memory) 134; ^ a storage interface 135; ^ an I / O interface 136 for reception of data to transmit, from an application; and ^ a power supply (not represented in FIG.1), e.g., a battery.

[0026] In accordance with an example, the power supply is external to the device. In each of mentioned memory, the word “register” used in the specification may correspond to area of small capacity (some bits) or to very large area (e.g., a whole program or large amount of received or decoded data). The ROM 133 comprises at least a program and parameters. The ROM 133 may store algorithms and instructions to perform techniques in accordance with present principles. When switched on, the CPU 132 uploads the program in the RAM and executes the corresponding instructions.

[0027] The RAM 134 comprises, in a register, the program executed by the CPU 132 and uploaded after switch-on of the device 130, input data in a register, intermediate data in different states of the method in a register, and other variables used for the execution of the method in a register.

[0028] Device 130 is linked, for example via bus 131 to a set of sensors 137 and to a set of rendering devices 138. Sensors 137 may be, for example, cameras, microphones, temperature sensors, Inertial Measurement Units, GPS, hygrometry sensors, IR or UV light sensors or wind sensors. Rendering devices 138 may be, for example, displays, speakers, vibrators, heat, fan, etc.

[0029] In accordance with examples, the device 130 is configured to implement a method according to the present principles, and belongs to a set comprising: ^ a mobile device; ^ a communication device; ^ a game device; ^ a tablet (or tablet computer); ^ a laptop; ^ a still picture camera; and ^ a video camera.

[0030] In XR applications, scene description is used to combine explicit and easy-to-parse description of a scene structure and some binary representations of media content. FIG. 2 shows an example of the syntax of a data stream encoding an extended reality scene description. FIG. 2 shows an example structure 210 of an XR scene description. The structure consists in a container which organizes the stream in independent elements of syntax. The structure may comprise a header part 220 which is a set of data common to every syntax element of the stream. For example, the header part comprises some of metadata about syntax elements, describing the nature and the role of each of them. The structure also comprises a payload comprising an element of syntax 230 and an element of syntax 240. Syntax element 230 comprises data representative of the media content items described in the nodes of the scene graph related to virtual elements. Images, meshes and other raw data may have been compressed according to a compression method. Element of syntax 240 is a part of the payload of the data stream and comprises data encoding the scene description as described according to the present principles.

[0031] FIG. 3 shows an example graph 310 of an extended reality scene description. In this example, the scene graph may comprise a description of real objects, for example ‘plane horizontal surface’ (that can be a table or a road) and a description of virtual objects 312, for example an animation of a car. Scene description is organized as an array of nodes. A node can be linked to child nodes to form a scene structure 311. A node can carry a description of a real object (e.g., a semantic description) or a description of a virtual object. In the example of FIG. 3, node 301 describes a virtual camera located in the 3D volume of the XR application. Node 302 describes a virtual car and comprises an index of a representation of the car, for example an index in an array of 3D meshes. Node 303 is a child of node 302 and comprises a description of one wheel of the car. The same way, it comprises an index to the 3D mesh of the wheel. The same 3D mesh may be used for several objects in the 3D scene as the scale, location and orientation of objects are described in the scene nodes. Scene graph 310 also comprises nodes that are a description of the spatial relation between the real objects and the virtual objects.

[0032] In time-based media streaming, the scene description itself can be time-evolving to provide the relevant virtual content for each sequence of a media stream. For instance, for advertising purpose, a virtual bottle can be displayed on a table during a video sequence where people are seated around the table. This kind of behavior can be achieved by relying on the framework defined in the Scene Description for MPEG media document.

[0033] Currently, the MPEG-I Scene Description framework uses “behavior” data to augment the time-evolving scene description and provides description of how a user can interact with the scene objects at runtime for immersive XR experiences. These behaviors are related to pre- defined virtual objects on which runtime interactivity is allowed for user specific XR experiences. These behaviors are also time-evolving and are updated through the existing scene description update mechanism.

[0034] FIG.4 shows an example of an extended reality scene description comprising behavior data, stored at scene level, describing how a user can interact with the scene objects, described at node level, at runtime for immersive XR experiences. When the XR application is started, media content items (e.g., meshes of virtual objects visible from the camera) are loaded, rendered and buffered to be displayed when triggered. For example, when a plane surface is detected in the real environment by sensors, the application displays the buffered media content item as described in related scene nodes. The timing is managed by the application according to features detected in the real environment and to the timing of the animation. A node of a scene graph may also comprise no description and only play a role of a parent for child nodes. FIG. 4 shows relationships between behaviors that are comprised in the scene description at the scene level and nodes that are components of the scene graph. Behaviors 410 are related to pre-defined virtual objects on which runtime interactivity is allowed for user specific XR experiences. Behavior 410 is also time-evolving and is updated through the scene description update mechanism.

[0035] A behavior comprises: - triggers 420 defining the conditions to be met for its activation; - a trigger control parameter defining logical operations between the defined triggers; - actions 430 to be proceeded when the triggers are activated; - an action control parameter defining the order of execution of the related actions; - a priority number enabling the selection of the behavior of highest priority in the case of competition between several behaviors on the same virtual object at the same time; - an optional interrupt action that specifies how to terminate this behavior when it is no longer defined in a newly received scene update; for instance, a behavior is no longer defined if a related object does not belong to the new scene or if the behavior is no longer relevant for this current media (e.g., audio or video) sequence.

[0036] Behavior 410 takes place at scene level. A trigger is linked to nodes and to the nodes’ child nodes. In the example of FIG.4, Trigger 1 is linked to nodes 1, 2 and 8. As Node 31 is a child of node 1, Trigger 1 is linked to node 31. Trigger 1 is also linked to node 14 as a child of node 8. Trigger 2 is linked to node 14. Indeed, a same node may be linked to several triggers. Trigger n is linked to nodes 5, 6 and 7. A behavior may comprise several triggers. For instance, a first behavior may be activated by trigger 1 AND trigger 2, AND being the trigger control parameter of the first behavior. A behavior may have several actions. For instance, the first behavior may perform Action m first and, then action 1, “first and then” being the action control parameter of the first behavior. A second behavior may be activated by trigger n and perform action 1 first and, then action 2, for example.

[0037] Often, the animation of an avatar body is driven by the motion of a skeleton. The skeleton is a directed acyclic graph of joint nodes linked by bone edges. The degrees of freedom of a skeleton are the 3D positions of its joints, the relative 3D rotations of each bone with respect to its parent bone in the graph, and the absolute position of its root joint. Joints are often mapped to physical human body joints such as elbow, ankle, and wrist as shown in FIG. 5. Note that the naming of joints may differ across skeleton representations. The root joint of the graph is usually chosen to be the pelvis, named “hips” in FIG.5.

[0038] The envelope of the avatar body is represented by a mesh. Example avatar mesh geometries are shown in FIG.6, where topologies of the reference avatar mesh models for the MPEG-I Scene Description standard are illustrated at various levels of detail (from left to right: high, medium and low). The animation of the envelope mesh is controlled by the motion of the skeleton joints through a process known as skinning. The most popular approach to skinning is Linear Blend Skinning (LBS). In this approach, the position of each vertex of the avatar mesh in a pose of the skeleton during an animation – hereafter referred to as an “animation pose” – is computed as a geometric transformation of the position of the same vertex in a known reference pose of the skeleton known as the “binding pose”. The “pose” of the skeleton is defined here as a particular instance of relative joint positions and bone rotations that defines the posture of the avatar body.

[0039] The skinning transformation applied to each of the mesh vertices is computed as a function of the 3D transformations undergone by a set of the skeleton joints attached to the considered vertex, between the binding pose and the animation pose. It depends on three sets of parameters: ^ A set of skinning weights, summing up to 1, defining which joints in the skeleton influence the position of the considered vertex and by how much. These weights are pre- determined as they do not depend on the animation pose. ^ A set of 3D “inverse bind” matrices defining the transformation of each joint from its world position in the binding pose of the skeleton to its local coordinate system in the skeleton (joint space). The “world position” refers the position of a point in the “world space” 3D coordinate system where the avatar and 3D virtual environment is represented. The local coordinate system of a joint is often aligned with the bone manipulated by the joint. Inverse-bind matrices can be precomputed prior to an animation as they depend only on the binding pose. ^ A set of 3D “joint-to-world” matrices defining the transformation of each joint from the above-defined joint space to the world space in the animation pose. The outputs of these transformations provide the positions of the skeleton joints in the animation pose.

[0040] Thus, given pre-determined skinning weights and inverse bind matrices, the positions of the mesh vertices in the frames of an animation can be obtained from the positions of the skeleton joints and the rotations of the corresponding bones at each frame. The positions of mesh vertices define the geometry of the avatar envelope. After texturing and lighting, the image of the avatar in a picture is obtained by rendering the textured mesh from a given camera viewpoint.

[0041] Morph targets provide another way to animate an avatar mesh, and particularly to sculpt facial expressions. They are essentially 3D deformations with respect to the geometry of a base mesh and are represented by the 3D offsets to the positions of each vertex of the base mesh. Typically, the base mesh models the avatar with a neutral face expression that shows no emotion. Morph targets are provided in addition to the base mesh. Each morph target is associated with a scalar weight. The desired shape of the mesh is obtained by adding to the vertex positions of the base mesh a linear combination of several morph target offsets weighted by their corresponding weights. In a typical setting, morph targets are mapped to deformations of the neutral face geometry incurred by the activation of facial muscles. Combining these deformations provides a convenient way for artists to sculpt the shape of the face mesh so that it expresses a specific emotion.

[0042] Avatars in the MPEG-I Scene Description standard

[0043] Avatars serve as user representations in gaming or virtual world platforms, allowing users to interact with their environment and with other users on the platform.

[0044] The glTF and MPEG-I Scene Description (SD) specifications provide standardized representations of 3D virtual worlds that allow interoperability across a variety of virtual environments proposed by these platforms. The glTF standard on which MPEG-I SD builds includes standardized descriptions of an avatar representation and an avatar animation model, including a skeleton of joints, skinning parameters and morph targets for sculpting facial expressions. The glTF skin object holds all the parameters needed to implement the mesh skinning process described above. In particular, theskin.joints array points to a hierarchy of node instances representing the skeleton joints. Each node stores the 3D transformation that defines its position and orientation with respect to its parent joint.

[0045] Further, amendment 2 of MPEG-I SD specifies a glTFnode for avatars by defining an MPEG_node_avatar extension. MPEG_node_avatar does not standardize a specific avatar representation scheme that would include a mesh model with a pre-determined topology, morph targets for facial animation, and a skeleton with its associated skinning parameters for the animation of the body. Instead, the representation scheme of an avatar is specified by the mandatorytype attribute ofMPEG_node_avatar, which holds a Uniform Resource Name indicating where the scheme is specified. A reference model that can be used by default is proposed in informative annex H of MPEG-I SD Amendment 2. In addition to the type attribute, the MPEG_node_avatar extension includes an array of Mapping instances that defines the semantics of body parts represented by sub-meshes of the full avatar mesh (e.g., “full_body / upper_body / arm_right / upper_arm_right”), and referenced by node instances. This allows interactivity triggers, e.g., detection of collisions with other scene elements, to be associated with specific body parts whose semantics is unambiguously defined.

[0046] For the sake of interoperability, immersive platforms featuring avatars interacting in a virtual environment will have to deal with a variety of avatar representations. Users of such platforms will want to bring their own avatars, perhaps created on other platforms or using external 3D asset creation tools. In particular, the skeletal animation model presented above, which is the standard way to animate virtual characters, will be instantiated with different skeleton topologies, i.e., count, connectivity and positions in the human body of the joints making up the skeleton graph, and morphologies, i.e., bone lengths given a predefined set of joint locations on the human body.

[0047] This also means that immersive platforms will have to deal with a variety of skeletal model semantics in the representations of avatars. The semantics of the skeletal model refers to the identification of the joints and their locations on the human body. The identification of a joint may be specified by a name, for instance “left knee joint”. For some joints, e.g., located on the spine, naming may be ambiguous. For instance, naming a joint “upper chest” does not unambiguously locate it on the human body. The location of such joints may be defined, for example, by localizing it on a reference detailed model of the human anatomy. In the considered example, the “upper chest” joint could be mapped to a specific vertebra on the spine.

[0048] For the purpose of animating a character, the semantics of the skeletal model is not required. As explained above, the world positions of the skeleton joints in the animation pose at each frame, together with the precomputed skinning weights and inverse bind matrices, are enough to compute the geometry of the avatar mesh at each frame of an animation. From this geometry and the associated texture, the avatar can be rendered in the image of the environment presented to the user display. The anatomical mappings of the joints are not needed in this process.

[0049] However, the implementation of some of the features of a virtual world platform may depend on the semantics of the skeletal model used to animate an avatar. For instance, the recognition of gestures or actions, or the detection of the proximity of specific body parts to scene elements, when it is based on the skeleton, will require identifying the joints relevant to these body parts and knowing their locations on the human body. The implementations of these features will likely be based on a predetermined reference skeletal model used by the platform. If these implementations are to be applied to user-defined avatars with their own skeleton model different from the reference skeleton model, adaptations of the algorithm designed for the reference skeletal model will be required. To this purpose, the semantics of the joints in the user-defined avatar representation will need to be specified to the platform.

[0050] To illustrate this requirement, consider as a first example the animation of an avatar walking in a 3D virtual environment. For realism, the application implementing the animation may want to have a sound emitted at each footstep of the avatar. In order to detect foot contact, following a standard approach as described in an article (see L. Mourot et al., “UnderPressure: Deep Learning for Foot Contact Detection, Ground Reaction Force Estimation and Footskate Cleanup”, in Computer Graphics Forum, Vol. 41 n°8, December 2022, pp. 195-206), the velocity and / or altitude with respect to the floor of foot joints of the skeleton will be thresholded at each animation frame. The thresholds will be set with respect to the semantics of the reference skeletal model used by the platform. For instance, this reference skeletal model may have a “toe” joint located at the distal phalangeal joint of the big toe and a “foot” joint located at the ankle. The thresholds used for foot contact detection will be adapted to the vertical distance of these joints to the ground, which depends on the bone lengths of the skeleton and the location of the joints in the human body. The custom skeletal model of a user-defined avatar may, however, have a different skeletal model semantics. For instance, the foot joints in this model may consist of a “heel” joint and a “toe” joint located at the base of the big toe. As a result, the foot detection thresholds computed for the reference skeleton semantics will not be applicable to the user-defined skeleton representation. An adaptation or a re- computation of the thresholds will require the platform to know the semantics of the user- defined skeletal model, specifically their precise location in the human body.

[0051] As a second example, consider a feature that automatically recognizes semantic gestures of the avatar, such as “raising the right arm to shake hands”, perhaps in order to trigger a social interaction with another avatar. This feature is especially useful when the animation of the avatar is controlled in real-time by a user rather than generated by activating a predefined animation. The implementation of this gesture recognition algorithm on a virtual world platform would typically rely on temporal sequences of joint positions, especially arm and hand joints. In this case, the gesture recognition algorithm will depend on the semantics of the reference skeletal model used by the platform, and specifically on the anatomical locations of the joints in the human body. If it is to be run on a custom avatar imported by a user, it will need to be adapted to the potentially different joint semantics of the skeletal model of the custom avatar. Hence, a specification of the joint semantics of the custom user avatar model will need to be made available to the platform. The platform could then train different gesture recognition algorithms on different skeleton joint semantics and select the algorithm suitable for the joint semantics of the user-defined skeleton based on the specification of this semantics.

[0052] Note that the semantics of the skeleton only provides information about the topology of the skeleton. For implementing features such as those described above, an application would also need to know its morphology, i.e., the skeleton bone lengths. This information can be derived from any animation pose of the skeleton as the norms of the translation vectors between connected joints in the skeletal graph.

[0053] In summary, to ensure interoperability of avatar representations across platforms there is a need for specifying the semantics of the skeleton joints used to animate user avatars, i.e., their precise localizations in the human body. Based on this specification, a platform can then adapt the implementation of some of the features it proposes, such as automated action or gesture recognition of an avatar, to the semantics of the skeleton model of the user avatar.

[0054] In order to solve the aforementioned issue, we propose to enrich the “MPEG_node_avatar” extension of MPEG-I SD with new attributes that describe the semantics of the skeletal model used for an avatar.

[0055] Specifically, we describe the semantics of skeleton joints of the avatar to be animated as a mapping to a highly-detailed skeleton model containing most or all of the joints and bones of human anatomy. This model will be referred to hereafter as the “reference skeleton”. An instance of such a reference skeleton is provided by the “LOA-4” version of the Web3D humanoid skeleton model, specified in section 4.8.5 of ISO / IEC 19774-1 2019 (Humanoid Animation (HAnim architecture)) and reproduced in FIG.7. It features detailed models of the hand and foot phalanges and joints for every vertebra of the spine.

[0056] Owing to the level of detail of the reference skeleton, each joint of a user-defined animation skeleton can be mapped to one of the joints of the reference skeleton. The identification and location of each joint in the user-defined skeleton on the human body is thereby accurately and unambiguously described.

[0057] A specification of the joint mapping representation that is compliant with the MPEG-I Scene Description format is provided below. However, it has to be understood that its meaning and use is generic. The proposed representation could be encoded in any other scene description format, such as XML or USD.

[0058] Proposed additions to the “MPEG_node_avatar” extension

[0059] The “MPEG_node_avatar” extension of the MPEG-I SD “node” object specifies the description of an avatar that is to be represented and animated in an application that consumes the scene description. The proposed additions enrich “MPEG_node_avatar” with attributes that define the semantics of the joints of the avatar skeleton. The semantics is described as a mapping of the avatar skeleton joints to corresponding joints of a pre-determined detailed reference skeleton model. The specification of joints in this reference skeleton model unambiguously determines their locations in the human body.

[0060] For the purpose of illustration, but without loss of generality, the specification of the avatar skeleton is assumed to conform to the MPEG-I SD standard. The skeleton joints are described as node instances in the “nodes” array of the scene description. The child to parent relationships that are encoded in the “children” attributes of the “node” instances define the structure of the skeleton graph. The skeleton joints are referenced as node indexes in the “joints” attribute of the “skin” object of the scene description attached to the avatar node.

[0061] According to one embodiment, the semantics of the joints of an avatar skeleton model is described by means of a new “skeleton_semantics” attribute that is added to the “MPEG_node_avatar” extension, as shown in Table 1. This attribute is optional. It is an instance of a new “Skeleton_Semantics” type, whose properties are shown in Table 2. In the Usage column, “M” stands for Mandatory and “O” for Optional. “skeleton_semantics” is needed only for application features that require the knowledge of the semantics of user-defined skeletons and can be ignored by applications that do not propose such features. If this attribute is not provided in a scene description, applications that require the knowledge of skeleton joint semantics for implementing some of their features may simply deactivate these features.

[0062] Within a “Skeleton_Semantics” instance, the mapping between the joints in a user-defined skeleton and in a reference skeleton is described by means of two attributes. A first attribute named “reference_skeleton” points to the specification of the “reference skeleton”. A second attribute named”joints_semantics” describes the joint mappings between the user-defined skeleton and the reference skeleton. Name Type Required Description skeleton_ Skeleton_ Semantics of joints of the skeleton used to No semantics Semantics animate the avatar Table 1: new “skeleton_semantics” attribute of the “MPEG_node_avatar” extension Name Type Usage Description reference_skeleton_repr Type of reference skeleton representation integer O esentation (see Table 3) if(reference_skeleton_r epresentation) == 0) { URN that uniquely identifies the avatar reference_skeletonstring Orepresentation that defines the reference skeleton model. Array of joint names in the reference skeleton corresponding to the joint names joints_semanticsstring[1-*] Oof the avatar skeleton referenced in the skin.joints array } if(reference_skeleton_r epresentation) == 1) { Array of indexes in the “nodes” array reference_skeleton integer[1-*] O of the nodes representing the hierarchy of joints of the reference skeleton Array of indexes in the “nodes” array of the nodes representing the joints of the joints_semanticsinteger[1-*] Oreference skeleton corresponding to the avatar skeleton joints in the skin.joints array } Table 2: description of the “Skeleton_Semantics” type properties. reference_skeleton_representation value Description 0 Representation of the reference skeleton as a URN 1 Representation of the reference skeleton as a hierarchy of joint nodes in the “nodes” array Table 3: semantics of the “reference_skeleton_representation” attribute

[0063] The example specification of these two attributes in conformance to the MPEG-I Scene Description standard is shown in Table 2. They can be represented in one of two formats. The two formats are mutually exclusive. The choice of the representation format is defined by the value of the ”reference_skeleton_representation” attribute, specified in Table 3.

[0064] If ”reference_skeleton_representation” is set to 0, “reference_skeleton” provides a Uniform Resource Name (URN) that points to a description of the reference skeleton model where each joint must be given a unique identifying name. Typically, this name matches the anatomical name of the joint in the human body, as illustrated in FIG.7. For disambiguation purposes, joints located in the left (resp. right) limbs of the body can be prefixed by “left_” or “l_” (resp. “right_” or “r”). The ”joints_semantics” attribute is an array of strings that provides, for each joint of the avatar skeleton, the name of the corresponding joint on the reference skeleton. The ”joints_semantics” array and the “skin.joints” array referencing the nodes of the avatar skeleton must have the same number of joint elements, taken in the same order. Alternatively stated, for any index i in these arrays, the ithelement of ”joints_semantics” must reference the same joint in the avatar skeleton as the ithelement of ”skin.joints”. For instance, assuming that skin.joints

[0014] , i.e., the element of the “skin.joints” array associated with the avatar skeleton with index 14, represents the left ankle joint, and that the location of this joint in the avatar body matches the joint named “l_talocrural” in the specification of the reference skeleton provided in the “reference_skeleton” URN, thenjoints_semantics

[0014] should be set to “l_talocrural”.

[0065] If ”reference_skeleton_representation” is set to 1, the reference skeleton is provided in the scene description as a hierarchy of nodes in the “nodes” array. The “reference_skeleton” attribute specifies the indexes of these nodes in the “nodes” array. Each reference skeleton “node” must be identified by a name, provided in the “name” attribute of the node, that unambiguously specifies its position on the human body. The same recommendations for joint naming as advised in the above for a URN representation of the reference skeleton apply. The ”joints_semantics” attribute is an array of indexes into the “nodes” array. The ithelement of ”joints_semantics” refers to the ithjoint of the avatar skeleton. It contains the index of the reference skeleton joint in the “nodes” array that maps to this avatar skeleton joint. The ”joints_semantics” array and the “skin.joints” array referencing the nodes of the avatar skeleton must have the same number of joint elements, taken in the same order. For instance, assume that skin.joints[5], i.e., the element of the “skin.joints” array associated with the avatar skeleton with index 5, represents the upper chest joint, and that the location of this joint in the avatar body matches the joint named “vt6” in the specification of the reference skeleton. Further assume that the “vt6” joint of the reference skeleton is instantiated as node 23 in the “nodes” array of the scene description. Then, node 23 in the “nodes” array should have a “name” attribute set to “vt6”, the “reference_skeleton” array should contain number 23, andjoints_semantics[5] should be set to 23.

[0066] As an alternative embodiment, the semantics of joints could also be specified in the existing “mappings” array of the “MPEG_node_avatar” extension. To this end, the (path, node) pairs representing the semantic mappings between sub-meshes of the avatar mesh and the corresponding semantic parts of the body could be extended by appending (path, node) pairs defining the semantics of the skeleton joints. In each such pair, the “node” would be defined as the index of the node representing the considered joint of the avatar skeleton in the “nodes” array of the scene description, and the”path” would be defined as the name of the reference skeleton joint corresponding to the considered joint. These reference skeleton joint names would be defined in an extra “reference_skeleton” attribute of “MPEG_node_avatar” with the same semantics as the “reference_skeleton” attribute in Table 2 when ”reference_skeleton_representation” is set to 0. Specifically, this “reference_skeleton” attribute would define a URN containing a description of the highly detailed reference skeleton model where each joint is given an unambiguous name. The “path” items in the (path,node) pairs defining the semantics of the joints would then be chosen among those names. The URN attribute can be optional by assuming a default URN is used if the attribute is not provided, where the default URN and corresponding reference skeleton description would be provided, e.g., in a normative annex in the standard.

[0067] As another alternative embodiment, the reference skeleton is specified in a normative annex of the standard rather than referenced as an attribute of eachSkeleton_Semantics instance. This is a variant of the reference_skeleton_representation=0 option described above in which thereference_skeleton attribute is removed.

[0068] Example MPEG-I SD description

[0069] FIG.8 illustrates an example scene description in MPEG-I SD format that is provided below in the JSON text. It describes a scene with 3 root nodes, with indexes 0, 1 and 11 in the “nodes” array.

[0070] Node 0, named “user_avatar”, describes the avatar of a user and is extended with the “MPEG_node_avatar” extension.

[0071] Node 1, named “reference_humanoid_root”, represents the root joint of the left leg of the reference skeleton represented in FIG. 7. For clarity, the other limbs of the reference skeleton have not been included in the scene description. Nodes 1 through 10 represent the hierarchy of joints in the highly detailed left leg model of FIG.7, wherein node n+1 is the child of node n for n in {1, 2, …, 9}, as indicated by the “children” attributes of the nodes. The relative pose of each node with respect to its parent is encoded by “translation” and “rotation” transform attributes. Each of these nodes is named according to the legend of FIG.7.

[0072] The skeleton of the avatar represented by node 0 is described in nodes 11 through 15, which form a second hierarchy of nodes wherein node 11 is the root and node n+1 is the child of node n for n in {11, 12, …, 14}. It should be understood that the numbering of nodes representing the skeleton of an avatar needs not be contiguous as is the case in the provided example. Like for the reference skeleton, for clarity, only the left leg joints have been included in the description. Each of the avatar joints is named in accordance with the avatar representation scheme specified by the “type” attribute of “MPEG_node_avatar”. In this example, the naming conforms to the MPEG-I SD reference avatar model represented in FIG.5.

[0073] The reference skeleton and avatar skeleton parts described in the example scene description are represented in FIG. 8A, and the mapping of joints is shown in FIG.8B, where the index of the node representing the joint in the array of nodes in the scene description is shown near each joint.

[0074] The geometry of the avatar represented by node 0 is modelled by mesh 0 in the “meshes” array and animated using skinning parameters described in skin 0 of the “skins” array. The “joints” attribute of skin 0 specify the joints used for skinning are referenced by nodes 11 through 15 in the “nodes” array, which form the skeleton of the avatar. The “skeleton” attribute points to the root of this joint hierarchy, i.e., node 11. The inverse bind matrices that transform the joint positions in bind pose to their local coordinate system are provided in a buffer referenced by accessor 5 in the “accessors” array, not included in the description for clarity.

[0075] The “JOINTS_0” (resp. “WEIGHTS_0”) attribute of the only primitive in the skinned avatar “mesh.primitives” array defines the indices of the joints in the “skin.joints” array (resp. the weights associated to these joints) that are used to animate each vertex of the mesh.

[0076] In accordance with one embodiment, the semantics of the avatar skeleton joints is specified using a new “skeleton_semantics” attribute in the “MPEG_node_avatar” extension. This attribute is an instance of the “Skeleton_Semantics” type defined in Table 2. Its “reference_skeleton_representation” attribute, set to 1 in this example, indicates that the semantics of joints in the reference skeleton is specified by a hierarchy of nodes in the “nodes” array of the description. The indexes of these nodes are provided by the “reference_skeleton” attribute of “skeleton_semantics”. In this example, they are the integers from 1 to 10. The semantics of the avatar skeleton joints is specified in the “joint_semantics” attribute of “skeleton_semantics” as the indexes of the reference skeleton joint nodes corresponding to each avatar joint. This specification is illustrated by two examples below.

[0077] As a first example, the first avatar joint for the animation is the node whose index is given by the first element in the “skin.joints” array, hence node 11. In the “nodes” array this is node named “avatar_pelvis_joint”. The corresponding joint in the reference skeleton is specified by the first index in “MPEG_node_avatar.skeleton_semantics.joint_semantics” array, hence node 1, which points in the “nodes” array to the joint named “reference_humanoid_root”. Hence, the semantics for the “avatar_pelvis_joint” is specified as the “reference_humanoid_root” joint of the reference skeleton, shown in FIG.7.

[0078] As a second example, the third avatar joint is the node whose index is given by the third element in the “skin.joints” array, hence joint 13. In the “nodes” array this is node named “avatar_l_lowerleg_joint”. The corresponding joint in the reference skeleton is specified by the third index in the “MPEG_node_avatar.skeleton_semantics.joint_semantics” array, hence node 3, which points in the “nodes” array to the joint named “reference_l_knee_joint”. Hence, the semantics for the “avatar_l_lowerleg_joint” is specified as the “reference_l_knee_joint” joint of the reference skeleton, shown in FIG.7. {   "scene": 0,    "scenes": [      {       "nodes": [0,1,11]     }    ],      "meshes": [      {       "name": "user_avatar_mesh",        "primitives": [          {           "attributes":              {               "POSITION": 0,               "TEXCOORD_0": 1,               "JOINTS_0": 2,               "WEIGHTS_0": 3             },           "indices": 4         }        ]     }    ],      "skins": [      {       "inverseBindMatrices": 5,       "skeleton": 11,       "joints": [11, 12, 13, 14, 15]        ]     }    ],      "nodes": [      {       "name": "user_avatar_node",        "skin": 0,        "mesh": 0,        "extensions": {         "MPEG_node_avatar": {           "isAvatar": true,        "type": “urn:my_avatar_type”,        "mappings": [],        "skeleton_semantics": {          "reference_skeleton_representation": 1,          "reference_skeleton": [1,2,3,4,5,6,7,8,9,10],         "joint_semantics": [1,2,3,4,10]        }     }   } },  {    "name": "reference_humanoid_root",    "children": [2]  },  {    "name": "reference_l_hip_joint",    "translation": [0.11, ‐0.02, 0.0],    "rotation": [0.991, 0.0, 0.0, ‐0.131],    "children": [3]  },  {    "name": "reference_l_knee_joint",    "translation": [0.0, ‐0.46, 0.0],    "rotation": [0.793, 0.0, 0.0, 0.609],    "children": [4]  },  {    "name": "reference_l_talocrural_joint",    "translation": [‐0.03, ‐0.34, 0.0],    "rotation": [0.999, 0.0, 0.0, ‐0.044],    "children": [5]  },  {    "name": "reference_l_calcaneuscuboid_joint",    "translation": [0.0, ‐0.03, 0.0],    "rotation": [1.0, 0.0, 0.0, 0.0],    "children": [6]  },  {    "name": "reference_l_talocalcaneonavicular_joint",   "translation": [‐0.01, ‐0.02, 0.0],    "rotation": [0.991, 0.0, 0.0, 0.131],    "children": [7]  },  {    "name": "reference_l_cuneonavicular_1_joint",    "translation": [‐0.02, ‐0.01, 0.0],    "rotation": [0.924, 0.0, 0.0, ‐0.353],    "children": [8]  },  {    "name": "reference_l_tarsometatarsal_1_joint",    "translation": [0.0, ‐0.01, 0.0],        "rotation": [0.793, 0.0, 0.0, 0.609],        "children": [9]      },       {        "name": "reference_l_metatarsophalangeal_1_joint",        "translation": [0.01, ‐0.05, 0.0],        "rotation": [0.966, 0.0, 0.0, 0.259],        "children":

[0010] },      {        "name": "reference_l_tarsal_interphalangeal_joint",        "translation": [0.0, ‐0.03, 0.0],        "rotation": [0.966, 0.0, 0.0, ‐0.259]     },      {        "name": "avatar_pelvis_joint",        "children":

[0012] },      {        "name": "avatar_l_upperleg_joint",        "translation": [0.10, ‐0.08, 0.0],        "rotation": [0.924, 0.0, 0.0, ‐0.383],        "children":

[0013] },      {        "name": "avatar_l_lowerleg_joint",        "translation": [0.03, ‐0.37, 0.0],        "rotation": [0.954, 0.0, 0.0, ‐0.301],        "children":

[0014] },      {        "name": "avatar_l_ankle_joint",        "translation": [0.04, ‐0.38, 0.0],        "rotation": [0.996, 0.0, 0.0, 0.087],        "children":

[0015] },      {        "name": "avatar_l_toe_joint",        "translation": [0.05, ‐0.015, 0.0],        "rotation": [0.866, 0.0, 0.0, 0.500]     }    ],     "extensionsUsed": [        "MPEG_node_avatar”    ],      "extensionsRequired": [        "MPEG_node_avatar"    ],    "asset": {        "version”: ["2.0"]    } }

[0079] If “reference_skeleton_representation” were set to 0, the reference skeleton would be specified as a string URN pointing to the description of an avatar representation scheme. This description would include a naming of the reference skeleton joints and one or more diagrams showing their localization on the human skeleton. The scene description excerpt below illustrates how the avatar node would be described in this configuration. Theskeleton_semantics.reference_skeleton attribute specifies a URN pointing to a description of the reference skeleton, “urn:mpeg:sd:reference_avatar_representation_scheme” in the example. Nodes 1 through 11 in the previous scene description, where “reference_skeleton_representation” is set to 1, would be removed. “MPEG_node_avatar.skeleton_semantics.joint_semantics” is now an array of strings encoding the semantics of the avatar skeleton joints as the names of their corresponding joints in the reference skeleton.   "nodes": [      {       "name": "user_avatar_node",        "skin": 0,        "mesh": 0,        "extensions": {         "MPEG_node_avatar": {           "isAvatar": true,            "type": “urn:my_avatar_type”,            "mappings": [],            "skeleton_semantics": {              "reference_skeleton_representation": 0,              "reference_skeleton":  “urn:mpeg:sd:reference_avatar_representation_scheme”,              "joint_semantics": [“reference_humanoid_root”,  “reference_l_hip_joint”, “reference_l_knee_joint”,  “reference_l_talocrural_joint”, “reference_l_tarsal_interphalangeal_1_joint”]            }         }       }     },

[0080] Parsing of the proposed new “MPEG_node_avatar” attributes

[0081] The scene description containing the specification of the avatar representation model including the semantics of the skeleton joints is encoded in some virtual world platform or asset creation tool then transmitted to the application where the avatar is to be rendered. It should be understood that the specification of the semantics of the skeleton joints of the avatar at the encoder side is the decision of the person that creates the avatar model. Any choice of joint locations is acceptable as long as it allows a proper animation of the vertices of the avatar mesh under standard human motion. After the animation skeleton of the avatar has been designed, the joints can be manually mapped to the joints of a highly detailed reference skeleton by visual comparison. This mapping can then be encoded in the scene description according to an embodiment.

[0082] FIG.9 illustrates a method of parsing the new attributes of the “MPEG_node_avatar” extension, according to an embodiment. Specifically, a decoder parses (910) the MPEG-I SD “nodes” array. The decoder checks whether there exists a node with an “MPEG_node_avatar” extension (920). If not, the scene description is parsed (925) without the extension; otherwise, the decoder continues to check whether the “MPEG_node_avatar” extension has “skeleton_semantics” attribute (930). If not, the scene description is parsed (935) without the attribute; otherwise, the decoder continues to check whether the “skeleton_semantics” has “reference_skeleton_representation” attribute (940). If yes, the decoder parses the “reference_skeleton_representation” attribute (950). If the value of “reference_skeleton_representation” is 0 (960), the decoder parses "reference_skeleton" as a string URN and "joint_semantics" as an array of string joint names in the URN description (970). If the value of “reference_skeleton_representation” is 1 (965), the decoder parses "reference_skeleton" and "joint_semantics" as arrays of node indexes in the “nodes” array of the scene description (980). If the value of “reference_skeleton_representation” is neither 0 nor 1, the decoder flags (990) an invalid attribute value.

[0083] Processing Model

[0084] At runtime an application consumes the scene description that contains an avatar “node” instance including an “MPEG_node_avatar” extension.

[0085] To animate avatars, the application will typically rely on a specific avatar representation format that includes a skinning framework with an application-specific skeleton topology. The “skeleton topology” is defined as the mapping of its joints to physical locations in the human body, together with the connectivity of the joints. This application-specific skeleton will be referred to hereafter as the “application reference skeleton”.

[0086] The avatar described in the scene description that is fed to the application, hereafter referred to as the “input avatar”, may have been designed outside of the platform that hosts the application. Thus, the topology of the skeleton used to animate the avatar will in general differ from the application reference skeleton.

[0087] Provided information on the mappings of the joints of the input avatar skeleton to the joints of a highly-detailed reference skeleton is transmitted in the scene description, according to an embodiment, the positions of the joints at any frame of the animation of the avatar can be used as landmarks by the application that receives the scene description to locate parts of the avatar body relative to the elements of its surrounding virtual environment. Among other benefits, the relative locations of these landmarks to walls, doors or furniture items in the environment may be computed by the application and leveraged to avoid collisions, or the relative locations of the hand joints of an avatar hand to the locations of the hand joints of another avatar may be computed by the application and used to automatically infer social behavior between the two avatars.

[0088] The application may provide features such as foot contact detection or gesture recognition that rely for their implementation on the locations of skeleton joints in the 3D world in which the avatar is represented. For instance, foot contact detection may be implemented by applying a threshold to the distance between the position in the 3D world of joints located in the foot of the avatar and the ground. Such implementations depend on the topology of the avatar skeleton as well as on its morphology, defined by the lengths of the bones connecting adjacent joints. The topology of the skeletons in the avatar animation sequences processed in the application features is the topology of the application reference skeleton. For this topology, the implementation of a feature can be designed to accommodate a variety of character morphologies, for instance by relying for the considered task on a neural network trained on a corpus of skeletons with varying morphologies.

[0089] To be able to accommodate a variety of avatar representation models whose skeletons do not share the topology of the application reference skeleton, the application may provide, for each feature, a set of implementations, each adapted to a skeleton topology within a “topology set”. In this setting, the application reference skeleton is not unique but may be selected among the set of skeletons defined in the “topology set”. This “topology set” will typically consist of skeleton topologies that are frequently used for character animation in the industry, such as, for instance, topologies represented in popular character animation datasets.

[0090] For the purpose of matching the skeleton topologies in the “topology set” to the skeleton topology of the input avatar, the semantics of the joints defining these skeleton topologies is assumed to be defined with respect to the same highly detailed “reference skeleton” that is used in the description of the input avatar and specified according to an embodiment. Advantageously, in the considered use case the reference skeleton is specified in a normative annex of the scene description standard that can be accessed by any avatar asset creation tool and any immersive virtual world platform.

[0091] With these prerequisites fulfilled, the description of the semantics of the input avatar animation skeleton can be used by the application to implement features that depend on this semantics following the processing pipeline described in FIG.10.

[0092] In particular, the application parses (1010) the scene description of the input avatar or a normative annex of the scene description standard, to retrieve the joint names and the topology of the reference skeleton that defines the semantics of the skeleton joints. In step 1020, the application parses the semantics of the input avatar provided in the scene description it is fed with. This involves mapping each joint of the input avatar skeleton to a joint of the highly detailed reference skeleton obtained in step 1010. In step 1030, for each “application reference skeleton” in the above-defined “topology set”, the application matches the joints of the considered application reference skeleton to the joints of the highly detailed reference skeleton parsed in step 1010.

[0093] In step 1040, based on the outcome of steps 1020 and 1030, the application selects the application reference skeleton in the topology set that best matches the avatar skeleton. The matching criterion could rely, for instance, on the minimization on the average joint distance between the considered application reference skeleton joints and the input avatar skeleton joints, after mapping to the reference skeleton joints. These distances could be defined, for instance, on an instantiation of the reference skeleton topology with the morphology of an average human adult.

[0094] Optionally in step 1050, if the joints of the selected application reference skeleton and the joints of the avatar skeleton do not match exactly, the topology of the input avatar skeleton can be translated to the topology of the selected application reference skeleton by a process known as motion retargeting. Motion retargeting at large refers to the task of transforming an input temporal sequence of animation skeletons with time-varying poses to an output sequence of animation skeletons with a different topology (joint semantics and connectivity) and morphology (bone lengths) as the input skeletons, while preserving the time-varying poses hence the motion content of the input sequence. Several motion retargeting algorithms have already been proposed in the scientific literature. One of such algorithms is described in a paper by L. Mourot, L. Hoyet, F. Le Clerc and P. Hellier titled “HuMoT: Human Motion Representation using Topology-Agnostic Transformers for Character Animation Retargeting”, published on the Internet as an arXiv document with reference arXiv:2305.18897. It relies on a neural network to retarget an input sequence of skeletons with a given topology and morphology, representing the motion of a character, to an output sequence of skeletons with a possibly different topology and a different morphology, but where the input character motion has been preserved. The neural network consists of an encoder network followed by a decoder network. The encoder network is conditioned on a template input skeleton that has the topology and morphology of the input skeleton sequence, in a predefined reference pose. In the same way, the decoder network is conditioned on a template output skeleton that has the topology and morphology of the target output skeletons, in the same predefined reference pose as the template input skeleton. The network is trained on a corpus of skeleton sequences with a large variety of skeletal topologies and morphologies, so that it can perform retargeting from an input skeleton sequence with any topology and morphology to an output skeleton sequence with any other skeleton topology and morphology. Running inference on such a network, or using any equivalent algorithm for topology- and morphology-generic skeletal motion retargeting, the application could, at any frame of the input animation sequence of the input avatar specified in the scene description that it consumes, map the skeleton of the input avatar at the current animation pose to a skeleton in the same animation pose but with the topology and morphology of the selected application reference skeleton. The application may then rely on the 3D positions of the joints of the latter selected application reference skeleton to run algorithms implementing features requiring knowledge of skeleton joint semantics, since these algorithms were developed under the assumption that the animation of the avatar is driven by this application reference skeleton.

[0095] Finally in step 1060, the application runs implementations of its features depending on the semantics of the avatar skeleton joints that are based on the mappings of the joints of the input avatar to the joints of the selected application reference skeleton in the topology set obtained in steps 1040 and, optionally, step 1050.

[0096] Various numeric values are used in the present application. The specific values are for example purposes and the aspects described are not limited to these specific values.

[0097] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.

[0098] The implementations and aspects described herein may be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed may also be implemented in other forms (for example, an apparatus or program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, cell phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end-users.

[0099] Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment.

[0100] Additionally, this application may refer to “determining” various pieces of information. Determining the information may include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.

[0101] Further, this application may refer to “accessing” various pieces of information. Accessing the information may include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.

[0102] Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information may include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.

[0103] It is to be appreciated that the use of any of the following “ / ”, “and / or”, and “at least one of”, for example, in the cases of “A / B”, “A and / or B” and least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.

[0104] As will be evident to one of ordinary skill in the art, implementations may produce a variety of signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal may be formatted to carry the bitstream of a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.

Claims

CLAIMS 1. A method, comprising: obtaining a reference skeleton model; obtaining a skeleton model for an avatar; mapping each joint of said skeleton model for said avatar to a corresponding joint in said reference skeleton model; and encoding information indicating said mapping for said avatar in scene description.

2. An apparatus, comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to: obtain a reference skeleton model; obtain a skeleton model for an avatar; map each joint of said skeleton model for said avatar to a corresponding joint in said reference skeleton model; and encode information indicating said mapping for said avatar in scene description.

3. A method, comprising: obtaining a reference skeleton model; decoding, from scene description, information indicating mapping of each joint of a skeleton model for an avatar to a corresponding joint in said reference skeleton model; and matching each joint of said avatar to said corresponding joint in said reference skeleton model based on said mapping.

4. An apparatus, comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to: obtain a reference skeleton model; decode, from scene description, information indicating mapping of each joint of a skeleton model for an avatar to a corresponding joint in said reference skeleton model; and match each joint of said avatar to said corresponding joint in said reference skeleton model based on said mapping.

5. The method of any one of claim 1 or 3, or the apparatus of claim 2 or 4, wherein said mapping is indicated by a name of said corresponding joint in said reference skeleton model.

6. The method of claim 5, or the apparatus of claim 5, wherein said reference skeleton model is indicated by a uniform resource name that points to a description of said reference skeleton model.

7. The method of claim 6, or the apparatus of claim 6, wherein said uniform resource name is encoded in said scene description.

8. The method of claim 6, or the apparatus of claim 6, wherein said uniform resource name and said reference skeleton model are obtained without said uniform resource name being indicated in said scene description.

9. The method of any one of claims 1, 3 and 5-8, or the apparatus of any one of claims 2 and 4-8, wherein said mapping is indicated by a node index of said corresponding joint in said reference skeleton model.

10. The method of any one of claims 6-9, further comprising, or the apparatus of any one of claims 6-9, wherein said one or more processors are further configured to perform: obtaining an array of indexes for said avatar, where the ithelement of said array of indexes contains an index of a corresponding joint in said reference skeleton model for the ithjoint in said skeleton model for said avatar, wherein said array of indexes is included in said scene description.

11. The method of any one of claims 6-10, or the apparatus of any one of claims 6- 10, wherein said reference skeleton model is represented by a graph of nodes representing a hierarchy of joints of said reference skeleton model.

12. The method of any one of claims 1, 3 and 5-11, or the apparatus of any one of claims 2 and 4-11, wherein said reference skeleton model is a highly-detailed skeleton that includes most or all of joints and bones of human anatomy.

13. The method of any one of claims 3 and 5-12, further comprising, or the apparatus of any one of claims 4-12, wherein said one or more processors are configured to perform: implementing a feature based on positions of skeleton joints driving animation of an avatar with a pre-defined avatar skeleton topology.

14. The method of claim 13, further comprising, or the apparatus of claim 13, wherein said one or more processors are further configured to perform: selecting an application reference skeleton for said skeleton model of said avatar.

15. The method of claim 14, further comprising, or the apparatus of claim 14, wherein said one or more processors are further configured to perform: obtaining a plurality of application reference skeletons defined with respect to said reference skeleton model, wherein an application reference skeleton that best matches said skeleton model of said avatar is selected.

16. The method of claim 14 or 15, further comprising, or the apparatus of claim 14 or 15, wherein said one or more processors are further configured to perform: retargeting an animation sequence of instances of said skeleton model of said avatar to translate topology of said skeleton model of said avatar to topology of said selected application reference skeleton.

17. The method of any one of claims 14-16, further comprising, or the apparatus of any one of claims 14-16, wherein said one or more processors are further configured to perform: obtaining a plurality of topologies, wherein said selected application reference skeleton corresponds to a topology of said plurality of topologies.

18. The method of any one of claims 14-17, further comprising, or the apparatus of any one of claims 14-17, wherein said one or more processors are further configured to perform: obtaining a distance between joints of a respective application reference skeleton and corresponding joints of said skeleton model of said avatar, after mapping to said reference skeleton model, wherein selection of said selected application reference skeleton is based on said distance.

19. A non-transitory computer readable medium comprising instructions which, when the instructions are executed by a computer, cause the computer to perform the method of any one of claims 1, 3 and 5-18.

Citation Information

Patent Citations

  • System and method for animating an avatar in a virtual world

    US20230410398A1

  • EP24305020A

Cited By

  • Volumetric capture through internet-of-things-based cameras

    US12563169B1