Avatar controllers in scene descriptions

ZA202608710APending Publication Date: 2026-09-30INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
ZA202608710
Authority / Receiving Office
ZA · ZA
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-27
Filing Date
2026-09-02
Publication Date
2026-09-30

AI Technical Summary

Technical Problem

The current MPEG-I Scene Description format does not clearly define how animations can be combined, particularly for modifying or animating the appearance of avatars, and lacks standardization in defining the meaning of animation names, making it difficult to combine multiple animations or understand their properties.

Method used

The proposed solution introduces an MPEG_controllers extension that defines controllers as high-level controllable animations, associating semantic tags with syntax structures to clearly identify their purpose and functionality, allowing for the combination and manipulation of avatar animations using controller sets and sets of controllers.

Benefits of technology

This approach enables clear and standardized manipulation of avatar animations, facilitating the exchange of avatar instances across heterogeneous environments and ensuring that applications understand the intended actions of controllers, thereby enhancing the flexibility and interoperability of avatar models in virtual scenes.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

NOT VISIBLE DUE TO STATUS OF PATENT
Need to check novelty before this filing date? Find Prior Art

Description

AVATAR CONTROLLERS IN SCENE DESCRIPTIONSCROSS-REFERENCE

[0001] This application claims the benefit of European Patent Application No. 24305458.2, filed 27 March 2024, the entire disclosure of which is incorporated herein by reference.BACKGROUND

[0002] The present disclosure relates to encoding for avatar 3D models in scene descriptions.

[0003] Various technologies are available for generating, processing, and rendering virtual three-dimensional (3D) scenes. The information characterizing a 3D scene, referred to as a scene description, can be time-dependent, allowing a 3D scene to change in a manner analogous to playback of a video. This kind of behavior can be achieved by relying on the framework defined in the Scene Description for MPEG media document, Information technology - Coded representation of immersive media - Part 14: Scene Description for MPEG media, ISO / IEC DIS 23090-14 :2021 (E). A scene update mechanism based on the JSON Patch protocol as defined in IETF RFC 6902 may be used to synchronize virtual content to MPEG media streams.

[0004] The current MPEG-I Scene Description (SD) format allows the storage of animations in “animations” assets in the gITF file. Each animation defines a named function that can modify a mesh given a time between zero and a maximum. A common use of these animations is to provide different actions for a mesh, for instance a character with walk, run and idle animations. Another is to provide animations that lead to different cinematics (e.g., like little movies, but rendered in real time).

[0005] The current animation encoding allows functions that modify one or more properties of a mesh, like the translation, rotation, or mesh target weights. More properties can be changed using extensions like KHR_animation_pointers. Each animation can have a name, but there is no standard to define the meaning of these names.SUMMARY

[0006] A method according to some embodiments comprises obtaining a runtime asset delivery file for a 3D scene, the file including: a plurality of first syntax structures, each first syntax structure defining an association between an input weight value and a corresponding transformation of at least one of a plurality of nodes in the scene; and at least one secondsyntax structure associating a respective semantic tag with at least one of the first syntax structures. The method further includes selecting at least one of the first syntax structures based on the associated semantic tag; and applying the transformation indicated by the selected first syntax structure.

[0007] An apparatus according to some embodiments comprises one or more processors, the apparatus being configured to perform at least: obtaining a runtime asset delivery file for a 3D scene, the file including a plurality of first syntax structures, each first syntax structure defining an association between an input weight value and a corresponding transformation of at least one of a plurality of nodes in the scene; and at least one second syntax structure associating a respective semantic tag with at least one of the first syntax structures. The apparatus is further configured to select at least one of the first syntax structures based on the associated semantic tag; and to apply the transformation indicated by the selected first syntax structure.

[0008] In some embodiments, the first syntax structures are controller structures in an MPEG_controllers extension.

[0009] In some embodiments, the file further includes a third syntax structure associated with a node in the scene, wherein the second syntax structure is included in the third syntax structure, and wherein the transformation corresponding to the first syntax structure is a transformation of a mesh in the node.

[0010] In some embodiments, the third syntax structure includes information indicating that the node represents an avatar. In some embodiments, the third syntax structure is an MPEG_node_avatar extension.

[0011] In some embodiments, the third syntax structure is an MPEG_node_controllers extension.

[0012] In some embodiments, the file further includes a fourth syntax structure associating a respective semantic tag with a set comprising a plurality of the first syntax structures.

[0013] In some embodiments, each of the first syntax structures has a respective index, and wherein the second syntax structure includes the index of at least one of the first syntax structures.

[0014] In some embodiments, the runtime asset delivery file is a gITF file.

[0015] In some embodiments, the semantic tag comprises a purpose property.

[0016] Some embodiments further include animating at least the node associated with the selected first syntax structure by applying changing weights to the transformation corresponding to the selected first syntax structure.

[0017] Some embodiments further include causing display of at least the node associated with the selected first syntax structure.

[0018] Some embodiments further include: obtaining a plurality of runtime asset delivery files, each of the runtime asset delivery files including: a respective node representing an avatar, at least one first syntax structure defining an association between an input weight value and a corresponding transformation of the respective avatar, and at least one second syntax structure associating a respective semantic tag with the first syntax structure; presenting a virtual environment including the avatars of the respective runtime asset delivery files; selecting an action to perform on a selected one of the avatars; using the semantic tags of the second syntax structures associated with the selected avatar, identifying a first syntax structure corresponding to the selected action; and applying the transformation indicated by the identified first syntax structure.

[0019] A computer-readable medium according to some embodiments stores a runtime asset delivery file for a 3D scene, the file including: a plurality of first syntax structures, each first syntax structure defining an association between an input weight value and a corresponding transformation of at least one of a plurality of nodes in the scene; and at least one second syntax structure associating a respective semantic tag with at least one of the first syntax structures.

[0020] A method according to some embodiments comprises encoding a runtime asset delivery file according to any of the embodiments described herein.

[0021] An apparatus according to some embodiments comprises one or more processors configured to encode, decode, and or parse a runtime asset delivery file according to any of the embodiments described herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0022] FIGs. 1A-1C illustrate an avatar in which a controller is used to move the avatar’s eyes from gazing to the avatar’s right (FIG. 1A), to gazing forward (FIG. 1 B), to gazing to the left (FIG. 1C).

[0023] FIGs. 2A-2C illustrate an avatar in which a controller is used to move the avatar’s eyes from gazing down (FIG. 2A), to gazing forward (FIG. 2B), to gazing up (FIG. 2C).

[0024] FIG. 3 is a flow chart illustrating a method of parsing a node according to an example embodiment.

[0025] FIG. 4 is a flow chart illustrating a method of parsing an avatar controller set according to an example embodiment.

[0026] FIG. 5 is a flow chart illustrating a method of parsing an avatar controller according to an example embodiment.

[0027] FIG. 6 is a schematic illustration of the logical structure of relevant syntax elements in a runtime asset delivery file.

[0028] FIG. 7 is a schematic illustration of the logical structure of relevant syntax elements in a runtime asset delivery file including semantic tags associated with controllers according to embodiments described herein.

[0029] FIG. 9 is a flow chart illustrating a method making use of several runtime asset delivery files according to an example embodiment.

[0030] FIG. 10 is a functional block diagram illustrating an example of an apparatus that may be used for encoding, decoding, and / or processing a scene description file according to example embodiments.DETAILED DESCRIPTION

[0031] The current scene description format does not define clearly how animations can be combined, particularly for modification or animation of the appearance of an avatar. For instance, if one wishes to use two animations of an avatar at a time, the current scene description format does not clearly define how the properties they change should be updated, starting with the order of animations.

[0032] This disclosure describes apparatus and methods for encoding an avatar with controllers, and how to use them. Controllers may be described as high-level controllable animations, e.g. functions that modify the avatar model given one or more weights.

[0033] Example embodiments may be implemented in an environment that makes use of the current MPEG-I Scene Description (SD) format together with an “MPEG_controllers” extension as described in European Patent Application No. EP 23306167.0, filed 10 July 2023. Example embodiments may further make use of an “MPEG_node_avatar” extension as described in European Patent Application No. EP 23305405.5, filed 24 March 2023.

[0034] In a conventional gITF system, an application can use an avatar with a controller contained in a gITF file if someone provides extra information. For instance, an application can load a gITF file with an avatar and controllers, and let the user define which controllers to use, and to what purpose. For instance, if a controller makes the avatar walk, the user must know or guess it.

[0035] One goal of avatar encoding in scene description is to allow for the exchange of avatar instance in heterogeneous environments. For example, one may wish to have the ability tobring one’s avatar into a virtual world and walk around. A user could want to change the expression or identity of its face. These features can be implemented using controllers. However, the application must know which controllers are related to what movement or expression, and what weights to use.

[0036] In the current description, the proposed format follows the gITF format and is compatible with the recent MPEG-I SD effort to extend gITF with MPEG-I SD extensions. Although example embodiments are described herein using gITF, it should be understood that other embodiments based on the same principles may be implemented using other formats (e.g. XML, USD, and the like).Overview of Avatar Signaling

[0037] In the current MPEG format, an avatar is not defined and can take any form, e.g., humanoid, and non-humanoid versions, or it can take the shape of any 3D object that has the propriety “isAvatar” set to “True”. Therefore, guiding applications to what to expect from the “MPEG_node_avatar” is helpful. In some embodiments it may be desirable for an avatar to respect some or all of the following features:• An avatar is represented by the node “MPEG_node_avatar,” and character representations outside this scope may not be qualified as an “avatar” under the MPEG-I SD standards.• The reconstruction and animation respect the supported primitives in the scene description for MPEG and / or other standard and animation formats.• An avatar may or not be explicitly represented by an avatar type property, which may be accompanied by the respective URI to the 3D asset.• If an avatar sets the avatar type property, the default may be an “MPEG_reference_avatar.” In some embodiments, the type of the avatar representation is provided as a URN that uniquely identifies the avatar representation scheme. The avatar representation scheme may define the format of all components that are used to reconstruct and animate the avatar.

[0038] In some embodiments, the avatar type is reconstructed, given some user inputs and retargeted onto a reference 3D avatar model, which may contain, for example a full bodybased mesh representation, skeleton and skinning weights, facial blend shapes and landmarks, and additional geometry including, eyes, jaws, teeth and tongue. Such embodiments allow a shared geometry and features representation for any humanoid character. Appropriate deformation allows for adaptation of the shape. This common base representation facilitates animation and representation of different morphologies or contexts.

[0039] In some embodiments, an extension such as an “MPEG_node_avatar” extension can be used to identify the avatar used. Given such an extension property, the client will know what avatar to reconstruct / render.

[0040] In some embodiments, a property such as isAvatar may be used with syntax and semantics such as the following:

[0041] In some embodiments, an MPEG_node_avatar extension may have a syntax and semantics as follows:

[0042] In some embodiments, the Mapping object may have the following syntax and semantics:Overview of Controllers and Animations

[0043] In gITF, animations are conventionally structured in JSON (JavaScript Object Notation) format as follows. An “animations” array includes one or more “animation” objects. Each “animation” object includes information identifying the component to be animated, called the “target,” and information identifying the corresponding data used to perform the animation, called a “sampler.” This connection between the target of the animation and the sampler is made in a “channel” object. The channel object further includes information indicating how the data affects the target of the animation. For example, where the target of the animation is amesh, the channel object may indicate whether the animation data is used to translate the mesh, to rotate the mesh, or to change the displacement of certain vertices of the mesh using “morph target weights” that have been defined for the mesh.

[0044] The “sampler” object that provides the data for the animation has “input” and “output” properties. Data associated with the “output” property indicates the transformation to be applied to the component of the scene. For example, where the transformation is a translation, the “output” data may provide the components of different vectors representing different displacements. Where the transformation is a rotation, the “output” data may provide the components of different rotation matrices. Where the transformation is applied to a morph target, the “output” data may provide different weights. The “input” property is used to help determine which of these output vectors, matrices, or weights is to be applied to the relevant component of the scene. The input property may correspond to a list of times. As an example, a sampler may identify an input with five different time values and an output with five different rotation matrices. Heuristically, these may be thought of as the columns of a table, in which a system performing an animation looks up the current time (e.g. the time since the beginning of the animation) in the “input” column and finds the appropriate rotation matrix in the “output” column. The times that are explicitly listed in the “input” column are referred to as the times of animation key frames. Interpolation is used when the current time is not one of the times listed explicitly in the “input,” with the type of interpolation (e.g. “step,” “linear,” or “cubicspline”) being indicated in the sampler object.

[0045] (To be more specific, the attributes of the “input” and “output” properties of a “sampler” object are not necessarily the relevant input or output data itself; instead, the attributes may be indices of objects called “accessors” that are used to access the relevant data, possibly through one or more additional layers of abstraction. Such details, however, are not necessary for the understanding of the principles described herein.)

[0046] A “controller” may be described as an object that is analogous to an animation but that transforms mesh properties given an input weight instead of an input time. For example, a controller object may include one or more sampler objects with respective inputs and outputs, and one or more channel objects that associate these samplers with corresponding mesh properties. Given an input weight, each sampler provides a corresponding transformation as its output, and that transformation may be applied to the appropriate mesh property identified by the channel objects. Controller objects provide greater flexibility than animation objects, at least in part because controller weights are not inherently functions of time (though controller weights can be animated as a function of time). Thus, the details of a controller-mediated transformation can be modified independently of the timeline of the scene animation as a whole.

[0047] The current MPEG_Controllers extension has the following features:• Controllers: each controller defines a function that changes mesh properties given a weight between a minimum and maximum value.• Controller weights: each value defines the weight of each controller.

[0048] Example embodiments may optionally include a component that defines a user interface for controllers that may be used to change one or more controller weights.

[0049] The "MPEG_controllers" property may be described as shown in Table 1.Table 1.

[0050] Table 1 describes features of the MPEG_controllers property, including a “ui” property. Note that in this and all other tables in the present disclosure, the column “required” (or an M for “mandatory”) indicates whether a feature is required or mandatory under the particular syntax illustrated in the table. An indication that a particular feature is “required” or “mandatory” for that particular syntax does not mean that the feature is a required feature of all embodiments. There may be other embodiments in which such features are not required.

[0051] In example embodiments, the number of items in "controllers" and "weights" is constrained to be the same.

[0052] In example embodiments, each controller in the "controllers" list of MPEG_controllers defines a function that changes one or more properties of a mesh. For example, in some embodiments, a controller changes the positions of the mesh vertices and thus models a deformation of the mesh.

[0053] The shape of the function associated to the controller is defined by a set of (input weight, output property change) values. The encoding of these values may be performed for compatibility with the keyframing scheme used in the corresponding animations (e.g. gITF animations). The keyframe timings pointed to by the samplers input attributes are mapped to the function input weights. The animated values pointed to by the channels target attributes are mapped to the output property change values. Thus, the animation framework (e.g. the gITF animation framework) is leveraged to shape up each controller as a possibly non-linear function of input weight control values.

[0054] In some embodiments, a controller is defined in a runtime asset delivery file. A controller is defined by an instance of the “controller” property. The controller property may include one or more of the attributes. The column “required” indicates attributes that are required to be included in at least some implementations of the controller property, though the same attributes may not be required in other implementations.Table 2.

[0055] In some embodiments, the Controller property has the same attributes as the standard Animation property of gITF, plus two additional ones: "animation" and "min".

[0056] In some embodiments, there are two modes for this property that depend on the definition of the "animation" property. If this one does not exist, then the controller directly defines the function, otherwise, it uses the one of an existing animation. In some implementations, the definition of a controller is constrained to be in one of the two modes, with other combinations leading to an error. In all modes, the "name" property can define the name of the controller. If it is not defined, the application may choose a name based on the index of the controller, like "controller_23" or "controller_0023"

[0057] In some embodiments, the "animation" property is not defined in the controller, in which case the controller directly defines its function. In this case, the "min" property is ignored. All the other properties ("channels", "samplers", "name", "extensions", "extras") may be usedaccording to those of a standard animation property. A usual animation parser can be run, except that time is considered as input weight values, and its minimum value can be negative.

[0058] In other embodiments, the "animation" property is defined in the controller. In such embodiments, the controller may not define its function but uses the one of an animation. In this case, the following properties may be ignored: "channels", "samplers". The properties “name”, “extensions”, and “extras”, however, may still be used when the animation property is defined.

[0059] The "animation" property defines the index of an animation in the “animations” asset. In this case, the function defined by the animation is used as a controller, except that time is considered as input weight values, and its minimum is shifted by the value of the "min" property. For instance, if "min" is -1.3, then animation time is decreased by -1.3, as if the animation was starting at -1.3 seconds. This property allows the definition of controllers with negative input weight values while referencing a gITF animation for which time cannot be negative.

[0060] In some embodiments, the "weights" attribute of MPEG_controllers contains an array of numbers (e.g. floating values) that define the weight associated to each controller. The weights may be provided in the same order as the order of the controllers in the "controllers" list: weights[0] is for controllers[0], weights[1] for controllers^], and so on.

[0061] These weights provide information regarding how the nodes targeted by the controllers change. In some embodiments, each weight is used to apply its corresponding controller, proceeding iteratively from the first one to the last one. For example, the function of controllers[0] may first be applied with weights[0] to update the nodes. Then, these updated nodes are modified by the function of controllers^] with weights[1], and so on. One can express these updates in the following way: sceneout= fn-i(wn-,fn-2(wn-2, ... f0(w0, scenein)')>) where ft is the function of controllers^], wtis weights[i], sceneinis the initial nodes and sceneoutis the updated nodes using the controllers and the weights.

[0062] In some embodiments, the controller changes happen at the same step as animations: they are applied on the initial nodes content, and before any transform in the other gITF properties (like the one in the “nodes” assets).

[0063] In some embodiments, controllers are not activated when gITF parsing is done; like animations, the application chooses whether or not to enable the controllers.

[0064] In some embodiments, controllers are not used with animations; the application may choose whether it enables controllers or one of the animations.

[0065] In some embodiments, controllers are animated using an extension such as KHR_animation_pointer, in which case an animation is defined with this extension to animate the controller weights (analogous to the use of morph target weights).Signaling with Controller Sets

[0066] In an example embodiment, a new "controllersets" list is added to the MPEG_node_avatar extension. An MPEG_node_avatar extension according to such embodiments may have the following syntax and semantics:Table 3 - The MPEG_node_avatar property with the new "controllersets" list.

[0067] In this and other tables used herein, an indication that a parameter is “required” indicates only that it is required by the syntax of the particular illustrated embodiment; there may be other embodiments in which the syntax does not require the parameter.

[0068] Each item in the “controllersets” list defines a controller set for the avatar and provides one or more semantic tags indicating how to use the controller set. An avatar controller set regroups several avatar controllers for a given purpose, for instance an avatar controller set can be dedicated to avatar movements (walk, run, idle, and the like), and another avatar controller set can be dedicated to facial expressions. Syntax elements in the scene description file may be used to associate semantic tags with different controllers and / or sets of controllers.

[0069] In some embodiments, avatar controller sets are replaced by a list of AvatarController (see below). These embodiments provide fewer features since there is a single set of avatar controllers, but they may use less memory.

[0070] In an example embodiment, an AvatarControllerSet property is defined as shown in Table 4.

[0071] In an example embodiment, the “description” property describes the avatar controller set and its usage. This is only illustrative, for example it can be used in a user interface to help the user understand what the controller set was intended for.

[0072] In an example embodiment, the “purpose” property contains a string that identifies the purpose of the avatar controller set. In some embodiments, the purpose property is constrained to follow predetermined values, defined by an application or a standard. In some embodiments, the type of this property is an enumeration (or “enum”) variable type. The “purpose” property and / or the “description” property may be used as semantic tags that provide information to allow a system to associate a controller set with a desired operation.

[0073] In an example embodiment, the “controllers” list defines the controllers to use and indicates how to use them, as described in greater detail below.

[0074] In an example embodiment, the “weights” property defines the default controller weights to use. The length of this array may be constrained to be equal to the number of items in “controllers”. These controller weights may replace the corresponding ones in the “weights” property of “MPEG_controllers”.

[0075] In an example embodiment, an AvatarController property may have the syntax and semantics as shown in Table 5.Table 5 - Example AvatarController property

[0076] In an example embodiment, the “description” property describes the avatar controller and its usage. This is only illustrative, for example it can be used in a user interface to help the user understand what the controller was intended for. The description may include human- readable text that is not necessarily constrained to a predetermined value.

[0077] In an example embodiment, the “purpose” property contains a string that identifies the purpose of the avatar controller. It may be constrained to follow predetermined values, defined by an application or a standard. In some embodiments, the type of this property is an enumeration (or “enum”) variable type. The “purpose” property and / or the “description” property may be used as semantic tags that provide information to allow a system to associate a controller with a desired operation. In a case where no “purpose” property is present for a particular controller, that controller may be assigned the “purpose” property of the corresponding controller set to which the controller belongs. Thus, semantic tags associated with a controller set may serve as default semantic tags for controllers within the set.

[0078] In an example embodiment, the “controller” property references a controller in the “controllers” list of the “MPEG_controllers” extension.

[0079] In an example embodiment, the “range” property defines the minimum and maximum values of the controller weight.

[0080] In some embodiments, controller set information as described above may be provided for non-avatar nodes instead of or in addition to providing such information for avatar nodes. As an example, controller sets may be added to gITF nodes by introducing a new “MPEG_node_controllers” extension. This allows the addition of controller sets to any gITF node. For the sake of clarity, such embodiments are described herein following the gITF format and with compatibility with the use of MPEG-I SD extensions. However, it should be understood that the principles described herein are not limited to the use of gITF and may be implemented using other coding formats (e.g. XML, USD, and the like).

[0081] In some embodiments, controller sets for nodes (which may or may not be avatar nodes) in a scene description file may be implemented using an "MPEG_node_controllers” extension. In an example embodiment, such an extension can be used to extend any gITF node. This extension may be used to define controller sets that an application can use to animate a mesh in the node. In example embodiments, the data in these controller sets includes one or more semantic tags that define the purpose of the controllers. Such semantic tags may be read by an application to identify which controller to use to implement a desired function.

[0082] In an example embodiment, the controller set is implemented as an “MPEG_node_controllers” extension having the following properties:Table 6 - Properties of the MPEG_node_controllers extension.

[0083] The “description” property describes the controller sets and their usage. This is only illustrative; for example, the description property can be used in a user interface to help the user understand the intended function of the controller set.

[0084] The “purpose” property contains a string that identifies the purpose of the controller sets. In some embodiments, the type of this property is an enumeration (or “enum”) variable type. The “purpose” property and / or the “description” property may be used as semantic tags that provide information to allow a system to associate a controller with a desired operation. In a case where no “purpose” property is present for a particular controller, that controller may be assigned the “purpose” property of the corresponding controller set to which the controller belongs. Thus, semantic tags associated with a controller set may serve as default semantic tags for controllers within the set.

[0085] Each item in the “controllersets” list defines a controller set for the node and how to use it. A controller set provides a group of one or more controllers for a given purpose. A controller may belong to more than one controller set. (And some controllers may not belong to any controller set.)

[0086] In some embodiments, controller sets are replaced by a list of “controller” properties as described below. Such embodiments may offer fewer features since there is no grouping of the controllers (other than their inclusion in the list), but such embodiments may use less memory.

[0087] In an example embodiment, the ControllerSet property is defined as follows:

[0088] The “description” property describes the controller set and its usage. This is only illustrative; for example it can be used in a user interface to help the user understand an intended function implemented by the controller set.

[0089] The “purpose” property in this embodiment contains a string that identifies the purpose of the controller set. It must follow predetermined values, defined by an application or a standard. In some embodiments, the type of this property is an enumeration (or “enum”) variable type. The “purpose” property and / or the “description” property may be used as semantic tags that provide information to allow a system to associate a controller with a desired operation. In a case where no “purpose” property is present for a particular controller, that controller may be assigned the “purpose” property of the corresponding controller set to which the controller belongs. Thus, semantic tags associated with a controller set may serve as default semantic tags for controllers within the set.

[0090] The “controllers” list in this embodiment defines the controllers to use and how to use them.

[0091] The “weights” property in this embodiment provides the default controller weights to use. The length of this array may be constrained to be equal to the number of items in “controllers”. These controller weights replace the corresponding ones in the “weights” property of “MPEG_controllers”.

[0092] In an example embodiment, the Controller property is defined as follows:Table 8 - The Controller property

[0093] The “description” property describes the controller and its usage. This is only illustrative, for example it can be used in a user interface to help the user understand what the controller was intended for.

[0094] The “purpose” property contains a string that identifies the purpose of the controller. In some embodiments, the purpose property is constrained to follow predetermined values, defined by an application or a standard. In some embodiments, the type of this property is an enumeration (or “enum”) variable type. The “purpose” property and / or the “description” property may be used as semantic tags that provide information to allow a system to associate a controller set with a desired operation.

[0095] The “controller” property in this embodiment references a controller in the “controllers” list of the “MPEG_controllers” extension.

[0096] The “range” property in this embodiment defines the minimum and maximum value the controller weight can get.

[0097] Examples of gITF schemas according to embodiments disclosed herein are provided below. Forthe sake of clarity, these examples only show parts of the gITF file related to avatars and controllers. In general, the file also includes information regarding a scene with one or more meshes representing an avatar.

[0098] In an example, it may be desirable to control the eyes of an avatar using two controllers. The first controller (index 0) controls the horizontal eye movement, and the second controller (index 1) controls the vertical eye movement. The controllers change the eyes rotations and the weights of two morph targets corresponding to eyelid movements, as shown in FIGs. 1A-1C and FIGs. 2A-2C.

[0099] The gITF file is configured with the following:• A head mesh (index 0) with four morph targets:o AU61_Eyes_turn_L: this morph target defines an animation for eyelids with a movement to the left. o AU61_Eyes_turn_R: this morph target defines an animation for eyelids with a movement to the right. o AU63_eyeUp: this morph target defines an animation for eyelids with a movement to the top. o AU63_eyeDown: this morph target defines an animation for eyelids with a movement to the bottom.• An eye mesh (index 1) for the left eye, centered around the origin.• An eye mesh (index 2) for the right eye, centered around the origin.

[0100] In this example, encoding is provided for these controllers (using “MPEG_controllers”) and avatar controllers that refer to them and give them a description and a purpose.

[0101] In an example embodiment, the controllers and the corresponding avatar controllers are encoded in the following way:{"extensions" : {"MPEG_controllers" : {"controllers" : [{" channels" : [{" sampler" : 0 ," target" : { "node" : 1 , "path" : "weights"}} , {" sampler" : 1 ," target" : { "node" : 2 , "path" : "rotation"}} , {" sampler" : 1 ," target" : { "node" : 3 , "path" : "rotation"}}" samplers" : [{" input" : 10 ," interpolation" : " CUBICSPLINE" ,"output" : 11 }, {"input" : 10,"interpolation" : "LINEAR", "output" : 12 }}, {"channels" : [{"sampler" : 0,"target" : { "node" : 1 , "path" : "weights" } }, {"sampler" : 1,"target" : { "node" : 2 , "path" : "rotation" } }, {"sampler" : 1,"target" : { "node" : 3 , "path" : "rotation" } }"samplers" : [{"input" : 14,"interpolation" : "CUBICSPLINE" , "output" : 15 }, {"input" : 14,"interpolation" : "LINEAR", "output" : 16 }}"weights" : [1.0, -1.0]}},"nodes" : [{"children" : [1, 2, 3] ,"extensions" : {"MPEG_node_avatar" : { "isAvatar" : true, "type" : "urn:mpeg: sd: 2023 : avatar" , "mappings" : [ {"path" : "full_body / upper_body / head" , "node" : 1}, {"path" : "full_body / upper_body / head / face / eye_right" , "node" : 2}, {"path" : "full_body / upper_body / head / face / eye_left" , "node" : 3}"controllersets" : [{"description" : "Update gaze", "purpose" : "gaze", "controllers" : [ {"description" : "Horizontal gaze" , "purpose" : "gaze : horizontal" , "controller" : 0}, { "description" : "Vertical gaze", "purpose" : "gaze : vertical" , "controller" : 1}}]}}}, {"mesh" : 0}, {"mesh" : 1 ,"translation" : [-2.53, 150.54, 9.70]}, {"mesh" : 2 , "translation" : [2.78, 150.54, 9.76] }}

[0102] Example embodiments use one or more semantic tags associated with a controller and / or set of controllers to identify an operation that may be performed on a node in a scene. Application software that is rendering and / or animating a scene may determine a function to be performed and may identify, based on the semantic tags, one or more controllers to be used to perform the function.

[0103] As for semantic tags applied to a set of controllers, in some embodiments, each semantic tag may be provided using a “purpose” property, which may be a property of AvatarControllerSet as described above. In example embodiments, the values of the purpose property are constrained to be values selected from a predetermined set of values defined by an application or a standard. The following table presents examples of such values:Table 9.

[0104] As for semantic tags applied to individual controllers (which may or may not be in a set of controllers), in some embodiments, a semantic tag is provided using the “purpose” property of AvatarController to define how an application can use an avatar controller. In example embodiments, the values of the purpose property are constrained to be values selected from a predetermined set of values defined by an application or a standard. The following table presents examples of such values:Table 10.

[0105] A scene description file according to embodiments disclosed herein may be parsed as follows. In an example embodiment, the gITF parser decodes all content (including “MPEG_controllers”) except the one from “MPEG_node_avatar” extension. Subsequently, it decodes the “MPEG_node_avatar” content found in the first node. It indicates that the node is an avatar ("isAvatar" is true), the type is the standard one ("type" is "urn:mpeg:sd:2023:avatar"), there is one part that corresponds to the head ("mappings" property).

[0106] The “controllersets” property of “MPEG_node_avatar” is parsed. In this example, there is a single AvatarControllerSet item. In this example, the avatar controller set can be described with the message “Update gaze” (“description” is “Update gaze”). As indicated by the encoded semantic tag, the purpose of this avatar controller set is to change the gaze of the avatar (“purpose” is “gaze”). The string “gaze” is an example of a value for the “purpose” property. The syntax and semantics of such a value may be defined by an application or a standard.

[0107] The parsing continues with the two items of the “controllersets” list.

[0108] The first item of the “controllersets” list is an avatar controller and can be described with the string “Horizontal gaze” (“description” is “Horizontal gaze”). The purpose of this avatar controller is to change the avatar gaze along the horizontal axis (“purpose” in “gaze:horizontal”). The “controller” property references the first controller of the “MPEG_controllers” extension. As a result, an application can use the first controller of the “MPEG_controllers” extension to change the avatar gaze along the horizontal axis.

[0109] The second item of the “controllersets” list is an avatar controller can be described with the string “Vertical gaze” (“description” is “Vertical gaze”). The purpose of this avatar controller is to change the avatar gaze along the vertical axis (“purpose” in “gaze:vertical”). The “controller” property references the second controller of the “MPEG_controllers” extension. As a result, an application can use the second controller of the “MPEG_controllers” extension to change the avatar gaze along the vertical axis.

[0110] Thus, using the avatar controllers’ data, an application can change the gaze of the avatar right / left and up / down.

[0111] A method for parsing an MPEG_node_avatar extension may proceed as shown in the flow chart of FIG. 3.

[0112] Parsing of a node in a gITF file begins at 501. At 502, the gITF loader decodes the standard properties and possible extensions. If it is determined at 503 that an MPEG_node_avatar extension is present, the loader parses it; otherwise, parsing ends. At 504, the "isAvatar" property is parsed. In embodiments where this property is required, an error may be signaled if it is absent. At 505, the "type" property is parsed. In embodiments where this property is required, an error may be signaled if it is absent. At 506, the "mappings" property is parsed. In embodiments where this property is required, an error may be signaled if it is absent.

[0113] At 507, a determination is made of whether the "controllersets" list is present. If so, it is parsed at 508. (FIG. 4 illustrates the parsing of each item of the "controllersets" list.)

[0114] At 509, the result of the data found in the MPEG_node_avatar extension (if any) is stored in a dedicated structure and should not change the scene.

[0115] An example method of parsing a "controllersets" list in MPEG_node_avatar is illustrated in the flow chart of FIG. 4.

[0116] The parsing of an AvatarControllerSet property in the gITF file begins at 601 . At 602, the "purpose" property is parsed. In embodiments where this property is required, an error may be signaled if it is absent. At 603, a determination is made of whether the “description” property is present; if so, it is parsed at 604. At 605, a determination is made of whether the “weights” property is present; if so, it is parsed at 606. At 607, each item of the “controllers” list is parsed as described in greater detail in FIG. 5. Parsing ends at 608.

[0117] FIG. 5 is a flow diagram illustrating example parsing of an item of the "controllers" list in an AvatarControllerSet.

[0118] Parsing of an AvatarController property begins at 701. At 702, the "controller" property is parsed. In embodiments where this property is required, an error may be signaled if it is absent. The value of “controller” may be constrained to be a valid index in the “controllers” list of “MPEG_controllers” extension. At 703, a determination is made ofwhetherthe “description” property is present; if so, it is parsed at 704. At 705, a determination is made of whether the “purpose” property is present; if so, it is parsed at 706. If “purpose” is not present, then the purpose of this controller is defined by the controller set that embeds it and its index in the controller set. For instance, if controller #2 has no purpose, then the “purpose” of the controllerset may define a default purpose for controller #2. At 707, a determination is made of whether the “range” property is present; if so, it is parsed at 708. Parsing ends at 709.

[0119] FIG. 6 schematically illustrates the structure of a 3D scene as represented in a runtime asset delivery file, e.g. a gITF file. The scene in this example includes four meshes. Three of the meshes, namely a left eye mesh 1202, a right eye mesh 1204, and a head mesh 1206, are associated with an avatar node by being referenced in an MPEG_node_avatar extension, as described above. For the sake of illustration, this scene also includes another node that is not an avatar node and is associated with a mesh 1208 referred to in FIG. 6 as a scene mesh. The scene of FIG. 6 includes four controllers defined in an MPEG_controllers extension as described above. One of the controllers 1210 is used to control vertical eye movement of the avatar, for example applying a rotation matrix to the left and right eyes while applying a morph target weight to appropriate targets (e.g. the eyelids) of the head mesh, with the appropriate rotation matrices and morph target weights being determined by the controller based on a weight input to that controller. Another one of the controllers 1212 is used to control horizontal eye movement of the avatar, for example applying a rotation matrix to the left and right eyes while applying a morph target weight to appropriate targets (e.g. the eyelids) of the head mesh, with the appropriate rotation matrices and morph target weights being determined by the controller based on a weight input to that controller.

[0120] Two additional controllers 1214, 1216 in FIG. 6 may apply appropriate transformations (e.g. translations, rotations, and / or morph target weights) to the scene mesh based on weight inputs. As described in greater detail above, the controllers may make use of samplers in determining an appropriate correspondence between input weights and output transformations applied to the meshes, but such details are not illustrated in FIG. 6.

[0121] One feature of the scene of FIG. 6 is that it does not include explicit information regarding the purpose of the different controllers. If the scene is represented using the conventional JSON format of gITF, then each controller is represented by an index indicating the position of that controller in an array of controllers within the MPEG_controllers extension, with the first controller in the array having an index of zero, the second having an index of one, and so on. Such an arrangement may be sufficient in a tightly controlled application environment, in which an application that processes the scene can be relied on to have information indicating which controller in the array applies which transformation to which mesh(es). Such information may indicate that the first controller, with index 0, controls the horizontal eye movement, and so on. However, this arrangement hinders the portability of the scene for use by other applications. For example, if an application has information only regarding the index of the different controllers (which it can determine implicitly from a gITF file) it may not have sufficient information to determine which controller to use to, for example,perform a horizontal eye movement. Applying a weight to the wrong controller could have an unintended effect on the scene by applying the wrong transformation and / or transforming the wrong mesh entirely.

[0122] FIG. 7 schematically illustrates an encoding of the scene of FIG. 6 with the addition of of avatar signaling with controller sets according to example embodiments as described herein. In the example of FIG. 7, the MPEG_node_avatar extension 1300 of the scene includes additional semantic tag information providing a description and purpose for the controllers associated with the avatar node. Specifically, the controller for horizontal eye movement is provided with the "description” property of "Horizontal gaze" and the “purpose” property of "gaze:horizontal.” Similarly, the controller for vertical eye movement is provided with the "description” property of "Vertical gaze" and the “purpose” property of "gaze:vertical.” In this example, the set of both controllers associated with the avatar is provided with semantic tags in the form of a "description" property of "Update gaze" and the "purpose" property of "gaze." In the example of FIG. 7, the semantic tags are organized in a hierarchy, with a more general semantic tag “gaze” being provided in a syntax structure 1302 associated with a set of controllers, and more specific semantic tags “gaze:vertical” and “gaze:horizontal” being provided by syntax structures 1304, 1306, respectively.

[0123] In some embodiments, the “description” and “purpose” properties of the controllers are provided explicitly in the MPEG_node_avatar extension, while the controllers are referenced within the MPEG_node_avatar extension by their index (e.g. representing their order within the MPEG_controllers extension that defines them). However, variations on this configuration are possible. For example, the “description” and “purpose” properties may be provided within the same data structure (such as a MPEG_controllers extension) used to define the controllers, or controllers may be defined in a data structure that provides avatar properties (such as in the MPEG_node_avatar extension). Some embodiments do not necessarily provide both a “description” and a “purpose” property. For example, some embodiments may provide only one property, e.g. the “purpose” property, which may be constrained to be selected from among a predetermined set of properties. Of course, properties such as “description” and “purpose” may have different names in different embodiments.

[0124] The principles described herein do not require semantic tags to be provided only by an extension associated with an avatar node. For example, the embodiment of FIG. 7 may be modified by replacing the MPEG_node_avatar extension 1300 with an MPEG_node_controllers extension or any other syntax structure that conveys semantic tags for controllers.

[0125] As illustrated in FIG. 8, in a method according to some embodiments, a runtime asset delivery file is obtained 802 for a 3D scene. The file may be obtained by a server or by a client device for the display and / or rendering of 3D scene information. In an example embodiments, the file includes a plurality of first syntax structures, which may be controllers in an MPEG_controllers structure. Each first syntax structure defines an association between an input weight value and a corresponding transformation of at least one of a plurality of nodes in the scene. The file further includes at least one second syntax structure, which may be a ControllerSet and / or one or more Controller properties as described above, among other possibilities. The second syntax structure associates a respective semantic tag (e.g. the “purpose” property described above, although other configurations are possible) with at least one of the first syntax structures. At 804, at least one of the first syntax structures is selected based on the associated semantic tag. At 806, the transformation indicated by the selected first syntax structure.

[0126] As a non-limiting example, the second syntax structure may associate the semantic tag “gaze:horizontal” with one of the first syntax structures. The selection of the first syntax structure may be based on this semantic tag. For example, in response to a determination by the client or server device that the horizontal direction of gaze should be adjusted for a mesh representing an avatar (or other eyed creature), the device in this example searches the second syntax structures to identify one of them that includes the tag “gaze:horizontal,” selects which first syntax structure is associated with that tag, and applies the transformation indicated by that first syntax structure (e.g. by applying varying input weights to that structure).

[0127] In some embodiments, the runtime asset delivery file further includes a third syntax structure, such as an MPEG_node_avatar or an MPEG_node_controllers structure described above, that is associated with a node in the scene. In some such embodiments, the second syntax structure is included in the third syntax structure. In such embodiments, the transformation corresponding to the first syntax structure may be a transformation of a mesh in that node.

[0128] In some embodiments, the third syntax structure includes information indicating that the node represents an avatar.

[0129] In some embodiments, the file further includes a fourth syntax structure, which may be a ControllerSet syntax structure as described above, that associates a respective semantic tag with a set comprising a plurality of the first syntax structures.

[0130] In some embodiments, each of the first syntax structures has a respective index, and wherein the second syntax structure includes the index of at least one of the first syntax structures.

[0131] In some embodiments, the runtime asset delivery file is a gITF file.

[0132] In some embodiments, the semantic tag comprises a “purpose” property as described above. The semantic tag and any other syntactic feature described herein may have different names in different embodiments.

[0133] In some embodiments, a client or server device may operate to animate at least the node associated with the selected first syntax structure by applying changing weights to the transformation corresponding to the selected first syntax structure.

[0134] Some embodiments further include causing display of at least the node associated with the selected first syntax structure. Such embodiments may include actually displaying a rendering of the node, e.g. on a display screen, providing a rendering of the node, to a display device, providing image or video data of the node to a display device, or other actions.

[0135] As shown in FIG. 9, in some embodiments, a device such as a client or server device obtains at 902 a plurality of runtime asset delivery files. These files may be received directly from different users 904, 906, 908, or they may be stored in a database (e.g. 910) where they may be associated with specific users or they may be selectable as options by more than one user. Each of the runtime asset delivery files includes a respective node representing an avatar 912, 914, 916, 918. Each of the runtime asset delivery files also includes at least one first syntax structure (e.g. a controller in an MPEG_controllers extension) defining an association between an input weight value and a corresponding transformation of the respective avatar, and at least one second syntax structure (e.g. a ControllerSet and / or a Controller property) that associates a respective semantic tag (e.g. the “purpose” property) with the first syntax structure.

[0136] At 920, the device presents a virtual environment including the avatars 912, 914, 916, 918 described in the respective runtime asset delivery files. The virtual environment may be, for example, a teleconference, a virtual world, or a virtual store, among other options.

[0137] At 922, the device selects an action to perform on a selected one of the avatars. This selection may be performed based at least in part on input from a user. As an example, user 912 may provide an input indicating that he wishes to turn his head.

[0138] At 924, using the semantic tags of the second syntax structures associated with the selected avatar, a first syntax structure is selected that corresponds to the selected action. For example, a first syntax structure may be associated with a semantic tag indicating that the corresponding transformation results in a head-turning motion. It may be the case, for example, that the different files from different users all include a controller that turns the avatar’s head, but those different controllers may have different indices in those different files.The use of semantic tags associated with those controllers makes it possible to identify the appropriate controller despite the lack of a standardized scheme of controller indices.

[0139] At 926, the transformation indicated by the identified first syntax structure is applied to the avatar.Use Case: Processing Workflow

[0140] When creating 3D assets, a common approach consists in splitting the procedure into several stages. Each stage focuses on a specific task, like working on the mesh modeling, the textures, the animation, the lights, etc. These stages are usually handled by different people or teams, or by the same person but at different times. To manage this, one solution is to store the result of each stage in a new file. Using the current gITF standard, one can already follow this approach for some tasks. However, this approach is not effective for other tasks, like complex animations of avatars.

[0141] Considering this context, one approach consists in using a collection of basic animations or controllers per model. For instance, faces are associated with a set of facial expressions; each one being encoded as an animation. During an early step of the workflow, artists must design these animations, in which case the current gITF standard is sufficient. During later stages, artists may combine these basic animations to create a final animated scene. In this last case, and if this is the last stage, the animations combinations can be baked into final animations into current gITF files. However, if this is not the last stage, and if another artist needs to make updates to the final animation (for instance, during the lighting stage), it is no longer possible with a standard approach: they must ask artists in the earlier stages to update animations and bake again.

[0142] Using the “MPEG_controllers” extension, it becomes possible to store information useful in the animation creation process in a gITF file. If this context, the basic animations are called controllers, and the combination of animations is then an animation of controllers. However, only using the standard and the “MPEG_controllers” extension, there is no information in the gITF file about the role of each controller. Then, when an artist gets the avatar from a previous stage, they (or the application) must have separate information indicating the purpose of each controller in the gITF file, because that file does not itself include such information. Moreover, without this information, different gITF files would need to have the same number of controllers and the same order. If, for some reason, a new controller has to be added to the current controllers list, all the processing chain must be updated. In complex environments involving many teams or even many compagnies, such an update can be very difficult or entirely unfeasible to execute.

[0143] Using proposed embodiments as described herein, the avatar controllers may be associated with meaningful information. For example, a set of purpose identifiers may be defined for each controller type (e.g., using the “purpose” property of AvatarController) . Alternatively or additionally, each controller may be described, giving information to the artist about the purpose of the controller (e.g., using the “description” property of AvatarController). Furthermore, example embodiments make it possible to provide different sets of controllers with the proposed extension. In such embodiments, each controller set may be dedicated to a specific purpose, like only modifying the identity or pose of the avatar. In some embodiments, controller sets may be used for versioning. This makes it possible for artists in the next stages to modify the controllers' animations themselves, without the need to guess the purpose of each controller and set.Use Case: Avatars in Virtual Worlds

[0144] Some example embodiments are useful in virtual worlds offered by different companies, governments, or users. With such embodiments, users can import their avatar into these virtual worlds, and they will work correctly in every virtual world that respects the same standard.

[0145] Many features may be expected from these avatars, like the ability to change their identity (face shape, nose size, etc.). To correctly change the identity, the application can use a controller set dedicated to identity. Using this controller set according to example embodiments, the user can change the identity inside the virtual world. Furthermore, the user can also update the avatar gITF file with new weights for this controller set and use it in another virtual world supporting the standard.

[0146] These mechanisms also work for other modifications, including live ones like expressions. In these cases, a controller set can be dedicated to these modifications, and used by the application to animate the avatar. Such controller sets can be updated: for example, if the user wants to change the way its avatar smiles, the controller dedicated to this expression can be modified. Again, it will update the avatar in the current virtual world, but also in others supporting the standard.Use Case: Avatars in Virtual Shops

[0147] Some embodiments may be used in cases where there are virtual shops handled by different companies, governments, or users. Using embodiments as described herein, users can import their avatar in these virtual shops, and they will work correctly in every virtual shop that respects the same standard.

[0148] People can buy clothing in a virtual shop. They first import their avatar in the shop, and then, they can try clothes, assuming their avatar has their body measurements. They canchange the pose of their avatar to see how the clothes match depending on situations. For example, an avatar controller can model different usual poses (stand, sit, etc.), helping them to see how a clothing can fit their body.Use Case: Avatars in Video Conferencing

[0149] In some embodiments, avatars can be used in video conferencing, where an avatar is used to represent the user. It is desirable for the avatar to be animated, in the first place with facial expressions that matches the user’s speech. So, at least one avatar controller set can be dedicated to facial expressions; however, several controllers set dedicated to facial expressions can also be put in the gITF file, each one for a specific mood: a facial expression sets when happy, one when sad, one when angry, etc. This principle can be repeated with no limit but imagination, like avatar standing, moves, funning poses, and the like.Example System Hardware

[0150] A device for the display and / or rendering of 3D scene information, together with its control electronics, may be implemented using a system such as the system of FIG. 10. FIG. 10 is a block diagram of an example of a system in which various aspects and embodiments are implemented. System 1000 can be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this document. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 1000, singly or in combination, can be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, the system 1000 is communicatively coupled to one or more other systems, or other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, the system 1000 is configured to implement one or more of the aspects described in this document.

[0151] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this document. Processor 1010 can include embedded memory, input output interface, and various other circuitries as known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device, and / or a non-volatile memory device). System 1000 includes a storage device 1040, which can include non-volatile memory and / or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory(EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash, magnetic disk drive, and / or optical disk drive. The storage device 1040 can include an internal storage device, an attached storage device (including detachable and non-detachable storage devices), and / or a network accessible storage device, as non-limiting examples.

[0152] System 1000 includes an encoder / decoder module 1030 configured, for example, to process data to provide an encoded or decoded scene, and the encoder / decoder module 1030 can include its own processor and memory. The encoder / decoder module 1030 represents module(s) that can be included in a device to perform the encoding and / or decoding functions. As is known, a device can include one or both of the encoding and decoding modules. Additionally, encoder / decoder module 1030 can be implemented as a separate element of system 1000 or can be incorporated within processor 1010 as a combination of hardware and software as known to those skilled in the art.

[0153] Program code to be loaded onto processor 1010 or encoder / decoder 1030 to perform the various aspects described in this document can be stored in storage device 1040 and subsequently loaded onto memory 1020 for execution by processor 1010. In accordance with various embodiments, one or more of processor 1010, memory 1020, storage device 1040, and encoder / decoder module 1030 can store one or more of various items during the performance of the processes described in this document. Such stored items can include, but are not limited to, the input scene description file, the decoded scene or portions of the decoded scene, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0154] In some embodiments, memory inside of the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other embodiments, however, a memory external to the processing device (for example, the processing device can be either the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory can be the memory 1020 and / or the storage device 1040, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of, for example, a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for coding and decoding operations, such as for MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also referred to as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part2), or WC (Versatile Video Coding, a new standard being developed by JVET, the Joint Video Experts Team).

[0155] The input to the elements of system 1000 can be provided through various input devices as indicated in block 1130. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples include composite video.

[0156] In various embodiments, the input devices of block 1130 have associated respective input processing elements as known in the art. For example, the RF portion can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) downconverting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the downconverted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs various of these functions, including, for example, downconverting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements can include inserting elements in between existing elements, such as, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.

[0157] Additionally, the USB and / or HDMI terminals can include respective interface processors for connecting system 1000 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, can be implemented, for example, within a separate input processing IC or within processor 1010 as necessary. Similarly, aspects of USB or HDMI interface processing can be implemented within separate interface ICs or within processor 1010 as necessary. The demodulated, error corrected, and demultiplexed stream is providedto various processing elements, including, for example, processor 1010, and encoder / decoder 1030 operating in combination with the memory and storage elements to process the datastream as necessary for presentation on an output device.

[0158] Various elements of system 1000 can be provided within an integrated housing, Within the integrated housing, the various elements can be interconnected and transmit data therebetween using suitable connection arrangement 1140, for example, an internal bus as known in the art, including the I nter-IC (I2C) bus, wiring, and printed circuit boards.

[0159] The system 1000 includes communication interface 1050 that enables communication with other devices via communication channel 1060. The communication interface 1050 can include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 1060. The communication interface 1050 can include, but is not limited to, a modem or network card and the communication channel 1060 can be implemented, for example, within a wired and / or a wireless medium.

[0160] Data is streamed, or otherwise provided, to the system 1000, in various embodiments, using a wireless network such as a Wi-Fi network, for example IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signal of these embodiments is received over the communications channel 1060 and the communications interface 1050 which are adapted for Wi-Fi communications. The communications channel 1060 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet for allowing streaming applications and other over- the-top communications. Other embodiments provide streamed data to the system 1000 using a set-top box that delivers the data over the HDMI connection of the input block 1130. Still other embodiments provide streamed data to the system 1000 using the RF connection of the input block 1130. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth network.

[0161] The system 1000 can provide an output signal to various output devices, including a display 1100, speakers 1110, and other peripheral devices 1120. The display 1100 of various embodiments includes one or more of, for example, a touchscreen display, an organic lightemitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 can be for a television, a tablet, a laptop, a cell phone (mobile phone), or other device. The display 1100 can also be integrated with other components (for example, as in a smart phone), or separate (for example, an external monitor for a laptop). The other peripheral devices 1120 include, in various examples of embodiments, one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disk player, a stereo system, and / or a lightingsystem. Various embodiments use one or more peripheral devices 1120 that provide a function based on the output of the system 1000. For example, a disk player performs the function of playing the output of the system 1000.

[0162] In various embodiments, control signals are communicated between the system 1000 and the display 1100, speakers 1110, or other peripheral devices 1120 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communications protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, the output devices can be connected to system 1000 using the communications channel 1060 via the communications interface 1050. The display 1100 and speakers 1110 can be integrated in a single unit with the other components of system 1000 in an electronic device such as, for example, a television. In various embodiments, the display interface 1070 includes a display driver, such as, for example, a timing controller (T Con) chip.

[0163] The display 1100 and speaker 1110 can alternatively be separate from one or more of the other components, for example, if the RF portion of input 1130 is part of a separate set- top box. In various embodiments in which the display 1100 and speakers 1110 are external components, the output signal can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0164] The system 1000 may include one or more sensor devices 1095. Examples of sensor devices that may be used include one or more GPS sensors, gyroscopic sensors, accelerometers, light sensors, cameras, depth cameras, microphones, and / or magnetometers. Such sensors may be used to determine information such as user’s position and orientation. Where the system 1000 is used as the control module for an extended reality display, the user’s position and orientation may be used in determining how to render image data such that the user perceives the correct portion of a virtual object or virtual scene from the correct point of view. In the case of head-mounted display devices, the position and orientation of the device itself may be used to determine the position and orientation of the user for the purpose of rendering virtual content. In the case of other display devices, such as a phone, a tablet, a computer monitor, or a television, other inputs may be used to determine the position and orientation of the user for the purpose of rendering content. For example, a user may select and / or adjust a desired viewpoint and / or viewing direction with the use of a touch screen, keypad or keyboard, trackball, joystick, or other input. Where the display device has sensors such as accelerometers and / or gyroscopes, the viewpoint and orientation used for the purpose of rendering content may be selected and / or adjusted based on motion of the display device.

[0165] The embodiments can be carried out by computer software implemented by the processor 1010 or by hardware, or by a combination of hardware and software. As a nonlimiting example, the embodiments can be implemented by one or more integrated circuits. The memory 1020 can be of any type appropriate to the technical environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. The processor 1010 can be of any type appropriate to the technical environment, and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.Further Embodiments

[0166] This disclosure describes a variety of aspects, including tools, features, embodiments, models, approaches, etc. Many of these aspects are described with specificity and, at least to show the individual characteristics, are often described in a manner that may sound limiting. However, this is for purposes of clarity in description, and does not limit the disclosure or scope of those aspects. Indeed, all of the different aspects can be combined and interchanged to provide further aspects. Moreover, the aspects can be combined and interchanged with aspects described in earlier filings as well.

[0167] The aspects described and contemplated in this disclosure can be implemented in many different forms. While some embodiments are illustrated specifically, other embodiments are contemplated, and the discussion of particular embodiments does not limit the breadth of the implementations. At least one of the aspects generally relates to encoding, decoding, and rendering of scene description information, and at least one other aspect generally relates to transmitting a file and / or bitstream generated or encoded. These and other aspects can be implemented as a method, an apparatus, a computer readable storage medium having stored thereon instructions for encoding, decoding, or rendering of scene description data according to any of the methods described, and / or a computer readable storage medium having stored thereon a bitstream generated according to any of the methods described.

[0168] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example,the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.

[0169] Various numeric values may be used in the present disclosure, for example. The specific values are for example purposes and the aspects described are not limited to these specific values.

[0170] Embodiments described herein may be carried out by computer software implemented by a processor or other hardware, or by a combination of hardware and software. As a nonlimiting example, the embodiments can be implemented by one or more integrated circuits. The processor can be of any type appropriate to the technical environment and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non-limiting examples.

[0171] When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of a corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of a corresponding method / process.

[0172] The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, ora programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end-users.

[0173] Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this disclosure are not necessarily all referring to the same embodiment.

[0174] Additionally, this disclosure may refer to “determining” various pieces of information. Determining the information can include one or more of, for example, estimating theinformation, calculating the information, predicting the information, or retrieving the information from memory.

[0175] Further, this disclosure may refer to “accessing” various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.

[0176] Additionally, this disclosure may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.

[0177] It is to be appreciated that the use of any of the following 7”, “and / or”, and “at least one of’, for example, in the cases of “A / B”, “A and / or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended for as many items as are listed.

[0178] Also, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments the encoder signals a particular one of a plurality of parameters for region-based filter parameter selection for de-artifact filtering. In this way, in an embodiment the same parameter is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of anyactual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun.

[0179] Implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal can be formatted to carry the bitstream of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.

[0180] We describe a number of embodiments. Features of these embodiments can be provided alone or in any combination, across various claim categories and types. Further, embodiments can include one or more of the following features, devices, or aspects, alone or in any combination, across various claim categories and types:• A bitstream or signal that includes one or more of the described syntax elements, or variations thereof.• A bitstream or signal that includes syntax conveying information generated according to any of the embodiments described.• Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements, or variations thereof.• Creating and / or transmitting and / or receiving and / or decoding according to any of the embodiments described.• A method, process, apparatus, medium storing instructions, medium storing data, or signal according to any of the embodiments described.

[0181] Note that various hardware elements of one or more of the described embodiments are referred to as “modules” that carry out (i.e., perform, execute, and the like) various functions that are described herein in connection with the respective modules. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memorydevices) deemed suitable for a given implementation. Each described module may also include instructions executable for carrying out the one or more functions described as being carried out by the respective module, and it is noted that those instructions could take the form of or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and / or the like, and may be stored in any suitable non-transitory computer- readable medium or media, such as commonly referred to as RAM, ROM, etc.

[0182] Although features and elements are described above in particular combinations, each feature or element can be used alone or in any combination with the other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, a read only memory (ROM), a random access memory (RAM), a register, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks, and digital versatile disks (DVDs). A processor in association with software may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

CLAIMS1. A method comprising: obtaining a runtime asset delivery file for a 3D scene, the file including: a plurality of first syntax structures, each first syntax structure defining an association between an input weight value and a corresponding transformation of at least one of a plurality of nodes in the scene; and at least one second syntax structure associating a respective semantic tag with at least one of the first syntax structures; selecting at least one of the first syntax structures based on the associated semantic tag; and applying the transformation indicated by the selected first syntax structure.

2. An apparatus comprising one or more processors, the apparatus being configured to perform at least: obtaining a runtime asset delivery file for a 3D scene, the file including: a plurality of first syntax structures, each first syntax structure defining an association between an input weight value and a corresponding transformation of at least one of a plurality of nodes in the scene; and at least one second syntax structure associating a respective semantic tag with at least one of the first syntax structures; selecting at least one of the first syntax structures based on the associated semantic tag; and applying the transformation indicated by the selected first syntax structure.

3. The method of claim 1 or the apparatus of claim 2, wherein the first syntax structures are controller structures in an MPEG_controllers extension.

4. The method of claim 1 or claim 3 as it depends from claim 1 , or the apparatus of claim 2 or claim 3 as it depends from claim 2, wherein the file further includes a third syntax structure associated with a node in the scene, wherein the second syntax structure is included in the third syntax structure, and wherein the transformation corresponding to the first syntax structure is a transformation of a mesh in the node.

5. The method of claim 4 as it depends from claim 1 , or the apparatus of claim 4 as it depends from claim 2, wherein the third syntax structure includes information indicating that the node represents an avatar.

6. The method of claim 4 as it depends from claim 1 , or the apparatus of claim 4 as it depends from claim 2, wherein the third syntax structure is an MPEG_node_avatar extension.

7. The method of claim 4 as it depends from claim 1 , or the apparatus of claim 4 as it depends from claim 2, wherein the third syntax structure is an MPEG_node_controllers extension.

8. The method of claim 1 or any of claims 3-7 as they depend from claim 1 , or the apparatus of claim 2 or any of claims 3-7 as they depend from claim 2, wherein the file further includes a fourth syntax structure associating a respective semantic tag with a set comprising a plurality of the first syntax structures.

9. The method of claim 1 or any of claims 3-8 as they depend from claim 1 , or the apparatus of claim 2 or any of claims 3-8 as they depend from claim 2, wherein each of the first syntax structures has a respective index, and wherein the second syntax structure includes the index of at least one of the first syntax structures.

10. The method of claim 1 or any of claims 3-9 as they depend from claim 1 , or the apparatus of claim 2 or any of claims 3-9 as they depend from claim 2, wherein the runtime asset delivery file is a gITF file.

11. The method of claim 1 or any of claims 3-10 as they depend from claim 1 , or the apparatus of claim 2 or any of claims 3-10 as they depend from claim 2, wherein the semantic tag comprises a purpose property.

12. The method of claim 1 or any of claims 3-11 as they depend from claim 1 , or the apparatus of claim 2 or any of claims 3-11 as they depend from claim 2, further comprising animating at least the node associated with the selected first syntax structure by applying changing weights to the transformation corresponding to the selected first syntax structure.

13. The method of claim 1 or any of claims 3-12 as they depend from claim 1 , or the apparatus of claim 2 or any of claims 3-12 as they depend from claim 2, further comprising causing display of at least the node associated with the selected first syntax structure.

14. The method of claim 1 or any of claims 3-13 as they depend from claim 1 , or the apparatus of claim 2 or any of claims 3-13 as they depend from claim 2, further comprising: obtaining a plurality of runtime asset delivery files, each of the runtime asset delivery files including: a respective node representing an avatar, at least one first syntax structure defining an association between an input weight value and a corresponding transformation of the respective avatar, and at least one second syntax structure associating a respective semantic tag with the first syntax structure; presenting a virtual environment including the avatars of the respective runtime asset delivery files; selecting an action to perform on a selected one of the avatars; using the semantic tags of the second syntax structures associated with the selected avatar, identifying a first syntax structure corresponding to the selected action; and applying the transformation indicated by the identified first syntax structure.

15. A computer-readable medium storing a runtime asset delivery file for a 3D scene, the file including: a plurality of first syntax structures, each first syntax structure defining an association between an input weight value and a corresponding transformation of at least one of a plurality of nodes in the scene; and at least one second syntax structure associating a respective semantic tag with at least one of the first syntax structures.