AÇÕES E COMPORTAMENTOS DE AVATAR EM AMBIENTES VIRTUAIS

BR112025019941A2Pending Publication Date: 2026-08-04INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
BR · BR
Patent Type
Applications
Current Assignee / Owner
INTERDIGITAL CE PATENT HOLDINGS SAS
Filing Date
2024-03-15
Publication Date
2026-08-04

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In one implementation, a new set of actions and behaviors associated with interactive spaces and avatars in 3D virtual environments are proposed. These actions and behaviors define the capabilities of an avatar in areas of interactivity. They can be used in MPEG-I Scene Description to support avatar social interactivity in 3D environments with corresponding time-based events. Generally, capabilities describe the allowed actions of an avatar following a trigger event in a region of interactivity. In the case "Disabilities" are defined for the user, information included in "Disabilities" will also impact the capabilities of a user avatar. The actions, for example, can include social action, restriction action, parental action, speech action, capabilities action and disabilities action. In one example, a new type of property (ACTION_SET_AVATAR) is included to the framework of MPEG_scene_activity. Under ACTION_SET_AVATAR, an object (avatarAction) is used to represent avatar-specific actions.
Need to check novelty before this filing date? Find Prior Art

Description

1 / 45 “AVATAR ACTIONS AND BEHAVIORS IN VIRTUAL ENVIRONMENTS” TECHNICAL FIELD

[001] The present modalities generally refer to digital human interaction in 3D virtual scenes, more particularly, to avatar actions and behaviors in virtual environments. BACKGROUND

[002] Extended reality (XR) is a technology that enables interactive experiences in which the real-world environment and / or video content is enhanced by virtual content, which can be defined in multiple sensory modalities, including visual, auditory, tactile, etc. During application runtime, the virtual content (3D content or audio / video file, for example) is rendered in real time in a way that is consistent with the user's context (environment, viewpoint, device, etc.). Scene graphs (such as those proposed by Khronos / glTF (Graphics Language Transmission Format) and its extensions defined in MPEG Scene Description or Apple / USDZ formats, for example) are one possible way to represent the content to be rendered. They combine a declarative description of the scene structure, linking objects from the real environment and virtual objects, on the one hand, and binary representations of the virtual content, on the other.Scene description structures ensure that timed media and corresponding relevant virtual content are available at any time during application rendering. Scene descriptions can also contain scene-level data, describing how a user can interact with scene objects at runtime for immersive XR experiences. SUMMARY

[003] According to one embodiment, a method is provided that comprises: obtaining, from a description for an extended reality scene, at least one parameter used to define one or more actions allowed for a node. Petition 870250084121, dated 09 / 18 / 2025, p. 48 / 108 2 / 45 of an avatar representing an avatar; activate a trigger for an action associated with said avatar node, where said action belongs to said one or more allowed actions; and initialize said action for said avatar node.

[004] According to another embodiment, a method is provided comprising: generating at least one parameter in a description for an extended reality scene to define one or more allowed actions for an avatar node that represents an avatar; associating a trigger with an action with the avatar node, where the action belongs to one or more allowed actions; and encoding the description for the extended reality scene.

[005] According to another embodiment, a device is provided comprising one or more processors and at least one memory, wherein said one or more processors are configured to: obtain, from a description of an extended reality scene, at least one parameter used to define one or more actions allowed for an avatar node representing an avatar; activate a trigger for an action associated with said avatar node, wherein said action belongs to one or more allowed actions; and initialize said action for said avatar node.

[006] According to another embodiment, a device is provided comprising one or more processors and at least one memory, wherein the processor(s) is / are configured to: generate at least one parameter in a description for an extended reality scene to define one or more allowed actions for an avatar node representing an avatar; associate a trigger with an action with the avatar node, wherein the action belongs to one or more allowed actions; and encode the description for the extended reality scene.

[007] One or more embodiments also provide a computer program comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method of Petition 870250084121, dated 09 / 18 / 2025, page 49 / 108 3 / 45 in accordance with any of the embodiments described herein. One or more of the present embodiments also provide a computer-readable storage medium having instructions stored therein for scene description processing in accordance with the methods described herein.

[008] One or more embodiments also provide a computer-readable storage medium having scene descriptions generated according to the methods described above. One or more embodiments also provide a method and apparatus for transmitting or receiving scene descriptions generated according to the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[009] Figure 1 shows an example of an XR processing mechanism architecture.

[010] Figure 2 shows an example of the syntax of a continuous data stream that encodes an extended reality scene description.

[011] Figure 3 shows an exemplary graphic of an extended reality scene description.

[012] Figure 4 shows an example of an extended reality scene description that includes behavioral data.

[013] Figure 5 illustrates an example of executing a hierarchical action, according to a modality.

[014] Figure 6 illustrates the generation of parameters with scene coding.

[015] Figure 7 illustrates a diagram of the encoder, according to a modality. DETAILED DESCRIPTION

[016] Various XR applications can be applied to different real or virtual contexts and environments. For example, in an industrial XR application, a virtual 3D content item (e.g., part A of an engine) is displayed when a Petition 870250084121, dated 09 / 18 / 2025, page 50 / 108 A 4 / 45 reference object (part B of an engine) is detected in the real environment by a camera attached to a head-mounted display device. The 3D content item is positioned in the real world with a defined position and scale relative to the detected reference object.

[017] For example, in an XR application for interior design, a 3D model of a piece of furniture is displayed when a specific catalog image is detected in the input camera view. The 3D content is positioned in the real world with a defined position and scale relative to the detected reference image. In another application, some audio file might start playing when the user enters an area near a church (real or virtually rendered in the extended real environment). In another example, an advertising jingle might play when the user sees a can of a specific soft drink in the real environment. In an outdoor game application, various virtual characters might appear, depending on the semantics of the scenario observed by the user.For example, bird characters are suitable for trees, so if the XR device's sensors detect real objects described by a semantic label "tree," birds can be added flying around the trees. In a complementary application implemented by smart glasses, a car noise can be initiated in the user's headset when a car is detected within the user's camera's field of view, to alert them to potential danger. Furthermore, the sound can be spatialized so that it comes from the direction where the car was detected.

[018] An XR application can also extend video content instead of a real environment. The video is displayed on a rendering device, and virtual objects described in the node tree are overlaid when timed events are detected in the video. In this context, the node tree comprises only descriptions of virtual objects. Petition 870250084121, dated 09 / 18 / 2025, page 51 / 108 5 / 45

[019] Figure 1 shows an example of an architecture of an XR processing mechanism 130 that can be configured to implement the methods described herein. A device according to the architecture of Figure 1 is linked to other devices via its bus 131 and / or via I / O interface 136.

[020] Device 130 comprises the following elements which are linked by a data and address bus 131: - a microprocessor 132 (or CPU), which is, for example, a DSP (or Digital Signal Processor); - a ROM (or Read-Only Memory) 133; - a RAM (or Random Access Memory) of 134; - a storage interface 135; - an I / O interface 136 for receiving data to be transmitted from an application; and - a power source (not shown in Figure 1), for example, a battery.

[021] According to one example, the power supply is external to the device. In each of the memories mentioned, the word “register” used in the specification can correspond to a small capacity area (a few bits) or a very large area (for example, an entire program or a large amount of received or decoded data). ROM 133 comprises at least one program and parameters. ROM 133 can store algorithms and instructions to perform techniques according to current principles. When switched, CPU 132 loads the program into RAM and executes the corresponding instructions.

[022] RAM 134 comprises, in a register, the program executed by CPU 132 and loaded after turning on device 130, input data in a register, intermediate data in different states of the method in a register and other variables used for the execution of the method in a register. Petition 870250084121, dated 09 / 18 / 2025, page 52 / 108 6 / 45

[023] Device 130 is linked, for example, via bus 131 to a set of sensors 137 and a set of rendering devices 138. Sensors 137 may be, for example, cameras, microphones, temperature sensors, Inertial Measurement Units, GPS, hygrometry sensors, infrared or ultraviolet light sensors, or wind sensors. Rendering devices 138 may be, for example, monitors, loudspeakers, vibrators, heaters, fans, etc.

[024] According to the examples, device 130 is configured to implement a method in accordance with these principles and belongs to a set comprising: - a mobile device; - a communication device; - a gaming device; - a tablet (or tablet computer); - a laptop; - a camera; - a video camera.

[025] In XR applications, scene description is used to combine an explicit and easily parsed description of a scene structure and some binary representations of media content. Figure 2 shows an example of the syntax of a continuous data stream that encodes an extended reality scene description. Figure 2 shows an example of structure 210 of an XR scene description. The structure consists of a container that organizes the stream into syntax-independent elements. The structure may comprise a header portion 220 which is a set of data common to each syntax element of the continuous stream. For example, the header portion comprises some metadata about syntax elements, describing the nature and function of each of them. The structure also Petition 870250084121, dated 09 / 18 / 2025, page 53 / 108 7 / 45 comprises a payload comprising a syntax element 230 and a syntax element 240. The syntax element 230 comprises representative data of the media content items described in the scene graph nodes related to the virtual elements. Images, meshes, and other raw data may have been compressed according to a compression method. The syntax element 240 is part of the data stream payload and comprises data that encodes the scene description as described according to the principles presented.

[026] Figure 3 shows an example of a 310 graph of an extended reality scene description. In this example, the scene graph can comprise a description of real objects, for example, 'flat horizontal surface' (which could be a table or a road) and a description of virtual objects 312, for example, an animation of a car. The scene description is organized as a matrix of nodes. A node can be linked to child nodes to form a scene structure 311. A node can carry a description of a real object (for example, a semantic description) or a description of a virtual object. In the example in Figure 3, node 301 describes a virtual camera located in the 3D volume of the XR application. Node 302 describes a virtual car and comprises an index of a representation of the car, for example, an index in a 3D mesh array. Node 303 is a child of node 302 and comprises a description of a car wheel. Similarly, it includes an index for the 3D mesh of the wheel.The same 3D mesh can be used for multiple objects in the 3D scene, as the scale, location, and orientation of the objects are described in the scene nodes. The 310 scene graph also includes nodes that describe the spatial relationship between real and virtual objects.

[027] In time-based streaming media, the scene description itself can evolve over time to provide the relevant virtual content for each sequence of a media stream. For example, for advertising purposes, a virtual bottle might be displayed on a table during a sequence of Petition 870250084121, dated 09 / 18 / 2025, page 54 / 108 8 / 45 video in which people are sitting around a table. This type of behavior can be achieved by relying on the structure defined in the Scene Description for MPEG media document.

[028] Currently, the MPEG-I scene description framework uses “behavior” data to augment the temporally evolving scene description and provides a description of how a user can interact with scene objects at runtime for immersive XR experiences. These behaviors relate to predefined virtual objects in which runtime interactivity is enabled for specific user XR experiences. These behaviors also evolve over time and are updated through the existing scene description update mechanism.

[029] Figure 4 shows an example of an extended reality scene description comprising behavioral data, stored at the scene level, describing how a user can interact with the scene objects, described at the node level, at runtime for immersive XR experiences. When the XR application is launched, media content items (e.g., meshes of virtual objects visible from the camera) are loaded, rendered, and buffered to be displayed when triggered. For example, when a flat surface is detected in the real environment by sensors, the application displays the buffered media content item as described in the related scene nodes. Timing is managed by the application according to the features detected in the real environment and the animation timing. A node in a scene graph may also contain no description at all and only play the role of parent-child nodes.Figure 4 shows relationships between behaviors that are understood in the scene description at the scene level and nodes that are components of the scene graph. Behaviors 410 are related to predefined virtual objects in the nodes. Petition 870250084121, dated 09 / 18 / 2025, page 55 / 108 9 / 45 which runtime interactivity is allowed for specific user XR experiences. The 410 behavior also evolves over time and is updated through the scene description update mechanism.

[030] A behavior comprises: - 420 shots defining the conditions to be met for its activation; - a trigger control parameter that defines logical operations between the defined triggers; - 430 actions to be processed when the shots are activated; - an action control parameter defining the order in which related actions are executed; - a priority number that allows selecting the highest priority behavior in case of competition between multiple behaviors on the same virtual object at the same time; - An optional stop action that specifies how to terminate this behavior when it is no longer defined in a newly received scene update; for example, a behavior will no longer be defined if a related object does not belong to the new scene or if the behavior is no longer relevant to this current media sequence (e.g., audio or video).

[031] Behavior 410 occurs at the scene level. A trigger is linked to nodes and to the child nodes of those nodes. In the example in Figure 4, trigger 1 is linked to nodes 1, 2, and 8. Since node 31 is a child of node 1, Trigger 1 is linked to node 31. Trigger 1 is also linked to node 14 as a child of node 8. Trigger 2 is linked to node 1. In fact, the same node can be linked to multiple triggers. Trigger n is linked to nodes 5, 6, and 7. A behavior can comprise multiple triggers. For example, a first behavior can be activated by trigger 1 AND trigger 2, E being the control parameter of the first behavior's trigger. A behavior can have multiple actions. For example, the first behavior can Petition 870250084121, dated 09 / 18 / 2025, page 56 / 108 10 / 45 perform Action m first and then action 1, where “first and then” is the control parameter for the first behavior. A second behavior can be triggered by performing action 1 first and then action 2, for example.

[032] Different formats can be used to represent the node tree. For example, the MPEG-I scene description structure using the Khronos glTF extension mechanism can be used for the node tree. In this example, an interactivity extension can be applied at the glTF scene level and is called MPEG_scene_interactivity. The corresponding semantics are provided in Table 1, where 'M' in the 'Usage' column indicates that the field is mandatory in an XR scene description format and 'O' indicates that the field is optional. TABLE 1 Name Type Usage Description triggers Arrangement M Contains the definition of all triggers used in that scene. actions Arrangement M Contains the definition of all actions used in that scene. behaviors Arrangement M Contains the definition of all behaviors used in that scene. A behavior is composed of a pair of (triggers, actions), trigger and action control parameters, a priority weighting, and an optional interrupt action.

[033] In this document, we present a new set of actions, for example, for MPEG-I Scene Description, to support avatar social interactivity in 3D environments with corresponding time-based events. These actions should be triggered by events between 3D objects in a virtual environment, such as dynamic objects (humanoid and non-humanoid characters, cars, Petition 870250084121, dated 09 / 18 / 2025, page 57 / 108 11 / 45 airplanes) or static objects (chairs, tables, plates). Current scene-level interactivity descriptions only support generic actions for a node in the scene, as illustrated in Table 2. However, they do not provide an avatar node with semantic information about human social behaviors, which is a common attribute in avatar and real-life user descriptions; for example, the act of walking is semantically described as "walking" or the ability to "walk," and not as a chain transformation of matrices describing the act of walking. At a higher level, it is better and more readable to describe actions with semantic information other than low-level, computer-readable 4x4 matrix multiplications. Therefore, this document presents high-level descriptions of actions that can have a significant impact on 3D social and interactive environments. Table 2: Types of actions available in MPEG_scene_interactivity Action Type Description “ACTION ACTIVATE” Adjust activation status of a node “ACTION TRANSFORM” Adjust transform for a node “ACTION BLOCK” Block the transform of a node “ACTION ANIMATION” Select and control an animation “ACTION MEDIA” Select and control media “ACTION MANIPULATE” Select a manipulation action “ACTION SET MATERIAL” Set new material for nodes “ACTION_SET_HAPTIC” Obtain haptic feedback from a set of nodes

[034] The proposed representation of user capabilities is intended to be compatible with scene description (CD) content and is primarily focused on the action and behavioral representation of an avatar in interactive regions. Details are provided below regarding the elements with their associated meaning, JSON (JavaScript Object Notation) encoding schemes, and how they can be used in Petition 870250084121, dated 09 / 18 / 2025, page 58 / 108 12 / 45 MPEG-I SD.

[035] In the current description, the proposed format follows the glTF format and is compatible with MPEG's current effort to extend glTF with MPEG extensions. However, the meaning and use are generic and can be encoded with any other formats, for example, XML and USD.

[036] Capabilities

[037] Here, we present the available actions and behaviors that an avatar can perform within an interactive region represented as a geometric primitive. The interactive region around an avatar can indicate which actions that avatar or 3D scene object is capable of performing, hence the definition of capabilities in the context of this document. As described earlier, behaviors are a set of conditions that will pair triggered events with specific actions and define temporal constraints on those conditions, allowing time-based events to occur in 3D virtual environments.

[038] Avatar Capability

[039] Here we present an illustrative example of capabilities associated with regions of interactivity in the context of social interactions between avatars and 3D objects in 3D virtual environments.

[040] Generally, capabilities describe the actions an avatar is allowed to take after a triggering event in an interactive region. For example, in a meeting room, viewer avatars can only use speech and upper body movements, such as gestures and head movements, or actions that describe the avatar's abilities (e.g., ability to run, walk, jump, talk, fly). If “Disabilities” are defined for the user, the information included in “Disabilities” will also impact the user's avatar capabilities.

[041] The following are some non-limiting examples of actions associated with an avatar in social environments. Petition 870250084121, dated 09 / 18 / 2025, page 59 / 108 13 / 45

[042] Social action corresponds to the user's social behavior. It can be generic (standard conversation and interactivity allowed) or defined and specific by the user. When interacting with another user, if a trigger is detected within an interactive region (given by proximity or collision triggers), that region will be designated for social interactions, allowing, for example, conversations between avatar users.

[043] Restricted action. In a scenario where permissions are required, for example, for reasons such as age, restrictions, access rights, a user's permitted movement can be limited. This action can also be used to limit the space in which the avatar can move.

[044] Parental action. This may limit interaction with permitted content to protect children and young adults.

[045] Speech action. This type of capability can be restricted to triggered events that only allow speech actions to be performed, for example, in a meeting room, viewer avatars can only use speech. This action will allow the use of a microphone or a pre-recorded media track. This action can be used in combination with the social action that allows certain types of interactivity, such as speech.

[046] Ability Action. This action lists the types of abilities allowed for avatars or 3D objects when in contact with a region that activates the actions. The different types of abilities should cover different types of activities, for example, but not limited to walking, flying, driving, talking. This list of actions will notify the engine and can be combined with other action modules, such as “Action_set_haptics” to enable haptic feedback or “Action_manipulate” to grasp objects. The goal is to create an action modifier that will restrict an avatar's animation to the list of provided ability actions. Petition 870250084121, dated 09 / 18 / 2025, pages 60 / 108 14 / 45

[047] Action for disabilities. Disabilities have the same effect as capabilities, although they are designed to inform the mechanism about the user's disabilities and, consequently, depending on the user's choice, this will impact the list of capability actions. This informative list is important for adapting the virtual environment to the individual needs of users. For example, users with hearing impairments should have visual cues instead of auditory cues.

[048] All provided actions can be used in combination with existing and newly introduced actions, if permitted, and can make use of existing interactive and animation tools from different fields, for example, using manipulators to perform a walking motion, using tactile manipulators to infer tactile feedback, or using a soundtrack / media for pre-recorded speech.

[049] Avatar Behavior

[050] Behavior is a set of parameters that defines the correspondence between actions and triggering events. This will couple the newly defined actions with collision, proximity, or user input triggers with a time-based event. Time-based behavior allows an interactivity region to have temporary actions and scale actions depending on the desired activity.

[051] Time-based actions

[052] Similar to time-based behavior, each action can define its own time period. This facilitates the individual definition of time for each action at the action level, rather than at the behavior level.

[053] Avatar Action and Behavior Patterns

[054] Next, we use MPEG-I as an example to illustrate the proposed actions and behaviors. In one modality, the actions and behaviors must respect the following requirements: Petition 870250084121, dated 09 / 18 / 2025, pp. 61 / 108 15 / 45 1. The representation of time-based actions and behaviors respects the primitives supported in the MPEG-I scene description and other available formats (XML, USD). 2. The interactive space is represented with a primitive, a trigger, an action, and a behavior label, which allows interactivity between an avatar representation, for example, individual body parts or an interactive avatar area, and 3D objects in the scene. 3. It allows multiple triggers for interaction with scenes, objects, and other avatars, and implements social, privacy, and interactive boundaries between any associated objects. 4. Time-based behaviors will determine the lifecycle of an action, and if not specifically defined, the event and action trigger time will be equal to the duration of the 3D scene.

[055] Action and Behavior in MPEG-I Scene Description

[056] In the MPEG-I scene description, we extend the “MPEG_scene_interactivity” element of the existing glTF node by adding the attributes described above.

[057] As the glTF MPEG interactivity extension allows triggering of events and behaviors in collision and proximity situations, the proposed extension contributes with an extension of new time-based actions and attributes for behaviors / actions for a trigger in the “MPEG_scene_interactivity” node. The generic implementation of the node can also be applied to the avatar representation, to add interactivity and time-based constraints to the avatar and its elements.

[058] We propose an extension that allows glTF models to use and interact with humanoid characters (avatars) and any other objects. We propose extending the action properties of the glTF scene element. Petition 870250084121, dated 09 / 18 / 2025, pp. 62 / 108 16 / 45 “MPEG_scene_interactivity” to define “ACTION_SET_AVATAR” which contains more avatar-related actions and time-based behavior constraints, as well as generic scene and node-level actions.

[059] Table 3 illustrates the new type of property added to the “MPEG_scene_interactivity” structure. Table 3: Description of the “MPEG_node_interactivity_action” extension. Name Description “ACTION_SET_AVATAR” Gets avatar-related actions on a set of nodes. Table 4 details the semantics of “ACTION, SET AVATAR”. Table 4: Object. Name Type Usage Description avatarAction object M Object that defines the type of actions of the avatar. The semantics are illustrated in Table 7.

[060] Under the new proposed “ACTION_SET_AVATAR”, we have an “avatarAction” object, which represents specific avatar actions. The semantic description is shown in Table 7. Table 5 illustrates the list of available specific avatar actions. Table 5: Types of actions. Action Type Description “Action Avatar Social” Adjust a node's action. “Action Avatar Restricted” Adjust a node's permissions. “Action_Avatar_Parental” Adjust parental and content usage permissions for a node. “Action Avatar Speech” Adjust a node's active speech.

[061] Table 6 illustrates the types of actions to be added at the scene level and general node of the interactivity structure. In the case where Petition 870250084121, dated 09 / 18 / 2025, pp. 63 / 108 17 / 45 “ACTION_SET_AVATAR” is not available or the framework does not implement any type of avatar, the system is still able to use the proposed action at the scene or node level. Table 6: Types of actions. Action Type Description “Action Social” Adjust a node's action. “Action Restricted” Adjust a node's permissions. “Action_Parental” Adjust a node's parental and content usage permissions. “Action Speech” Adjust a node's active speech. “Action Capabilities” Adjust a node's capabilities. “Action Disabilities” Adjust a node's disabilities.

[062] The semantics of the new proposed actions are provided in Table 7. Table 7: Semantic description of new action properties.______ Name Type Usage Description type chain M Defines the type of social action (Table 2, Table 5, Table 6). child number The index of a child action to be executed after the condition is met. The default value “-1” means there is no child node. duration number The duration of an action in seconds. The default value is “-1” for infinite duration, greater than “0” to define the action's lifespan if it is a time-based event, otherwise it will be ignored. If “0” the action is canceled and not executed. Petition 870250084121, dated 09 / 18 / 2025, pp. 64 / 108 18 / 45 Extension string: Extended attributes for any action that may not be included or that may be specific to the application. This makes it easier for the content creator to extend, for example, social actions or capabilities to other types of actions. If(type == “Action Social”){ authorized array M One or more elements from Table 8 that define the types of social actions. nodes array M indices of the nodes in the array to apply “Social Parameters” (Table 8) listed in the “authorized” field.} If(type == “Action Restricted”)! permissionjd string M Unique string identifier that restricts interaction between nodes without equal permission id. nodes array M indices of the nodes in the array to which “Restricted Parameters” will be applied, i.e., the nodes whose permission IDs should be checked / applied.} If(type == “Action_ _Parental”){ Petition 870250084121, dated 09 / 18 / 2025, pp. 65 / 108 19 / 45 age number M An element from Table 9 that defines the minimum age recommendation for users, given the content of the node list. descriptors array M One or more elements from Table 10 that add additional explicit semantics to the content present in the node list. nodes array M indices of the nodes in the array to apply “Parental Parameters”, for example, as described as age and descriptors.} If(type == “Action Speech”){ microphone Boolean O Indicates whether the user uses a microphone type as an audio input device. “0” is False and “1” is True. The default value is 0. media String O URI (Uniform Resource Identifier) ​​for a media track to play a pre-recorded audio file. nodes array M indices of the nodes in the array to allow “Speech” media.} If(type == “Action _Capabilities”){ Petition 870250084121, dated 09 / 18 / 2025, pp. 66 / 108 20 / 45 capabilities array M One or more elements from Table 11 define the capabilities of an avatar / object. nodes array M indices of the nodes in the array to apply “capabilities parameters”.} If(type == “Action Disabilities”)! disabilities array M One or more elements from Table 12 that define the disabilities of an avatar / object. nodes array M indices of the nodes in the array to apply “Disability Parameters”.}

[063] Table 8 illustrates the types of actions when the action type is Action_Social. Table 8: Types of social actions. Social Action Description “conversation” Allow social conversations with users and enable the “Action Speech” functionality if it is not already enabled. “interaction” Allow interaction between users.

[064] Table 9 defines the minimum age recommendation for users, given the content of the node list. Table 9: Types of age levels. Age Description 3 Content suitable for all ages. 7 Content with scenes or sounds that may be frightening for younger children. Petition 870250084121, dated 09 / 18 / 2025, pp. 67 / 108 21 / 45 12 Content featuring violence with unrealistic graphic characters. 16 Content featuring violence that mimics reality. 18 Content developed for adults only.

[065] Table 10 describes additional explicit semantics of the content present in the list of nodes. Table 10: Type of parental descriptors. Parental Descriptors Description “violence” Contains depictions of violence. “bad language” Contains inappropriate language. “fear” Contains images or sounds that may be frightening or scary. “gambling” Contains elements that encourage or teach gambling. “sex” Contains sexual posture. “drugs” Contains illustrations of the use of illicit drugs, alcohol, or tobacco. “discrimination” The game contains depictions of ethnic, religious, nationalist, or other stereotypes that may encourage hatred. “in-game_ .purchases” Offers the option to purchase digital services.

[066] Table 11 defines the capabilities of an avatar / object. Table 11: Semantics of capabilities.______________________ Abilities Description “walk” The ability to walk. “run” The ability to run. “jump” The ability to jump. Petition 870250084121, dated 09 / 18 / 2025, pages 68 / 108 22 / 45 “fly” The ability to fly. “swim” The ability to swim. “climb” The ability to go over, climb, ascend and descend objects, such as climbing onto a chair, climbing stairs, descending stairs, scaling a wall, etc. “grasp” The ability to grasp objects with hand-like representations. “manipulate” The ability to interact with and alter the spatial position of 3D objects using collision or proximity detectors. “ride” The ability to pilot a vehicle or animal, such as motorcycles or horses. “drive” The ability to use a vehicle, such as cars or trucks, etc. “pilot” The ability to pilot a vehicle, such as ships or airplanes.

[067] Table 12 defines the deficiencies of an avatar / object. Table 12: Semantics of the deficiencies. Disabilities Description “Cerebral palsy” A group of disorders that affect a person's ability to move and maintain balance. “Spinal cord injuries” Spinal cord injury indicates damage to any part of the spinal cord or nerves at the end of the spinal canal. Results in permanent loss of strength, sensation, and function (mobility and sensitivity). “Amputation” Indicates the removal of part or all of a body part that is covered by the skin. “Musculoskeletal injuries” Refers to damage to the muscular or skeletal systems, usually caused by strenuous activities. Petition 870250084121, dated 09 / 18 / 2025, pp. 69 / 108 23 / 45 "Hearing loss" refers to the loss of hearing ability. This avatar will need visual cues and text-to-speech replacement for guidance. "Vision impairment" refers to the loss or inability to see. This avatar will primarily need audio cues for guidance.

[068] Table 13 illustrates the semantic description of the new behavior property (the duration of a behavior). Table 13: Semantic description of new behavioral properties. Name Type Usage Description Duration Number The duration in seconds of a behavior. The default value is "-1" for infinite duration, greater than "0" to set the action's lifespan if it's a time-based event; otherwise, it will be ignored.

[069] Examples of glTF scheme

[070] The following glTF is an example of an instantiation of an “MPEG_scene_interactivity” action extension on clients that support “MPEG_scene_interactivity”. Each example illustrates a simplistic scenario, assuming that nodes or node avatars have metadata available to enable permission flags or capabilities.

[071] Note that a large number of instantiations are possible, depending on the application. Here we give several examples for illustrative purposes. Example 1 { Scene: 0, we:[ ________{____________________________________________________________________________________________________________________ Petition 870250084121, dated 09 / 18 / 2025, pp. 70-108 24 / 45 mesh : 0, name : Box_Yellow, translation : [ -10, 0, ] }, { mesh : 1, name : Box_Red, translation : [ 10, 0, ] }, { Mesh: 2, Name: avatar, Translation: [ 3, 0, ], extensions: { MPEG node avatar: { Petition 870250084121, dated 09 / 18 / 2025, pp. 71 / 108 25 / 45 isAvatar: True} }}, ], scenes: [ { extensions: { MPEG_scene_interactivity: { triggers: [ { type : TRIGGER_PROXIMITY, distanceLowerLimit : 0.0, distanceUpperLimit : 1.0, nodes : [0,2]}, { type : TRIGGER_PROXIMITY, distanceLowerLimit : 0,0, distanceUpperLimit : 1,0, nodes : [1,2]} ], actions: [ { type: ACTION RESTRICTED, Petição 870250084121, de 18 / 09 / 2025, pág. 72 / 108 26 / 45 activationStatus: 0, nodes: [2], permission_id: 123654789}, { type: ACTION_SOCIAL, nodes: [2], authorised: [conversation, interaction]} ], behaviors: [ { triggers: [0], actions: [0], triggersCombinationControl: #1, triggersActivationControl: TRIGGER_ACTIVATE_FIRST_ENTER, actionsControl: 0, priority: 1}, { triggers: [1], actions: [1], triggersCombinationControl: #1, triggersActivationControl: TRIGGER_ACTIVATE_FIRST_ENTER, actionsControl: 0, priority: 1} Petição 870250084121, de 18 / 09 / 2025, pág. 73 / 108 27 / 45 ] }} } ] }

[072] In Example 1, there are three nodes, and the indexed node 2 with the name “avatar” represents an avatar because it has an extension “isAvatar” set to “True”.

[073] In this example, when a node representing an avatar approaches within 0.0 and 1.0 units of node 0 (“Box_Yellow”) or 1 (“Box_Red”), proximity triggering is activated and an action is performed. In this example, we illustrate two behaviors, and each example is defined in the behaviors section. Each behavior will link a trigger and an action. Behavior 0 checks if the “avatar” node is close to the “Box_Yellow” node. This is set by the “nodes” field in proximity triggering (nodes: [0,2], which represent the indices of the “Box_Yellow” and “avatar” nodes), and if this condition is “True” the action with index “0” is initialized. This refers to “ACTION_RESTRICTED”, and this action will check if the avatar has the necessary permission to enter or interact with this node.

[074] The second behavior has exactly the same condition as the first, but the trigger is in the “Box_Red”, and the action is to enable “conversation” and “interaction” within the box. This behavior will occur when the object enters the predefined proximity for the first time. To disable actions or permissions, a different behavior needs to be set. Example 2 { scene: 0, Petition 870250084121, dated 09 / 18 / 2025, pp. 74 / 108 28 / 45 nodes:[ { mesh : 0, name : Box_Yellow, translation : [ -10, 0, ] }, { mesh : 2, name : avatar, translation : [ 3, 0, ], extensions: { MPEG_node_avatar: { isAvatar: True} }}, ], scenes: [ _J_________________ Petition 870250084121, dated 09 / 18 / 2025, pp. 75 / 108 29 / 45 extensions: { MPEG_scene_interactivity: { triggers: [ { type : TRIGGER_PROXIMITY, distanceLowerLimit : 0.0, distanceUpperLimit : 1.0, nodes : [0,1]} ], actions: [ { type: ACTION_PARENTAL, activationStatus: 0, nodes: [1], age: 3, Descriptors: [in-game_purchases]} ], behaviors: [ { triggers: [0], actions: [0], triggersCombinationControl: #1, triggersActivationControl: TRIGGER_ACTIVATE_FIRST_ENTER, actionsControl: 0, priority: 1 Petition 870250084121, dated 09 / 18 / 2025, pp. 76 / 108 30 / 45 ]} }} ]}

[075] In Example 2, there are two nodes, and the indexed node 1 named “avatar” represents an avatar. The authorization and control of each node must be handled on the engine side and not on the scene description side. When a node representing an avatar approaches within 0.0 and 1.0 units of distance from node 0 (“Box_Yellow”), proximity triggering is activated and an action is executed.

[076] In this scenario, we illustrate an example that defines a single behavior. The behavior will link a trigger and an action. Behavior 0 checks if the “avatar” node is near the “Box_Yellow” node. This is set by the “nodes” field in the proximity trigger (nodes: [0,1], which represent the indices of the “Box_Yellow” and “avatar” nodes), and if this condition is “True” the action with index “0” is initialized. This refers to “ACTION_PARENTAL”, and this action will signal to the user the content type and minimum age required to check if the avatar has the necessary permission to enter or interact with this node.

[077] This behavior will occur when the object first enters the vicinity. To disable actions or permissions, a different behavior needs to be configured. Example 3 { scene: 0, Petition 870250084121, dated 09 / 18 / 2025, pp. 77 / 108 31 / 45 nodes:[ { mesh : 0, name : Box_Yellow, translation : [ -10, 0, ] }, { mesh : 2, name : avatar, translation : [ 3, 0, ], extensions: { MPEG_node_avatar: { isAvatar: True} }}, ], scenes: [ _J_________________ Petição 870250084121, de 18 / 09 / 2025, pág. 78 / 108 32 / 45 extensions: { MPEG_scene_interactivity: { triggers: [ { type : TRIGGER_PROXIMITY, distanceLowerLimit : 0,0, distanceUpperLimit : 1,0, nodes : [0,1]} ], actions: [ { type: ACTION_SPEECH, activationStatus: 0, nodes: [1], microphone: True, duration: 180} ], behaviors: [ { triggers: [0], actions: [0], triggersCombinationControl: #1, TRIGGER_ACTIVATE_FIRST_ENTER, actionsControl: 0, triggersActivationControl: Petição 870250084121, de 18 / 09 / 2025, pág. 79 / 108 33 / 45 priority: 1} ]} }} ]}

[078] In Example 3, there are two nodes, and the indexed node 1 named “avatar” represents an avatar. The authorization and control of each node must be handled on the engine side and not on the scene description side. When a node representing an avatar approaches within 0.0 and 1.0 units of node 0 (“Box_Yellow”), proximity triggering is activated and an action is performed.

[079] In this scenario, we illustrate an example that defines a single behavior. The behavior will link a trigger and an action. Behavior 0 checks if the “avatar” node is close to the “Box_Yellow” node. This is set by the “nodes” field in the proximity trigger (nodes: [0,1], which represent the indices of the “Box_Yellow” and “avatar” nodes), and if this condition is “True” the action with index “0” is initialized. This refers to “ACTION_SPEECH”, and this action will signal to the application that this node avatar can use the microphone for a duration of 180 seconds. Example 4 { scene: 0, nodes:[ { mesh : 0, Petição 870250084121, de 18 / 09 / 2025, pág. 80 / 108 34 / 45 name : Box_Yellow, translation : [ -10, 0, ] }, { mesh : 2, name : avatar, translation : [ 3, 0, ], extensões: { MPEG_node_avatar: { isAvatar: True} }}, ], scenes: [ { extensions: { MPEG_scene_interactivity: { triggers: [ Petição 870250084121, de 18 / 09 / 2025, pág. 81 / 108 35 / 45 { type : TRIGGER_PROXIMITY, distanceLowerLimit : 0,0, distanceUpperLimit : 0,0, nodes : [0,1]} ], actions: [ { type: ACTION_CAPABILITIES, activationStatus: 0, nodes: [1], capabilities: [climb,ride,fly], duration: 240}, ], behaviors: [ { triggers: [0], actions: [0], triggersCombinationControl: #1, triggersActivationControl: TRIGGER_ACTIVATE_FIRST_ENTER, actionsControl: 0, priority: 1}, ____________________________]__________________________________________________________________________________________________________________________________________________________ Petition 870250084121, dated 09 / 18 / 2025, pp. 82-108 36 / 45 }} ]}

[080] Example 4 is similar to Example 3. The difference is that the node actions are triggered based on contact (lower limit = upper limit = 0,0). Additionally, in “ACTION_CAPABILITIES”, it defines new capabilities for the avatar (climb, walk, and fly) when proximity triggering is activated for a duration of 240 seconds (instead of 180 seconds in Example 3). Example 5 { scene: 0, we:[ { mesh : 0, name : Box_Yellow, translation : [ -10, 0, ] }, { Mesh: 2, Name: avatar, Petição 870250084121, de 18 / 09 / 2025, pág. 83 / 108 37 / 45 translation : [ 3, 0, ], extensions: { MPEG_node_avatar: { isAvatar: True} }}, ], scenes: [ { extensions: { MPEG_scene_interactivity: { triggers: [ { type : TRIGGER_PROXIMITY, distanceLowerLimit : 0,0, distanceUpperLimit : 1,0, nodes : [0,1]} ], actions: [ { type: ACTION DISABILITIES, Petição 870250084121, de 18 / 09 / 2025, pág. 84 / 108 38 / 45 activationStatus: 0, nodes: [1], disabilities: [Hearing loss]}, ], behaviors: [ { triggers: [0], actions: [0], triggersCombinationControl: #1, triggersActivationControl: TRIGGER_ACTIVATE_FIRST_ENTER, actionsControl: 0, priority: 1}, ] }} } ] }

[081] Example 5 is similar to Example 4. The difference is that in “ACTION_DISABILITIES”, it signals the disabilities available in the “Box_Yellow”, which notifies users that interactivity with this region will take “Hearing loss” into account and display the appropriate visual cues. Example 6 Petition 870250084121, dated 09 / 18 / 2025, pp. 85 / 108 39 / 45 { scene: 0, we:[ { mesh : 0, name : Box_Yellow, translation : [ -10, 0, ] }, { Mesh: 2, Name: avatar, Translation: [ 3, 0, ], extensions: { MPEG_node_avatar: { isAvatar: True} }}, ]_____________________________________________________________________________________ Petition 870250084121, dated 09 / 18 / 2025, pp. 86 / 108 40 / 45 scenes: [ { extensions: { MPEG_scene_interactivity: { triggers: [ { type : TRIGGER_PROXIMITY, distanceLowerLimit : 0,0, distanceUpperLimit : 1,0, nodes : [0,1]} ], actions: [ { type: ACTION_SET_AVATAR, activationStatus: 0, avatarAction: { type: ACTION_AVATAR_DISABILITIES, nodes: [1], disabilities: [Hearing loss]} } ], behaviors: [ { triggers: [0], actions: [0], Petição 870250084121, de 18 / 09 / 2025, pág. 87 / 108 41 / 45 triggersCombinationControl: #1, triggersActivationControl: TRIGGER_ACTIVATE_FIRST_ENTER, actionsControl: 0, priority: 1}, ] }} } ] }

[082] Example 6 is similar to Example 5. The difference lies in the use of “ACTION_SET_AVATAR” to specify “ACTION_AVATAR_DISABILITIES”. The “ACTION_SET_AVATAR” flag indicates that node 1 is an avatar node and the action is avatar-specific (Disability). Specifically, it flags the disabilities available in the “Box_Yellow”, notifying users that interactivity with this region will take into account “Hearing_Loss” and display appropriate visual cues.

[083] Figure 5 illustrates an example of executing a hierarchical action according to a modality. This example illustrates how actions can affect the activation of subsequent actions. As illustrated in Figure 5, for each trigger (510), we evaluate (520) whether the trigger activation conditions (e.g., proximity) are met in each scene update. If the trigger condition does not meet the conditions, the processing model continues to the next scene update without changes to the trigger or activation of actions. On the other hand, if the Petition 870250084121, dated 09 / 18 / 2025, pages 88 / 108 If 42 / 45 trigger conditions are met, the trigger is activated (550) and the action is initiated (560) if the action conditions (e.g., permission) are met (540). Once the action is initialized (560), we evaluate whether the action has child actions (530), if so, they are also evaluated (540) and initialized (560) if the conditions are met. Once all actions and the actions of their dependent children (560) are initialized, the application continues to the next scene update (570).

[084] Figure 6 illustrates the generation of scene-encoded parameters by an encoder (610), which uses a scene description file format as input and outputs a scene-representative encoded data format. In particular, Figure 7 illustrates the encoding diagram for the encoder, according to a modality. In particular, for the extended node “Interactivity” (710) of the node “Scene” (705), there are “Behaviors” (720) defining links between “Triggers” (730) and “Actions” (740). These “Actions” (740) are encoded (1) if a “Node” node (750, 760) is seen as an avatar (e.g., using the extended attribute “is_avatar”) and (2) if triggers are activated by the avatar (770). As a result, parameters such as “Action_Parental()” (780) and “Action_Speech()” (790) are generated.

[085] Several numerical values ​​are used in this application. The specific values ​​are for example purposes and the aspects described are not limited to those specific values.

[086] Several methods are described here, and each of the methods comprises one or more steps or actions to achieve the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. In addition, terms such as “first”, “second”, etc. may be used in various ways to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Petition 870250084121, dated 09 / 18 / 2025, page 89 / 108 43 / 45 decoding”. The use of such terms does not imply an ordering of the modified operations, unless specifically required. Therefore, in this example, the first decoding does not need to be performed before the second decoding and may occur, for example, before, during, or at a time overlapping with the second decoding.

[087] The implementations and aspects described herein can be implemented in, for example, a method or a process, a device, a software program, a continuous data stream, or a signal. Even if discussed only in the context of a single implementation form (for example, discussed only as a method), the implementation of the attributes discussed can also be implemented in other forms (for example, a device or program). A device can be implemented, for example, in appropriate hardware, software, and firmware. Methods can be implemented in, for example, a device, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device.Processors also include communication devices, for example, computers, cell phones, personal / portable digital assistants (PDAs), and other devices that facilitate the communication of information between end users.

[088] The reference to “a modality” or “a modality” or “an implementation” or “an implementation”, as well as other variations thereof, means that a particular attribute, structure, characteristic, and so forth described in connection with the modality are included in at least one modality. Thus, the appearances of the phrase “in a modality” or “in a modality” or “in an implementation” or “in an implementation”, as well as any other variations, that appear in various places throughout this application are not necessarily all referring to the same modality. Petition 870250084121, dated 09 / 18 / 2025, pp. 90-108 44 / 45

[089] In addition, this request may refer to the “determination” of various information. Determining the information may include one or more of the following: for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.

[090] In addition, this request may refer to “access” to various pieces of information. Access to information may include one or more of the following: for example, receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[091] Furthermore, this request may refer to the “receiving” of various pieces of information. Receiving, as well as “accessing,” is intended to be a broad term. Receiving information may include one or more of the following: for example, accessing the information or retrieving it (e.g., from memory). In addition, “receiving” is typically involved, in one form or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, deleting the information, calculating the information, determining the information, predicting the information, or estimating the information.

[092] It should be understood that the use of any of the following “ / ”, “and / or”, and “at least one of”, for example, in the cases of “A / B”, “A and / or B” and “at least one of A and B”, is intended to cover the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B and / or C” and “at least one of A, B and C”, such formulation is intended to cover the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and second listed options. Petition 870250084121, dated 09 / 18 / 2025, pp. 91-108 45 / 45 (A and B) only, or selecting the first and third listed options (A and C) only, or selecting the second and third listed options (B and C) only, or selecting all three options (A, B, and C). This can be extended, as is clear to anyone versed in the technique and related techniques, to as many items as are listed.

[093] As will be evident to anyone skilled in the art, implementations can produce a variety of formatted signals to carry information that can, for example, be stored or transmitted. The information may include, for example, instructions for performing a method or data produced by one of the described implementations. For example, a signal may be formatted to carry the bitstream of a described mode. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using a radio frequency portion of the spectrum) or as a baseband signal. The formatting may include, for example, encoding a continuous data stream and modulating a carrier with the encoded continuous data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of wired or wireless links, as is known.The signal can be stored on a processor-readable medium. Petition 870250084121, dated 09 / 18 / 2025, pp. 92 / 108

Claims

1 / 5 CLAIMS 1. A method, CHARACTERIZED in that it comprises: obtaining, from a scene description for an extended reality scene, at least one parameter used to define one or more allowed actions for an avatar node representing an avatar; activating a trigger for an action associated with said avatar node, wherein said action belongs to one or more allowed actions; and initializing said action for said avatar node.

2. Method, according to claim 1, CHARACTERIZED in that one or more permitted actions include at least one of the following types: - adjusting the action of said avatar node, - adjusting restrictions of said avatar node, - adjusting parental and content usage permissions of said avatar node, - adjusting the permitted speech activity of said avatar node, - adjusting the capabilities of said avatar node, and - adjusting deficiencies of said avatar node.

3. Method, according to claim 1, CHARACTERIZED in that at least one parameter indicates the capabilities of the avatar.

4. Method according to claim 3, or apparatus defined in claim 3, CHARACTERIZED in that said capabilities include at least one of the following: - the ability to walk, - the ability to run, - the ability to jump, - the ability to fly, - the ability to swim, - the ability to pass over, climb, scale, descend objects, - the ability to hold objects with the hands as representations, Petition 870250084121, dated 09 / 18 / 2025, p. 104 / 108 2 / 5 - the ability to interact with and alter the spatial position of 3D objects using collision or proximity detectors, - the ability to ride a vehicle or animal, - the ability to use a vehicle, and - the ability to pilot a vehicle.

5. Method, according to claim 1, CHARACTERIZED in that at least one parameter indicates deficiencies of said avatar.

6. Method, according to claim 1, CHARACTERIZED in that at least one parameter indicates a content type of a list of nodes.

7. Method, CHARACTERIZED by the fact that it comprises: generating at least one parameter in a scene description for an extended reality scene to define one or more allowed actions for an avatar node representing an avatar; associating a trigger with an action with said avatar node, where said action belongs to one or more allowed actions; and encoding said scene description for said extended reality scene.

8. Method, according to claim 7, CHARACTERIZED in that said one or more permitted actions include at least one of the following types: - adjusting the action of said avatar node, - adjusting restrictions of said avatar node, - adjusting parental and content usage permissions of said avatar node, - adjusting the permitted speech activity of said avatar node, - adjusting the capabilities of said avatar node, and - adjusting the deficiencies of said avatar node.

9. Method, according to claim 7, CHARACTERIZED in that said at least one parameter indicates the capabilities of said avatar. Petition 870250084121, dated 09 / 18 / 2025, pp. 105 / 108 3 / 5 10. Method, according to claim 7, CHARACTERIZED in that said at least one parameter indicates the deficiencies of said avatar.

11. Device, CHARACTERIZED by the fact that it comprises one or more processors and at least one memory, wherein said processor(s) is / are configured to: obtain, from a scene description for an extended reality scene, at least one parameter used to define one or more actions allowed for an avatar node representing an avatar; activate a trigger for an action associated with said avatar node, wherein said action belongs to one or more allowed actions; and initiate said action for said avatar node.

12. Device, according to claim 11, CHARACTERIZED in that the aforementioned one or more permitted actions include at least one of the following types: - adjusting the action of said avatar node, - adjusting restrictions of said avatar node, - adjusting parental and content usage permissions of said avatar node, - adjusting the permitted speech activity of said avatar node, - adjusting the capabilities of said avatar node, and - adjusting the deficiencies of said avatar node.

13. Device, according to claim 11, CHARACTERIZED in that at least one parameter indicates the capabilities of said avatar.

14. Apparatus, according to claim 11, CHARACTERIZED in that said capabilities include at least one of the following: - the ability to walk, - the ability to run, - the ability to jump, - the ability to fly, Petition 870250084121, dated 09 / 18 / 2025, p. 106 / 108 4 / 5 - the ability to swim, - the ability to pass over, climb, ascend and descend objects, - the ability to hold objects with manual representations, - the ability to interact with and alter the spatial position of 3D objects using collision or proximity detectors, - the ability to pilot a vehicle or animal, - the ability to use a vehicle and - the ability to pilot a vehicle.

15. Device, according to claim 11, CHARACTERIZED in that at least one parameter indicates deficiencies of said avatar.

16. Device according to claim 11, CHARACTERIZED in that at least one parameter indicates a content type of a list of nodes.

17. Device, CHARACTERIZED by the fact that it comprises one or more processors and at least one memory, wherein said processor(s) is / are configured to: generate at least one parameter in a scene description for an extended reality scene to define one or more allowed actions for an avatar node representing an avatar; associate a trigger with an action with said avatar node, wherein said action belongs to one or more allowed actions; and encode said scene description for said extended reality scene.

18. Device according to claim 17, CHARACTERIZED in that said one or more permitted actions include at least one of the following types: - adjusting the action of said avatar node, - adjusting restrictions of said avatar node, - adjusting parental and content usage permissions of said avatar node, Petition 870250084121, dated 09 / 18 / 2025, pp. 107 / 108 5 / 5 - adjusting the permitted speech activity of said avatar node, - adjusting the capabilities of said avatar node, and - adjusting the deficiencies of said avatar node.

19. Device according to claim 17, CHARACTERIZED in that at least one parameter indicates the capabilities of said avatar.

20. Device according to claim 17, CHARACTERIZED in that at least one parameter indicates avatar deficiencies.

21. Device, according to claim 17, CHARACTERIZED in that at least one parameter indicates a content type of a list of nodes. Petition 870250084121, dated 09 / 18 / 2025, pp. 108 / 108