Avatar actions and behavior in a virtual environment
The integration of semantic avatar actions and behaviors into XR environments through MPEG-I scene descriptions and glTF format addresses the lack of detailed avatar interactions, enabling more immersive and adaptable XR experiences.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- INTERDIGITALCE PATENT HLDG SAS
- Filing Date
- 2024-03-15
- Publication Date
- 2026-04-10
AI Technical Summary
Current extended reality (XR) technologies lack the ability to provide detailed semantic descriptions of avatar actions and behaviors in 3D virtual environments, particularly in social interactions, limiting immersive and interactive experiences.
The introduction of a new set of actions and behaviors for MPEG-I scene descriptions that support social interactions of avatars, using high-level semantic descriptions for avatar actions and behaviors, triggered by events in the virtual environment, and integrated with the glTF scene graph format to enhance interactivity.
Enables more immersive and interactive XR experiences by allowing avatars to perform human-readable actions and behaviors, such as walking, talking, and interacting, based on semantic triggers, enhancing user-specific interactions and adapting to individual user capabilities and disabilities.
Smart Images

Figure 2026511176000001_ABST
Abstract
Description
Technical Field
[0001] This embodiment generally relates to digital human interaction within a 3D virtual scene, and more particularly, to the actions and behaviors of avatars in a virtual environment.
Background Art
[0002] Extended Reality (XR) is a technology that enables an interactive experience in which the real-world environment and / or video content is enhanced by virtual content that can be defined across multiple sensory modalities including vision, hearing, touch, etc. During the execution of an application, the virtual content (e.g., 3D content or audio / video files) is rendered in real time in a way that matches the user's context (environment, viewpoint, device, etc.). A scene graph (e.g., as proposed by Khronos / glTF (Graphics Language Transmission Format), and its extensions defined in the MPEG scene description format or Apple / USDZ) is a way to represent the content to be rendered. They combine, on the one hand, a declarative description of the scene structure that links objects in the real environment with virtual objects, and on the other hand, a binary representation of the virtual content. The scene description framework ensures that the timed media and the corresponding associated virtual content are available at any time during the rendering of the application. Scene descriptions can also carry scene-level data that describes how a user can interact with scene objects at runtime for an immersive XR experience.
Summary of the Invention
[0003] According to one embodiment, a method is provided which includes: obtaining from a description of an augmented reality scene at least one parameter used to define one or more permitted actions for an avatar node representing an avatar; activating a trigger for an action associated with the avatar node, wherein the action belongs to the one or more permitted actions; and initiating the action for the avatar node.
[0004] According to another embodiment, a method is provided which includes generating at least one parameter in a description for an augmented reality scene to define one or more permitted actions for an avatar node representing an avatar, associating a trigger for an action with an avatar node, wherein the action belongs to the one or more permitted actions, and encoding a description for the augmented reality scene.
[0005] According to another embodiment, a device is provided comprising one or more processors and at least one memory, wherein the one or more processors are configured to take from a description of an augmented reality scene at least one parameter used to define one or more permitted actions for an avatar node representing an avatar, and to activate a trigger for an action associated with the avatar node, wherein the action belongs to the one or more permitted actions and initiates the action for the avatar node.
[0006] According to another embodiment, a device is provided comprising one or more processors and at least one memory, wherein the one or more processors are configured to perform the following: generate at least one parameter in a description of an augmented reality scene to define one or more permitted actions for an avatar node representing an avatar; associate a trigger for an action with the avatar node, wherein the action belongs to the one or more permitted actions; and encode the description of the augmented reality scene.
[0007] One or more embodiments also provide a computer program that, when executed by one or more processors, includes instructions causing one or more processors to perform a method according to any embodiment described herein. One or more embodiments also provide a computer-readable storage medium storing instructions for processing a scene description according to the method described herein.
[0008] One or more embodiments also provide a computer-readable storage medium storing scene descriptions generated according to the method described above. One or more embodiments also provide a method and apparatus for transmitting or receiving scene descriptions generated according to the method described herein. [Brief explanation of the drawing]
[0009] [Figure 1] This figure shows an exemplary architecture of an XR processing engine. [Figure 2] This figure shows an example of the syntax for a data stream that encodes an augmented reality scene description. [Figure 3] This figure shows an illustrative graph of an augmented reality scene description. [Figure 4] This figure shows an example of an augmented reality scene description that includes behavioral data. [Figure 5] An example of executing a hierarchical action according to one embodiment is shown. [Figure 6]This demonstrates the generation of parameters using scene coding. [Figure 7] This figure shows an encoder according to one embodiment. [Modes for carrying out the invention]
[0010] Various XR applications can be applied to different contexts and real or virtual environments. For example, in industrial XR applications, a virtual 3D content item (e.g., engine part A) is displayed when a reference object (engine part B) is detected in the real environment by a camera attached to a head-mounted display device. The 3D content item is placed in the real world at a defined position and scale relative to the detected reference object.
[0011] For example, in an XR application for interior design, a 3D model of furniture is displayed when a given image from a catalog is detected within the input camera view. The 3D content is placed in the real world at a defined position and scale relative to the detected reference image. In another application, an audio file can start playing when the user enters an area near a church (either real or virtually rendered in an augmented reality environment). In yet another example, an ad jingle file may play when the user sees a given can of soda in the real world. In an outdoor game application, various virtual characters may appear depending on the semantics of the landscape observed by the user. For example, since bird characters are suitable for trees, if the sensors of the XR device detect a real object described by the semantic label "tree," birds flying around a tree can be added. In a companion application implemented with smart glasses, if a car is detected within the user's camera's field of view, the sound of a car may be emitted into the user's headset to warn the user of a potential hazard. Furthermore, the sound may be spatially adjusted to sound as if it is coming from the direction in which the vehicle was detected.
[0012] XR applications can also extend video content rather than the real environment. The video is displayed on a rendering device, and virtual objects described within a node tree are overlaid when timed events are detected within the video. In such a context, the node tree contains only descriptions of virtual objects.
[0013] Figure 1 shows an exemplary architecture of an XR processing engine 130 that may be configured to carry out the methods described herein. A device according to the architecture of Figure 1 is connected to other devices via a bus 131 and / or an I / O interface 136.
[0014] Device 130 comprises the following elements, which are linked together by a data and address bus 131: - Microprocessor 132 (or CPU), for example, DSP (Digital Signal Processor), - ROM(Read Only Memory)133, - RAM (Random Access Memory) 134, - Storage interface 135, - I / O interface 136 for receiving data sent from the application. - Power source (not shown in Figure 1), for example, a battery.
[0015] For example, the power supply is located outside the device. In each of the memories mentioned herein, the term “register” as used herein may refer to a small area (a few bits) or a very large area (e.g., an entire program or a large amount of received or decoded data). ROM 133 contains at least the program and parameters. ROM 133 may also store algorithms and instructions for performing techniques based on the principles of the present invention. When power is applied, the CPU 132 uploads the program to RAM and executes the corresponding instructions.
[0016] RAM134 contains, in its registers, a program executed by CPU132 and uploaded after device 130 is powered on, input data in its registers, intermediate data for different states of the method in its registers, and other variables used to execute the method in its registers.
[0017] Device 130 is linked to a set of sensors 137 and a set of rendering devices 138, for example, via a bus 131. The sensors 137 may be, for example, a camera, microphone, temperature sensor, inertial measurement unit, GPS, humidity sensor, infrared or ultraviolet sensor, or wind sensor. The rendering devices 138 may be, for example, a display, speaker, vibrator, heater, fan, etc.
[0018] According to some examples, device 130 is configured to execute a method based on this principle and belongs to a set including the following. - Mobile device, - Communication device, - Gaming device, - Tablet (or tablet computer), - Laptop, - Still camera, - Video camera.
[0019] In an XR application, a scene description is used to combine an explicit and syntactically easy-to-analyze description of the scene structure with some binary representations of media content. FIG. 2 shows an example of the syntax of a data stream encoding an extended reality scene description. FIG. 2 shows an exemplary structure 210 of an XR scene description. This structure is composed of containers that organize the stream into independent elements of syntax. This structure can include a header portion 220 that is a set of data common to all syntax elements of the stream. For example, the header portion includes some metadata regarding syntax elements that describe the nature and role of each syntax element. This structure also includes a payload including syntax element 230 and syntax element 240. Syntax element 230 includes data representing media content items described in the nodes of the scene graph related to virtual elements. Images, meshes, and other raw data may be compressed according to a compression method. Syntax element 240 is part of the payload of the data stream and includes data encoding the scene description explained according to this principle.
[0020] Figure 3 shows an exemplary graph 310 of an extended reality scene description. In this example, the scene graph can include descriptions of real objects such as, for example, a "horizontal plane" (which can be a table or a road), and descriptions of virtual objects 312 such as, for example, an animation of a car. The scene description is organized as an array of nodes. The nodes can be linked to child nodes to form a scene structure 311. A node can have a description of a real object (e.g., a semantic description) or a description of a virtual object. In the example of Figure 3, node 301 describes a virtual camera located within the 3D space of the XR application. Node 302 describes a virtual car and includes an index of the representation of the car, e.g., an index within an array of 3D meshes. Node 303 is a child node of node 302 and includes a description of one wheel of the car. Similarly, it includes an index to the 3D mesh of the wheel. Since the scale, position, and orientation of the object are described in the scene node, the same 3D mesh can be used for several objects within the 3D scene. The scene graph 310 also includes nodes that are descriptions of the spatial relationships between real objects and virtual objects.
[0021] In time-based media streaming, by changing the scene description itself over time, virtual content related to each sequence of the media stream can be provided. For example, for advertising purposes, a virtual bottle can be displayed on a table during a video sequence where people are sitting around the table. This kind of behavior can be achieved by relying on the framework defined in the scene description of the MPEG media document.
[0022] Currently, the MPEG-I scene description framework extends time-varying scene descriptions using "behavior" data to provide a description of how users can interact with scene objects at runtime for immersive XR experiences. These behaviors are associated with predefined virtual objects for which runtime interactivity is permitted for user-specific XR experiences. These behaviors also change over time and are updated through existing scene description update mechanisms.
[0023] Figure 4 shows an example of an augmented reality scene description that includes scene-level stored behavior data, describing how a user can interact with node-level described scene objects when running an immersive XR experience. When an XR application starts, media content items (e.g., meshes of virtual objects visible to the camera) are loaded, rendered, and buffered so that they are displayed when triggered. For example, if a sensor detects a plane in the real environment, the application displays the buffered media content item as described in the relevant scene node. Timing is managed by the application according to the timing of features detected in the real environment and animations. Nodes in the scene graph may not contain descriptions and may only act as parents of child nodes. Figure 4 shows the relationship between the behaviors included in the scene-level scene description and the nodes that are components of the scene graph. Behavior 410 relates to a predefined virtual object that enables runtime interactivity for a user-specific XR experience. Behavior 410 also changes over time and is updated via a scene description update mechanism.
[0024] The behavior includes the following elements: - Trigger 420 defines the conditions that must be met for the behavior to be activated. - Trigger control parameters that define logical operations between defined triggers, - Action 430 that is processed when the trigger is activated, - Action control parameters that define the execution order of related actions. - A priority number that allows the selection of the highest priority behavior when several behaviors conflict simultaneously on the same virtual object. - An optional interruption action that specifies how to terminate this behavior if it becomes undefined in a newly received scene update. For example, if the relevant object is no longer present in the new scene, or if this behavior is no longer relevant to the current media (audio or video) sequence, the behavior is no longer defined.
[0025] Behavior 410 is executed at the scene level. Triggers are linked to nodes and their child nodes. In the example in Figure 4, trigger 1 is linked to nodes 1, 2, and 8. Since node 31 is a child of node 1, trigger 1 is linked to node 31. Trigger 1 is also linked to node 14 as a child of node 8. Trigger 2 is linked to node 1. In fact, the same node may be linked to several triggers. Trigger n is linked to nodes 5, 6, and 7. A single behavior may contain several triggers. For example, the first behavior may be activated by trigger 1 AND trigger 2, where AND is the trigger control parameter for the first behavior. A single behavior may have several actions. For example, the first behavior may first perform action m and then action 1, where "first and then" is the action control parameter for the first behavior. The second behavior is activated by trigger n, and can, for example, perform action 1 first, then action 2.
[0026] Node trees can be represented using different formats. For example, an MPEG-I scene description framework using the Khronos glTF extension mechanism can be used for a node tree. In this example, the interactivity extension is applied at the glTF scene level and is called "MPEG_scene_interactivity". The corresponding semantics are provided in Table 1, where "M" in the "Usage" column indicates that the field is required in the XR scene description format, and "O" indicates that the field is optional.
[0027] [Table 1]
[0028] This specification introduces a new set of actions, for example, for MPEG-I scene descriptions, to support the social interaction of avatars in a 3D environment with corresponding time-based events. These actions are triggered by events between 3D objects in the virtual environment, such as dynamic objects (humanoid and non-humanoid characters, cars, airplanes) or static objects (chairs, tables, plates). Current interactivity descriptions at the scene level only support general actions for nodes in a scene, as shown in Table 2. However, these descriptions do not provide avatar nodes with semantic information about human social behavior, which is a common attribute in the descriptions of avatars and real-world users. For example, the action of walking is semantically described as the word "walk" or the ability to "walk," not as a chain of matrix transformations describing the action of walking. At a higher level, describing actions using semantic information rather than low-level computer-readable 4x4 matrix multiplication is clearer and more human-readable. Therefore, this specification presents high-level action descriptions that can have a significant impact in 3D social and interactive environments.
[0029] [Table 2]
[0030] The proposed representation of user capabilities is intended to be compatible with scene description (SD) content and primarily focuses on representing avatar actions and behaviors in the interactive domain. Below, we provide details on the relevant semantic elements, JSON (JavaScript Object Notation) encoding schemes, and how they can be used within MPEG-I SD.
[0031] In its current description, the proposed format follows the glTF format and is compatible with the current MPEG effort to extend glTF with MPEG extensions. However, its meaning and use are general-purpose and can be coded in any other format, such as XML and USD.
[0032] Capability
[0033] Here, we introduce the available actions and behaviors that an avatar can perform within an interactive region represented as a geometric primitive. The interactive region surrounding an avatar can indicate what actions such an avatar or 3D scene object can perform, and thus can define capabilities in the context of this specification. As previously mentioned, behavior is a set of conditions that combine triggered events with specific actions, define the temporal constraints of such conditions, and enable time-based events to occur in a 3D virtual environment.
[0034] Avatar Capability
[0035] Here, we present examples of capabilities related to interactivity in the context of social interaction between avatars and 3D objects within a 3D virtual environment.
[0036] Generally, capabilities describe the permitted actions of an avatar following a trigger event in the realm of interactivity. For example, in a conference room, spectating avatars are only permitted to use actions that describe speech, as well as upper body movements such as gestures and head movements, or the avatar's capabilities (e.g., the ability to run, walk, jump, talk, and fly). If "disabilities" are defined for a user, the information contained in the "disabilities" also affects the capabilities of the user's avatar.
[0037] Below are some non-exclusive examples of actions associated with avatars in a social environment.
[0038] Social action corresponds to a user's social behavior. This can be generic (allowing default conversation and interactivity) or specific, as defined by the user. When interacting with another user, if a trigger is detected within the interactive area (provided by proximity or collision triggers), this area is designated for social interaction, allowing, for example, conversation between avatar users.
[0039] Restricted action. For example, in scenarios where permission is required for reasons such as age, restrictions, or access rights, you can restrict the user's permitted movement. This action can also be used to restrict the space in which an avatar is allowed to move.
[0040] Parental action allows you to restrict interaction with permitted content to protect children and young adults.
[0041] Speech action. This type of ability can be restricted to trigger events that only allow the speech action to be performed; for example, in a conference room, spectator avatars may only be allowed to speak. This action allows the use of a microphone or a pre-recorded media track. This action can be used in combination with social actions that grant permission for certain types of interactivity, such as speaking.
[0042] Capabilities action. This action lists the types of abilities that an avatar or 3D object is allowed to perform when it comes into contact with an area where the action is activated. Different types of abilities should cover different types of activities, such as walking, flying, driving, and talking, but not limited to these. Such a list of actions is communicated to the engine and can be combined with other action modules, such as "Action_set_haptics" to enable haptic feedback, or "Action_manipulate" for grasping objects. The purpose is to generate an action modifier that restricts the avatar's animation to the provided list of capabilities actions.
[0043] Disabilities actions. Disabilities have the same effect as abilities, but are designed to notify the engine of a user's disability, and as a result, affect the list of ability actions depending on the user's choice. This information list is important for adapting the virtual environment to the needs of individual users. For example, a user with a hearing impairment should have visual cues instead of audio cues.
[0044] All actions provided can be used in combination with existing and newly introduced actions, where permitted, and can leverage existing interactive and animation tools in various fields. For example, a manipulator can be used to perform "walking" movements, and a haptic manipulator can be used to infer sound / media tracks for tactile feedback or pre-recorded speech.
[0045] Avatar Behavior
[0046] Behavior is a set of parameters that define the matching between actions and trigger events. This combines newly defined actions, which have collision, proximity, or user input triggers, with time-based events. Time-based behavior makes it possible to set temporary actions in interactive areas and schedule actions in response to desired activities.
[0047] Time-based actions
[0048] Similar to time-based behavior, each action can define its own time span. This makes it easier to define the time for each action individually at the action level rather than the behavior level.
[0049] Avatar Action and Behavior Standards
[0050] The proposed actions and behaviors will be described below using MPEG-I as an example. In one embodiment, the actions and behaviors should comply with the following requirements: 1. The representation of actions and time-based behavior conforms to primitives supported in MPEG-I scene descriptions and other available formats (XML, USD). 2. The interactive space is represented by primitives, triggers, actions, and behavior labels, which enables interaction between avatar representations, such as individual body parts or interactive areas of the avatar, and 3D objects in the scene. 3. Enable multiple interaction triggers with scenes, objects, and other avatars, and implement social, private, and interactive boundaries between any associated objects. 4. Time-based behavior determines the lifecycle of an action; unless otherwise specified, the event trigger and action duration are equal to the duration of the 3D scene.
[0051] Actions and Behaviors in MPEG-I Scene Description
[0052] By adding the aforementioned attributes to the MPEG-I scene description, the existing glTF node "MPEG_scene_interactivity" element is extended.
[0053] The MPEG interactivity glTF extension enables event triggering and behavior in collision and proximity situations; therefore, the proposed extension contributes to the extension of new actions and time-based attributes for trigger behavior / actions in the node "MPEG_scene_interactivity". This generic node implementation can also be applied to avatar representations to add interactions and time-based constraints to the avatar and its respective elements.
[0054] The applicant proposes an extension that enables glTF models to use and interact with humanoid characters (avatars) and any other objects. The applicant proposes extending the glTF scene element "MPEG_scene_interactivity" action property to define "ACTION_SET_AVATAR," which includes more avatar-related actions and time-based behavior constraints, as well as general scene and node-level actions.
[0055] Table 3 shows the new type of property added to the "MPEG_scene_interactivity" framework.
[0056] [Table 3]
[0057] [Table 4]
[0058] Under the newly proposed "ACTION_SET_AVATAR," there is an object called "avatarAction" that represents avatar-specific actions. Its semantic description is shown in Table 7. Table 5 shows a list of available avatar-specific actions.
[0059] [Table 5]
[0060] Table 6 shows the types of actions that can be added at the scene level and general node level of the interactive framework. If "ACTION_SET_AVATAR" is not available or the framework does not implement any type of avatar, the system can still use the actions proposed at the scene level or node level.
[0061] [Table 6]
[0062] A newly proposed semantic description of the action is provided in Table 7.
[0063] [Table 7-1]
[0064] [Table 7-2]
[0065] Table 8 shows the action types when the action type is Action_Social.
[0066] [Table 8]
[0067] Table 9 defines the minimum age recommendation for a given user with a list of node contents.
[0068] [Table 9]
[0069] Table 10 describes the additional explicit semantics of the content present in the list of nodes.
[0070] [Table 10]
[0071] Table 11 defines the capabilities of the avatar / object.
[0072] [Table 11]
[0073] Table 12 defines the defects of avatars / objects.
[0074] [Table 12]
[0075] Table 13 shows the semantic description of the new behavior property (duration of behavior).
[0076] [Table 13]
[0077] Examples of glTF schemes
[0078] The following glTFs are examples of instantiations of the "MPEG_scene_interactivity" action extension in a client that supports "MPEG_scene_interactivity". Each example shows a simplified scenario that assumes the node or node avatar has metadata available to grant permission or capability flags.
[0079] Note that a large number of instances can be created depending on the application. Here are some examples for illustrative purposes.
[0080] [Table 14-1]
[0081] [Table 14-2]
[0082] In Example 1, there are three nodes, and the node at index 2, which has the name "avatar," represents an avatar because it has the extension "isAvatar" set to "True."
[0083] In this example, when the node representing the avatar comes within a distance unit of 0.0 to 1.0 of either node 0 ("Box_Yellow") or node 1 ("Box_Red"), the proximity trigger is activated and an action is performed. This example demonstrates two behaviors, each defined in the Behavior section. Each behavior links a trigger to an action. Behavior 0 verifies whether the node "avatar" is close to the node "Box_Yellow". This is set by the "nodes" field in the proximity trigger (nodes:[0,2] representing the node indices for "Box_Yellow" and "avatar"), and if this condition is "True", the action with index "0" is invoked. This is "ACTION_RESTRICTED", and this action verifies whether the avatar has the necessary permissions to enter or interact with this node.
[0084] The second behavior has exactly the same conditions as the first behavior, but the trigger is on "Box_Red," and the action enables "conversation" and "interaction" within the box. This behavior occurs when the object first enters a given proximity. Different behaviors must be set to disable the action or permission.
[0085] [Table 15-1]
[0086] [Table 15-2]
[0087] In Example 2, there are two nodes, with the node at index 1, named "avatar," representing the avatar. Authentication and control of each node must be handled on the engine side, not on the scene description side. When the node representing the avatar comes within a distance unit of 0.0 to 1.0 of node 0 ("Box_Yellow"), the proximity trigger is activated and an action is performed.
[0088] This scenario provides an example of defining a single behavior. A behavior links a trigger and an action. Behavior 0 verifies whether the node "avatar" is close to the node "Box_Yellow". This is set by the "nodes" field in the proximity trigger (nodes: [0,1] representing the indices of the "Box_Yellow" and "avatar" nodes), and if this condition is "True", the action with index "0" is initiated. This is "ACTION_PARENTAL", and this action signals the user the type of content and minimum age required to verify whether the avatar is allowed to enter this node or interact with this node, including the necessary permissions.
[0089] This behavior occurs when an object first enters a neighborhood. Different behavior must be configured to disable the action or permission.
[0090] [Table 16-1]
[0091] [Table 16-2]
[0092] In Example 3, there are two nodes, with the node at index 1, named "avatar," representing the avatar. Authentication and control of each node must be handled on the engine side, not on the scene description side. When the node representing the avatar comes within a distance unit of 0.0 to 1.0 of node 0 ("Box_Yellow"), the proximity trigger is activated and an action is executed.
[0093] This scenario illustrates an example of defining a single behavior. A behavior links a trigger and an action. Behavior 0 verifies whether the node "avatar" is in proximity to the node "Box_Yellow". This is set by the "nodes" field in the proximity trigger (nodes: [0,1] representing the indices of the "Box_Yellow" and "avatar" nodes), and if this condition is "True", the action with index "0" is initiated. This indicates "ACTION_SPEECH", and this action signals to the application that this node avatar can use the microphone for a duration of 180 seconds.
[0094] [Table 17-1]
[0095] [Table 17-2]
[0096] Example 4 is similar to Example 3. The difference is that actions on nodes are triggered based on contact (lower limit = upper limit = 0.0). Also, in "ACTION_CAPABILITIES", a proximity trigger is activated for a duration of 240 seconds (instead of 180 seconds in Example 3) to set new avatar abilities (climb, ride, and fly).
[0097] [Table 18-1]
[0098] [Table 18-2]
[0099] Example 5 is similar to Example 4. The difference is that "ACTION_DISABILITIES" signals available disabilities in "Box_Yellow" and notifies the user that interactivity with this area will take "Hearing_loss" into consideration and show appropriate visual cues.
[0100] [Table 19-1]
[0101] [Table 19-2]
[0102] Example 6 is similar to Example 5. The difference is that it uses "ACTION_SET_AVATAR" to specify "ACTION_AVATAR_DISABILITIES". The "ACTION_SET_AVATAR" signaling indicates that node 1 is an avatar node and the action is an avatar-specific action (disability). Specifically, it signals the available disabilities in "Box_Yellow" and informs the user that interactivity with this area will take "Hearing_loss" into consideration and show the appropriate visual queuing.
[0103] Figure 5 shows an example of hierarchical action execution according to one embodiment. This example shows how an action may affect the activation of subsequent actions. As shown in Figure 5, for each trigger (510), each time the scene is updated, it is evaluated whether the trigger activation condition (e.g., proximity) is met (520). If the conditional trigger does not meet the condition, the processing model proceeds to the next scene update without changing the trigger or activating the action. On the other hand, if the trigger condition is met, the trigger is activated (550), and if the action condition (e.g., permission) is met (540), the action is started (560). Once the action is started (560), it is evaluated whether the action has child actions (530), and if it does, the child actions are also evaluated (540), and if their conditions are met, they are started (560). Once all actions and their dependent child actions have started (560), the application proceeds to the next scene update (570).
[0104] Figure 6 shows the generation of parameters using scene encoding by an encoder (610) that takes a scene description file format as input and outputs an encoded data format representing the scene. In particular, Figure 7 shows the encoding of the encoder according to one embodiment. Specifically, for the extension node "Interactivity" (710) of the "Scene" node (705), there is "Behaviours" (720) that defines the link between "Trigger" (730) and "Actions" (740). These "Actions" (740) are encoded when (1) the "Node" nodes (750, 760) are recognized as avatars (for example, using the extension attribute "is_avatar"), and (2) the trigger is activated by the avatar (770). As a result, parameters such as "Action_Parental()" (780) and "Action_Speech()" (790) are generated.
[0105] Various numerical values are used in this application. Certain values are illustrative, and the embodiments described are not limited to these specific values.
[0106] Various methods are described herein, each of which includes one or more steps or actions to implement the described method. The order and / or use of any particular steps and / or actions can be modified or combined unless a specific order of steps or actions is required for the proper operation of the method. Furthermore, terms such as “first,” “second,” etc., may be used in various embodiments to modify elements, components, steps, operations, etc., such as “first decryption” and “second decryption.” The use of such terms does not imply a modified order of operations unless specifically requested. Therefore, in this example, the first decryption does not need to be performed before the second decryption, but may, for example, be performed before the second decryption, during the second decryption, or during a period overlapping with the second decryption.
[0107] The embodiments and aspects described herein can be implemented, for example, as methods or processes, apparatus, software programs, data streams, or signals. Even if an embodiment of a described feature is described only in the context of a single form (for example, discussed only as a method), it can still be implemented in other forms (for example, apparatus or programs). Apparatus can be implemented, for example, as appropriate hardware, software, and firmware. Methods can be implemented in apparatus such as processors, which generally refer to processing devices, including, for example, computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices, such as computers, mobile phones, portable / personal digital assistants (PDAs), and other devices that facilitate the communication of information between end users.
[0108] The phrases "one embodiment," "an embodiment," "one implementation," or "an implementation," as well as references to other variations thereof, mean that certain features, structures, characteristics, etc., described in relation to an embodiment are included in at least one embodiment. Therefore, the phrases "in one embodiment," "in one embodiment," "in one implementation," or "in an implementation," appearing in various places throughout this application, as well as any other variations, do not necessarily all refer to the same embodiment.
[0109] Furthermore, this application may refer to "determining" various types of information. Determining information may include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from memory.
[0110] Furthermore, this application may refer to "accessing" various types of information. Accessing information may include, for example, one or more of the following: receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0111] Furthermore, this application may refer to "receiving" various types of information. Receiving is intended to be a broad term, similar to "accessing." Receiving information may include, for example, accessing information or retrieving information (for example, from memory) one or more of these. Moreover, "receiving" typically involves, in some way, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information during the operation.
[0112] For example, in the cases of "A / B", "A and / or B", and "at least one of A and B", it should be understood that the use of " / ", "and / or", and "at least one of" is intended to encompass the selection of only the first enumerated option (A), or only the second enumerated option (B), or both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such phrasing is intended to encompass the selection of only the first enumerated option (A), or only the second enumerated option (B), or only the third enumerated option (C), or only the first and second enumerated options (A and B), or only the first and third enumerated options (A and C), or only the second and third enumerated options (B and C), or all three options (A, B, and C). This can be extended to the same number of items listed, as will be obvious to those skilled in the art in this and related fields.
[0113] As will be apparent to those skilled in the art, embodiments can generate a variety of signals formatted to carry information that can be stored or transmitted. This information may include, for example, instructions for performing a method or data generated by one of the embodiments described. For example, a signal may be formatted to carry a bitstream of the embodiment described. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is well known. The signal may be stored in a processor-readable medium.
Claims
1. From the augmented reality scene description, obtain at least one parameter used to define one or more permitted actions for the avatar node representing the avatar, Activating a trigger for an action associated with the avatar node, wherein the action belongs to one or more permitted actions, To initiate the aforementioned action for the aforementioned avatar node, A method that includes this.
2. To define one or more permitted actions for an avatar node representing an avatar, generate at least one parameter in the description of the augmented reality scene, Associating a trigger for an action with the avatar node, wherein the action belongs to one or more permitted actions, Encoding the description of the augmented reality scene, A method that includes this.
3. A device comprising one or more processors and at least one memory, The one or more processors described above are From the augmented reality scene description, obtain at least one parameter used to define one or more permitted actions for the avatar node representing the avatar, Activating a trigger for an action associated with the avatar node, wherein the action belongs to one or more permitted actions, To initiate the aforementioned action for the aforementioned avatar node, A device configured to perform [a certain action].
4. A device comprising one or more processors and at least one memory, The one or more processors described above are To define one or more permitted actions for an avatar node representing an avatar, generate at least one parameter in the description of the augmented reality scene, Associating a trigger for an action with the avatar node, wherein the action belongs to one or more permitted actions, Encoding the description of the augmented reality scene, A device configured to perform [a certain action].
5. The one or more permitted actions described above are of the following types: - Setting the action of the aforementioned avatar node, - Setting restrictions on the aforementioned avatar node, - Setting parental and content usage permissions for the aforementioned avatar node, - Setting the permitted speech actions for the avatar node, - Setting the capabilities of the aforementioned avatar node, - Setting a failure in the aforementioned avatar node, A method according to claim 1 or 2, or the apparatus according to claim 3 or 4, comprising at least one of the above.
6. The method according to any one of claims 1, 2, and 5, or the apparatus according to any one of claims 3 to 5, wherein the at least one parameter indicates the capabilities of the avatar.
7. The aforementioned ability is, - The ability to walk, - The ability to run, - The ability to jump, - The ability to fly, - Swimming ability, - The ability to overcome, climb onto, ascend to, and descend objects, - The ability to hold objects using hands and other means, - The ability to interact with 3D objects using collision or proximity detectors and change their spatial position, - The ability to ride a vehicle or animal, - The ability to use vehicles, - The ability to operate a vehicle, A method according to claim 6, or the apparatus according to claim 6, comprising at least one of the above.
8. The method according to any one of claims 1, 2, and 5 to 7, or the apparatus according to any one of claims 3 to 7, wherein the at least one parameter indicates a defect in the avatar.
9. The aforementioned malfunction is, - Cerebral palsy and, - Spinal cord injury and, - Cutting, - Musculoskeletal injuries, - Hearing impairment and, - Visual impairment and, A method according to claim 8, or the apparatus according to claim 8, comprising at least one of the above.
10. The method according to any one of claims 1, 2, and 5 to 9, or the apparatus according to any one of claims 3 to 9, wherein the at least one parameter indicates a limitation of the avatar.
11. The method according to any one of claims 1, 2, and 5 to 10, or the apparatus according to any one of claims 3 to 10, wherein the at least one parameter indicates the minimum recommended age for the content of the list of nodes.
12. The method according to any one of claims 1, 2, and 5 to 11, or the apparatus according to any one of claims 3 to 11, wherein the at least one parameter indicates the content type of the list of nodes.
13. The aforementioned content type is, - Violence and, - Inappropriate language and, - Fear, - Gambling and, - Sexual content, - Drugs and, - Discrimination and, - In-game purchases and, A method according to claim 12, or the apparatus according to claim 12, comprising at least one of the above.
14. A non-temporary computer-readable medium that, when executed by a computer, includes instructions causing the computer to perform the method according to any one of claims 1, 2, and 5 to 13.