Trigger activation mechanism in temporal evolution scene description

By introducing behavioral information and interaction logic into the XR scene description, the problem of lack of runtime interaction mechanism in the prior art is solved, and an immersive and personalized XR experience is achieved.

CN119948532APending Publication Date: 2025-05-06INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380068376.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-23
Filing Date
2023-09-18
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing Extended Reality (XR) scenario description framework lacks a mechanism to describe how users interact with scene objects at runtime, resulting in an immersive user-specific XR experience.

Method used

By introducing behavior information into the scene description, the logical relationship between triggers, actions and behavior is defined, allowing the activation of activities at runtime based on the user's interaction conditions to realize interaction with virtual objects or light sources.

Benefits of technology

It realizes runtime interaction in XR applications, allowing users to dynamically interact with virtual objects or light sources, and enhances the personalization and interactivity of the immersive XR experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948532A_ABST
    Figure CN119948532A_ABST
Patent Text Reader

Abstract

Some embodiments of a method may include obtaining scene description data for a 3D scene, the scene description data may include scene element information describing each of a plurality of scene elements in a scene, trigger information describing at least one trigger condition, action information describing an action to be performed on one or more scene elements associated with the action, and behavior information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions; and in response to determining that at least one of the trigger conditions has produced a result and the activation information indicates the result to excite the at least one trigger, performing the at least one action on the at least first node associated with the at least one action.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to European patent application No. EP22306405.6 filed on September 23, 2022, the entire contents of which are incorporated herein by reference.

[0003] This application incorporates by reference in their entirety the following applications: European Patent Application No. EP22305024.6, filed on January 12, 2022, entitled “METHODS AND DEVICES FOR INTERACTIVE RENDERING OF A TIME-EVOLVING EXTENDED REALITY SCENE” (the “’024 Application”); International Application No. PCT / EP2023 / 065281, filed on June 7, 2023 (the ’281 Application); and European Patent Application No. EP22305880.1, filed on June 16, 2022, entitled “SYSTEMS AND METHODS FOR PROVIDING INTERACTIVITY WITH LIGHT SOURCES IN A SCENEDESCRIPTION” (the “’880 Application”). Background Art

[0004] This section is intended to introduce the reader to various aspects of the art, which may be related to various aspects of the present principles described and / or claimed below. It is believed that this discussion is helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present principles. Therefore, it should be understood that these statements should be read from this perspective, and not as an admission of prior art.

[0005] Extended Reality (XR) is a technology that enables interactive experiences in which the real world environment and / or video content is augmented by virtual content, which can be defined across multiple sensory modalities, including vision, hearing, touch, etc. The virtual content (e.g., 3D content or audio / video files) is rendered in real time during the runtime of the application in a manner consistent with the user context (environment, viewpoint, device, etc.). Scene graphs (such as those proposed by Khronos / glTF and its extensions defined in the MPEG scene description format or Apple / USDZ) are possible ways to represent the content to be rendered. They combine, on the one hand, a declarative description of the scene structure linking real environment objects and virtual objects, and on the other hand, a binary representation of the virtual content. Although such a scene description framework ensures that timed media and the corresponding related virtual content are available at any time during the rendering of the application, there is no description of how the user can interact with the scene objects at runtime for an immersive XR experience.

[0006] There is a lack of XR systems that can employ XR scene descriptions that include metadata describing how users can interact with scene objects at runtime and how these interactions can be updated during runtime of the XR application. Summary of the invention

[0007] Embodiments described herein include methods used in video encoding and decoding (collectively referred to as "decoding").

[0008] A first example method according to some embodiments may include: acquiring scene description data of a 3D scene, wherein the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with at least one action, and behavior information, wherein the behavior information may include: a trigger list of at least one trigger, an action list of at least at least one action; trigger combination information and activation information, the trigger combination information indicating a combination of a first trigger condition corresponding to a first trigger in the trigger list and other trigger conditions corresponding to other triggers in the trigger list, the activation information indicating when to perform at least one action according to a result of the trigger combination information; and performing at least one action on one or more scene elements associated with at least one action in response to the following logic: (i) a combination of a first trigger condition corresponding to a first trigger in the trigger list and other trigger conditions corresponding to other triggers in the trigger list has produced a result, and (ii) activation information indicating that the result triggers the first trigger and the other triggers.

[0009] For some embodiments of the first example method, the trigger list may include one trigger, and there are no other triggers in the trigger list, the trigger combination information may include a unary operator operating on the one trigger, and the activation information may indicate when to execute the unary operator on the one trigger.

[0010] For some embodiments of the first example method, the trigger list may include at least two triggers, the trigger combination information may indicate how to combine the at least two trigger conditions, and the activation information may indicate when to perform at least one action according to a result of the trigger combination information.

[0011] For some embodiments of the first example method, performing the first action may include performing the first action one or more times based on the activation information.

[0012] For some embodiments of the first example method, determining a combination of a trigger condition of at least one trigger in the trigger list and trigger conditions of other triggers in the trigger list may include performing a logical operation on at least one of the trigger conditions.

[0013] For some embodiments of the first example method, determining a combination of a trigger condition of at least one trigger in the trigger list and trigger conditions of other triggers in the trigger list may include performing a logical OR operation of at least two trigger conditions.

[0014] For some embodiments of the first example method, at least a first one of the trigger conditions is a visibility condition that is satisfied when a specified scene element is visible to a specified camera node.

[0015] For some embodiments of the first example method, at least a first one of the trigger conditions is a proximity condition that is satisfied when a distance from a user camera to a specified scene element is within a specified boundary.

[0016] For some embodiments of the first example method, at least a first one of the trigger conditions is a user input condition that is satisfied when a specified user interaction is detected.

[0017] For some embodiments of the first example method, at least a first one of the trigger conditions is a timing condition that is satisfied during a specified time period.

[0018] For some embodiments of the first example method, at least a first one of the trigger conditions is a collider condition that is satisfied in response to detecting a conflict between specified scene elements.

[0019] For some embodiments of the first example method, the action information may describe at least two actions, and the behavior information may include information indicating, for at least one behavior, an order in which at least two of the at least two actions are to be performed.

[0020] For some embodiments of the first example method, the action information may describe at least two actions, and the behavior information may include information indicating that at least two of the at least two actions are to be performed simultaneously.

[0021] For some embodiments of the first example method, at least one of the scene elements in the scene is a virtual object.

[0022] Some embodiments of the first example method may further include rendering a 3D scene according to the scene description data operated by the first action.

[0023] For some embodiments of the first example method, the trigger information may include an array of two or more triggers in the 3D scene.

[0024] For some embodiments of the first example method, the action information may include an array of two or more actions in the 3D scene.

[0025] For some embodiments of the first example method, the behavior information may include an array of two or more behaviors in the 3D scene.

[0026] For some embodiments of the first example method, the scene description data may be provided in a JSON format.

[0027] For some embodiments of the first example method, the scene description data may be provided in a GLTF format.

[0028] A first example method / apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions which, when executed by the processor, are operable to cause the apparatus to perform any of the methods shown above.

[0029] A second example method / apparatus according to some embodiments may include: acquiring scene description data of a 3D scene, wherein the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with at least one action, and behavior information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions; and performing a first action on at least a first node associated with the first action in response to the following logical determination: (i) a combination of at least one of the trigger conditions has been satisfied, and (ii) a first action of at least one action is associated with the combination of the trigger conditions through the behavior information.

[0030] For some embodiments of the second example method, the logical determination may further include determining that a trigger associated with the at least one trigger condition is activated.

[0031] For some embodiments of the second example method, the logical determination may further include determining an activation state of a combination of at least one trigger condition, and performing the first action may include performing the first action one or more times based on the activation state.

[0032] For some embodiments of the second example method, determining the combination of at least one triggering condition may include performing a logical operation on at least one of the triggering conditions.

[0033] For some embodiments of the second example method, determining the combination of at least one trigger condition may include performing a logical OR operation on at least two trigger conditions.

[0034] For some embodiments of the second example method, at least a first of the at least one triggering condition is a visibility condition that is satisfied when the specified scene element is visible to the specified camera node.

[0035] For some embodiments of the second example method, at least a first trigger condition of the at least one trigger condition is a proximity condition that is satisfied when a distance from a user camera to a specified scene element is within a specified boundary.

[0036] For some embodiments of the second example method, at least a first trigger condition of the at least one trigger condition is a user input condition that is satisfied when a specified user interaction is detected.

[0037] For some embodiments of the second example method, at least a first trigger condition of the at least one trigger condition is a timing condition that is satisfied during a specified time period.

[0038] For some embodiments of the second example method, at least a first trigger condition of the at least one trigger condition is a collider condition satisfied in response to detecting a conflict between specified scene elements.

[0039] For some embodiments of the second example method, the behavior information may identify at least one behavior, the behavior information for each of the at least one behavior identifying a trigger associated with one of the at least one trigger condition and one of the at least one action.

[0040] For some embodiments of the second example method, the action information may describe at least two actions, and the behavior information may include information indicating, for at least one behavior, an order in which at least two of the at least two actions are to be performed.

[0041] For some embodiments of the second example method, the action information may describe at least two actions, and the behavior information may include information indicating that at least two of the at least two actions are to be performed simultaneously.

[0042] For some embodiments of the second example method, at least one of the scene elements in the scene is a virtual object.

[0043] Some embodiments of the second example method may further include rendering a 3D scene according to the scene description data operated by the first action.

[0044] For some embodiments of the second example method, the trigger information may include an array of two or more triggers in the 3D scene.

[0045] For some embodiments of the second example method, the action information may include an array of two or more actions in the 3D scene.

[0046] For some embodiments of the second example method, the behavior information may include an array of two or more behaviors in the 3D scene.

[0047] For some embodiments of the second example method, the scene description data may be provided in a JSON format.

[0048] For some embodiments of the second example method, the scene description data may be provided in a GLTF format.

[0049] A second example method / apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions which, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0050] According to a third example method of some embodiments, which is a method for rendering an extended reality scene relative to a user in a timed environment, the method may include: obtaining a description of the extended reality scene, the description may include: a scene tree of linked nodes, the nodes describing timed objects, virtual objects or relationships between objects; behavior data items, the behavior data items may include: at least one trigger control parameter, the trigger control parameter is a description of conditions associated with one or more triggers; activation conditions associated with the trigger control parameters; at least one action, which is an action that describes the processing performed by the extended reality engine on the object described by the node of the scene tree; and applying the action of the behavior data item to the associated object under the condition that a logical combination of at least one of the triggers that trigger the behavior data item and the activation conditions associated with the trigger control parameters of the behavior data item are satisfied.

[0051] For some embodiments of the third example method, the logical combination of at least one trigger may include a logical operation on at least one of the conditions associated with the one or more triggers.

[0052] For some embodiments of the third example method, the logical combination of at least one trigger may include a logical OR operation of at least two conditions associated with the one or more triggers.

[0053] A third example method / apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0054] According to a fourth example method of some embodiments, which is a method for updating a first description of an extended reality scene including behavior data items with a second description of the extended reality scene at runtime, the method may include: for each ongoing behavior data item of the first description, if the ongoing behavior data item is not applicable to the second description: if there is an interruption action for the ongoing application in the first description, processing the interruption action; stopping the ongoing behavior; and applying the second description.

[0055] According to a fifth example apparatus of some embodiments, which is a device for rendering an extended reality scene relative to a user in a timed environment, the apparatus may include: a memory associated with a processor, the processor being configured to: obtain a description of the extended reality scene, the description including: a scene tree of linked nodes, the nodes describing timed objects, virtual objects, or relationships between objects; behavior data items, the behavior data items may include: at least one trigger, the trigger being a description of a condition; the trigger being activated when its condition is detected in the timed environment; and at least one action, the action being a description of processing performed by the extended reality engine on an object described by a node of the scene tree; and applying the action of the behavior to the associated object under the condition that the trigger of the behavior data item is activated.

[0056] For some embodiments of the fifth example device, the processor is further configured to: when obtaining a description of an extended reality scene, attribute an activation state set to false to at least one trigger of the description; when a condition of at least one trigger is met for the first time, set the activation state of the trigger to true; and when the condition of at least one trigger is met, activate the trigger.

[0057] For some embodiments of the fifth example apparatus, the processor is further configured to: when a condition of at least one trigger is met, if the activation state of the trigger is set to true, activate the trigger only if the description of the trigger authorizes a second activation.

[0058] According to a sixth example device of some embodiments, which is a device for updating a first description of an extended reality scene including behavior data items with a second description of the extended reality scene at runtime, the device may include: a memory associated with a processor, the processor being configured to: for each ongoing behavior data item of the first description, if the ongoing behavior data item is not applicable to the second description: if the ongoing behavior data item has an interruption action for the ongoing application in the first description, process the interruption action; stop the ongoing behavior; and apply the second description.

[0059] A seventh example apparatus according to some embodiments may include one or more processors configured to perform any of the methods listed above.

[0060] An eighth example apparatus according to some embodiments may include a computer readable medium including instructions for causing one or more processors to perform any of the methods listed above.

[0061] For some embodiments of the eighth example apparatus, the computer-readable medium is a non-transitory storage medium.

[0062] A tenth example apparatus according to some embodiments may include a computer program product including instructions which, when executed by one or more processors, cause the one or more processors to perform any of the methods listed above.

[0063] An eleventh example apparatus according to some embodiments may include: a signal including scene description data of a 3D scene, wherein the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the action, and behavior information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions.

[0064] A twelfth example method / apparatus according to some embodiments may include a computer-readable medium comprising scene description data for a 3D scene, wherein the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the action, and behavior information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions.

[0065] In other embodiments, encoder and decoder devices are provided to perform the methods described herein. The encoder or decoder device may include a processor configured to perform the methods described herein. The device may include a computer-readable medium (e.g., a non-transitory medium) storing instructions for performing the methods described herein. In some embodiments, the computer-readable medium (e.g., a non-transitory medium) stores a video encoded using any of the methods described herein.

[0066] One or more of the embodiments further provides a computer-readable storage medium on which instructions for performing bidirectional optical flow encoding or decoding of video data according to any of the above methods are stored. The embodiment further provides a computer-readable storage medium on which a bitstream generated according to the above method is stored. The embodiment further provides a method and apparatus for sending a bitstream generated according to the above method. The embodiment further provides a computer program product including instructions for performing any of the described methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1A is a schematic side view illustrating an example waveguide display that may be used with extended reality (XR) applications in accordance with some embodiments.

[0068] Figure 1Bis a schematic side view illustrating example alternative display types that may be used with extended reality applications in accordance with some embodiments.

[0069] Figure 1C is a schematic side view illustrating example alternative display types that may be used with extended reality applications in accordance with some embodiments.

[0070] Figure 1D is a system diagram illustrating a collection of example interfaces for a system according to some embodiments.

[0071] Figure 1E is a system diagram illustrating a collection of example interfaces for scene description in accordance with some embodiments.

[0072] Figure 2 is a system diagram illustrating a collection of example interfaces of an MPEG-1 node hierarchy for elements supporting scene interactivity in accordance with some embodiments.

[0073] Figure 3 is a block diagram showing an example of the logical relationship between trigger information (describing triggers 1 to n), action information (describing actions 1 to m), and behavior information (describing the relationship between the trigger and the action) according to some embodiments, where the trigger and the action may refer to one or more nodes in a scene description (such as a hierarchical scene graph).

[0074] Figure 4 is a schematic plan view illustrating example relationships of objects described in an augmented reality scene according to some embodiments.

[0075] Figure 5A is a data structure illustrating a collection of example triggers for an augmented reality scenario description in accordance with some embodiments.

[0076] Figure 5B is a data structure illustrating a collection of example actions describing an augmented reality scenario in accordance with some embodiments.

[0077] Figure 5C is a data structure illustrating a collection of example behaviors for augmented reality scenario descriptions according to some embodiments.

[0078] Figure 5D is a data structure illustrating a collection of example supplemental information for an augmented reality scene description in accordance with some embodiments.

[0079] Figure 6 is a data syntax diagram illustrating an example syntax for a data stream for encoding an extended reality (XR) scene description in accordance with some embodiments.

[0080] Figure 7 is a flow chart illustrating an example process according to some embodiments.

[0081] Figure 8 is a flow diagram illustrating an example process for activating a trigger or combination of triggers in accordance with some embodiments.

[0082] Fig. 9 is a flow diagram illustrating an example process for an extended reality scene with a trigger mechanism for rendering updates in accordance with some embodiments.

[0083] Fig.10 is a flow diagram illustrating an example process for trigger activation in a time-evolving scene description in accordance with some embodiments.

[0084] The entities, connections, arrangements, etc. depicted in and described in connection with the various figures are presented as examples and not as limitations. Therefore, any and all statements or other indications as to what a particular figure "depicts," what a particular element or entity in a particular figure "is" or "has," and any and all similar statements—which in isolation and out of context may be read as absolute and therefore limiting—may only be properly read as constructively following a clause such as "In at least one embodiment, ..." For brevity and clarity of presentation, this implicit leading clause is not repeated in the detailed description. DETAILED DESCRIPTION

[0085] Figure 1A 1 is a schematic side view illustrating an example waveguide display that may be used with extended reality (XR) applications according to some embodiments. An image is projected by an image generator 102. The image generator 102 may project an image using one or more of a variety of techniques. For example, the image generator 102 may be a laser beam scanning (LBS) projector, a liquid crystal display (LCD), a light emitting diode (LED) display (including an organic LED (OLED) or micro-LED (μLED) display), a digital light processor (DLP), a liquid crystal on silicon (LCoS) display, or other type of image generator or light engine.

[0086] Light representing an image 112 generated by the image generator 102 is coupled into the waveguide 104 by the diffraction coupler 106. The coupler 106 diffracts the light representing the image 112 into one or more diffraction orders. For example, the light ray 108 (which is one of the light rays representing a portion of the bottom of the image) is diffracted by the coupler 106, and one of the diffraction orders 110 (e.g., the second order) is at an angle that enables propagation through the waveguide 104 by total internal reflection. The image generator 102 displays the image as directed by the control module 124, which operates to render image data, video data, point cloud data, or other displayable data.

[0087] At least a portion of the light 110 that has been coupled into the waveguide 104 by the diffractive incoupler 106 is coupled out of the waveguide by the diffractive outcoupler 114. At least some of the light coupled out of the waveguide 104 replicates the angle of incidence of the light coupled into the waveguide. For example, in the illustration, outcoupled light rays 116a, 116b, and 116c replicate the angle of the incoupled light ray 108. Because the light exiting the outcoupler replicates the direction of the light entering the incoupler, the waveguide substantially replicates the original image 112. The user's eye 118 can focus on the replicated image.

[0088] exist Figure 1A In the example of , the coupler 114 couples out only a portion of the light at each reflection, thereby allowing a single input beam (such as beam 108) to generate multiple parallel output beams (such as beams 116a, 116b, and 116c). In this way, at least some of the light originating from each portion of the image may reach the user's eye even if the eye is not perfectly aligned with the center of the coupler. For example, if the eye 118 were to move downward, beam 116c could enter the eye even if beams 116a and 116b did not move downward, so the user could still perceive the bottom of the image 112 despite the position shift. Therefore, the coupler 114 partially operates as an exit pupil expander in the vertical direction. The waveguide may also include one or more additional exit pupil expanders ( Figure 1A (not shown) to expand the exit pupil in the horizontal direction.

[0089] In some embodiments, the waveguide 104 is at least partially transparent with respect to light originating from outside the waveguide display. For example, at least some of the light 120 from a real-world object (such as object 122) passes through the waveguide 104, allowing the user to see the real-world object when using the waveguide display. When the light 120 from the real-world object also passes through the diffraction grating 114, there will be multiple diffraction orders, and therefore multiple images. In order to minimize the visibility of multiple images, it is desirable that the zero-order diffraction (without deviation 114) has a large diffraction efficiency for the light 120 and the zero-order, while the higher diffraction orders are lower in energy. Therefore, in addition to expanding and coupling out the virtual image, the coupler 114 is preferably configured to pass the zero-order of the real image. In such an embodiment, the image displayed by the waveguide display can appear superimposed on the real world.

[0090] Figure 1B1 is a schematic side view illustrating an example alternative display type that can be used with an extended reality application according to some embodiments. In an XR head mounted display device 130, a control module 132 controls a display 134 (which can be an LCD) to display an image. The head mounted display includes a partially reflective surface 136 that reflects (and in some embodiments, both reflects and focuses) an image displayed on the LCD to make the image visible to the user. The partially reflective surface 136 also allows at least some external light to pass through, thereby allowing the user to see their surroundings.

[0091] Figure 1C 1 is a schematic side view illustrating an example alternative display type that may be used with an extended reality application in accordance with some embodiments. In an XR head mounted display device 140, a control module 142 controls a display 144 (which may be an LCD) to display an image. The image is focused by one or more lenses of a display optics 146 to make the image visible to the user. Figure 1C In some such embodiments, external camera 148 may be used to capture images of the external environment and display such images on display 144 along with any virtual content that may also be displayed.

[0092] The embodiments described herein are not limited to any particular type or structure of XR display device.

[0093] Figure 1D is a system diagram illustrating a collection of example interfaces for a system according to some embodiments. An extended reality display device together with its control electronics may use a system such as Figure 1D System 150 can be implemented as a device including various components described below, and is configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 150 can be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components, either individually or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of system 150 are distributed over multiple ICs and / or discrete components. In various embodiments, system 150 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more aspects described in this document.

[0094] The system 150 includes at least one processor 152 configured to execute instructions loaded therein for implementing, for example, various aspects described in this document. The processor 152 may include embedded memory, input-output interfaces, and various other circuits known in the art. The system 150 includes at least one memory 154 (e.g., a volatile memory device and / or a non-volatile memory device). As a non-limiting example, the system 150 may include a storage device 158, which may include a non-volatile memory and / or a volatile memory, including but not limited to an electrically erasable programmable read-only memory (EEPROM), a read-only memory (ROM), a programmable read-only memory (PROM), a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a flash memory, a magnetic disk drive, and / or an optical disk drive. As a non-limiting example, the storage device 158 may include an internal storage device, an attached storage device (including removable and non-removable storage devices), and / or a network accessible storage device.

[0095] The system 150 includes an encoder / decoder module 156, which is configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 156 may include its own processor and memory. The encoder / decoder module 156 represents a module(s) that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. In addition, the encoder / decoder module 156 may be implemented as a separate element of the system 150, or may be incorporated into the processor 152 as a combination of hardware and software known to those skilled in the art.

[0096] Program code to be loaded onto the processor 152 or encoder / decoder 156 to perform various aspects described in this document may be stored in the storage device 158 and subsequently loaded onto the memory 154 for execution by the processor 152. According to various embodiments, one or more of the processor 152, memory 154, storage device 158, and encoder / decoder module 156 may store one or more of various items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0097] In some embodiments, memory internal to the processor 152 and / or encoder / decoder module 156 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 152 or the encoder / decoder module 152) is used for one or more of these functions. The external memory may be a memory 154 and / or a storage device 158, such as a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store, for example, an operating system for a television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG refers to Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard developed by JVET (Joint Video Experts Group)).

[0098] As shown in block 172, input to the elements of system 150 may be provided through various input devices. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives RF signals transmitted over the air, for example, by a broadcaster, (ii) a component (COMP) input terminal (or a collection of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 1C Other examples not shown include composite video.

[0099] In various embodiments, the input device of block 172 has associated corresponding input processing elements as known in the art. For example, the RF portion may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal, or band limiting a signal to a band of frequencies), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower band of frequencies to select a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF portion of various embodiments includes: one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF portion may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near baseband frequency) or down-converting to baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal transmitted by wired (for example cable) medium, and filter to the desired frequency band by filtering, down-conversion and again to perform frequency selection. Various embodiments rearrange the order of above-mentioned (and other) elements, remove some in these elements, and / or add other elements of similar or different functions. Adding element can include inserting element between existing element, for example inserting amplifier and analog-to-digital converter. In various embodiments, the RF part comprises antenna.

[0100] In addition, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 150 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomom error correction, can be implemented as needed, for example, in a separate input processing IC or in the processor 152. Similarly, aspects of USB or HDMI interface processing can be implemented as needed in a separate interface IC or in the processor 152. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 152 and the encoder / decoder 156, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on an output device.

[0101] The various elements of system 150 may be provided within an integrated housing in which the various elements may be interconnected and transmit data between them using a suitable connection arrangement 174, such as an internal bus known in the art, including an inter-IC (I2C) bus, wiring, and printed circuit boards.

[0102] The system 150 includes a communication interface 160 that enables communication with other devices via a communication channel 162. The communication interface 160 may include, but is not limited to, a transceiver configured to send and receive data through the communication channel 162. The communication interface 160 may include, but is not limited to, a modem or a network card, and the communication channel 162 may be implemented, for example, within a wired and / or wireless medium.

[0103] In various embodiments, data is streamed or otherwise provided to the system 150 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of these embodiments are received through a communication channel 162 and a communication interface 160 suitable for Wi-Fi communication. The communication channel 162 of these embodiments is typically connected to an access point or router that provides access to an external network (including the Internet) to allow streaming applications and other over-the-top communications. Other embodiments provide streaming data to the system 150 using a set-top box that delivers data through an HDMI connection of an input block 172. Other embodiments provide streaming data to the system 150 using an RF connection of an input block 172. As described above, various embodiments provide data in a non-streaming manner. In addition, various embodiments use a wireless network other than Wi-Fi, such as a cellular network or a Bluetooth network.

[0104] The system 150 can provide output signals to various output devices, including a display 176, a speaker 178, and other peripherals 180. The display 176 of various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. The display 176 can be used for a television, a tablet computer, a laptop computer, a mobile phone, or other devices. The display 176 can also be integrated with other components (for example, as in a smart phone), or it can be separate (for example, an external monitor for a laptop computer). In various examples of embodiments, other peripherals 180 include one or more of a stand-alone digital video disk (or digital versatile disk) (DVR, both terms), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more peripherals 180 that provide functions based on the output of the system 150. For example, a disk player performs the function of playing the output of the system 150.

[0105] In various embodiments, control signals are transmitted between the system 150 and the display 176, speaker 178, or other peripheral device 180 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. Output devices can be communicatively coupled to the system 1000 via dedicated connections through corresponding interfaces 164, 166, and 168. Alternatively, the output devices can be connected to the system 150 using a communication channel 162 via the communication interface 160. The display 176 and speaker 178 can be integrated into a single unit with other components of the system 150 in an electronic device (e.g., a television). In various embodiments, the display interface 164 includes a display driver, such as, for example, a timing controller (T_Con) chip.

[0106] For example, if the RF portion of input 172 is part of a separate set-top box, the display 176 and speaker 178 may optionally be separate from one or more other components. In various embodiments where the display 176 and speaker 178 are external components, the output signal may be provided via a dedicated output connection (including, for example, an HDMI port, a USB port, or a COMP output).

[0107] The system 150 may include one or more sensor devices 168. Examples of sensor devices that can be used include one or more GPS sensors, gyroscope sensors, accelerometers, light sensors, cameras, depth cameras, microphones, and / or magnetometers. Such sensors can be used to determine information such as the position and orientation of the user. In the case where the system 150 is used as a control module (e.g., control modules 124, 132) of an extended reality display, the position and orientation of the user can be used to determine how to render image data so that the user perceives the correct part of the virtual object or virtual scene from the correct viewpoint. In the case of a head-mounted display device, the position and orientation of the device itself can be used to determine the position and orientation of the user for the purpose of rendering virtual content. In the case of other display devices (such as phones, tablets, computer monitors, or televisions), other inputs can be used to determine the position and orientation of the user for the purpose of rendering content. For example, the user can use a touch screen, a keypad or keyboard, a trackball, a joystick, or other inputs to select and / or adjust the desired viewpoint and / or viewing direction. Where the display device has sensors such as an accelerometer and / or a gyroscope, the viewpoint and orientation for purposes of rendering content may be selected and / or adjusted based on the motion of the display device.

[0108] The embodiments may be performed by computer software implemented by the processor 152 or by hardware or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 154 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology (such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples). The processor 152 may be of any type suitable for the technical environment and may encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture as non-limiting examples.

[0109] XR scene description framework

[0110] The present principles generally relate to the field of rendering of extended reality scene descriptions and extended reality rendering. This document is also understood in the context of formatting and playback of extended reality applications when rendered on an end-user device such as a mobile device or a head mounted display (HMD).

[0111] In XR applications, scene descriptions are used to combine an explicit and easily parsable description of the scene structure and some binary representation of the media content.

[0112] In time-based media streams, the scene description itself can be time-evolving to provide relevant virtual content for each sequence of the media stream. For example, for advertising purposes, a virtual bottle can be shown during a video sequence of people drinking.

[0113] This behavior can be achieved by relying on the framework defined in the scene description of MPEG media documents (Information technology-Coded representation of immersive media-Part14:Scene Description for MPEGmedia, ISO / IEC DIS23090-14:2021(E)). A scene update mechanism based on the JSON patch protocol defined in IETF RFC 6902 can be used to synchronize virtual content to MPEG media streams.

[0114] Figure 1E is a system diagram illustrating a collection of example interfaces for scene description according to some embodiments. Figure 1E In , scene descriptions are stored as items in glTF.json 194. The example scene description has three video tracks 182, 183, 184, an audio track 185 and a JSON patch update track 186 in the ISOBMFF file configuration 181 according to some embodiments. Figure 1EThe example in shows video samples 187, 188, 189, 190 within video tracks 182, 183, 184 for an example configuration. The example JSON patch update track 186 has multiple sample update patches 191, 192, 193. The example ISOBMFF file configuration 181 shows gltf buffers.bin 195 connected to the item gltf.json 194.

[0115] Although the MPEG-1 scene description framework ensures that timed media and corresponding related virtual content are available at any time, it does not provide a description of how users can interact with scene objects at runtime for immersive XR experiences. Therefore, user-specific XR experiences for consuming immersive media are not supported.

[0116] Example embodiments as described herein can be used to provide a scene description that includes a virtual object or light source, but even if the virtual object or light source is available, the virtual object or light source is not necessarily displayed or rendered. In some embodiments, one or more of the following aspects can be considered when determining whether to display a virtual object or light source.

[0117] Spatial aspects may be considered when determining whether to display a virtual object or light source. For example, a virtual object or light source may not be displayed if the user environment is not suitable (e.g., the user is too far away from the location of the rendered timed media) or if the user is not looking in the right direction, or if the virtual object should be displayed on a specific area of ​​the user (e.g., above his left hand which has not been detected).

[0118] The time aspect may be considered when determining whether to display a virtual object or light source. For example, if the user is not yet ready or wants to trigger the display of the object themselves (e.g., using a specific gesture), the virtual object or light source may not be displayed until an appropriate trigger is detected.

[0119] In some embodiments, it is specified in the scene description which objects or light sources the user is allowed to manipulate or interact with through potential haptic feedback.

[0120] Runtime interactivity

[0121] Figure 2 is a system diagram illustrating a collection of example interfaces of an MPEG-1 node hierarchy for elements supporting scene interactivity in accordance with some embodiments. Figure 2 An exemplary MPEG-1 node hierarchy 200 is shown. In accordance with the present principles, except for Figure 3In addition to the node tree described, behavioral metadata items (referred to herein as "behaviors") are added to the scene description. In an example embodiment, the time-evolving scene description is enhanced by adding information identifying behaviors. These behaviors can be associated with predefined virtual objects, on which runtime interactivity is allowed for user-specific XR experiences.

[0122] In some embodiments, these behaviors evolve over time. In such embodiments, the behaviors can be updated through an existing scene description update mechanism.

[0123] In an example embodiment, behavior is characterized by one or more of the following properties:

[0124] One or more triggers that define the conditions to be met for activation.

[0125] Trigger control parameters that define the logical operation between the defined triggers.

[0126] The action to be performed in response to the activation of a trigger.

[0127] Action control parameters Action control parameters define the execution order of the defined actions.

[0128] A priority number that enables the highest priority behavior to be selected if several behaviors occur simultaneously on the same virtual object.

[0129] An optional abort action that specifies how to terminate a behavior when it is no longer defined in a newly received scene update. For example, a behavior may be no longer defined if the associated object has been removed or if the behavior is no longer relevant to the current media (e.g., audio or video) sequence.

[0130] By adding these behaviors, you can define time-dependent user interactivity in immersive content for XR experiences.

[0131] When the second scene description is received, some behaviors of the first scene description may be "ongoing", i.e. they are triggered and their actions are running. The second scene description can be provided as update metadata, i.e. metadata describing the differences between the first scene description and the second scene description. The second scene description includes a node tree describing objects that are the same or different from the objects of the first scene description. The objects of the node tree of the first scene description may no longer exist in the second description. If objects related to the running actions of ongoing behaviors are missing in the second scene description, then these ongoing behaviors are no longer applicable. Similarly, if an ongoing behavior is not defined in the second description, then the ongoing behavior is no longer applicable. The interrupt action field describes how to correctly interrupt a running action on an ongoing behavior.

[0132] Figure 3 is a block diagram showing an example of the logical relationship between trigger information (describing triggers 1 to n), action information (describing actions 1 to m), and behavior information (describing the relationship between the trigger and the action) according to some embodiments, where the trigger and the action may refer to one or more nodes in a scene description (such as a hierarchical scene graph).

[0133] In XR applications, scene descriptions are used to combine an explicit and easily parsable description of the scene structure and some binary representation of the media content. The above section describes the action mechanism of scene descriptions. These actions are associated with predefined virtual objects, on which runtime interactivity is allowed for user-specific XR experiences. Figure 3 An example behavioral mechanism structure 300 is illustrated. Figure 3 The example structure 300 shows an example trigger structure 304 and an example action structure 306 within an example behavior structure 302 . Figure 3 Also shown is an example node 308 with connections to specific example triggers and actions.

[0134] Figure 4 4 is a schematic plan view showing an example relationship of an extended reality scene description object according to some embodiments. In this example, the scene graph 400 includes a description of a real object 412, such as a "planar horizontal surface" (which can be a table or floor or board) and a description of a virtual object 414, such as an animation of a walking character. The scene graph node 414 is associated with a media content item 416, which is an encoding of data for rendering and displaying a walking character (e.g., as a textured animated 3D mesh). The scene graph 400 also includes a node 410, which is a description of the spatial relationship between the real object described in node 412 and the virtual object described in node 414. In this example, node 410 describes the spatial relationship that makes the character walk on a planar surface. When the XR application is started, the media content item 416 is loaded, rendered, and buffered to be displayed when triggered. When the sensor (or the camera of some embodiments) detects a planar surface in the real environment, the application displays the buffered media content item, as described in node 410. The timing is managed by the application according to the timing of the features and animations detected in the real environment. A scene graph node may also contain no description and act only as a parent node to its child nodes.

[0135] XR applications are diverse and can be applied to different contexts and real or virtual environments. For example, in an industrial XR application, a virtual 3D content item (e.g., part A of an engine) is displayed when a camera mounted on a head-mounted display device detects a reference object (part B of an engine) in the real environment. The 3D content item is positioned in the real world at a defined position and scale relative to the detected reference object.

[0136] For example, in an XR application for interior design, when a given image from a catalog is detected in the input camera view, a 3D model of furniture is displayed. The 3D content is positioned in the real world at a position and scale defined relative to the detected reference image. In another application, when the user enters an area near a church (real or virtually rendered in the extended real environment), some audio files may start playing. In another example, when the user sees a can of a given soda in the real environment, an advertising jingle file can be played. In an outdoor gaming application, various virtual characters can appear depending on the semantics of the scenery observed by the user. For example, a bird character is suitable for a tree, so if the sensors of the XR device detect a real object described by the semantic label 'tree', a bird can be added to fly around the tree. In a companion application implemented by smart glasses, when a car is detected in the field of view of the user's camera, a car noise can be emitted in the user's headphones to warn him of potential danger; in addition, the sound can be spatialized so that it arrives from the direction from which the car was detected.

[0137] XR applications can also augment video content instead of the real environment. The video is displayed on a rendering device, and when a timed event is detected in the video, the virtual objects described in the node tree are overlaid. In such a context, the node tree only includes virtual object descriptions.

[0138] Example embodiments are described with reference to the scope of the MPEG-1 scene description framework using the Khronos glTF extension mechanism, which supports additional scene description features, such as node trees. However, the principles described herein are not limited to a particular scene description framework.

[0139] In an example embodiment, the glTF scene description is extended to support interactivity. The interactivity extension is applied at the glTF scene level and is called MPEG_scene_interactivity. The corresponding semantics are provided in Table 1.

[0140]

[0141] Table 1: Semantics of an example MPEG_scene_interactivity extension

[0142] In Table 1 and other semantic tables described herein, the "Usage" column indicates that "M" represents a "mandatory" feature and "O" represents an "optional" feature. However, such a feature may be "mandatory" or "optional" solely based on the particular proposed syntax. Features marked as "mandatory" are not necessarily required features to implement an application. For example, in some embodiments, there are features marked as "mandatory" to meet the expectations of a particular type of parsing and rendering software; however, in other embodiments, the feature may be optional or may be omitted entirely, with the corresponding functionality being implemented using default values ​​or not being implemented at all, without departing from the scope of the present disclosure.

[0143] Figures 5A to 5D Non-limiting examples of extended reality scene descriptions in accordance with some embodiments are shown.

[0144] exist FIG. 5A to FIG. 5D In the first example presented in , a virtual 3D object is continuously displayed and transformed during a media sequence. Once the user's left hand is detected, the virtual 3D object is placed on the user's left hand and continuously follows it.

[0145] As an example of an interactive virtual object according to some embodiments, a virtual 3D advertisement object may be continuously displayed and transformed during a defined period (e.g., between 20 seconds and 40 seconds) in an MPEG media sequence. In this example, once the user's left hand is detected, the virtual 3D object is placed on the user's left hand, and the virtual 3D object continuously follows the user's hand.

[0146] In this example, two behaviors are defined to support this interactivity scenario. The first behavior has the following parameters:

[0147] The first trigger, which is associated with the time sequence of the MPEG media between 20 seconds and 40 seconds, is activated (ACTIVATE_ON) as long as the condition is met.

[0148] Two sequential actions to enable and transform a virtual 3D object (node ​​0).

[0149] The second line has the following parameters:

[0150] The trigger combination having a second trigger associated with the detection of the user's left hand and a third trigger associated with the non-detection of the user's right hand is activated (ACTIVATE_ON) as long as the condition is met.

[0151] A single action that places a virtual object (node ​​0) on the user's left hand.

[0152] Both behaviors define the same interrupt action of disabling a virtual object (node ​​0). Since both behaviors affect the same virtual object (node ​​0), a higher priority is set to the second behavior associated with the user gesture (left hand gesture) to execute the desired interactivity scene. In this example, the desired behavior can be implemented using scene interactivity information, which can be provided in JSON format, such as Figures 5A to 5D The example shown in .

[0153] For some embodiments, scene descriptions for MPEG media may be used to combine scene structures with binary representations of media content. See “Information Technology - Coded Representation of Immersive Media - Part 14: Scene Description for MPEG Media,” ISO / IEC DIS 23090-14:20 21(E).

[0154] While the MPEG-1 scene description framework ensures that timed media and corresponding related virtual content are available at any time, it does not describe how a user interacts with scene objects at runtime for an immersive XR experience. European Patent Application No. EP22305024 (filed on January 12, 2022) ("'024") discusses enhancing time-evolving scene descriptions by adding behaviors. These behaviors are associated with predefined virtual objects, allowing runtime interactivity on predefined virtual objects for user-specific XR experiences.

[0155] A behavior consists of a collection of triggers that define the conditions to be met for activation and a collection of actions to be performed when the triggers are activated. Triggers control how parameters are defined to allow logical operations between triggers and activation policies.

[0156] Regarding the activation policy, a single Boolean flag (ActivateOnce) is specified for each trigger and indicates whether (i) the trigger is activated every time the condition is met or (ii) the trigger is activated only once after the condition is met. Table 2 provides an example syntax for a trigger configured in this manner.

[0157] This activation mechanism may not be well suited when a set of triggers for a behavior references multiple triggers: depending on how those triggers are combined by logical operations (e.g., AND, OR, NOT, or other logical operators), the activation flag may not be coherent between all triggers in the following combination: a proximity trigger combined with a visibility trigger (logically AND'ed), where both ActivateOnce flags are set to true. The combination of triggers may never occur. In one scenario, when a proximity trigger is to be activated, a visibility trigger may be activated only once and never again. For example, an initial visibility trigger may be activated only once and never again even when the proximity trigger is later activated multiple times.

[0158] The combination of triggers can be a combination of each condition, rather than the activation state of each trigger. In addition, the activation flag as defined in the application '024 may not allow for additional situations that may occur, such as:

[0159] Activates the trigger when the stated condition is no longer met

[0160] Activate the trigger only n times during the entire timeline of the scene, or once each time the condition is met (or not met). The value of n can range from 1 to infinity.

[0161] According to some embodiments, the application defines a new activation flag at the behavior level rather than the trigger level. The flag specifies the activation mechanism for the combination of triggers by defining several scenarios for activating the triggers. The conditions of each reference trigger are evaluated and combined according to the trigger control parameters. The control parameters allow the description of a combination of logical operations between the referenced triggers. It can be a string that describes the combination. The result of the evaluation (satisfied or not) and the activation flag tell the presentation engine when to execute the action of the behavior. The application introduces new data in the behavior-based interactive scene description. These data fields are used by the runtime processing model to control the behavior.

[0162] The application is described in detail in the context of the MPEG-1 Scene Description Framework using the Khronos glTF extension mechanism ("Khronos Group") to support additional scene description features.This application modifies the MPEG_scene_interactivity gltf extension defined in the '024 application.

[0163] This modification affects the definitions of triggers (in tables 2 and 3) and actions (in tables 5 and 6):

[0164] ·Removal of parameters from ActivateOnce trigger

[0165] ●Create new activation behavior parameters.

[0166] Modify the triggersControl parameter so that multiple operations (e.g., AND, OR, NOT, or other operators) can be used for a combination of triggers. It can be a string that combines the trigger index and the logical operation in the trigger array. For example:

[0167] The "#" tag indicates the trigger index, the "&" indicates the logical AND operation, the "|" indicates the logical OR operation, the "~" indicates the NOT operation, and the brackets group certain operations. Such a syntax can give the following string: "#1&~#2|(#3 )"

[0168] The default empty string can be interpreted as a logical OR between all triggers.

[0169] An example process for activating a trigger or combination of triggers according to some embodiments is described in Fig.10 and discussed further below.

[0170] The parameters of the example trigger structure are shown in Table 2. Table 2 shows the activateOnce Boolean value, which indicates how often the trigger is activated when the trigger's condition is met. This Boolean is removed in Table 3.

[0171]

[0172]

[0173]

[0174]

[0175] Table 2: First example trigger semantics

[0176] In some embodiments, instead of using string parameters associated with the Khronos OpenXR Interaction Profile path syntax to define gestures for user body parts and USER_INPUT triggers, other syntax formats may be used, such as a syntax format for representation of haptic objects where an array of vertices (geometric model) and a binary mask (body part mask) are used to specify where haptic effects should be applied.

[0177] An updated trigger structure according to some embodiments is shown in Table 3. Table 3 removes the activateOnce Boolean value shown in Table 2. The updated example trigger structure described below in Table 3 corresponds to Figure 5A .

[0178]

[0179]

[0180]

[0181]

[0182] Table 3: Second example trigger semantics

[0183] Example parameters of the action structure are shown in Table 4. The example action structure described below in Table 4 corresponds to Figure 5B .

[0184]

[0185]

[0186]

[0187]

[0188]

[0189] Table 4: Action semantics

[0190] Example parameters of the behavior structure are shown in Table 5. Table 5 shows a triggersControl enumeration, which combines multiple triggers using a logical OR operator when the enumeration is 0, and combines multiple triggers using a logical AND operator when the enumeration is 1.

[0191]

[0192]

[0193] Table 5: First example behavior semantics

[0194] According to some embodiments, example parameters of the updated behavior structure are shown in Table 6. Table 6 shows an activation enumeration, which lists the various activation states of how and when various triggers need to be activated. Table 6 also shows an activator control (triggersControl) string, which uses a string value to indicate the logical operator to be applied when combining multiple triggers. The updated example behavior structure described below in Table 6 corresponds to Figure 5C .

[0195]

[0196]

[0197]

[0198] Table 6: Second example behavior semantics

[0199] exist FIG. 5A to FIG. 5D In the first example presented in , a virtual 3D object is continuously displayed and transformed during a media sequence. Once the user's left hand is detected, the virtual 3D object is placed on the user's left hand, and the media sequence continuously follows the virtual 3D object.

[0200] Figure 5A is a data structure illustrating a collection of example triggers for an augmented reality scenario description in accordance with some embodiments. Figure 5A A data structure 502 is shown with a header indicating that interactivity metadata belongs to a scene description. Three triggers for two behaviors are described for the example. Triggers can be listed in the behavior field (such as Figure 5C Listing them in separate arrays allows a method or apparatus to use the same trigger for several behaviors. The parameters of an example trigger structure are shown in Table 3. The "activateOnce" field shown in the data structure of Table 2 is not in Tables 3 and Figure 5A in the data structure. Figure 5A Corresponding to Table 3. Returning to the example listed above, Figure 5A The first trigger is a time sequence of MPEG media between 20 seconds and 40 seconds. Figure 5A The second trigger is user input associated with the user's left hand. Figure 5A The third trigger is user input associated with the user's right hand.

[0201] Figure 5B is a data structure illustrating a collection of example actions describing an extended reality scenario according to some embodiments. Data structure 504 has an "actions" field that includes: a description of the three actions required to perform the two behaviors of the illustrative example. The first action, which enables the object at node 0 in the node tree, has index 0 because it is the first action in the action array. The second action at node 0, which places the object on the user's left hand, has index 1, and in this example, the third action, which transforms the object at node 0 according to the transformation matrix, has index 2. The fourth action, which disables the object at node 0 with index 3, is an interruption action common to both behaviors. Figure 5B Corresponds to Table 4.

[0202] Figure 5C is a data structure illustrating a collection of example behaviors described in an augmented reality scenario according to some embodiments. Data structure 506 has a "behavior" field that includes descriptions of two behaviors of illustrative examples. The list of triggers and actions is represented by Figure 5A The flip-flop array and Figure 5BThe interrupt action of the two behaviors refers to the fourth action with index number 3 in the action array. The priority 2 of the second behavior is higher than the first behavior with priority 1. Since the two behaviors apply to the same node 0 of the node tree, if the two behaviors are active at the same time, the second behavior is selected. Example parameters of the behavior structure are shown in Table 6. Table 6 and Figure 5C The "activation" field shown in Table 5 is not in the data structure of Table 5. The "triggersControl" field shown in the data structure of Table 5 is in Tables 6 and Figure 5C The data structure is changed into a string describing the set of logical operations to be performed on the list of triggers. Figure 5C Corresponding to Table 6. For the first behavior (index 0), the triggersControl field indicates that only the first trigger is used. The trigger is activated as long as the condition is met, which means that the time series is between 20 seconds and 40 seconds. When the first behavior is applicable, the first action (activate / enable the virtual object) and the third action (transform the virtual object) are performed. For the second behavior, the trigger condition is met when the left hand is detected and the right hand is not detected. When the second behavior is applicable, the second action (place the virtual object at the user's left hand) is performed.

[0203] Figure 5D It is a data structure illustrating a collection of example information of an extended reality scene description according to some embodiments. Data structure 508 has a "node" field including an example description of a node. For example, in a scene, a node can be named. Triggers, actions, and behaviors can also have unique id numbers or unique names. Therefore, when a scene description is updated, it is straightforward to detect whether an ongoing behavior or node belongs to a new scene description.

[0204] Figure 6 is a data syntax diagram illustrating an example syntax for a data stream for encoding an extended reality (XR) scene description in accordance with some embodiments. The structure resides in a container that organizes the stream into independent syntax elements. The structure may include a header portion 602, which is a set of data common to each syntax element of the stream. For example, the header portion includes some metadata about the syntax elements, describing the properties and role of each of the syntax elements. The structure also includes a payload, which includes syntax element 604 and syntax element 606. Syntax element 604 includes data representing a media content item described in a node of a scene graph associated with a virtual element. Images, meshes, and other raw data may have been compressed according to a compression method. Syntax element 606 is part of the payload of the data stream and includes data for encoding the scene description, such as information about the Figures 5A to 5D as described.

[0205] Figure 7is a flowchart illustrating a method performed according to some embodiments. Scene description data for a 3D scene is obtained. The scene description data may be in GLTF format or other formats. Processing 700 may include obtaining 702 scene description data, the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the action, and behavior information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions. In some embodiments, an action is performed on a scene element associated with a corresponding node in a hierarchical scene description graph. Alternatively or additionally, an action may be performed on a scene element that is not associated with a particular node (e.g., an animation or MPEG media element included in the scene description).

[0206] In some embodiments, characteristics of user interactivity are monitored 704 to detect whether a trigger condition is met. The monitored characteristics may include, for example, the position of a user camera or viewpoint (as determined by user input and / or by sensors such as gyroscopes and accelerometers, among other possibilities), user gestures (as detected by a camera, wrist-worn or hand-held accelerometer, or other sensor), or other user input.

[0207] In response to determining 706 that at least one of the triggering conditions has been satisfied and at least a first one of the actions is associated with the triggering condition via the behavior information, the first action is performed 708 on at least a first scene element associated with the first action.

[0208] In some embodiments, the action results in a modification of one or more scene elements in a scene description of a 3D scene. In some embodiments, the method includes rendering the 3D scene according to the changed scene description. In other embodiments, the modified scene description is provided to a separate renderer for rendering the 3D scene according to conventional 3D rendering techniques. The 3D scene can be displayed on any display device, such as the display device described herein or other display devices. In some embodiments, the 3D scene can be displayed as an overlay with a real-world scene using an optical perspective or video perspective display. It should be noted that the display of the 3D scene as mentioned herein includes using a 2D display device to display a 2D projection of the 3D scene.

[0209] Figure 8 is a flow chart illustrating an example process for activating a trigger or combination of triggers according to some embodiments. During runtime, the application iterates over each defined behavior and Fig.10 The procedure shown checks the implementation of the relevant triggers.

[0210] As described above, according to some embodiments, the application defines a new activation flag at the behavior level rather than the trigger level. See Table 5. The flag specifies the activation mechanism for the combination of triggers by defining several scenarios for activating the triggers. The conditions of each reference trigger are evaluated and combined according to the trigger control parameters. The control parameters allow the description of a combination of logical operations between the referenced triggers. It can be a string that describes the combination. The result of the evaluation (satisfied or not) and the activation flag tell the presentation engine when to execute the action of the behavior. The application introduces new data in the behavior-based interactive scene description. These data fields are used by the runtime processing model to control the behavior.

[0211] The application is described in detail in the context of the MPEG-1 Scene Description Framework using the Khronos glTF extension mechanism ("Khronos Group") to support additional scene description features.This application modifies the MPEG_scene_interactivity gltf extension defined in the '024 application.

[0212] This modification affects the definitions of triggers (in tables 2 and 3) and actions (in tables 5 and 6):

[0213] ●Removal of parameters from ActivateOnce trigger

[0214] ●Create new activation behavior parameters.

[0215] Modify the triggersControl parameter so that multiple operations (e.g., AND, OR, NOT, or other operators) can be used for a combination of triggers. It can be a string that combines the trigger index and the logical operation in the trigger array. For example:

[0216] The "#" tag indicates the trigger index, the "&" indicates the logical AND operation, the "|" indicates the logical OR operation, the "~" indicates the NOT operation, and the brackets group certain operations. Such a syntax can give the following string: "#1&~#2|(#3 )"

[0217] The default empty string can be interpreted as a logical OR between all triggers.

[0218] Return to Figure 8 Discussion, Figure 8 Represents the state of a combination of triggers ({T}xxxx). A combination of conditions ({conditions}) can be evaluated consecutively: each condition can be evaluated and all evaluation results are combined after the triggersControl parameter. Depending on the state and value of the activation flag, the referenced action is executed:

[0219] S1: Once the condition is met for the first time (used when Activate = ACTIVATE_FIRST_ENTER)

[0220] S2: Once each time the condition is met (used when Activate = ACTIVATE_EACH_ENTER)

[0221] S3: As long as the conditions are met (used when Activate=ACTIVATE_ON)

[0222] S4: Once the condition is no longer met for the first time (used when Activate = ACTIVATE_FIRST_EXIT)

[0223] S5: Once each time the condition is not met (used when Activate = ACTIVATE_EACH – EXIT)

[0224] S6: As long as the condition is not met (used when Activate = ACTIVATE_OFF)

[0225] After the scene is loaded 814, the behaviors are loaded and each combination of triggers is evaluated and initialized to either state {T} closed 810 or state {T} open 804, and the referenced actions are executed (S6 case 826 or S3 case 820, respectively).

[0226] From state {T}Close 810, when a combination of conditions 812 are met, the trigger enters state {T}Enter 802, and the referenced action is executed: only the first time (S1 case 816) or every time (S2 case 818).

[0227] Entering from state {T}, the trigger automatically enters state {T} at 804 and executes the referenced action (S3 case 820).

[0228] As long as the condition 806 is met, the trigger is in state {T}on 804 and the referenced action is executed (S3 case 820).

[0229] From state {T}on 804, when condition 806 is no longer met, the trigger enters state {T}exit 808, and the referenced action is executed: only the first time (S4 case 822) or every time (S5 case 824).

[0230] Exiting from state {T}, the trigger automatically enters state {T} closed 810 and executes the referenced action (S6 case 826).

[0231] Alternatively, the S1 case 816 or the S4 case 822 may be replaced by a (S1-n) case or a (S4-n) case, respectively, which activates the trigger when the trigger enters state {T} enter 802 or state {T} exit 808, respectively, for the first n times. The value of n may be specified in the parameters of the additional behavior.

[0232] Alternatively, the diagram of triggers can be applied outside the mechanism of the above-mentioned behavior. A combination of triggers can be specified for a scenario, and S1 to S6 can be considered as events that are triggered to execute predefined callback functions.

[0233] The MPEG_scene_interactivity extension semantics may be added to the MPEG-1 SD standard.

[0234] Fig. 9 9 is a flow chart illustrating an example process of an extended reality scene for rendering an updated trigger mechanism according to some embodiments. For some embodiments of process 900, scene description data of a 3D scene is obtained 902. The scene description data may include trigger information, action information, behavior information. For some embodiments, the scene description data may include the following Tables 3, 4, and 6 and Figure 5A , 5B , 5C and 5D. The method may also include monitoring 904 one or more trigger conditions associated with the action. The one or more trigger conditions associated with the action may be logically combined 906. For example, in the case of one trigger, the combination of triggers may not be performed, but for example a NOT operation may be performed on the trigger according to some embodiments. In another example, the first trigger state may be inverted (by a NOT logical operator operation), and the inverted first trigger may be logically ORed with the second trigger state. The method may then determine 908 whether the logical combination of one or more triggers produces a true result. If the logical combination produces a true result, the method may determine 910 whether the activation state allows the associated action to be executed. If so, the associated action may be executed 912 on the associated node. For example, the activation state may be ACTIVATE_EACH_ENTER. If the device executing the method is currently in the S2 state (such as with respect to Figure 8 As described), whenever the logical combination of one or more triggers produces a true result, the associated action can be executed. If the activation state does not allow the action to be executed, the method can return to monitoring one or more trigger conditions associated with the action.

[0235] For some embodiments, a check may be performed to determine that the logical combination has produced a result (which may be a true or false result). If a result of the logical combination is produced, a check may be performed to determine whether the activation state allows the action to be performed. If the activation state allows the action to be performed, the action may be performed on the node.

[0236] Fig.10 is a flow chart illustrating an example process for a trigger activation mechanism in a time-evolving scene description according to some embodiments. For some embodiments, the example process 1000 may include obtaining 1002 scene description data for a 3D scene. For some embodiments, the scene description data may include scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the action, and behavior information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions. For some embodiments, in response to 1004 the following logic determines that the example process may also include performing the first action on at least a first node associated with the first action: (i) a combination of at least one of the trigger conditions has been satisfied, and (ii) associating at least a first action of the actions with the combination of the trigger conditions via the behavior information.

[0237] Although methods and systems according to some embodiments are generally discussed in the context of extended reality (XR), some embodiments may be applicable to any XR context, such as, for example, a virtual reality (VR) / mixed reality (MR) / augmented reality (AR) context. In addition, although the term "head mounted display (HMD)" is used herein according to some embodiments, for some embodiments, some embodiments may be applicable to a wearable device (which may or may not be attached to the head) capable of, for example, XR, VR, AR, and / or MR.

[0238] An example method according to some embodiments may include: acquiring scene description data of a 3D scene, wherein the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the action, and behavior information, wherein the behavior information may include: a list of at least one trigger, trigger combination information indicating how at least one trigger condition corresponding to a trigger in the list is combined with other trigger conditions corresponding to other triggers in the list, and activation information indicating when to execute a combination of at least one trigger in the list with other triggers in the list; and determining that a first action associated with a first action in at least one action may be performed in response to the following logic: (i) a combination of the trigger condition of at least one trigger in the list with the trigger condition of other triggers in the list has produced a true result, and (ii) activation information indicating that a combination of at least one trigger in the list with other triggers in the list is to be executed.

[0239] For some embodiments of the example method, the list may include one trigger and no other triggers of the list exist, the trigger combination information may include a unary operator that operates on the one trigger, and the activation information may indicate when to execute the unary operator on the one trigger.

[0240] For some embodiments of the example method, the list may include at least two triggers, the trigger combination information may indicate how to combine the at least two triggers, and the activation information may indicate when to execute the combination of the at least two triggers of the list.

[0241] For some embodiments of the example method, performing the first action may include performing the first action one or more times based on the activation information.

[0242] For some embodiments of the example method, determining a combination of a triggering condition of at least one trigger in the list and triggering conditions of other triggers in the list may include performing a logical operation on at least one of the triggering conditions.

[0243] For some embodiments of the example method, determining a combination of a triggering condition of at least one trigger in the list and a triggering condition of other triggers in the list may include performing a logical OR operation of at least two triggering conditions.

[0244] For some embodiments of the example method, at least a first of the trigger conditions may be a visibility condition that is satisfied when a specified scene element is visible to a specified camera node.

[0245] For some embodiments of the example method, at least a first of the trigger conditions may be a proximity condition that is satisfied when a distance from a user camera to a specified scene element is within specified boundaries.

[0246] For some embodiments of the example method, at least a first of the trigger conditions may be a user input condition that is satisfied when a specified user interaction is detected.

[0247] For some embodiments of the example method, at least a first of the trigger conditions may be a timing condition that is satisfied during a specified time period.

[0248] For some embodiments of the example method, at least a first of the trigger conditions may be a collider condition satisfied in response to detecting a conflict between specified scene elements.

[0249] For some embodiments of the example method, the action information may describe at least two actions, and the behavior information may include information indicating, for at least one behavior, an order in which at least two of the at least two actions are to be performed.

[0250] For some embodiments of the example method, the action information may describe at least two actions, and the behavior information may include information indicating that at least two of the at least two actions are to be performed simultaneously.

[0251] For some embodiments of the example method, at least one of the scene elements in the scene may be a virtual object.

[0252] In some embodiments, the example method may further include rendering a 3D scene according to the scene description data operated by the first action.

[0253] For some embodiments of the example method, the trigger information may include an array of two or more triggers in the 3D scene.

[0254] For some embodiments of the example method, the action information may include an array of two or more actions in the 3D scene.

[0255] For some embodiments of the example method, the behavior information may include an array of two or more behaviors in the 3D scene.

[0256] For some embodiments of the example methods, the scene description data may be provided in JSON format.

[0257] For some embodiments of the example methods, the scene description data may be provided in a GLTF format.

[0258] Another example method according to some embodiments may include: obtaining scene description data of a 3D scene, wherein the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the action, and behavior information, wherein the behavior information may associate at least one of the trigger conditions with at least one of the actions; and determining that a first action may be performed on at least a first node associated with the first action in response to the following logic: (i) a combination of at least one of the trigger conditions has been satisfied, and (ii) at least a first action of the actions is associated with the combination of the trigger conditions through the behavior information.

[0259] For some embodiments of another example method, the logical determination may further include determining that at least one trigger is activated.

[0260] For some embodiments of another example method, the logical determination may further include determining an activation state of a combination of at least one trigger, and performing the first action may include performing the first action one or more times based on the activation state of the combination.

[0261] For some embodiments of another example method, determining the combination of at least one triggering condition may include performing a logical operation on at least one of the triggering conditions.

[0262] For some embodiments of another example method, determining the combination of at least one trigger condition may include performing a logical OR operation of at least two trigger conditions.

[0263] For some embodiments of another example method, at least a first one of the trigger conditions may be a visibility condition that is satisfied when a specified scene element is visible to a specified camera node.

[0264] For some embodiments of another example method, at least a first one of the trigger conditions may be a proximity condition that is satisfied when a distance from a user camera to a specified scene element is within a specified boundary.

[0265] For some embodiments of another example method, at least a first one of the trigger conditions may be a user input condition that is satisfied when a specified user interaction is detected.

[0266] For some embodiments of another example method, at least a first one of the trigger conditions may be a timing condition that is satisfied during a specified time period.

[0267] For some embodiments of another example method, at least a first one of the trigger conditions may be a collider condition satisfied in response to detecting a conflict between specified scene elements.

[0268] For some embodiments of another example method, the behavior information may identify at least one behavior, and the behavior information for each behavior may identify at least one of the triggers and at least one of the actions.

[0269] For some embodiments of another example method, the action information may describe at least two actions, and the behavior information may include information indicating, for at least one behavior, an order in which at least two of the at least two actions are to be performed.

[0270] For some embodiments of another example method, the action information may describe at least two actions, and the behavior information may include information indicating that at least two of the at least two actions are to be performed simultaneously.

[0271] For some embodiments of another example method, at least one of the scene elements in the scene may be a virtual object.

[0272] In some embodiments, another example method may further include: rendering a 3D scene according to the scene description data operated by the first action.

[0273] For some embodiments of another example method, the trigger information may include an array of two or more triggers in the 3D scene.

[0274] For some embodiments of another example method, the action information may include an array of two or more actions in the 3D scene.

[0275] For some embodiments of another example method, the behavior information may include an array of two or more behaviors in the 3D scene.

[0276] For some embodiments of another example method, the scene description data may be provided in a JSON format.

[0277] For some embodiments of another example method, the scene description data may be provided in a GLTF format.

[0278] Another example method for rendering an extended reality scene relative to a user in a timed environment according to some embodiments may include: obtaining a description of the extended reality scene, which may include: a scene tree of linked nodes, the nodes describing timed objects, virtual objects, or relationships between objects; behavior data items, which may include: at least one trigger control parameter, which is a description of conditions associated with one or more triggers; activation conditions associated with the trigger control parameters; at least one action, which is a description of the processing performed by the extended reality engine on the objects described by the nodes of the scene tree; and applying the action of the behavior to the associated object under the condition that a logical combination of at least one of the triggers that trigger the behavior data item and the activation conditions associated with the trigger control parameters of the behavior are satisfied.

[0279] For some embodiments of another example method, the logical combination of at least one trigger may include: a logical operation on at least one of the trigger conditions.

[0280] For some embodiments of another example method, the logical combination of at least one trigger may include a logical OR operation of at least two trigger conditions.

[0281] Another example method according to some embodiments for updating a first description of an extended reality scene with a second description of the extended reality scene (which may include behavior data items and a second description of the extended reality scene) at runtime may include: for each ongoing behavior data item of the first description, if the ongoing behavior data item is not applicable to the second description: if the ongoing behavior data item has an interrupt action for the ongoing application in the first description, processing the interrupt action; stopping the ongoing behavior; and applying the second description.

[0282] An example device for rendering an extended reality scene relative to a user in a timed environment according to some embodiments may include: a memory associated with a processor, the processor being configured to: obtain a description of the extended reality scene, the description may include: a scene tree of linked nodes, the nodes describing timed objects, virtual objects, or relationships between objects; behavioral data items, the behavioral data items may include: at least one trigger, the trigger being a description of a condition; the trigger being activated when its condition is detected in the timed environment; and at least one action, the action being a description of processing performed by the extended reality engine on an object described by a node of the scene tree; and applying the action of the behavior to the associated object under the condition that the trigger of the behavior data item is activated.

[0283] For some embodiments of the example device, the processor may be further configured to: when obtaining a description of an extended reality scene, attribute an activation state set to false to at least one trigger of the description; when a condition of at least one trigger is met for the first time, set the activation state of the trigger to true; and when the condition of at least one trigger is met, activate the trigger.

[0284] For some embodiments of the example device, the processor may be further configured to: when a condition of at least one trigger is met, if the activation state of the trigger is set to true, activate the trigger only if the description of the trigger authorizes a second activation.

[0285] Another example device for updating a first description of an extended reality scene including behavior data items with a second description of the extended reality scene at runtime according to some embodiments may include: a memory associated with a processor, the processor being configured to: for each ongoing behavior data item of the first description, if the ongoing behavior data item is not applicable to the second description: if the ongoing behavior data item has an interruption action for the ongoing application in the first description, process the interruption action; stop the ongoing behavior; and apply the second description.

[0286] An example apparatus according to some embodiments may comprise one or more processors configured to perform the method of any of the preceding claims.

[0287] Example computer readable media according to some embodiments may include instructions for causing one or more processors to perform any of the methods listed above.

[0288] For some embodiments of the example computer-readable medium, the computer-readable medium may be a non-transitory storage medium.

[0289] An example computer program product according to some embodiments may include instructions that, when executed by one or more processors, cause the one or more processors to perform any of the methods listed above.

[0290] According to an example signal of scene description data including a 3D scene in some embodiments, the scene description data may include: scene element information describing each of multiple scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the action, and behavior information, wherein the behavior information may associate at least one of the trigger conditions with at least one of the actions.

[0291] According to another example computer-readable medium including scene description data of a 3D scene in some embodiments, the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the action, and behavior information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions.

[0292] A first example method according to some embodiments may include: acquiring scene description data of a 3D scene, wherein the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with at least one action, and behavior information, wherein the behavior information may include: a trigger list of at least one trigger, an action list of at least at least one action; trigger combination information and activation information, the trigger combination information indicating a combination of a first trigger condition corresponding to a first trigger in the trigger list and other trigger conditions corresponding to other triggers in the trigger list, the activation information indicating when to perform at least one action according to a result of the trigger combination information; and performing at least one action on one or more scene elements associated with at least one action in response to the following logic: (i) a combination of a first trigger condition corresponding to a first trigger in the trigger list and other trigger conditions corresponding to other triggers in the trigger list has produced a result, and (ii) activation information indicating that the result triggers the first trigger and the other triggers.

[0293] For some embodiments of the first example method, the trigger list may include one trigger, and there are no other triggers in the trigger list, the trigger combination information may include a unary operator operating on the one trigger, and the activation information may indicate when to execute the unary operator on the one trigger.

[0294] For some embodiments of the first example method, the trigger list may include at least two triggers, the trigger combination information may indicate how to combine the at least two trigger conditions, and the activation information may indicate when to perform at least one action according to a result of the trigger combination information.

[0295] For some embodiments of the first example method, performing the first action may include performing the first action one or more times based on the activation information.

[0296] For some embodiments of the first example method, determining a combination of a trigger condition of at least one trigger in the trigger list and trigger conditions of other triggers in the trigger list may include performing a logical operation on at least one of the trigger conditions.

[0297] For some embodiments of the first example method, determining a combination of a trigger condition of at least one trigger in the trigger list and trigger conditions of other triggers in the trigger list may include performing a logical OR operation of at least two trigger conditions.

[0298] For some embodiments of the first example method, at least a first one of the trigger conditions is a visibility condition that is satisfied when a specified scene element is visible to a specified camera node.

[0299] For some embodiments of the first example method, at least a first one of the trigger conditions is a proximity condition that is satisfied when a distance from a user camera to a specified scene element is within a specified boundary.

[0300] For some embodiments of the first example method, at least a first one of the trigger conditions is a user input condition that is satisfied when a specified user interaction is detected.

[0301] For some embodiments of the first example method, at least a first one of the trigger conditions is a timing condition that is satisfied during a specified time period.

[0302] For some embodiments of the first example method, at least a first one of the trigger conditions is a collider condition that is satisfied in response to detecting a conflict between specified scene elements.

[0303] For some embodiments of the first example method, the action information may describe at least two actions, and the behavior information may include information indicating, for at least one behavior, an order in which at least two of the at least two actions are to be performed.

[0304] For some embodiments of the first example method, the action information may describe at least two actions, and the behavior information may include information indicating that at least two of the at least two actions are to be performed simultaneously.

[0305] For some embodiments of the first example method, at least one of the scene elements in the scene is a virtual object.

[0306] Some embodiments of the first example method may further include rendering a 3D scene according to the scene description data operated by the first action.

[0307] For some embodiments of the first example method, the trigger information may include an array of two or more triggers in the 3D scene.

[0308] For some embodiments of the first example method, the action information may include an array of two or more actions in the 3D scene.

[0309] For some embodiments of the first example method, the behavior information may include an array of two or more behaviors in the 3D scene.

[0310] For some embodiments of the first example method, the scene description data may be provided in a JSON format.

[0311] For some embodiments of the first example method, the scene description data may be provided in a GLTF format.

[0312] A first example method / apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions which, when executed by the processor, are operable to cause the apparatus to perform any of the methods shown above.

[0313] A second example method / apparatus according to some embodiments may include: acquiring scene description data of a 3D scene, wherein the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with at least one action, and behavior information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions; and performing a first action on at least a first node associated with the first action in response to the following logical determination: (i) a combination of at least one of the trigger conditions has been satisfied, and (ii) a first action of at least one action is associated with the combination of the trigger conditions through the behavior information.

[0314] For some embodiments of the second example method, the logical determination may further include determining that a trigger associated with the at least one trigger condition is activated.

[0315] For some embodiments of the second example method, the logical determination may further include determining an activation state of a combination of at least one trigger condition, and performing the first action may include performing the first action one or more times based on the activation state.

[0316] For some embodiments of the second example method, determining the combination of at least one triggering condition may include performing a logical operation on at least one of the triggering conditions.

[0317] For some embodiments of the second example method, determining the combination of at least one trigger condition may include performing a logical OR operation on at least two trigger conditions.

[0318] For some embodiments of the second example method, at least a first of the at least one triggering condition is a visibility condition that is satisfied when the specified scene element is visible to the specified camera node.

[0319] For some embodiments of the second example method, at least a first trigger condition of the at least one trigger condition is a proximity condition that is satisfied when a distance from a user camera to a specified scene element is within a specified boundary.

[0320] For some embodiments of the second example method, at least a first trigger condition of the at least one trigger condition is a user input condition that is satisfied when a specified user interaction is detected.

[0321] For some embodiments of the second example method, at least a first trigger condition of the at least one trigger condition is a timing condition that is satisfied during a specified time period.

[0322] For some embodiments of the second example method, at least a first trigger condition of the at least one trigger condition is a collider condition satisfied in response to detecting a conflict between specified scene elements.

[0323] For some embodiments of the second example method, the behavior information may identify at least one behavior, the behavior information for each of the at least one behavior identifying a trigger associated with one of the at least one trigger condition and one of the at least one action.

[0324] For some embodiments of the second example method, the action information may describe at least two actions, and the behavior information may include information indicating, for at least one behavior, an order in which at least two of the at least two actions are to be performed.

[0325] For some embodiments of the second example method, the action information may describe at least two actions, and the behavior information may include information indicating that at least two of the at least two actions are to be performed simultaneously.

[0326] For some embodiments of the second example method, at least one of the scene elements in the scene is a virtual object.

[0327] Some embodiments of the second example method may further include rendering a 3D scene according to the scene description data operated by the first action.

[0328] For some embodiments of the second example method, the trigger information may include an array of two or more triggers in the 3D scene.

[0329] For some embodiments of the second example method, the action information may include an array of two or more actions in the 3D scene.

[0330] For some embodiments of the second example method, the behavior information may include an array of two or more behaviors in the 3D scene.

[0331] For some embodiments of the second example method, the scene description data may be provided in a JSON format.

[0332] For some embodiments of the second example method, the scene description data may be provided in a GLTF format.

[0333] A second example method / apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0334] According to a third example method of some embodiments, which is a method for rendering an extended reality scene relative to a user in a timed environment, the method may include: obtaining a description of the extended reality scene, the description may include: a scene tree of linked nodes, the nodes describing timed objects, virtual objects or relationships between objects; behavior data items, the behavior data items may include: at least one trigger control parameter, the trigger control parameter is a description of conditions associated with one or more triggers; activation conditions associated with the trigger control parameters; at least one action, which is an action that describes the processing performed by the extended reality engine on the object described by the node of the scene tree; and applying the action of the behavior data item to the associated object under the condition that a logical combination of at least one of the triggers that trigger the behavior data item and the activation conditions associated with the trigger control parameters of the behavior data item are satisfied.

[0335] For some embodiments of the third example method, the logical combination of at least one trigger may include a logical operation on at least one of the conditions associated with the one or more triggers.

[0336] For some embodiments of the third example method, the logical combination of at least one trigger may include a logical OR operation of at least two conditions associated with the one or more triggers.

[0337] A third example method / apparatus according to some embodiments may include: a processor; and a non-transitory computer readable medium storing instructions which, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0338] According to some embodiments, a fourth example method is a method for updating a first description of an extended reality scene including behavior data items with a second description of the extended reality scene at runtime. The fourth example method may include, for each ongoing behavior data item of the first description, if the ongoing behavior data item is not applicable to the second description: if the ongoing behavior data item has an interruption action for the ongoing application in the first description, processing the interruption action; stopping the ongoing behavior; and applying the second description.

[0339] According to a fifth example device of some embodiments, which is a device for rendering an extended reality scene relative to a user in a timed environment, the device may include: a memory associated with a processor, the processor being configured to: obtain a description of the extended reality scene, the description including: a scene tree of linked nodes, the nodes describing timed objects, virtual objects, or relationships between objects; behavior data items, the behavior data items may include: at least one trigger, the trigger is a description of a condition; the trigger is activated when its condition is detected in the timed environment; and at least one action, the action is a description of the processing performed by the extended reality engine on the object described by the node of the scene tree; and under the condition that the trigger of the behavior data item is activated, applying the action of the behavior to the associated object.

[0340] For some embodiments of the fifth example device, the processor is further configured to: when obtaining a description of an extended reality scene, attribute an activation state set to false to at least one trigger of the description; when a condition of at least one trigger is met for the first time, set the activation state of the trigger to true; and when the condition of at least one trigger is met, activate the trigger.

[0341] For some embodiments of the fifth example apparatus, the processor is further configured to: when a condition of at least one trigger is met, if the activation state of the trigger is set to true, activate the trigger only if the description of the trigger authorizes a second activation.

[0342] A sixth example device according to some embodiments, which is a device for updating a first description of an extended reality scene including behavior data items with a second description of the extended reality scene at runtime, may include: a memory associated with a processor, the processor being configured to: for each ongoing behavior data item of the first description, if the ongoing behavior data item is not applicable to the second description: if there is an interrupt action for the ongoing application in the first description, process the interrupt action; stop the ongoing behavior; and apply the second description.

[0343] A seventh example apparatus according to some embodiments may include one or more processors configured to perform any of the methods listed above.

[0344] An eighth example apparatus according to some embodiments may include a computer readable medium including instructions for causing one or more processors to perform any of the methods listed above.

[0345] For some embodiments of the eighth example apparatus, the computer-readable medium is a non-transitory storage medium.

[0346] A tenth example apparatus according to some embodiments may include a computer program product including instructions which, when executed by one or more processors, cause the one or more processors to perform any of the methods listed above.

[0347] An eleventh example apparatus according to some embodiments may include: a signal including scene description data of a 3D scene, wherein the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the action, and behavior information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions.

[0348] A twelfth example method / apparatus according to some embodiments may include a computer-readable medium comprising scene description data for a 3D scene, wherein the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the action, and behavior information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions.

[0349] According to another exemplary computer-readable medium including scene description data of a 3D scene in some embodiments, the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the action, and behavior information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions.

[0350] The present disclosure describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are specifically described, and at least in order to illustrate individual features, are usually described in a manner that may sound restrictive. However, this is for the purpose of describing clearly, and does not limit the disclosure or scope of those aspects. In fact, all different aspects can be combined and interchanged to provide further aspects. In addition, these aspects can also be combined and interchanged with the aspects described in the earlier applications.

[0351] The aspects described and contemplated in this disclosure may be implemented in many different forms. Although some embodiments are specifically shown, other embodiments are contemplated, and the discussion of specific embodiments does not limit the breadth of implementation. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects may be implemented as methods, devices, computer-readable storage media having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having stored thereon a bitstream generated according to any of the described methods.

[0352] In this disclosure, the terms "reconstruction" and "decoding" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image", "picture" and "frame" may be used interchangeably. Typically, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.

[0353] The terms HDR (high dynamic range) and SDR (standard dynamic range) generally convey to one of ordinary skill in the art a specific value of dynamic range. However, additional embodiments are also intended in which references to HDR are understood to mean "higher dynamic range" and references to SDR are understood to mean "lower dynamic range". Such additional embodiments are not constrained by any specific value of dynamic range that may often be associated with the terms "high dynamic range" and "standard dynamic range".

[0354] Various methods are described herein, and each method includes one or more steps or actions for realizing the described method.Unless the correct operation of the method requires the steps or actions of a specific order, the order and / or use of specific steps and / or actions can be modified or combined.In addition, terms such as "first", "second" etc. can be used to modify elements, components, steps, operations, etc. in various embodiments, such as, for example, "first decoding" and "second decoding".Unless specifically needed, the use of these terms does not mean the sequencing of the operation to modification.Therefore, in this example, the first decoding does not need to be performed before the second decoding, and can, for example, occur before, during, or in a time period overlapping with the second decoding.

[0355] For example, various numerical values ​​may be used in the present disclosure. The specific values ​​are for example purposes, and the described aspects are not limited to these specific values.

[0356] The embodiments described herein may be performed by computer software implemented by a processor or other hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The processor may be of any type suitable for the technical environment, and may include one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture as non-limiting examples.

[0357] Various implementations involve decoding. "Decoding" as used in the present disclosure may encompass, for example, all or part of the processing performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such processing includes one or more of the processing typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processing also or alternatively includes processing performed by the decoder of the various embodiments described in the present disclosure, such as extracting a picture from a tiled (packed) picture, determining the upsampling filter to be used and then upsampling the picture, and flipping the picture back to its intended orientation.

[0358] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Based on the context of the particular description, it will be clear whether the phrase "decoding process" is intended to refer specifically to a subset of operations or generally to a broader decoding process.

[0359] Various embodiments relate to encoding. In a manner similar to the above discussion of "decoding", "encoding" as used in the present disclosure may encompass, for example, all or part of the processing performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processing includes one or more processes typically performed by an encoder, such as partitioning, differential encoding, transforms, quantization, and entropy encoding. In various embodiments, such processing also or alternatively includes processing performed by the encoder of the various embodiments described in the present invention.

[0360] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Based on the context of the particular description, it will be clear whether the phrase "encoding process" is intended to refer specifically to a subset of operations or generally to a broader encoding process.

[0361] Various embodiments relate to rate-distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is usually considered, often given constraints on computational complexity. Rate-distortion optimization is usually formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. There are different methods to solve the rate-distortion optimization problem. For example, these methods can be based on extensive testing of all coding options, including all considered modes or coding parameter values, and a complete evaluation of their coding costs and the associated distortion of the reconstructed signal after encoding and decoding. A faster approach can also be used to save coding complexity, in particular, to calculate approximate distortion based on a prediction or prediction residual signal rather than on a reconstructed signal. A mixture of these two approaches can also be used, such as by using approximate distortion only for some possible coding options and using full distortion for other coding options. Other methods only evaluate a subset of possible coding options. More generally, many methods use any of a variety of techniques to perform optimization, but the optimization is not necessarily a complete evaluation of both coding cost and associated distortion.

[0362] When the figures are presented as flow charts, it should be understood that they also provide block diagrams of corresponding apparatuses. Similarly, when the figures are presented as block diagrams, it should be understood that they also provide flow charts of corresponding methods / processes.

[0363] The embodiments and aspects described herein can be implemented in, for example, methods or processes, devices, software programs, data streams, or signals. Even if only discussed in the context of a single implementation form (e.g., discussed only as a method), the implementation of the features discussed can also be implemented in other forms (e.g., devices or programs). The device can be implemented in, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as, for example, a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.

[0364] Reference to "one embodiment" or "an embodiment" or "an implementation" or "implementation" and other variations thereof means that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" and any other variations thereof appearing in various places throughout this disclosure are not necessarily all referring to the same embodiment.

[0365] Additionally, the present disclosure may involve “determining” various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.

[0366] Additionally, the present disclosure may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0367] Additionally, the present disclosure may involve "receiving" various information. Like "accessing," receiving is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, during operations such as, for example, storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, "receiving" is generally involved in one way or another.

[0368] It should be understood that, for example, in the case of "A / B," "A and / or B," and "at least one of A and B," the use of any of the following " / ," "and / or," and "at least one of" is intended to cover selection of only the first listed option (A), or only the second listed option (B), or both options (A and B). As another example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such wording is intended to cover selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). This can be extended to as many items as listed.

[0369] In addition, as used herein, the word "signal" refers to a corresponding decoder indicating something in addition to other things. For example, in some embodiments, the encoder signals a specific one of the multiple parameters of the region-based filter parameter selection for de-artifact filtering. In this way, in an embodiment, the same parameters are used at both the encoder side and the decoder side. Therefore, for example, the encoder can send (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. On the contrary, if the decoder already has specific parameters and other parameters, signaling can be used without sending (implicit signaling) to simply allow the decoder to know and select specific parameters. By avoiding the transmission of any actual function, bit saving is achieved in various embodiments. It should be understood that signaling can be completed in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the verb form of the word "signal" is related to the above, the word "signal" can also be used as a noun in this article.

[0370] Embodiments may generate various signals formatted to carry information that may, for example, be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described embodiments. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor readable medium.

[0371] We have described a number of embodiments. Features of these embodiments may be provided individually or in any combination across various claim categories and types. In addition, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination across various claim categories and types:

[0372] • A bitstream or signal comprising one or more of the described syntax elements, or variations thereof.

[0373] • A bitstream or signal comprising syntax conveying information generated according to any of the described embodiments.

[0374] • Creating and / or sending and / or receiving and / or decoding a bitstream or signal comprising one or more of the described syntax elements or variants thereof.

[0375] • Creating and / or sending and / or receiving and / or decoding according to any of the described embodiments.

[0376] ●A method, process, apparatus, medium storing instructions, medium storing data, or signal according to any of the described embodiments.

[0377] Note that one or more of the various hardware elements in the described embodiments are referred to as "modules" that perform (i.e., perform, execute, etc.) the various functions described herein in conjunction with the corresponding modules. As used herein, a module includes hardware that is considered suitable for a given implementation (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices). Each described module may also include executable instructions for performing one or more functions described as being performed by the corresponding module, and it should be noted that those instructions may take the form of hardware (i.e., hard-wired) instructions, firmware instructions, software instructions, and / or the like, or include hardware (i.e., hard-wired) instructions, firmware instructions, software instructions, and / or the like, and may be stored in any suitable non-transitory computer-readable medium or media, such as what is commonly referred to as RAM, ROM, and the like.

[0378] Although the features and elements are described above in specific combinations, each feature or element may be used alone or in any combination with other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memory, semiconductor storage devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor associated with the software may be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.

[0379] Note that one or more of the various hardware elements in the described embodiments are referred to as "modules" that implement (i.e., perform, execute, etc.) the various functions described herein in conjunction with the corresponding modules. As used herein, a module includes hardware that a person skilled in the relevant art deems suitable for a given implementation (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices). Each described module may also include executable instructions for performing one or more functions described as being performed by the corresponding module, and it should be noted that those instructions may take the form of hardware (i.e., hard-wired) instructions, firmware instructions, software instructions, and / or the like, or include hardware (i.e., hard-wired) instructions, firmware instructions, software instructions, and / or the like, and may be stored in any suitable non-transitory computer-readable medium or media, such as what is commonly referred to as RAM, ROM, and the like.

[0380] Although the features and elements are described above in specific combinations, it will be understood by those skilled in the art that each feature or element may be used alone or in any combination with other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memory, semiconductor storage devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor associated with the software may be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A method comprising: Get the scene description data of the 3D scene, The scene description data includes: describing scene element information for each of a plurality of scene elements in the scene, Trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the at least one action, and Behavioral information, The behavior information includes: a trigger list of at least one trigger, an action list of at least the at least one action; trigger combination information indicating a combination of a first trigger condition corresponding to a first trigger in the trigger list and other trigger conditions corresponding to other triggers in the trigger list, and activation information indicating when to perform the at least one action according to a result of the trigger combination information; and The at least one action is performed on one or more scene elements associated with the at least one action in response to the following logic determination: (i) the combination of the first trigger condition corresponding to the first trigger in the trigger list and the other trigger conditions corresponding to the other triggers in the trigger list has produced the result, and (ii) the activation information indicating that the result triggers the first trigger and the other triggers.

2. The method according to claim 1, in, the trigger list includes one trigger, and no other triggers in the trigger list exist, The trigger combination information includes a unary operator for operating the one trigger, and The activation information indicates when to execute the unary operator on the one trigger.

3. The method according to claim 1, in, The trigger list includes at least two triggers, The trigger combination information indicates how to combine the at least two trigger conditions, and The activation information indicates when to execute the at least one action according to a result of the trigger combination information.

4. The method according to any one of claims 1 to 3, wherein performing the first action comprises: The first action is performed one or more times based on the activation information.

5. The method according to any one of claims 1 to 4, wherein: Determining a combination of the triggering condition of the at least one trigger in the trigger list and the triggering conditions of other triggers in the trigger list includes: performing a logical operation on at least one of the triggering conditions.

6. The method according to any one of claims 1 to 4, wherein: Determining a combination of the triggering condition of the at least one trigger in the trigger list and the triggering conditions of other triggers in the trigger list includes: performing a logical OR operation of at least two triggering conditions.

7. The method according to any one of claims 1 to 6, wherein: At least a first one of the trigger conditions is a visibility condition that is satisfied when a specified scene element is visible to a specified camera node.

8. The method according to any one of claims 1 to 7, wherein: At least a first one of the trigger conditions is a proximity condition that is satisfied when a distance from a user camera to a specified scene element is within a specified boundary.

9. The method according to any one of claims 1 to 8, wherein: At least a first trigger condition among the trigger conditions is a user input condition that is satisfied when a specified user interaction is detected.

10. The method according to any one of claims 1 to 9, wherein: At least a first one of the trigger conditions is a timing condition that is satisfied during a specified time period.

11. The method according to any one of claims 1 to 10, wherein: At least a first one of the trigger conditions is a conflictor condition satisfied in response to detecting a conflict between specified scene elements.

12. The method according to any one of claims 1 to 11, in, The action information describes at least two actions, and The behavior information includes information indicating, for at least one behavior, an order in which at least two of the at least two actions are to be performed.

13. The method according to any one of claims 1 to 11, in, The action information describes at least two actions, and The behavior information includes information indicating that at least two of the at least two actions are to be executed simultaneously.

14. The method according to any one of claims 1 to 13, wherein: At least one of the scene elements in the scene is a virtual object.

15. The method according to any one of claims 1 to 14, further comprising: The 3D scene is rendered according to the scene description data operated by the first action.

16. The method according to any one of claims 1 to 15, wherein: The trigger information includes an array of two or more triggers in the 3D scene.

17. The method according to any one of claims 1 to 16, wherein: The action information includes an array of two or more actions in the 3D scene.

18. The method according to any one of claims 1 to 17, wherein: The behavior information includes an array of two or more behaviors in the 3D scene.

19. The method according to any one of claims 1 to 18, wherein: The scene description data is provided in JSON format.

20. The method according to any one of claims 1 to 19, wherein: The scene description data is provided in GLTF format.

21. An apparatus comprising: processor; as well as A non-transitory computer readable medium storing instructions which, when executed by the processor, are operable to cause the apparatus to perform the method according to any one of claims 1 to 20.

22. A method comprising: Get the scene description data of the 3D scene, The scene description data includes: describing scene element information for each of a plurality of scene elements in the scene, Trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the at least one action, and Behavioral information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions; and The first action is performed on at least a first node associated with the first action in response to the following logic determination: (i) a combination of at least one of the trigger conditions has been satisfied, and (ii) a first action of the at least one action is associated with the combination of the trigger conditions through the behavior information.

23. The method according to claim 22, wherein: The logical determination also includes determining that a trigger associated with the at least one trigger condition is activated.

24. The method according to any one of claims 22-23, in, The logical determination further comprises determining an activation state of a combination of the at least one trigger condition, and The performing of the first action includes: performing the first action once or multiple times based on the activation state.

25. The method according to any one of claims 22 to 24, wherein: Determining the combination of the at least one triggering condition includes performing a logical operation on at least one of the triggering conditions.

26. The method according to any one of claims 22 to 24, wherein: Determining the combination of the at least one trigger condition includes performing a logical OR operation on the at least two trigger conditions.

27. The method according to any one of claims 22 to 26, wherein: At least a first trigger condition of the at least one trigger condition is a visibility condition that is satisfied when a specified scene element is visible to a specified camera node.

28. The method according to any one of claims 22 to 27, wherein: At least a first trigger condition of the at least one trigger condition is a proximity condition that is satisfied when a distance from a user camera to a specified scene element is within a specified boundary.

29. The method according to any one of claims 22 to 26, wherein: At least a first trigger condition of the at least one trigger condition is a user input condition that is satisfied when a specified user interaction is detected.

30. The method according to any one of claims 22 to 26, wherein: At least a first trigger condition of the at least one trigger condition is a timing condition satisfied during a specified time period.

31. The method according to any one of claims 22 to 26, wherein: At least a first trigger condition of the at least one trigger condition is a conflictor condition satisfied in response to detecting a conflict between specified scene elements.

32. The method according to any one of claims 22 to 31, wherein: The behavior information identifies at least one behavior, and the behavior information of each of the at least one behavior identifies a trigger associated with one of the at least one trigger condition and one of the at least one action.

33. The method according to any one of claims 22 to 32, in, The action information describes at least two actions, and The behavior information includes information indicating, for at least one behavior, an order in which at least two of the at least two actions are to be performed.

34. The method according to any one of claims 22 to 32, in, The action information describes at least two actions, and The behavior information includes information indicating that at least two of the at least two actions are to be executed simultaneously.

35. The method according to any one of claims 22 to 34, wherein: At least one of the scene elements in the scene is a virtual object.

36. The method according to any one of claims 22 to 35, further comprising: The 3D scene is rendered according to the scene description data operated by the first action.

37. The method according to any one of claims 22 to 36, wherein: The trigger information includes an array of two or more triggers in the 3D scene.

38. The method according to any one of claims 22 to 37, wherein: The action information includes an array of two or more actions in the 3D scene.

39. The method according to any one of claims 22 to 38, wherein: The behavior information includes an array of two or more behaviors in the 3D scene.

40. The method according to any one of claims 22 to 39, wherein: The scene description data is provided in JSON format.

41. The method according to any one of claims 22 to 40, wherein: The scene description data is provided in GLTF format.

42. An apparatus comprising: processor; as well as A non-transitory computer readable medium storing instructions which, when executed by the processor, are operable to cause the apparatus to perform the method of any one of claims 22 to 41.

43. A method for rendering an extended reality scene relative to a user in a timed environment, the method comprising: Obtain a description of the extended reality scene, the description comprising: A scene tree of linked nodes describing timed objects, virtual objects, or relationships between objects; Behavior data items, the behavior data items comprising: at least one trigger control parameter, the trigger control parameter being a description of a condition associated with one or more triggers; an activation condition associated with said trigger control parameter; at least one action that is a description of a process performed by an extended reality engine on an object described by a node of the scene tree; and Upon triggering a logical combination of at least one of the triggers of the behavior data item and satisfying the activation condition associated with the trigger control parameter of the behavior data item, applying the action of the behavior data item to an associated object.

44. The method of claim 43, wherein the logical combination of at least one trigger comprises a logical operation on at least one of the conditions associated with the one or more triggers.

45. The method of claim 43, wherein: The logical combination of the at least one trigger includes a logical OR operation of at least two conditions associated with the one or more triggers.

46. ​​An apparatus comprising: processor; as well as A non-transitory computer readable medium storing instructions which, when executed by the processor, are operable to cause the apparatus to perform the method of any one of claims 43 to 45.

47. A method for updating a first description of an extended reality scene including behavior data items with a second description of the extended reality scene at runtime, the method comprising, for each ongoing behavior data item of the first description, if the ongoing behavior data item is not applicable to the second description: If an interruption action for the ongoing application exists in the first description, processing the interruption action; Stop the ongoing conduct; and The second description is applied.

48. A device for rendering an extended reality scene relative to a user in a timed environment, the device comprising a memory associated with a processor, the processor configured to: Obtain a description of the extended reality scene, the description comprising: A scene tree of linked nodes describing timed objects, virtual objects, or relationships between objects; Behavior data items, including: At least one trigger, which is a description of a condition; A trigger is activated when its condition is detected in the timing environment; and at least one action, an action being a description of processing performed by an extended reality engine on an object described by a node of the scene tree; and Under the condition that the trigger of the behavior data item is activated, the action of the behavior is applied to the associated object.

49. The apparatus of claim 48, wherein: The processor is further configured to: When obtaining a description of the augmented reality scene, attributing an activation state set to false to at least one trigger of the description; When a condition of the at least one trigger is met for the first time, setting the activation state of the trigger to true; as well as When a condition of the at least one trigger is met, the trigger is activated.

50. The apparatus of claim 49, wherein: The processor is further configured to, when a condition of the at least one trigger is met, if an activation state of the trigger is set to true, activate the trigger only if the description of the trigger authorizes a second activation.

51. An apparatus for updating, at runtime, a first description of an extended reality scene including a behavior data item with a second description of the extended reality scene, the apparatus comprising a memory associated with a processor, the processor being configured to: For each ongoing behavior data item of the first description, if the ongoing behavior data item is not applicable to the second description: If an interruption action for the ongoing application exists in the first description, processing the interruption action; Stop the ongoing conduct; and The second description is applied.

52. An apparatus comprising one or more processors configured to perform the method of any one of claims 1-20, 22-41, 43-45, and 47.

53. A computer-readable medium comprising instructions for causing one or more processors to perform the method of any one of claims 1-20, 22-41, 43-45, and 47.

54. The computer readable medium of claim 53, wherein: The computer readable medium is a non-transitory storage medium.

55. A computer program product comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1-20, 22-41, 43-45, and 47.

56. A signal comprising scene description data of a 3D scene, wherein: The scene description data includes: describing scene element information for each of a plurality of scene elements in the scene, Trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the action, and Behavior information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions.

57. A computer readable medium comprising scene description data of a 3D scene, wherein: The scene description data includes: describing scene element information for each of a plurality of scene elements in the scene, Trigger information describing at least one trigger condition, action information describing at least one action to be performed on one or more scene elements associated with the action, and Behavior information, wherein the behavior information associates at least one of the trigger conditions with at least one of the actions.