Generic condition triggering in virtual environment

By introducing behavioral information and conditional triggering mechanisms into the extended reality scene description, the problem of insufficient user interactivity is solved, enabling users to have a specific immersive experience and dynamic rendering, and enhancing the real-time interactive capabilities of virtual objects.

CN121794656APending Publication Date: 2026-04-03INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing extended reality technologies are inadequate in terms of user interactivity and immersive experience, and cannot effectively support real-time interaction and dynamic rendering between users and virtual objects.

Method used

By introducing behavioral information, including triggering conditions and action information, into the scene description data, a condition-triggered mechanism is implemented, allowing actions to be updated and executed at runtime, thereby enhancing the interactivity between users and virtual objects.

Benefits of technology

It enables users to have a specific immersive extended reality experience, supports real-time interaction and dynamic rendering, and improves the user's ability to interact with the virtual environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121794656A_ABST
    Figure CN121794656A_ABST
Patent Text Reader

Abstract

Some embodiments of an example method may include obtaining scene description data for a three-dimensional (3D) scene, where the scene description data includes behavioral information including: trigger information describing at least one trigger condition, and action information describing an action performed on a scene element in the 3D scene, where the action information includes action information describing an action performed on the scene element in the 3D scene; the at least one triggering condition comprises condition triggering; and in response to determining that the at least one trigger condition has occurred, performing the action on the scene element.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims the benefit of European patent application No. EP23306182, filed on July 11, 2023, entitled “GENERICCONDITIONAL TRIGGER IN VIRTUAL ENVIRONMENTS”, which is incorporated herein by reference in its entirety. Background Technology

[0002] Extended Reality (XR) is a technology that enables interactive experiences by enhancing real-world environments and / or video content with virtual content that can be defined across multiple sensory modalities, including visual, auditory, tactile, and other sensory modalities. During application runtime, virtual content (e.g., 3D content or audio / video files) is rendered in real-time in a manner consistent with the user's context (environment, perspective, device, etc.). Scene graphs (e.g., scene graphs proposed by Khronos / glTF and their extensions defined in the MPEG scene description format or Apple / USDZ) are possible ways to represent the content to be rendered. They combine, on the one hand, an explanatory description of the scene structure linking real-world objects and virtual objects, and on the other hand, a binary representation of the virtual content. Summary of the Invention

[0003] The embodiments described herein include methods used in video encoding and decoding (collectively, “code processing”).

[0004] An example method according to some embodiments may include: obtaining scene description data of a three-dimensional (3D) scene, wherein the scene description data includes behavioral information, the behavioral information including: trigger information describing at least one triggering condition, and action information describing an action performed on a scene element in the 3D scene, wherein the at least one triggering condition includes a conditional trigger; and wherein the trigger information includes a path to an attribute, wherein the attribute is used in conjunction with the conditional trigger; and in response to determining that the at least one triggering condition has occurred, performing the action on the scene element.

[0005] In some embodiments of the example method, the at least one triggering condition includes a test of the attribute.

[0006] An example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the device to perform the method of any of the above-listed claims.

[0007] An example method according to some embodiments may include: obtaining scene description data of a three-dimensional (3D) scene, wherein the scene description data includes behavioral information, the behavioral information including: trigger information describing at least one triggering condition, and action information describing an action performed on a scene element in the 3D scene, wherein the at least one triggering condition may include a conditional trigger; and performing the action on the scene element in response to determining that the at least one triggering condition has occurred.

[0008] In some embodiments of the example method, the at least one conditional trigger may include a path to a storage location, and the trigger information describing the at least one triggering condition is stored in the storage location.

[0009] Some embodiments of the example method may also include updating the triggering information at runtime.

[0010] Some embodiments of the example method may also include updating the triggering information before runtime.

[0011] Some embodiments of the example method may also include updating the condition trigger at runtime.

[0012] Some embodiments of the example method may also include updating the condition trigger before runtime.

[0013] In some embodiments of the example method, the triggering information may include one or more parameters, and the at least one triggering condition may be based on at least one of the one or more parameters.

[0014] In some embodiments of the example method, at least one of the one or more parameters includes a value field and a comparator field.

[0015] In some embodiments of the example method, the triggering information further includes node information indicating one or more nodes, and the method further includes: searching in a memory structure corresponding to the one or more nodes for a field that matches the value field; and updating the conditional trigger based on the matched field.

[0016] In some embodiments of the example method, the triggering information further includes a descriptor field, and the method further includes using the descriptor field in conjunction with the one or more parameters to determine the at least one triggering condition.

[0017] Some embodiments of the example method may further include: detecting a second condition; and updating the triggering information in response to detecting the second condition.

[0018] In some embodiments of the example method, the triggering information may also include at least one metadata field.

[0019] In some embodiments of the example method, the at least one triggering condition is based on the at least one metadata field.

[0020] In some embodiments of the example method, the conditional triggering may include two or more conditions.

[0021] In some embodiments of the example method, the at least one condition trigger includes a path to the parameter to be compared, and the at least one trigger condition includes the parameter to be compared.

[0022] In some embodiments of the example method, the at least one triggering condition includes a test of the attribute.

[0023] An example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0024] Another example method according to some embodiments may include handling conditional triggers associated with a three-dimensional (3D) scene, wherein the conditional triggers may include a mechanism for updating one or more triggering conditions associated with the conditional triggers.

[0025] An example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the device to perform the methods listed above.

[0026] Another example method according to some embodiments may include: obtaining scene description data of a three-dimensional (3D) scene, wherein the scene description data includes behavioral information, the behavioral information including: trigger information describing at least one triggering condition, and action information describing an action performed on a scene element in the 3D scene, wherein the at least one triggering condition includes a conditional trigger; and wherein the trigger information includes a path to an attribute, wherein the attribute is used in conjunction with the conditional trigger; and in response to determining that the at least one triggering condition has occurred, performing the action on the scene element.

[0027] In some embodiments of yet another example method, the at least one triggering condition includes a test of the attribute.

[0028] Another example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0029] In another embodiment, an encoder device and a decoder device are provided to perform the methods described herein. The encoder device or decoder device may include a processor configured to perform the methods described herein. The device may include a computer-readable medium (e.g., a non-transitory medium) storing instructions for performing the methods described herein. In some embodiments, the computer-readable medium (e.g., a non-transitory medium) stores video encoded using any of the methods described herein.

[0030] One or more of these embodiments also provide a computer-readable storage medium having instructions stored thereon for performing bidirectional optical flow, encoding or decoding video data according to any of the methods described above. This embodiment also provides a computer-readable storage medium having a bitstream generated according to the methods described above stored thereon. This embodiment also provides a method and apparatus for transmitting a bitstream generated according to the methods described above. This embodiment also provides a computer program product including instructions for performing any of the described methods. Attached Figure Description

[0031] Figure 1A This is a schematic side view illustrating an example waveguide display that can be used with extended reality (XR) applications according to some embodiments.

[0032] Figure 1B This is a schematic side view illustrating an example alternative display type that can be used with extended reality applications according to some embodiments.

[0033] Figure 1C This is a schematic side view illustrating an example alternative display type that can be used with extended reality applications according to some embodiments.

[0034] Figure 1D This is a system diagram illustrating a set of example interfaces for a system according to some embodiments.

[0035] Figure 1E This is a system diagram illustrating a set of example interfaces for a scene description (stored as an item in glTF.json) in an ISOBMFF file, three video tracks, an audio track, and a JSON patch update track, according to some embodiments.

[0036] Figure 2 This is a system diagram illustrating a set of example interfaces for an MPEG-I node hierarchy of elements supporting scene interactivity, according to some embodiments.

[0037] Figure 3 This is a block diagram illustrating an example of the logical relationship between trigger information (describing triggers 1 to n), action information (describing actions 1 to m), and behavior information (describing the relationship between triggers and actions) according to some embodiments, wherein triggers and actions can refer to one or more nodes in a scene description, such as a hierarchical scene graph.

[0038] Figure 4 This is a schematic plan view illustrating example relationships between objects describing an extended reality scene according to some embodiments.

[0039] Figure 5 This is a flowchart illustrating an example preprocessing step triggered by conditions according to some embodiments.

[0040] Figure 6 This is a flowchart illustrating an example preprocessing step triggered by conditions according to some embodiments.

[0041] Figure 7 This is a flowchart illustrating an example preprocessing step triggered by conditions according to some embodiments.

[0042] Figures 8A to 8B A list of code forms an example code structure illustrating basic scene elements according to some embodiments.

[0043] Figure 9 This is a list of code examples illustrating the structure of condition-triggered metadata according to some embodiments.

[0044] Figures 10A to 10B A list of code forms an example code structure that illustrates condition-triggered events according to some embodiments.

[0045] Figure 11 This is a flowchart illustrating an example process for handling condition-triggered events according to some embodiments.

[0046] The entities, connections, arrangements, etc., depicted and described in conjunction with the various figures are presented by way of example rather than limitation. Therefore, any and all statements or other indications regarding what a particular figure “depicts,” what a particular element or entity in a particular figure “is” or “has,” and any and all similar statements that might be interpreted in isolation and out of context as absolute and therefore limiting, should only be correctly interpreted as constructively beginning with a clause such as “In at least one embodiment, …”. For the sake of simplicity and clarity, this implicit leading clause will not be repeated in the detailed description. Annoying . Detailed Implementation

[0047] Figure 1A This is a schematic side view illustrating an example waveguide display that can be used with extended reality (XR) applications according to some embodiments. An image is projected by an image generator 102. The image generator 102 can use one or more of a variety of techniques for projecting images. For example, the image generator 102 can be a laser beam scanning (LBS) projector, a liquid crystal display (LCD), a light-emitting diode (LED) display (including organic LED (OLED) or micro LED (µLED) displays), a digital light processor (DLP), a liquid crystal on silicon (LCoS) display, or other types of image generators or light engines.

[0048] The light representing the image 112 generated by the image generator 102 is coupled into the waveguide 104 by a diffraction input coupler 106. The input coupler 106 diffracts the light representing the image 112 into one or more diffraction orders. For example, ray 108, which is part of the bottom of the image, is diffracted by the input coupler 106, and one of the diffraction orders 110 (e.g., second order) is at an angle that allows it to propagate through the waveguide 104 by total internal reflection. The image generator 102 displays the image according to the instructions of the control module 124, which operates to render image data, video data, point cloud data, or other displayable data.

[0049] At least a portion of the light 110, already coupled into waveguide 104 by diffraction input coupler 106, is coupled out of the waveguide by diffraction output coupler 114. At least some of the light coupled out of waveguide 104 replicates the incident angle of the light coupled into the waveguide. For example, in the illustration, output coupled rays 116a, 116b, and 116c replicate the angle of input coupled ray 108. Because the light leaving the output coupler replicates the direction of the light entering the input coupler, the waveguide essentially replicates the original image 112. The user's eye 118 can focus on the replicated image.

[0050] exist Figure 1AIn the example, output coupler 114 outputs only a portion of the coupled light on each reflection, allowing a single input beam (such as beam 108) to generate multiple parallel output beams (such as beams 116a, 116b, and 116c). In this way, even if the eye is not perfectly aligned with the center of the output coupler, at least some of the light from each part of the image may reach the user's eye. For example, if eye 118 moves downwards, beam 116c may enter the eye even if beams 116a and 116b do not, so the user can still perceive the bottom of image 112 despite the positional shift. Therefore, output coupler 114 partially functions as an outgoing pupil dilator in the vertical direction. The waveguide may also include one or more additional outgoing pupil dilators (…). Figure 1A (Not shown in the image) to expand the exit pupil in the horizontal direction.

[0051] In some embodiments, waveguide 104 is at least partially transparent relative to light originating from outside the waveguide display. For example, at least some of the light 120 from a real-world object (such as object 122) passes through waveguide 104, allowing the user to see the real-world object when using the waveguide display. Multiple diffraction orders and thus multiple images will exist as the light 120 from the real-world object also passes through diffraction grating 114. To minimize the visibility of multiple images, it is desirable that the zeroth-order diffraction (not deflected by 114) has high diffraction efficiency for both light 120 and the zeroth order, while higher diffraction orders have lower energy. Therefore, in addition to expanding and output coupling the virtual image, output coupler 114 is preferably configured to allow the zeroth order of the real image to pass through. In such embodiments, the image displayed by the waveguide display may appear to be superimposed on the real world.

[0052] Figure 1B This is a schematic side view illustrating example alternative display types that can be used with extended reality applications according to some embodiments. In the XR head-mounted display device 130, a control module 132 controls a display 134 (which may be an LCD) to display images. The head-mounted display includes a partially reflective surface 136 that reflects (and in some embodiments, both reflects and focuses) the image displayed on the LCD to make the image visible to the user. The partially reflective surface 136 also allows at least some external light to pass through, thereby allowing the user to see their surroundings.

[0053] Figure 1C This is a schematic side view illustrating example alternative display types that can be used with extended reality applications according to some embodiments. In the XR head-mounted display device 140, a control module 142 controls a display 144 (which may be an LCD) to display an image. The image is focused by one or more lenses of a display optics 146 to make the image visible to the user. Figure 1C In this example, external light does not reach the user's eyes directly. However, in some such embodiments, an external camera 148 may be used to capture images of the external environment and display such images on a display 144 along with any virtual content that may also be displayed.

[0054] The embodiments described herein are not limited to any particular type or structure of XR display device.

[0055] Figure 1D This is a system diagram illustrating a set of example interfaces for a system according to some embodiments. Interfaces such as... Figure 1D The system shown implements an extended reality display device and its control electronics. System 150 may be embodied as an apparatus including the various components described below and configured to perform one or more aspects described in this document. Examples of such apparatus include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 150 may be embodied individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 150 are distributed across multiple ICs and / or discrete components. In various embodiments, system 150 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more aspects described in this document.

[0056] System 150 includes at least one processor 152 configured to execute instructions loaded thereon to implement various aspects described herein, such as those described. Processor 152 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 150 includes at least one memory 154 (e.g., a volatile memory device and / or a non-volatile memory device). System 150 may include a storage device 158, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 158 may include internal storage, attached storage (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0057] System 150 includes an encoder / decoder module 156 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 156 may include its own processor and memory. The encoder / decoder module 156 represents a module that can be included in a device to perform encoding and / or decoding functions. It is well known that a device can include one or both encoding and decoding modules. Alternatively, the encoder / decoder module 156 may be implemented as a separate element of system 150, or it may be incorporated within processor 152 as a combination of hardware and software known to those skilled in the art.

[0058] Program code to be loaded onto processor 152 or encoder / decoder 156 to execute the various aspects described herein may be stored in storage device 158 and subsequently loaded onto memory 154 for execution by processor 152. According to various embodiments, one or more of processor 152, memory 154, storage device 158, and encoder / decoder module 156 may store one or more of various items during the execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operational logic processing.

[0059] In some embodiments, the memory within processor 152 and / or encoder / decoder module 156 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, external memory (e.g., processor 152 or encoder / decoder module 152) is used for one or more of these functions. External memory may be memory 154 and / or storage device 158, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG stands for Moving Picture Experts Group; MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Various Video Coding, i.e., a new standard developed by JVET (Joint Video Experts Group)).

[0060] Inputs to the components of system 150 can be provided through various input devices, as indicated in box 172. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 1C Other examples not shown include composite video.

[0061] In various embodiments, the input device of block 172 has corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting the signal band to a band), (ii) down-converting the selected signal, (iii) further band-limiting to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired data packet stream. The RF section in various embodiments includes one or more elements performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners performing various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering again to the desired frequency band. Various embodiments rearrange the order of the above (and other) components, remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as insert amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0062] Additionally, USB and / or HDMI terminals may include corresponding interface processors for connecting system 150 to other electronic devices via USB and / or HDMI connections. It should be understood that aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within processor 152, as needed. Similarly, aspects of USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within processor 152. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 152 and an encoder / decoder 156 operating in conjunction with memory and storage elements, to process the data stream as needed for presentation on an output device.

[0063] Various components of system 150 can be housed within an integrated housing. Within the integrated housing, the various components can be interconnected and transmit data between them using suitable connection means 174 (e.g., internal buses known in the art, including inter-IC (I2C) buses, wiring, and printed circuit boards).

[0064] System 150 includes a communication interface 160 that enables communication with other devices via a communication channel 162. The communication interface 160 may include, but is not limited to, a transceiver configured to send and receive data via the communication channel 162. The communication interface 160 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 162 may be implemented, for example, within a wired and / or wireless medium.

[0065] In various embodiments, data is streamed to or otherwise provided to system 150 using a wireless network, such as a Wi-Fi network, for example, IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). In these embodiments, the Wi-Fi signal is received via a communication channel 162 and a communication interface 160 adapted for Wi-Fi communication. The communication channel 162 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other top-level communications. Other embodiments use a set-top box to provide streaming data to system 150, delivering data via an HDMI connection to input box 172. Still other embodiments use an RF connection to input box 172 to provide streaming data to system 150. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0066] System 150 can provide output signals to various output devices, including display 176, speaker 178, and other peripheral devices 180. Display 176 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 176 can be used in televisions, tablets, laptops, mobile phones, or other devices. Display 176 can also be integrated with other components (e.g., in a smartphone) or standalone (e.g., an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 180 include one or more of a standalone digital video disc (or digital versatile disc) (DVR, for both terms), an optical disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 180 based on the output of system 150 to provide functionality. For example, an optical disc player performs the function of playing the output of system 150.

[0067] In various embodiments, signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention are used to transmit control signals between system 150 and display 176, speaker 178, or other peripheral devices 180. Output devices may be communicatively coupled to system 1000 via dedicated connections through corresponding interfaces 164, 166, and 168. Alternatively, output devices may be connected to system 150 via communication interface 160 using communication channel 162. Display 176 and speaker 178 may be integrated into a single unit along with other components of system 150 in an electronic device, such as a television. In various embodiments, display interface 164 includes a display driver, such as a timing controller (TCon) chip.

[0068] For example, if the RF input section 172 is part of a separate set-top box, the display 176 and speaker 178 can alternatively be separate from one or more of the other components. In various embodiments where the display 176 and speaker 178 are external components, the output signal can be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0069] System 150 may include one or more sensor devices 168. Examples of sensor devices that may be used include one or more GPS sensors, gyroscope sensors, accelerometers, light sensors, cameras, depth cameras, microphones, and / or magnetometers. Such sensors can be used to determine information such as the user's position and orientation. Where system 150 is used as a control module (such as control modules 124, 132) for an extended reality display, the user's position and orientation can be used to determine how to render image data so that the user perceives the correct portion of a virtual object or scene from the correct perspective. In the case of a head-mounted display device, the position and orientation of the device itself can be used to determine the user's position and orientation for rendering virtual content. In the case of other display devices (such as telephones, tablets, computer monitors, or televisions), other inputs can be used to determine the user's position and orientation for rendering content. For example, a user can select and / or adjust the desired viewpoint and / or viewing direction by using a touchscreen, keypad or keyboard, trackball, joystick, or other inputs. Where the display device has sensors such as accelerometers and / or gyroscopes, the viewpoint and orientation can be selected and / or adjusted based on the movement of the display device for rendering content.

[0070] The embodiments can be executed by processor 152 or by computer software implemented by hardware or a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. As a non-limiting example, memory 154 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 152 can be of any type suitable for the technical environment and can encompass one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0071] Scene description architecture for XR This principle generally relates to the fields of rendering extended reality scene descriptions and extended reality rendering. This document can also be understood in the context of extended reality applications being formatted and played back when rendered on end-user devices such as mobile devices or head-mounted displays (HMDs).

[0072] In XR applications, scene descriptions are used to combine a clear and easily parsed description of the scene structure with some binary representation of the media content.

[0073] In time-based media streaming, the scene description itself can evolve over time to provide relevant virtual content for each sequence of the media stream. For example, for advertising purposes, a virtual bottle could be displayed in a video sequence of people drinking water.

[0074] This behavior can be achieved by relying on the framework defined in the following document: Scene Description for MPEG media document, Information technology – Coded representation of immersive media – Part 14: Scene Description for MPEG media, ISO / IEC DIS23090-14:2021 (E). A scene update mechanism based on the JSON Patch protocol, as defined in IETF RFC 6902, can be used to synchronize virtual content to the MPEG media stream.

[0075] Figure 1E This is a system diagram illustrating a set of example interfaces for a scene description (stored as an item in glTF.json) in an ISOBMFF file according to some embodiments, three video tracks, an audio track, and a JSON patch update track. Figure 1E This is an example ISOBMFF file 190, and other elements can be used in such files.

[0076] While the MPEG-I scene description framework ensures that timed media and corresponding virtual content are available at all times, it does not describe how users can interact with scene objects at runtime to obtain an immersive XR experience. Therefore, user-specific XR experiences for consuming immersive media are not supported.

[0077] The example embodiments described herein can be used to provide a scene description that includes virtual objects or light sources, but even if virtual objects or light sources are available, they are not necessarily displayed or rendered. In some embodiments, one or more of the following aspects may be considered when determining whether to display a virtual object or light source.

[0078] Spatial factors can be considered when determining whether to display virtual objects or lights. For example, virtual objects or lights may not be displayed if the user's environment is unsuitable (e.g., the user is too far from the rendered timing media), or if the user is not looking in the correct direction, or if the virtual object should be displayed in a specific area of ​​the user (e.g., above their left hand, which has not yet been detected).

[0079] Timing can be considered when determining whether to display virtual objects or light sources. For example, if the user is not yet ready or wants to trigger the display of an object themselves (e.g., using a specific gesture), the virtual object or light source may not be displayed until an appropriate trigger is detected.

[0080] In some embodiments, the scene description specifies which objects or light sources the user is allowed to manipulate or interact with via potential haptic feedback.

[0081] Runtime interactivity Figure 2 This is a system diagram illustrating a set of example interfaces for an MPEG-I node hierarchy of elements supporting scene interactivity, according to some embodiments. Based on this principle, besides... Figure 2 The MPEG-I node hierarchy 200 and such as about Figure 4 In addition to the described node tree, behavioral metadata items (referred to herein as "behaviors") are added to the scene description. In the example embodiment, the temporal evolution scene description is enhanced by adding information identifying behaviors. These behaviors may be associated with predefined virtual objects that allow runtime interactivity to obtain a user-specific XR experience.

[0082] In some embodiments, these behaviors are time-varying. In such embodiments, the behaviors can be updated using an existing scene description update mechanism.

[0083] In the example embodiment, the behavior is characterized by one or more of the following properties: • Define one or more triggers that must be met for activation.

[0084] • Define the trigger control parameters for the logical operations between the defined triggers.

[0085] • Actions performed in response to triggered activation.

[0086] • Define the action control parameters for the execution order of the defined actions.

[0087] • When multiple actions occur simultaneously on the same virtual object, the priority number of the highest priority action can be selected.

[0088] • An optional interrupt action to specify how to terminate a behavior when it is no longer defined in a newly received scene update. For example, the behavior is no longer defined if the relevant object has been removed, or if the behavior is no longer relevant to the current media (e.g., audio or video) sequence.

[0089] By adding these behaviors, time-related user interactivity in immersive content for XR experiences can be defined.

[0090] When the second scene description is received, some of the behaviors in the first scene description may be "in progress," meaning they have been triggered and their actions are running. The second scene description can be provided as update metadata, that is, metadata describing the differences between the first and second scene descriptions. The second scene description includes a node tree of objects that can be common objects or different objects from the first scene description. Objects in the node tree of the first scene description may no longer exist in the second description. If the second scene description lacks objects related to the running actions of an ongoing behavior, then those ongoing actions are no longer applicable. Similarly, if an ongoing behavior is not defined in the second description, then that ongoing behavior is no longer applicable. The Interrupt Action field describes how to properly interrupt the ongoing actions of an ongoing behavior.

[0091] Figure 3 This is a block diagram illustrating an example of the logical relationship between trigger information (describing triggers 1 to n), action information (describing actions 1 to m), and behavior information (describing the relationship between triggers and actions) according to some embodiments, wherein triggers and actions can refer to one or more nodes in a scene description, such as a hierarchical scene graph.

[0092] In XR applications, scene descriptions are used to combine a well-defined and easily parsed description of the scene structure with some binary representation of the media content. The preceding sections described the action mechanisms used for scene descriptions. These behaviors are related to predefined virtual objects that allow runtime interactivity to achieve a user-specific XR experience. Figure 3 The structure 300 of the example behavior mechanism is shown. The example structure 300 shows example trigger information 302 and action information 304. Within the trigger information 302 are example triggers 1 (306), 2 (308), ..., n (310). Within the action information 304 are example triggers 1 (312), 2 (314), ..., n (316). Triggers 306, 308, 310 and actions 312, 314, 316 are shown as having example relationships with various nodes 318.

[0093] Figure 4This is a schematic plan view illustrating example relationships between objects describing an extended reality scene according to some embodiments. In this example, scene graph 400 includes descriptions of real-world objects 412, such as a “flat horizontal surface” (which could be a table, floor, or plate), and descriptions of virtual objects 414, such as an animation of a walking character. Scene graph node 414 is associated with media content item 416, which is an encoding of data used to render and display the walking character (e.g., as a textured, animated 3D mesh). Scene graph 400 also includes node 410, which describes the spatial relationships between the real-world objects described in node 412 and the virtual objects described in node 414. In this example, node 410 describes the spatial relationships that allow the character to walk on a flat surface. When an XR application is launched, media content item 416 is loaded, rendered, and buffered for display upon triggering. When a sensor (or a camera in some embodiments) detects a flat surface in the real environment, the application displays the buffered media content item as described in node 410. Timing is managed by the application based on the timing of features detected in the real environment and the animation. Nodes in a scene graph may not include descriptions, but instead simply act as the parent node of child nodes.

[0094] XR applications are diverse and can be applied to various backgrounds and real or virtual environments. For example, in industrial XR applications, when a camera mounted on a head-mounted display detects a reference object (part B of an engine) in a real environment, a virtual 3D content item (e.g., part A of the engine) is displayed. The 3D content item is positioned in the real world with a position and scale defined relative to the detected reference object.

[0095] For example, in an XR application for interior design, a 3D model of furniture is displayed when a given image from a catalog is detected in the input camera view. The 3D content is positioned in the real world with a location and scale defined relative to the detected reference image. In another application, some audio files can start playing when a user enters an area near a church (whether real or virtually rendered in an extended real-world environment). In yet another example, an advertising jingle file can play when a user sees a given can of soda in a real-world environment. In outdoor gaming applications, various virtual characters can appear based on the semantics of the scenery observed by the user. For example, bird characters are suitable for trees, so if the XR device's sensors detect a real object described by the semantic tag "tree," birds flying around the trees can be added. In a companion application implemented with smart glasses, when a car is detected within the user's camera's field of view, car noise can be emitted through the user's headphones to warn them of potential danger; furthermore, the sound can be spatialized so that it travels from the direction of the detected car.

[0096] XR applications can also enhance video content, rather than the real environment. The video is displayed on a rendering device, and when a timed event is detected in the video, virtual objects described in the node tree are overwritten. In this context, the node tree only includes descriptions of the virtual objects.

[0097] The example implementation is described with reference to the scope of the MPEG-I scene description framework and using the Khronos glTF extension mechanism, which supports additional scene description features such as node trees. However, the principles described herein are not limited to any particular scene description framework.

[0098] In the example embodiment, the glTF scene description is extended to support interactivity. This interactivity extension is applied at the glTF scene level and is referred to as MPEG_scene_interactivity. The corresponding semantics are provided in Table 1.

[0099]

[0100] Table 1: Semantics of the example MPEG_scene_interactivity extension In Table 1 and other semantic tables described herein, the “Purpose” column indicates that “M” represents a “mandatory” feature and “O” represents a “optional” feature. However, depending on the syntax of a particular proposal, such features may only be “mandatory” or “optional.” A feature marked as “mandatory” is not necessarily a feature required to implement the application. For example, in some embodiments, there are features marked as “mandatory” to meet the expectations of a particular type of parsing and rendering software; however, in other embodiments, the feature may be optional, or may be omitted entirely, without departing from the scope of this disclosure, where default values ​​are used to implement the corresponding functionality or the corresponding functionality is not implemented at all.

[0101] Extended Reality (XR) is a technology that enables interactive experiences by enhancing real-world environments and / or video content with virtual content that can be defined across multiple sensory modalities, including visual, auditory, tactile, and other sensory modalities. During application runtime, virtual content (e.g., 3D content or audio / video files) is rendered in real-time in a manner consistent with the user's context (environment, perspective, device, etc.). Scene graphs (e.g., scene graphs proposed by Khronos / glTF and their extensions defined in the MPEG Scene Description format or Apple / USDZ) are possible ways to represent the content to be rendered. They combine, on the one hand, an explanatory description of the scene structure linking real-world objects and virtual objects, and on the other hand, a binary representation of the virtual content. While such scene description frameworks ensure that timed media and corresponding associated virtual content are available at any time during application rendering, there is no description of how condition-based interactive events are handled within an MPEG-I Scene Description (SD) environment.

[0102] This application discusses 3D scene and object interaction within an immersive environment. It describes a set of generic conditional triggers that allow the use of additional information to activate triggers. This set of triggers can be used with MPEG-I Scene Description (SD) to support interactivity in a 3D environment based on generic metadata information present in nodes and corresponding time-based events. Current scene-level interactivity support only supports generic triggers for any node within the scene, as shown in Table 2. Therefore, problems arise if nodes contain information that can be used to trigger interactive events, because the current interactivity framework of MPEG-I SD does not allow such conditional event-based signaling. See also MPEG extension .

[0103]

[0104] Table 2: List of events currently available for triggering MPEG_scene_interactivity In the interactivity framework, a "behavior" is a set of conditions that pair triggered events with specific actions and describe the temporal constraints of such conditions, thus allowing time-based events to occur in the 3D virtual environment. An "action" is a modifier for a 3D node that may affect its spatial position, materials, media controls, or haptic feedback. A "trigger" is an event that occurs between a node or the user. Table 2 shows a list of triggers currently available for MPEG_scene_interactivity, which can be... MPEG extension It can be found in Table 8.2-3.

[0105] The following sections introduce new extensions that allow glTF models to use other 3D objects and interact with them through metadata information present in nodes. These sections introduce new conditional triggering. These additions, described in the "MPEG_scene_interactivity" section, can be applied at the scene level and, if needed, can be extended to the node level, for example, in some implementations.

[0106] Conditional triggering The following sections detail the meaning of the elements and their associations, the JSON encoding scheme, and how they can be used within MPEG-I SD.

[0107] This format follows the glTF format and is compatible with current MPEG work to extend the glTF format with MPEG extensions. However, the meaning and use are "general" and can be encoded in other formats, such as Extensible Markup Language (XML) and Universal Scene Description (USD).

[0108] Table 3 introduces generic conditional triggers (e.g., “TRIGGER_conditional”), which are added at the same hierarchical level as the original trigger list shown in Table 2 above. New generic conditional triggers can be added to the list shown in Table 2 (see Table 4). For some embodiments, generic conditional triggers may include a “generic” path to one or more parameters. Trigger conditions can be configured via such parameters stored at a location in the generic path. Therefore, for some embodiments, the specific parameters used for triggering may not necessarily be specified in the table, but can be configured and / or updated at runtime. For some embodiments, a set of default parameters may be specified or stored at startup (or before startup) and can be updated later. For example, a separate condition may occur, and the parameters stored at the generic path may be updated. Many trigger parameters are possible according to the embodiments disclosed and contemplated herein. For example, for some embodiments, specific trigger parameters may be updated to follow the evolution of standards or specifications.

[0109]

[0110] Table 3: Conditional Triggering Semantics In some implementations, the attribute "Descriptor" can be a string or an array of strings. The attribute "Descriptor" compares multiple schemes against a given value in the "Parameter" attribute.

[0111]

[0112] Table 4: Trigger list for MPEG_scene_interactivity plus general conditional triggers Conditional preprocessing Figure 5 This is a flowchart illustrating an example preprocessing step triggered by conditions according to some embodiments. Figure 6 This is a flowchart illustrating an example preprocessing step triggered by conditions according to some embodiments. Figure 7 This is a flowchart illustrating an example preprocessing step triggered by conditions according to some embodiments. Figure 5 , Figure 6 and Figure 7 This demonstrates the preprocessing for conditional triggering used in the immersive and interactive systems discussed above. This preprocessing is performed to validate the behavior and triggering constructors.

[0113] Figure 5 , Figure 6 and Figure 7 Example procedures 500, 600, and 700 are shown respectively, for applications that can parse each trigger present in the 502, 602, and 702 behaviors and evaluate their conditions. For Figure 5 , Figure 6 and Figure 7 Execute the trigger check for 504, 604, and 704 errors.

[0114] for Figure 5 , Figure 6 and Figure 7 If a conditional trigger is encountered, checks 506, 608, and 716 are performed to see if the nodes listed in the trigger contain a descriptor, which is the path to the node value to be compared. If the current trigger is not a conditional trigger, processing continues to handle other triggers. If no descriptor is found in any node, errors 508, 610, and 718 can be indicated. This indication can be used by the model to signal to the application (or any other engine handling such interaction models) that if the node list has not changed, the trigger and current behavior will not be validated and handled. This allows the engine to optimize its computation and ignore or correct behaviors that might never be completed during the resolution phase. If all nodes contain descriptors and are therefore conditional trigger extension nodes, processing (e.g., resolution of the trigger) continues 510, 612, and 720. For some embodiments, Figures 5 to 7 The limitations shown can be applied to Table 3 above.

[0115] for Figure 6 and Figure 7 If the current trigger is not a conditional trigger, processing continues to handle other triggers such as 606 and 706. For Figure 7 If the current trigger is not a conditional trigger, then in some embodiments, an "enumeration" (enumeration) can be performed to jump to the appropriate unconditional trigger box 708, 710, 712, 714 (wherein) Figure 7 Some example triggers are listed below.

[0116] At runtime, the processing model can remain unchanged from the original interaction model. For example, if a condition is not met, the application continues. If the behavior being evaluated satisfies all trigger conditions, its action will be initiated in some embodiments. In some embodiments, if one of the triggers does not satisfy one or more of the conditions, the application continues to evaluate scene updates until all trigger conditions are satisfied. In some embodiments, if one of the triggers does not satisfy one or more of the conditions, the application continues to evaluate each scene update until all trigger conditions are satisfied.

[0117] glTF mode example The following glTF modes are instantiated as examples (not exhaustive) triggered by clients supporting "MPEG_scene_interactivity". Depending on the application, numerous instantiations are possible, and for illustrative purposes, several examples are given in the following sections.

[0118] Figures 8A to 8B A list of code forms an example code structure illustrating basic scene elements according to some embodiments. Figure 8A and Figure 8B The example code lists 800 and 850 together show some basic scene elements. There are two nodes. The first node is an extension of the avatar node (…). Figure 8A The camera in lines 17 to 34 Figure 8A (lines 3 to 35 in the text), and the second node is a sphere node ( Figure 8A (Lines 36 to 44). These elements are related to Figure 9 A to Figure 10B Use the example code list shown.

[0119] Figure 9 This is a list of code examples illustrating the structure of condition-triggered metadata according to some embodiments. This list of example code 900 illustrates a structure based on avatars (…). Figure 8A Lines 20-32) describe the associated metadata information and how conditional triggering can be used to trigger interactions with objects in the scene. In this example, the defined behavior uses proximity triggering (…). Figure 9 Lines 7 to 12) and conditional triggering ( Figure 9 (Lines 13-23) Both begin the action. The combo triggers when the avatar enters within 0.0 and 1.0 units of the red sphere and the avatar (or the person using the avatar) is 18 years of age or older.

[0120] In this example, there exists from Figure 8A and Figure 8BThe basic scene shown has two nodes. The first node is a camera with an avatar node extension, and the second node is a sphere node used to demonstrate interaction. These objects are examples of possible object types in the scene.

[0121] This example will trigger the instantiation of `trigger_conditional`. The behavior class will instantiate the behavior that triggers the `trigger_proximity`, for example, `trigger_proximity` and `trigger_conditional`, so the two objects (e.g., avatar and sphere) can have a two-step interaction. In some embodiments, this means triggering a proximity trigger before triggering a second trigger. Figure 9 (Lines 7 to 12).

[0122] In this example, the value being compared ("18") is specified via a conditional trigger. Figure 9 Line 18) and the comparator used ("greaterThanOrEqualTo") Figure 9 Line 19) Both. The attribute looked up in the node is determined by a conditionally triggered naming convention ("node.MPEG_node_avatar.metadata.age") ( Figure 9 The action is specified in line 15. In this example, since node 0 (the avatar) has a value greater than or equal to the value being compared ("18"), an activation action ("ACTION_ACTIVATE") is initiated, and the corresponding interaction is applied.

[0123] Figures 10A to 10B A list of code examples illustrating the structure of conditional triggering according to some embodiments is provided. This second list of example code 1000, 1050 illustrates how conditional triggering can be used with multiple parameters of different types of comparisons to limit interaction. In this example, the action associated with the behavior is triggered based on proximity triggers and multiple conditions in the avatar's metadata information, but these conditions are, of course, merely examples.

[0124] In this example, there exists from Figure 8A and Figure 8B The basic scene shown has two nodes. The first node is a camera with an avatar node extension, and the second node is a sphere node used to demonstrate interaction. These objects are examples of possible object types in the scene.

[0125] This example will trigger the instantiation of `trigger_conditional`. The behavior class will instantiate the behavior of the trigger paired with `trigger_proximity`, for example, `trigger_proximity` and `trigger_conditional`, so the two objects (e.g., avatar and sphere) can have a two-step interaction. In some embodiments, this means triggering a proximity trigger before triggering a second trigger. Figure 10A (Lines 7 to 12).

[0126] This example is Figure 9 The example shown is similar. The difference is that multiple properties can be compared using a single conditional trigger. Figure 10A Lines 15 to 24 and Figure 10B (Lines 1 to 33). For this example, the "descriptor" property is an array with indices aligned with the array of "parameters" properties.

[0127] In this example, conditional triggering is based on a comparison of multiple parameters, but each of these parameters is generally defined using a PATH name relative to the pattern. See, for example, [link to example]. Figure 10A See the line “node.MPEG_node_avatar.metadata.capability” in the code to get an example of this name. Figure 10A and Figure 10B The examples illustrate the complexity and flexibility of new conditional triggers. In some embodiments, conditional triggers can be defined entirely within the schema, rather than in the syntax table defining the trigger.

[0128] Figure 11 This is a flowchart illustrating an example process for handling conditional triggering according to some embodiments. For some embodiments, example process 1100 may include obtaining 1102 scene description data of a three-dimensional (3D) scene. For some embodiments of example process 1000, the scene description data may include 1104 behavioral information, which includes: trigger information describing at least one triggering condition, and action information describing an action to be performed on a scene element in the 3D scene, wherein the at least one triggering condition includes a conditional trigger. For some embodiments, example process 1100 may also include performing 1106 an action on a scene element in response to determining that at least one triggering condition has occurred.

[0129] While methods and systems according to some embodiments are generally discussed in the context of extended reality (XR), some embodiments can be applied to any XR context, such as virtual reality (VR) / mixed reality (MR) / augmented reality (AR) contexts. Additionally, although the term "head-mounted display (HMD)" is used herein according to some embodiments, for some embodiments, some embodiments can be applied to wearable devices capable of supporting, for example, XR, VR, AR, and / or MR (which may or may not be attached to the head).

[0130] An example method according to some embodiments may include: obtaining scene description data of a three-dimensional (3D) scene, wherein the scene description data includes behavioral information, the behavioral information including: trigger information describing at least one triggering condition, and action information describing an action performed on a scene element in the 3D scene, wherein the at least one triggering condition includes a conditional trigger; and wherein the trigger information includes a path to an attribute, wherein the attribute is used in conjunction with the conditional trigger; and in response to determining that the at least one triggering condition has occurred, performing the action on the scene element.

[0131] In some embodiments of the example method, the at least one triggering condition includes a test of the attribute.

[0132] An example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the device to perform the method of any of the above-listed claims.

[0133] An example method according to some embodiments may include: obtaining scene description data of a three-dimensional (3D) scene, wherein the scene description data includes behavioral information, the behavioral information including: trigger information describing at least one triggering condition, and action information describing an action performed on a scene element in the 3D scene, wherein the at least one triggering condition may include a conditional trigger; and performing the action on the scene element in response to determining that the at least one triggering condition has occurred.

[0134] In some embodiments of the example method, the at least one conditional trigger may include a path to a storage location, and the trigger information describing the at least one triggering condition is stored in the storage location.

[0135] Some embodiments of the example method may also include updating the triggering information at runtime.

[0136] Some embodiments of the example method may also include updating the triggering information before runtime.

[0137] Some embodiments of the example method may also include updating the condition trigger at runtime.

[0138] Some embodiments of the example method may also include updating the condition trigger before runtime.

[0139] In some embodiments of the example method, the triggering information may include one or more parameters, and the at least one triggering condition may be based on at least one of the one or more parameters.

[0140] In some embodiments of the example method, at least one of the one or more parameters includes a value field and a comparator field.

[0141] In some embodiments of the example method, the triggering information further includes node information indicating one or more nodes, and the method further includes: searching in a memory structure corresponding to the one or more nodes for a field that matches the value field; and updating the conditional trigger based on the matched field.

[0142] In some embodiments of the example method, the triggering information further includes a descriptor field, and the method further includes using the descriptor field in conjunction with the one or more parameters to determine the at least one triggering condition.

[0143] Some embodiments of the example method may further include: detecting a second condition; and updating the triggering information in response to detecting the second condition.

[0144] In some embodiments of the example method, the triggering information may also include at least one metadata field.

[0145] In some embodiments of the example method, the at least one triggering condition is based on the at least one metadata field.

[0146] In some embodiments of the example method, the conditional triggering may include two or more conditions.

[0147] In some embodiments of the example method, the at least one condition trigger includes a path to the parameter to be compared, and the at least one trigger condition includes the parameter to be compared.

[0148] In some embodiments of the example method, the at least one triggering condition includes a test of the attribute.

[0149] An example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0150] Another example method according to some embodiments may include handling conditional triggers associated with a three-dimensional (3D) scene, wherein the conditional triggers may include a mechanism for updating one or more triggering conditions associated with the conditional triggers.

[0151] An example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the device to perform the methods listed above.

[0152] Another example method according to some embodiments may include: obtaining scene description data of a three-dimensional (3D) scene, wherein the scene description data includes behavioral information, the behavioral information including: trigger information describing at least one triggering condition, and action information describing an action performed on a scene element in the 3D scene, wherein the at least one triggering condition includes a conditional trigger; and wherein the trigger information includes a path to an attribute, wherein the attribute is used in conjunction with the conditional trigger; and in response to determining that the at least one triggering condition has occurred, performing the action on the scene element.

[0153] In some embodiments of yet another example method, the at least one triggering condition includes a test of the attribute.

[0154] Another example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the device to perform any of the methods listed above.

[0155] This disclosure describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described in a specific manner and are generally described in a way that may sound limiting, at least for the purpose of illustrating the various features. However, this is for the purpose of clarity of description and does not limit the disclosure or scope of these aspects. In fact, all the different aspects can be combined and interchanged to provide other aspects. Furthermore, these aspects can also be combined and interchanged with aspects described in previous documents.

[0156] The aspects described and contemplated in this disclosure can be implemented in many different forms. While some embodiments are specifically shown, other embodiments are contemplated, and the discussion of particular embodiments does not limit the breadth of implementations. At least one aspect relates generally to video encoding and decoding, and at least one other aspect relates generally to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions thereon stored thereon for encoding or decoding video data according to any of the methods, and / or computer-readable storage media having bitstreams generated according to any of the methods stored thereon.

[0157] In this disclosure, the terms “reconstruction” and “decoding” are used interchangeably, as are the terms “pixel” and “sample”, and the terms “image,” “picture,” and “frame.” Typically, but not necessarily, the term “reconstruction” is used on the encoder side, while “decoding” is used on the decoder side.

[0158] The terms HDR (High Dynamic Range) and SDR (Standard Dynamic Range) generally convey specific values ​​of dynamic range to those skilled in the art. However, it is also intended to include additional embodiments in which references to HDR should be understood as meaning "higher dynamic range" and references to SDR should be understood as meaning "lower dynamic range." Such additional embodiments are not bound by any specific values ​​of dynamic range that may generally be associated with the terms "high dynamic range" and "standard dynamic range."

[0159] This document describes various methods, each of which includes one or more steps or actions to implement the described method. The order and / or use of specific steps and / or actions can be modified or combined unless the correct operation of the method requires a specific order of steps or actions. Additionally, in various embodiments, terms such as "first," "second," etc., may be used to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." Unless specifically required, the use of such terms does not imply an ordering of the modified operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding, but can occur, for example, before, during, or in a time period overlapping with the second decoding.

[0160] For example, various numerical values ​​may be used in this disclosure. Specific values ​​are for illustrative purposes, and the aspects described are not limited to these specific values.

[0161] The embodiments described herein can be executed by computer software implemented by a processor or other hardware or a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. As a non-limiting example, the processor can be of any type suitable for the technical environment and can encompass one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0162] Various implementations involve decoding. As used in this disclosure, "decoding" can encompass all or part of a process performed, for example, on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also includes, or alternatively includes, processes performed by a decoder of the various implementations described in this disclosure, such as extracting images from stitched (packed) images, determining an upsampling filter to use and then upsampling the images, and flipping the images back to their intended orientation.

[0163] As another example, in one embodiment, "decoding" refers only to entropy decoding; in another embodiment, "decoding" refers only to differential decoding; and in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Based on the specific context of the description, it will become clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to refer to the broader decoding process.

[0164] Various implementations involve encoding. Similar to the discussion above regarding “decoding,” “encoding” as used in this disclosure can include, for example, all or part of the processing performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as partitioning, differential coding, transform, quantization, and entropy coding. In various embodiments, such processes also include, or alternatively include, processes performed by a decoder of the various implementations described in this disclosure.

[0165] As another example, in one embodiment, "encoding" refers only to entropy encoding; in another embodiment, "encoding" refers only to differential encoding; and in yet another embodiment, "encoding" refers to a combination of entropy encoding and differential encoding. Based on the specific context of the description, it will become clear whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to refer to a broader encoding process.

[0166] Various embodiments involve rate distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often due to constraints on computational complexity. Rate distortion optimization is generally formulated as minimizing a rate distortion function, which is a weighted sum of rate and distortion. Different approaches exist to address the rate distortion optimization problem. For example, these approaches can be based on extensive testing of all encoding options, including all considered modes or encoding parameter values, where a comprehensive evaluation of the encoding cost and associated distortion of the reconstructed signal is performed after encoding and decoding. Faster methods can also be used to save encoding complexity, particularly by calculating approximate distortion based on prediction or prediction of the residual signal rather than the reconstructed signal. A hybrid of these two approaches can also be used, such as by using approximate distortion only for some possible encoding options and full distortion for others. Other methods evaluate only a subset of possible encoding options. More generally, many methods employ any of a variety of techniques to perform optimization, but optimization is not necessarily a comprehensive evaluation of encoding cost and associated distortion.

[0167] When a diagram is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a diagram is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / device.

[0168] The implementations and aspects described herein can be implemented, for example, in methods or processes, devices, software programs, data streams, or signals. Even if discussed only in the context of a single implementation (e.g., discussed only as a method), the implementation of the features in question can also be implemented in other forms (e.g., devices or programs). Devices can be implemented, for example, with appropriate hardware, software, and firmware. Methods can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0169] The references to "an embodiment" or "an embodiment" or "an implementation" or "an implementation," and other variations thereof, mean that a particular feature, structure, characteristic, etc., described in connection with that embodiment is included in at least one embodiment. Therefore, the phrases "in an embodiment" or "in an embodiment" or "in an implementation" or "in an implementation," and any other variations appearing throughout this disclosure, do not necessarily refer to the same embodiment.

[0170] Additionally, this disclosure may relate to “determining” various types of information. Determining information may include one or more of the following: for example, estimated information, calculated information, predicted information, or information retrieved from memory.

[0171] Furthermore, this disclosure may relate to “accessing” various types of information. Accessing information may include one or more of the following: for example, receiving information, retrieving information (e.g., retrieving from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0172] Additionally, this disclosure may relate to "receiving" various types of information. Like "access," the intent to receive is a broad term. Receiving information may include one or more of the following: for example, accessing information or retrieving information (e.g., retrieving from memory). Furthermore, "receiving" is generally referred to in one or more ways during operation, such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0173] It should be understood that, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As yet another example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” this wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or selecting all three options (A, B, and C). This can be extended to as many as the items listed.

[0174] Furthermore, as used herein, the term “signaling” specifically refers to instructing the corresponding decoder to do something. For example, in some embodiments, the encoder signals a particular parameter among several parameters for region-based filter parameter selection for artifact removal filtering. In this way, in embodiments, the same parameter is used at both the encoder and decoder sides. Thus, for example, the encoder can send (explicitly signal) a specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling can be used without transmitting (implicitly signaling) to simply allow the decoder to know and select the specific parameter. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntactic elements, flags, etc., are used to send information to the corresponding decoder. Although the verb form of the word “signal” has been referred to above, the word “signal” can also be used as a noun herein.

[0175] Implementations can generate various signals that are formatted to carry information, such as information that can be stored or transmitted. The information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is well known that signals can be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.

[0176] Several embodiments are described. Features of these embodiments may be provided individually or in any combination across various claim classes and types. Furthermore, embodiments may include one or more of the following features, means, or aspects individually or in any combination across various claim classes and types: • A bitstream or signal that includes one or more of the described syntactic elements or their variants.

[0177] • A bitstream or signal, which includes syntax for conveying information generated according to any embodiment of the described embodiments.

[0178] • Create and / or send and / or receive and / or decode bit streams or signals, which include one or more of the described syntactic elements or their variants.

[0179] • Create and / or send and / or receive and / or decode according to any of the described embodiments.

[0180] • A method, process, apparatus, medium for storing instructions, medium for storing data, or signal according to any of the described embodiments.

[0181] It should be noted that the various hardware elements, one or more of which are described in the embodiments, are referred to as “modules”, which implement (i.e., perform, execute, etc.) the various functions described herein in connection with the respective modules. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more memory devices) that a person skilled in the art would consider suitable for a given implementation. Each described module may also include executable instructions for performing one or more functions described as being performed by the respective module, and it should be noted that these instructions may take the form of hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc., and may be stored in any suitable non-transitory computer-readable medium such as RAM, ROM, etc.

[0182] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented in a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM discs and digital multifunction disks (DVDs). The processor associated with the software can be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A method comprising: Obtain scene description data for a three-dimensional (3D) scene. The scene description data includes behavioral information, which includes: Trigger information describing at least one trigger condition, and Action information describing actions performed on scene elements in the 3D scene. Wherein, the at least one triggering condition includes conditional triggering; and The triggering information includes the path to the attribute. The attribute is used in conjunction with the conditional trigger; and In response to determining that at least one of the triggering conditions has occurred, the action is performed on the scene element.

2. The method according to claim 1, in, The at least one triggering condition includes a test of the attribute.

3. An apparatus comprising: processor; as well as A memory that stores instructions that, when executed by the processor, are operable to cause the device to perform the method as described in any one of claims 1 to 2.

4. A method comprising: Obtain scene description data for a three-dimensional (3D) scene. The scene description data includes behavioral information, which includes: Trigger information describing at least one trigger condition, and Action information describing actions performed on scene elements in the 3D scene. Wherein, the at least one triggering condition includes conditional triggering; and In response to determining that at least one of the triggering conditions has occurred, the action is performed on the scene element.

5. The method according to claim 4, in, The at least one conditional trigger includes a path to the storage location, and The triggering information describing the at least one triggering condition is stored in the storage location.

6. The method according to any one of claims 4 to 5, further comprising updating the triggering information at runtime.

7. The method according to any one of claims 4 to 5, further comprising updating the triggering information before runtime.

8. The method according to any one of claims 4 to 7, further comprising updating the condition trigger at runtime.

9. The method according to any one of claims 4 to 7, further comprising updating the condition trigger before runtime.

10. The method according to any one of claims 4 to 9, in, The triggering information includes one or more parameters, and The at least one triggering condition is based on at least one of the one or more parameters.

11. The method according to claim 10, wherein, At least one of the one or more parameters includes a value field and a comparator field.

12. The method according to any one of claims 10 to 11, in, The triggering information also includes node information indicating one or more nodes, and The method further includes: Search the memory structure corresponding to the one or more nodes for a field that matches the value field; and The condition trigger is updated based on the matched field.

13. The method according to any one of claims 4 to 12, in, The triggering information also includes a descriptor field, and The method further includes using the descriptor field in conjunction with the one or more parameters to determine the at least one triggering condition.

14. The method according to any one of claims 4 to 13, further comprising: The second condition for detection; as well as In response to detecting the second condition, the triggering information is updated.

15. The method according to any one of claims 4 to 14, wherein, The triggering information includes at least one metadata field.

16. The method according to any one of claims 4 to 15, wherein, The at least one triggering condition is based on the at least one metadata field.

17. The method according to any one of claims 4 to 16, wherein, The conditional triggering includes two or more conditions.

18. The method according to any one of claims 4 to 17, in, The at least one conditional trigger includes a path to the parameter to be compared, and The at least one triggering condition includes the parameter to be compared.

19. The method according to any one of claims 4 to 18, in, The at least one triggering condition includes a test of the attribute.

20. An apparatus comprising: processor; as well as A memory that stores instructions that, when executed by the processor, are operable to cause the device to perform the method as described in any one of claims 4 to 19.

21. A method comprising handling conditional triggers associated with a three-dimensional (3D) scene, wherein, The conditional triggering includes a mechanism for updating one or more triggering conditions associated with the conditional triggering.

22. An apparatus comprising: processor; as well as A memory that stores instructions that, when executed by the processor, are operable to cause the device to perform the method of claim 21.

23. A method comprising: Obtain scene description data for a three-dimensional (3D) scene. The scene description data includes behavioral information, which includes: Trigger information describing at least one trigger condition, and Action information describing actions performed on scene elements in the 3D scene. Wherein, the at least one triggering condition includes conditional triggering; and The triggering information includes the path to the attribute. The attribute is used in conjunction with the conditional trigger; and In response to determining that at least one of the triggering conditions has occurred, the action is performed on the scene element.

24. The method according to claim 23, in, The at least one triggering condition includes a test of the attribute.

25. An apparatus comprising: processor; as well as A memory that stores instructions that, when executed by the processor, are operable to cause the device to perform the method as described in any one of claims 23 to 24.