Universal avatar triggering in virtual environment

By introducing behavioral information and interactive mechanisms into extended reality technology, the problem of insufficient user interactivity in existing technologies is solved, enabling users to have a specific immersive interactive experience and enhancing the integration of virtual content with the real environment.

CN121794031APending Publication Date: 2026-04-03INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing extended reality technologies have failed to effectively support specific interactive experiences for users, especially in the combination of virtual content and the real environment, where immersive interactivity is lacking.

Method used

By introducing behavioral information, including triggering conditions and action information, into the scene description data, the interactive behavior between the user and virtual objects is defined using the MPEG-I scene description framework and glTF configuration, and the triggering and execution of these behaviors are realized through encoder and decoder devices.

Benefits of technology

It enables users to have specific interactive experiences in extended reality environments, enhances the integration of virtual content with the real environment, and provides immersive interactivity and dynamic responsiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121794031A_ABST
    Figure CN121794031A_ABST
Patent Text Reader

Abstract

Some embodiments of a method may include obtaining scene description data for a three-dimensional (3D) scene, where the scene description data includes behavioral information including: trigger information describing at least one trigger condition, and action information describing an action to be performed on a scene element in the 3D scene, wherein the at least one trigger condition corresponds to an avatar in the 3D scene; and in response to determining that the at least one trigger condition has occurred, performing an action on the scene element.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications This application claims the benefit of European Patent Application No. EP23306183 entitled “GENERIC AVATAR TRIGGER IN VIRTUAL ENVIRONMENTS”, filed on July 11, 2023, which is incorporated herein by reference in its entirety. Background Technology

[0002] Extended Reality (XR) is a technology that enables interactive experiences where real-world environments and / or video content are enhanced by virtual content, which can be defined across multiple sensory modalities, including visual, auditory, tactile, and others. During application runtime, the virtual content (e.g., 3D content or audio / video files) is rendered in real-time in a manner consistent with the user's context (environment, point of view, device, etc.). Scene graphs (such as those proposed by Khronos / glTF and their extensions defined in the MPEG scene description format or Apple / USDZ) are possible ways to represent the content to be rendered. They combine, on the one hand, a declarative description of the scene structure linking real-world objects and virtual objects, and on the other hand, a binary representation of the virtual content. Summary of the Invention

[0003] The embodiments described herein include methods used in video encoding and decoding (collectively, “coding”).

[0004] Example methods according to some embodiments may include: obtaining scene description data of a three-dimensional (3D) scene, wherein the scene description data includes behavioral information, the behavioral information including: trigger information describing at least one trigger condition, and action information describing an action to be performed on a scene element in the 3D scene, wherein the at least one trigger condition corresponds to an avatar in the 3D scene, and wherein the at least one trigger condition tests a specific characteristic of the avatar; and performing an action on the scene element in response to determining that at least one trigger condition has occurred.

[0005] Example apparatuses according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0006] Example methods according to some embodiments may include: obtaining scene description data of a three-dimensional (3D) scene, wherein the scene description data includes behavioral information, the behavioral information including: trigger information describing at least one trigger condition, and action information describing an action to be performed on a scene element in the 3D scene, wherein the at least one trigger condition corresponds to an avatar in the 3D scene; and performing an action on the scene element in response to determining that at least one trigger condition has occurred.

[0007] For some embodiments of the example method, at least one triggering condition corresponds to a generic avatar trigger.

[0008] For some embodiments of the example method, at least one triggering condition corresponds to a combination of a general avatar triggering condition and a specific avatar triggering condition.

[0009] For some embodiments of the example method, at least one triggering condition corresponds to a specific avatar triggering condition.

[0010] For some embodiments of the example method, at least one triggering condition corresponds to a general avatar trigger that points to one or more specific avatar triggering conditions.

[0011] For some embodiments of the example method, specific avatar triggering conditions are selected from the group including: social actions, avatar permission, age restrictions, media device conditions, avatar abilities, and avatar disabilities.

[0012] For some embodiments of the example method, the triggering information may include: generic avatar trigger, comparator field, and node information.

[0013] For some embodiments of the example method, at least one triggering condition is based on a generic avatar trigger, a comparator field, and node information.

[0014] For some embodiments of the example method, the trigger information may include at least one metadata field.

[0015] For some embodiments of the example method, at least one triggering condition corresponds to at least one metadata field.

[0016] For some embodiments of the example method, at least one metadata field corresponds to an avatar.

[0017] For some embodiments of the example method, at least one metadata field corresponds to a second avatar in the 3D scene.

[0018] For some embodiments of the example method, at least one metadata field corresponds to a scene element that is separate from the avatar.

[0019] For some embodiments of the example method, at least one triggering condition corresponds to the avatar's social action.

[0020] For some embodiments of the example method, at least one triggering condition corresponds to a social action of a second avatar in the 3D scene.

[0021] For some embodiments of the example method, at least one triggering condition corresponds to an age restriction associated with the avatar.

[0022] For some embodiments of the example method, at least one triggering condition corresponds to a content restriction associated with the avatar.

[0023] For some embodiments of the example method, at least one triggering condition corresponds to an ability associated with the avatar.

[0024] For some embodiments of the example method, at least one triggering condition corresponds to a disability associated with the avatar.

[0025] For some embodiments of the example method, the scene description data is compatible with MPEG-I scene description (SD).

[0026] For some implementations of the example methods, the scene description data is compatible with the glTF configuration.

[0027] Some embodiments of the example method may further include: determining that the node associated with at least one triggering condition does not support MPEG node avatars; and transmitting the node configuration error to another device.

[0028] For some embodiments of the example method, at least one triggering condition corresponds to the entire avatar trigger specified at the scene level in the 3D scene.

[0029] For some embodiments of the example method, at least one triggering condition corresponds to a sub-avatar trigger specified at the scene level of the 3D scene.

[0030] Example apparatuses according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0031] In additional embodiments, encoder and decoder devices are provided to perform the methods described herein. The encoder or decoder device may include a processor configured to perform the methods described herein. The device may include a computer-readable medium (e.g., a non-transitory medium) storing instructions for performing the methods described herein. In some embodiments, the computer-readable medium (e.g., a non-transitory medium) stores video encoded using any of the methods described herein.

[0032] One or more of these embodiments also provide a computer-readable storage medium storing instructions for performing bidirectional optical flow, encoding, or decoding video data according to any of the methods described above. This embodiment also provides a computer-readable storage medium storing a bitstream generated according to the methods described above. This embodiment further provides a method and apparatus for transmitting a bitstream generated according to the methods described above. This embodiment also provides a computer program product including instructions for performing any of the methods described above. Attached Figure Description

[0033] Figure 1A This is a schematic side view of an example waveguide display that can be used with extended reality (XR) applications according to some embodiments.

[0034] Figure 1B This is a schematic side view illustrating an example alternative display type that can be used with extended reality applications according to some embodiments.

[0035] Figure 1C This is a schematic side view illustrating an example alternative display type that can be used with extended reality applications according to some embodiments.

[0036] Figure 1D This is a system diagram illustrating a set of example interfaces of a system according to some embodiments.

[0037] Figure 1E This is a system diagram illustrating a set of example interfaces for a scene description (stored as an item in glTF.json) in an ISOBMFF file, three video tracks, one audio track, and one JSON patch update track, according to some embodiments.

[0038] Figure 2 This is a system diagram illustrating a set of example interfaces of an MPEG-I node hierarchy that supports scene interactivity according to some embodiments.

[0039] Figure 3 This is a block diagram illustrating an example of the logical relationship between trigger information (describing triggers 1 to n), action information (describing actions 1 to m), and behavior information (describing the relationship between triggers and actions) according to some embodiments, wherein triggers and actions may refer to one or more nodes in a scene description (such as a hierarchical scene graph).

[0040] Figure 4 This is a schematic plan view illustrating example relationships of objects describing an extended reality scene according to some embodiments.

[0041] Figure 5 This is a flowchart illustrating an example preprocessing step triggered by an avatar according to some embodiments.

[0042] Figures 6A-6B A list of code forms an example code structure illustrating the basic scene elements according to some embodiments.

[0043] Figures 7A-7B A list of code forms an example code structure illustrating the independent triggering of avatar-related information according to some embodiments.

[0044] Figures 8A-8B A list of codes is formed that illustrates example code structures (e.g., avatar triggering conditions) according to some embodiments.

[0045] Figure 9 This is a list of code illustrating an example code structure for external avatar-triggered metadata according to some embodiments.

[0046] Figure 10 This is a flowchart illustrating an example process for handling avatar triggering according to some embodiments.

[0047] Figure 11 This is a flowchart illustrating an example process for handling avatar triggering according to some embodiments.

[0048] Figure 12 This is a flowchart illustrating an example process for handling avatar triggering according to some embodiments.

[0049] Entities, connections, arrangements, and such depictions in various figures, as well as those described in conjunction with various figures, are presented by way of example rather than by way of limitation. Therefore, any and all statements or other indications regarding what a particular figure “depicts,” what a particular element or entity in a particular figure “is” or “has,” and any and all similar statements (which may be understood in isolation and out of context as absolute and therefore limiting) may only be correctly understood as being preceded by a clause such as “In at least one embodiment, …”. For the sake of brevity and clarity, this implicit introductory clause is not tiresomely repeated in the detailed description. Detailed Implementation

[0050] Figure 1AThis illustration shows a schematic side view of an example waveguide display that can be used with extended reality (XR) applications according to some embodiments. The image is projected by an image generator 102. The image generator 102 can project the image using one or more of a variety of technologies. For example, the image generator 102 can be a laser beam scanning (LBS) projector, a liquid crystal display (LCD), a light-emitting diode (LED) display (including organic LED (OLED) or micro LED (µLED) displays), a digital light processor (DLP), a liquid crystal on silicon (LCoS) display, or other types of image generators or light engines.

[0051] The light representing image 112 generated by image generator 102 is coupled into waveguide 104 via diffraction in-coupler 106. In-coupler 106 diffracts the light representing image 112 into one or more diffraction orders. For example, ray 108 representing a portion of the bottom of the image is diffracted by in-coupler 106, and one of the diffraction orders 110 (e.g., the second order) is at an angle capable of propagating through waveguide 104 via total internal reflection. Image generator 102 displays the image according to instructions from control module 124, which operates to render image data, video data, point cloud data, or other displayable data.

[0052] At least a portion of the light 110, already coupled into waveguide 104 by diffraction-in coupler 106, is coupled out of the waveguide by diffraction-out coupler 114. At least some of the light coupled out of waveguide 104 replicates the angle of incidence of the light coupled into the waveguide. For example, in the illustration, out-coupled rays 116a, 116b, and 116c replicate the angle of the input coupled ray 108. Because the light leaving the out-coupler replicates the direction of the light entering the in-coupler, the waveguide essentially replicates the original image 112. The user's eye 118 can focus on the replicated image.

[0053] exist Figure 1AIn the example, the out-coupler 114 outputs only a portion of the coupled light, where each reflection allows a single input beam (such as beam 108) to generate multiple parallel output beams (such as beams 116a, 116b, and 116c). Thus, even if the user's eye is not perfectly aligned with the center of the out-coupler, at least some of the light originating from each part of the image may reach the user's eye. For example, if eye 118 moves downwards, beam 116c may enter the eye even if beams 116a and 116b do not, so the user can still perceive the bottom of image 112 despite the positional shift. Therefore, the out-coupler 114 partially functions as an exit pupil expander in the vertical direction. The waveguide may also include one or more additional exit pupil expanders (…). Figure 1A (not shown in the image) to expand the exit pupil in the horizontal direction.

[0054] In some embodiments, waveguide 104 is at least partially transparent to light originating outside the waveguide display. For example, at least some of the light 120 from a real-world object (such as object 122) passes through waveguide 104, allowing the user to see the real-world object when using the waveguide display. Since the light 120 from the real-world object also passes through diffraction grating 114, there will be multiple diffraction orders, and therefore multiple images. To minimize the visibility of multiple images, it is desirable that the zeroth-order diffraction (without the bias of 114) has a large diffraction efficiency for both light 120 and the zeroth order, while higher diffraction orders are lower in energy. Therefore, in addition to extending and coupling virtual images, the out-coupler 114 is preferably configured to allow the zeroth order of the real image to pass through. In such embodiments, the image displayed by the waveguide display may appear to be superimposed on the real world.

[0055] Figure 1B This is a schematic side view illustrating an example alternative display type that can be used with extended reality applications according to some embodiments. In the XR head-mounted display device 130, a control module 132 controls a display 134 (which may be an LCD) to display images. The head-mounted display includes a partially reflective surface 136 that reflects (and in some embodiments, both reflects and focuses) the image displayed on the LCD to make the image visible to the user. The partially reflective surface 136 also allows at least some external light to pass through, thereby allowing the user to see their surroundings.

[0056] Figure 1CThis is a schematic side view illustrating an example alternative display type that can be used with extended reality applications according to some embodiments. In an XR head-mounted display device 140, a control module 142 controls a display 144 (which may be an LCD) to display an image. The image is focused by one or more lenses of a display optics 146 to make the image visible to the user. Figure 1C In this example, the external light does not reach the user's eyes directly. However, in some such embodiments, the external camera 148 can be used to capture images of the external environment and display such images on the display 144 along with any virtual content that may also be displayed.

[0057] The embodiments described herein are not limited to any particular type or structure of XR display device.

[0058] Figure 1D This is a system diagram illustrating a set of example interfaces of a system according to some embodiments. Interfaces such as... Figure 1D Systems such as 1000 are used to implement extended reality display devices and their control electronics. System 150 can be implemented as a device including the various components described below and configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptops, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 150 can be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 150 are distributed across multiple ICs and / or discrete components. In various embodiments, system 150 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input ports and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described in this document.

[0059] System 150 includes at least one processor 152 configured to execute instructions loaded therein for implementing aspects such as those described in this document. Processor 152 may include embedded memory, input / output interfaces, and various other circuitry as known in the art. System 150 includes at least one memory 154 (e.g., a volatile memory device and / or a non-volatile memory device). System 150 may include a storage device 158, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 158 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0060] System 150 includes an encoder / decoder module 156 configured to, for example, process data to provide encoded or decoded video, and the encoder / decoder module 156 may include its own processor and memory. Encoder / decoder module 156 represents one or more modules that can be included in a device to perform encoding and / or decoding functions. It is well known that a device may include one or both encoding and decoding modules. Additionally, encoder / decoder module 156 may be implemented as a separate element of system 150, or may be incorporated into processor 152 as a combination of hardware and software as known to those skilled in the art.

[0061] Program code to be loaded onto processor 152 or encoder / decoder 156 to execute the aspects described in this document may be stored in storage device 158 and subsequently loaded onto memory 154 for execution by processor 152. According to various embodiments, one or more of processor 152, memory 154, storage device 158, and encoder / decoder module 156 may store one or more items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing equations, formulas, operations, and operational logic.

[0062] In some embodiments, the memory within processor 152 and / or encoder / decoder module 156 is used to store instructions and provide working memory for processing during encoding or decoding. However, in other embodiments, external memory (e.g., processor 152 or encoder / decoder module 152) is used for one or more of these functions. External memory may be memory 154 and / or storage device 158, such as volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory (such as RAM) is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG stands for Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC stands for High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Various Video Coding, a new standard developed by the Joint Video Experts Group JVET).

[0063] Inputs to the components of system 150 can be provided through various input devices as indicated in block 172. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) component (COMP) input terminals (or a set of COMP input terminals); (iii) universal serial bus (USB) input terminals; and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 1C Other examples not shown include composite video.

[0064] In various embodiments, the input device of block 172 has associated corresponding input processing elements as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal, or limiting a signal band to a band), (ii) down-converting the selected signal, (iii) re-band-limiting the signal to a narrower band to select (e.g.,) a signal band that may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section in various embodiments includes one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF section may include tuners that perform various functions among these functions, including, for example, down-converting a received signal to a lower frequency (e.g., intermediate frequency or near-baseband frequency) or down-converting it to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. Various embodiments rearrange the order of the components described above (and others), remove some of these components, and / or add other components that perform similar or different functions. Adding components may include inserting components between existing components, such as, for example, inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0065] Additionally, the USB and / or HDMI terminals may include corresponding interface processors for connecting system 150 to other electronic devices across USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented as needed, for example, within a separate input processing IC or within processor 152. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed, either within a separate interface IC or within processor 152. Demodulation, error correction, and demultiplexing streams are provided to various processing elements, including, for example, processor 152 and encoder / decoder 156, which operate in conjunction with memory and storage elements to process the data streams as needed for presentation on an output device.

[0066] Various components of system 150 can be provided within an integrated housing in which various components can be interconnected and transmit data therebetween using a suitable connection arrangement 174 (e.g., internal buses as known in the art, including inter-IC (I2C) buses, wiring and printed circuit boards).

[0067] System 150 includes a communication interface 160 that enables communication with other devices via a communication channel 162. The communication interface 160 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 162. The communication interface 160 may include, but is not limited to, a modem or network interface card (NIC), and the communication channel 162 may be implemented, for example, within a wired and / or wireless medium.

[0068] In various embodiments, a wireless network such as Wi-Fi (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)) is used to stream or otherwise provide data to system 150. In these embodiments, Wi-Fi signals are received via a communication channel 162 and a communication interface 160 suitable for Wi-Fi communication. The communication channel 162 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to system 150, delivering data via an HDMI connection to input block 172. Still other embodiments use an RF connection to input block 172 to provide streaming data to system 150. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0069] System 150 can provide output signals to various output devices, including display 176, speaker 178, and other peripheral devices 180. Display 176 in various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. Display 176 can be used in televisions, tablet computers, laptop computers, cellular phones (mobile phones), or other devices. Display 176 can also be integrated with other components (e.g., as in a smartphone) or separate (e.g., an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 180 include one or more of a stand-alone digital video disc (or digital universal disc) (DVR, for both terms), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 180 that provide functionality based on the output of system 150. For example, a disk player performs the function of playing the output of system 150.

[0070] In various embodiments, signaling (such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols enabling device-to-device control with or without user intervention) is used to transmit control signals between system 150 and display 176, speaker 178, or other peripheral devices 180. Output devices can be communicatively coupled to system 1000 via dedicated connections through corresponding interfaces 164, 166, and 168. Alternatively, output devices can be connected to system 150 via communication interface 160 using communication channel 162. In electronic devices (such as, for example, televisions), display 176 and speaker 178 can be integrated into a single unit with other components of system 150. In various embodiments, display interface 164 includes a display driver, such as, for example, a timing controller (TCon) chip.

[0071] For example, if the RF portion of input 172 is part of a separate set-top box, then display 176 and speaker 178 can alternatively be separated from one or more other components. In various embodiments where display 176 and speaker 178 are external components, output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0072] System 150 may include one or more sensor devices 168. Examples of sensor devices that may be used include one or more GPS sensors, gyroscope sensors, accelerometers, light sensors, cameras, depth cameras, microphones, and / or magnetometers. Such sensors can be used to determine information such as the user's position and orientation. Where system 150 is used as a control module (such as control modules 124, 132) for an extended reality display, the user's position and orientation can be used to determine how image data is rendered so that the user perceives the correct portion of a virtual object or scene from the correct viewpoint. In the case of a head-mounted display device, the device's own position and orientation can be used to determine the user's position and orientation for the purpose of rendering virtual content. In the case of other display devices such as telephones, tablets, computer monitors, or televisions, other inputs can be used to determine the user's position and orientation for the purpose of rendering content. For example, the user can use a touchscreen, keypad or keyboard, trackball, joystick, or other inputs to select and / or adjust the desired viewpoint and / or viewing direction. When the display device has sensors such as accelerometers and / or gyroscopes, the viewpoint and orientation used for rendering content can be selected and / or adjusted based on the movement of the display device.

[0073] The embodiments may be executed by computer software implemented by processor 152, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. As a non-limiting example, memory 154 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, such as optical storage devices, magnetic storage devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 152 may be of any type suitable for the technical environment and may encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0074] Scene description framework for XR This principle typically relates to the realm of extended reality scene description and rendering. It is also understood within the context of formatting and playback of extended reality applications when rendering on end-user devices such as mobile devices or head-mounted displays (HMDs).

[0075] In XR applications, scene descriptions are used to combine a clear and easily parsed description of the scene structure with some binary representation of the media content.

[0076] In time-based media streaming, the scene description itself can evolve over time to provide relevant virtual content for each sequence of the media stream. For example, a virtual bottle could be displayed during a video sequence of people drinking alcohol for advertising purposes.

[0077] This behavior can be achieved by relying on a framework defined in the scene description of the MPEG media document, namely Information technology – Coded representation of immersive media – Part 14: Scene Description for MPEG media, ISO / IEC DIS 23090-14:2021 (E). A scene update mechanism based on a JSON patching protocol as defined in IETF RFC 6902 can be used to synchronize virtual content with the MPEG media stream.

[0078] Figure 1E This is a system diagram illustrating a set of example interfaces for a scene description (stored as an item in glTF.json) in an ISOBMFF file, three video tracks, one audio track, and one JSON patch update track, according to some embodiments. Figure 1E This is an example ISOBMFF file 190, and other elements can be found in such a file.

[0079] While the MPEG-I scene description framework ensures that timed media and corresponding virtual content are available at all times, it does not provide a description of how users can interact with scene objects at runtime to achieve an immersive XR experience. Therefore, it does not support user-specific XR experiences using immersive media.

[0080] The example embodiments described herein may be used to provide a scene description including virtual objects or light sources, but even if available, they do not necessarily display or render the virtual objects or light sources. In some embodiments, one or more of the following aspects may be considered when determining whether to display virtual objects or light sources.

[0081] Spatial factors can be considered when determining whether to display virtual objects or lights. For example, virtual objects or lights may not be displayed if the user environment is unsuitable (e.g., the user is too far from the rendered timing media), or if the user is not looking in the correct direction, or if the virtual object should be displayed in a specific area of ​​the user (e.g., above their undetected left hand).

[0082] When determining whether to display a virtual object or light source, timing can be considered. For example, if the user is not yet ready or wants to trigger the display of the object themselves (e.g., using a specific gesture), the virtual object or light source may not be displayed until an appropriate trigger is detected.

[0083] In some embodiments, the scene description specifies which objects or light sources the user is allowed to manipulate or interact with via potential haptic feedback.

[0084] Runtime interactivity Figure 2 This is a system diagram illustrating a set of example interfaces of an MPEG-I node hierarchy supporting scene interactivity according to some embodiments. According to this principle, besides... Figure 2 The MPEG-I node hierarchy 200 and such as about Figure 4 In addition to the described node tree, behavioral metadata items (referred to herein as "behaviors") are added to the scene description. In the example embodiment, the scene description, which evolves over time, is enhanced by adding information identifying behaviors. These behaviors may be associated with predefined virtual objects on which interactivity is permitted for a user-specific XR experience when running.

[0085] In some embodiments, these behaviors evolve over time. In such embodiments, the behaviors can be updated using an existing scene description update mechanism.

[0086] In the example embodiment, the behavior is characterized by one or more of the following properties: • One or more triggers, which define the conditions that must be met for activation.

[0087] • Trigger control parameters define the logical operations between defined triggers.

[0088] • Actions implemented in response to activation.

[0089] • Action control parameters define the order in which the defined actions are executed.

[0090] • Priority numbering enables the selection of the highest priority behavior when several behaviors occur simultaneously on the same virtual object.

[0091] • Optional interrupt action, specifying how to terminate the behavior if it is no longer defined in a newly received scene update. For example, the behavior is no longer defined if the relevant object has been removed, or if the behavior is no longer associated with the current media (e.g., audio or video) sequence.

[0092] By adding these behaviors, time-related user interactivity can be defined in immersive content for XR experiences.

[0093] When the second scene description is received, some of the behaviors in the first scene description may be "in progress," meaning they have been triggered and their actions are running. The second scene description can be provided as updated metadata (i.e., metadata describing the differences between the first and second scene descriptions). The second scene description includes a node tree describing objects that may be the same as or different from those in the first scene description. Objects in the node tree of the first scene description may no longer exist in the second description. If objects related to the running actions of an ongoing behavior are missing in the second scene description, then those ongoing behaviors are no longer applicable. Similarly, if an ongoing behavior is not defined in the second description, then the ongoing behavior is no longer applicable. The interrupt action field describes how to properly interrupt the running actions of an ongoing behavior.

[0094] Figure 3 This is a block diagram illustrating an example of the logical relationship between trigger information (describing triggers 1 to n), action information (describing actions 1 to m), and behavior information (describing the relationship between triggers and actions) according to some embodiments, wherein triggers and actions may refer to one or more nodes in a scene description (such as a hierarchical scene graph).

[0095] In XR applications, scene descriptions are used to combine a clear and easily parsed description of the scene structure with some binary representation of the media content. The above section describes the action mechanism of scene descriptions. These behaviors are associated with predefined virtual objects, on which interactivity is permitted for a user-specific XR experience. Figure 3The diagram illustrates the structure 300 of the example behavior mechanism. Example structure 300 shows example trigger information 302 and action information 304. Within trigger information 302 are example triggers 1 (306), 2 (308), ..., n (310). Within action information 304 are example triggers 1 (312), 2 (314), ..., n (316). Triggers 306, 308, 310 and actions 312, 314, 316 are shown as having example relationships with each node 318.

[0096] Figure 4 This is a schematic plan view illustrating example relationships between objects describing an extended reality scene according to some embodiments. In this example, scene graph 400 includes descriptions of real-world objects 412, such as a “flat horizontal surface” (which could be a table, floor, or plate), and descriptions of virtual objects 414, such as an animation of a walking character. Scene graph node 414 is associated with media content item 416, which is an encoding (e.g., as a textured animated 3D mesh) of data used to render and display the walking character. Scene graph 400 also includes node 410, which describes the spatial relationships between the real-world objects described in node 412 and the virtual objects described in node 414. In this example, node 410 describes the spatial relationships that allow the character to walk on a flat surface. When an XR application is started, media content item 416 is loaded, rendered, and buffered for display upon triggering. When a flat surface is detected in the real-world environment by a sensor (or, in some embodiments, a camera), the application displays the buffered media content item, as described in node 410. Timing is managed by the application based on features detected in the real-world environment and the timing of animations. Nodes in the scene graph may also omit descriptions and simply serve as parent nodes for child nodes.

[0097] XR applications are diverse and can be applied to various contexts and real or virtual environments. For example, in industrial XR applications, when a reference object (part B of the engine) is detected in a real environment by a camera mounted on a head-mounted display, a virtual 3D content item (e.g., part A of the engine) is displayed. The 3D content item is positioned in the real world with a position and scale defined relative to the detected reference object.

[0098] For example, in an XR application for interior design, a 3D model of furniture is displayed when a given image from a catalog is detected in the input camera view. The 3D content is positioned in the real world with a location and scale defined relative to the detected reference image. In another application, some audio files might start playing when a user enters an area near a church (realistically or virtually rendered in an extended reality environment). In yet another example, an advertising jingle sound file could play when a user sees a given can of soda in a real-world environment. In outdoor gaming applications, various virtual characters may appear based on the semantics of the scene observed by the user. For example, bird characters are suitable for trees, so if the XR device's sensors detect a real-world object described by the semantic label "tree," birds flying around a tree could be added. In a companion application implemented with smart glasses, car noise could be played in the user's headphones when a car is detected within the user's camera's field of view to warn them of potential danger; furthermore, the sound could be spatialized so that it travels from the direction of the detected car.

[0099] XR applications can also enhance video content rather than the real-world environment. The video is displayed on a rendering device, and when a timed event is detected in the video, virtual objects described in the node tree are overwritten. In this context, the node tree only includes descriptions of the virtual objects.

[0100] Referring to an example embodiment of the range description of an MPEG-I scene description framework using the Khronos glTF extension mechanism, which supports additional scene description features such as node trees, the principles described herein are not limited to any particular scene description framework.

[0101] In the example implementation, the glTF scene description is extended to support interactivity. The interactivity extension is applied at the glTF scene level and is referred to as MPEG_scene_interactivity. The corresponding semantics are provided in Table 1. Table 1 illustrates the current top-level extension of the “MPEG_scene_interactivity” framework, which includes the “triggers,” “actions,” and “behaviors” features. See document ISO / IEC 23090-14, CDAM 2: Support for Haptics, Augmented Reality, Avatars, Interactivity, MPEG-I Audio, and Lighting, ISO / IEC JTC 1 / SC 29 / WG 03 N00797 (“MPEG Extension”). Table 1: Semantics of the example MPEG_scene_interactivity extension.

[0102] In Table 1 and other semantic tables described herein, the “Usage” column indicates “M” for “mandatory” features and “O” for “optional” features. However, such features may be “mandatory” or “optional” only depending on the specific proposed grammar. Features marked as “mandatory” are not necessarily essential for implementing the application. For example, in some embodiments, features marked as “mandatory” exist to meet the expectations of a particular type of parsing and rendering software; however, in other embodiments, the feature may be optional, or the feature may be omitted entirely, wherein the corresponding functionality is implemented using default values ​​or is not implemented at all without departing from the scope of this disclosure.

[0103] Extended Reality (XR) is a technology that enables interactive experiences where real-world environments and / or video content are augmented with virtual content, which can be defined across multiple sensory modalities, including visual, auditory, tactile, and other sensory modalities. During application runtime, the virtual content (e.g., 3D content or audio / video files) is rendered in real-time in a manner consistent with the user context (environment, viewpoint, device, etc.). Scene graphs (such as those proposed by Khronos / glTF and their extensions defined in formats such as MPEG Scene Description or Apple / USDZ) are possible ways to represent the content to be rendered. They combine, on the one hand, a declarative description of the scene structure linking real-world objects and virtual objects, and on the other hand, a binary representation of the virtual content. While such scene description frameworks ensure that timed media and corresponding associated virtual content are available at any time during application rendering, they do not describe how avatars can interact with scene objects in an MPEG-I Scene Description (SD) environment.

[0104] This application discusses 3D scene and object interaction in immersive environments. A set of triggers is described herein that enables interactivity between avatars and 3D scene objects, and allows avatar metadata to activate the trigger. This set of triggers can be used with MPEG-I Scene Description (SD) to support avatar interactivity in 3D environments. Current scene-level interactivity support only supports generic triggers for any node in the scene, as illustrated in Table 2. Therefore, a problem arises if nodes contain information that can be used to trigger interactive events, because the current interactivity framework of MPEG-I SD does not allow such avatar-based event-based signaling. With increasing interest in avatars and additional information at the node and scene levels, such an avatar mechanism is needed. See MPEG Extensions. Table 2: List of events currently available for triggering MPEG_scene_interactivity.

[0105] In the interactivity framework, "Behaviors" are a set of conditions that pair triggering events with specific actions and describe the temporal constraints of such conditions, allowing time-based events to occur in the 3D virtual environment. "Actions" are modifications to 3D nodes that can affect their spatial location, their materials, media controls, or haptic feedback. "Triggers" are events that occur between nodes or users. Table 2 shows a list of triggers currently available for MPEG_scene_interactivity, which can be found in Table 8.2-3 of the MPEG extension.

[0106] The following sections present new extensions that allow glTF models to use and interact with avatars and other 3D objects. These sections introduce new triggers that can be combined with actions specific to the avatar interactivity framework. These additions can be applied at the scene level in the "MPEG_scene_interactivity" section and can be extended to the node level if needed.

[0107] Avatar Trigger The following sections detail the elements, their associated meanings, JSON encoding schemes, and how they can be used in MPEG-ISD. These sections focus on the triggering representation of avatars and scenes, 3D objects, and user interactivity. User representation is compatible with SD content.

[0108] This format follows the glTF format and is compatible with current MPEG efforts to extend the glTF format using MPEG extensions. However, its meaning and purpose are "general" and can be encoded in other formats, such as Extensible Markup Language (XML) and Universal Scene Description (USD).

[0109] Two example configurations compatible with MPEG-I SD are shown below. The first configuration adds a top-level avatar trigger and several subtypes under the top-level avatar trigger. The first configuration is detailed in Tables 3, 4, and 5. For some embodiments, this first configuration can be specified as the entire avatar trigger at the scene level of the 3D scene (Table 2).

[0110] The second configuration adds a dedicated avatar trigger at the same level as the current trigger shown in Table 2. This second configuration is detailed in Tables 6 and 7. In some embodiments, this second configuration can be specified as a sub-avatar trigger at the scene level (Table 2) of the 3D scene. While these sections are independent of each other, they can complement one another. For example, the “TRIGGER_AVATAR” trigger presented in Table 3 can point to a scheme of triggers presented in other sections, such as “TRIGGER_AVATAR_SOCIAL”.

[0111] Configuration 1: Triggering the ultimate form of MPEG-I SD This section introduces the new “TRIGGER_AVATAR” trigger, which is added to the list of original triggers shown in Table 2 above. Table 3 includes the “TRIGGER_AVATAR” trigger as an extension of the MPEG-I Scene Description Interactivity Framework. For example, Table 8.2-3 of the MPEG extensions can be updated as shown in Table 3 below.

[0112] The new “generic” avatar triggers shown below are introduced at the same hierarchical level as the triggers presented in Table 2. Table 3: Trigger list for MPEG_scene_interactivity plus generic avatar triggers.

[0113] Table 4 illustrates the semantics of the TRIGGER_AVATAR type. In some embodiments, the “avatarTrigger” string may indicate details of a specific avatar trigger, such as specific parameters that vary depending on the avatar trigger type. Table 4: Semantics of the “TRIGGER_AVATAR” type.

[0114] Table 5 illustrates the subtypes of the "avatarTrigger" attribute. These attribute subtypes provide example schemas used and referenced by the "avatarTrigger" trigger type, demonstrating the flexibility and versatility of "generic" avatar triggers. Table 5 introduces the subtypes of avatars that can be applied to avatars interacting within interactive areas. Table 5 also introduces specialized avatar triggers at the same hierarchical level as the triggers presented in Table 2. Table 5: Avatar Trigger Subtypes of the “avatarTrigger” Property in Table 4.

[0115] Configuration 2: Dedicated top-level avatar trigger for MPEG-I SD Table 6 describes specific incarnation triggers, which are added to the list of original triggers shown in Table 2. Table 6: Trigger list for MPEG_scene_interactivity with dedicated avatar.

[0116] Table 7 presents the semantic descriptions of avatar triggers and action properties. Table 7: Semantic description of the nature of the new action.

[0117] Trigger preprocessing Figure 5 This is a flowchart illustrating an example preprocessing step triggered by an avatar according to some embodiments. Figure 5 Example process 500 shows how avatar triggers can be processed in immersive and interactive systems via preprocessing to verify the construction of behaviors and triggers. Figure 5 This demonstrates the processing model used to resolve previously presented avatar triggers.

[0118] The application can parse each trigger present in behavior 502 and perform an avatar trigger check 504. If a trigger that is not an avatar trigger exists, the process can process another (e.g., a non-avatar) trigger 506. If an avatar trigger is encountered, a check 508 is performed to see if the node listed in the trigger is "MPEG_node_avatar". If none of the nodes are avatar extension nodes, an error can be indicated 512. Such an indication can signal to the application (or any other engine handling such an interactive model) that if the list of nodes has not changed, the trigger and the current behavior will not be validated for processing. This allows the engine to optimize its computation and ignore or correct behaviors that may never be completed at the parsing stage. If all nodes are avatar extension nodes, processing (e.g., parsed triggers) continues 510. For some embodiments, Figure 5 The limitations shown can be applied to Tables 3 through 7 above.

[0119] At runtime, the processing model can remain unchanged from the original interactive model. For example, if a certain condition is not met, the application continues. In some embodiments, if all triggering conditions are met for the evaluated behavior, its action will be initiated. In some embodiments, if one or more conditions are not met for one of the triggers, the application continues to evaluate scene updates until all triggering conditions are met. In some embodiments, if one or more conditions are not met for one of the triggers, the application continues to evaluate each scene update until all triggering conditions are met.

[0120] Avatar Action List Tables 8 through 13 introduce example avatar-related metadata parameters. These parameters are merely examples of the numerous possibilities for avatar-related metadata. Many of these example parameters are used in the examples shown later in the tables. Table 8: Types of metadata. Table 9: Types of social actions. Table 10: Types of PEGI age classes. Table 11: Types of PEGI content descriptors. Table 12: Capability Semantics. Table 13: Disability semantics.

[0121] glTF Scheme Example The following glTF schemes are example (non-exhaustive) instances triggered by clients that support "MPEG_scene_interactivity". Numerous instantiations are possible, depending on the application, and several examples are given in the following sections for illustrative purposes.

[0122] Figures 6A-6B A list of code forms an example code structure illustrating the basic scene elements according to some embodiments. Figure 6A and 6B Together, sample code lists 600 and 650 show some basic scene elements. There are two nodes. The first node is an avatar node extension ( Figure 8A The camera in lines 17 to 34 Figure 8A (lines 3 to 35), and the second node is a sphere node ( Figure 8A (Lines 36 to 44). These elements are related to Figures 7A to 9 Use the example code list shown in B together.

[0123] Figures 7A-7B A list of code examples illustrating the independently triggered sample code structures based on some embodiments of avatar-related information is provided. In this list of sample code 700 and 750, there are... Figure 6A and 6BThe basic scene shown has two nodes. The first node is a camera with an avatar node extension, and the second node is a sphere node used for graphical interaction. These objects are examples of the types of objects that can exist in the scene. Triggering TRIGGER_ACTIVATE_FIRST_ENTER indicates how and when a combination of triggers is activated. Triggering TRIGGER_ACTIVATE_FIRST_ENTER corresponds to the case where an action is initiated once when one or more of the trigger conditions are met for the first time.

[0124] This example instantiates the following triggers: `trigger_proximity`, `trigger_social`, `trigger_restricted`, `trigger_parental`, `trigger_speech`, `trigger_capabilities`, and `trigger_disabilities`. The behavior class instantiates the behavior of each of the triggers paired with `trigger_proximity`, such as `trigger_proximity` and `trigger_social`; `trigger_proximity` and `trigger_restricted`, allowing two objects (e.g., an avatar and a sphere) to have a two-step interaction. In some embodiments, this means proximity triggering ("..."). Figure 8B The 7th to 12th lines are triggered before the second trigger is triggered.

[0125] Each action listed in this example is initiated by a proximity condition, followed by a comparison condition that compares the triggered value against the avatar node's metadata. For example, the first action initiates both a proximity and a social trigger. The social trigger has "authorized" and "comparator" properties, and uses the value of the "comparator" property to compare the value of the "authorized" property. The node avatar (0) is compared against these properties, and if they are valid, the action is initiated, and the application continues until another interactive signaling is received.

[0126] In this example, the avatar node is allowed to interact with node 1 (index 1) when the proximity between two nodes is between 0 and 1 cm (triggered via "TRIGGER_AVATAR_SOCIAL"). The interaction can be defined by the implementing application. For example, node 0 (which can be an avatar) can pick up node 1 (which can be a sphere) and manipulate the sphere in 3D space.

[0127] This example also specifies restricted and parental triggers, which relate to the ability to interact with the object. For example, if such parameters are not met, the application may not even allow social interaction to occur. For instance, the password might not match, or the object (the sphere) might contain content (such as sound or display content) that is inappropriate for the avatar's age.

[0128] Triggers grant the avatar the ability to "walk." The application allows the avatar user to move / displace within a 3D environment while holding a sphere object (node ​​1). Disability triggers ("TRIGGER_AVATAR_DISABILITIES") adapt to the avatar user's disabilities. For example, the sphere object can activate visual cues in the scene, allowing the avatar user to walk and manipulate the surrounding 3D scene without audio / sound.

[0129] Figures 8A-8B A list of code examples illustrating sample code structures (e.g., avatar triggering conditions) according to some embodiments is provided. For example code lists 800 and 850, similar to... Figures 7A-7B The code list shown has elements from... Figure 6A and 6B The basic scene shown has two nodes. The first node is a camera with an avatar node extension, and the second node is a sphere node used for graphical interaction. These objects are examples of the types of objects that may exist in the scene.

[0130] This example instantiates the following triggers: `trigger_proximity` and `trigger_avatar`. The behavior class instantiates the behavior of the trigger paired with `trigger_proximity`, such as `trigger_proximity` and `trigger_avatar`, allowing two objects (e.g., an avatar and a sphere) to have a two-step interaction. In some embodiments, this means triggering proximity (…). Figure 8B (Lines 7 to 12) are triggered before the second trigger is triggered.

[0131] Figure 8A and 8B The example shown is the same as Figure 7A and 7B The example is similar, except that the comparator uses the external avatar metadata parameter "schema.avatar" to evaluate the node avatar triggering conditions. This operation can correspond to comparing multiple attributes simultaneously using a single comparator.

[0132] In the current example, if the avatar satisfies the avatar trigger ( Figure 8B The other conditions listed below (lines 13-28) will trigger when the avatar enters the proximity trigger ( Figure 8BThe behavior of a single listed node is first triggered when any of the nodes specified in lines 7 to 12 are within a proximity of [0.0 to 1.0 distance].

[0133] according to Figure 8A and 8B In the example shown, the avatar triggers behavior only if: (1) the avatar is authorized to interact; (2) the avatar meets the listed permission IDs; (3) the avatar is associated with microphone input; (4) the avatar has a "colorblind" disability; (5) the avatar is able to walk; (6) the avatar is able to express admiration and joy in animations; (7) the avatar is able to engage in social interactions; and (8) the avatar is associated with a PEGI age of 18 and is authorized to consume content labeled "bad_language," "violence," and "fear." If the avatar does not meet any of these properties or conditions, the avatar may not trigger the actions listed in the behavior list. This example illustrates setting various metadata settings in the associated object to meet the example criteria listed above. For some embodiments, the node object may not necessarily be an avatar node, although the metadata of such a node may be related to the example shown above.

[0134] Figure 9 This is a list of code illustrating an example code structure for external avatar triggering metadata according to some embodiments. For this example code list 900, similar to the code list shown above, there are... Figure 6A and 6B The basic scene shown has two nodes. The first node is a camera with an avatar node extension, and the second node is a sphere node used for graphical interaction. These objects are examples of the types of objects that may exist in the scene.

[0135] This example instantiates the following triggers: `trigger_proximity` and `trigger_avatar`. The behavior class instantiates the behavior of the trigger paired with `trigger_proximity`, such as `trigger_proximity` and `trigger_avatar`, allowing two objects (e.g., an avatar and a sphere) to have a two-step interaction. In some embodiments, this means triggering proximity (…). Figure 8B (Lines 7 to 12) are triggered before the second trigger is triggered.

[0136] Figure 9 The example shown is the same as Figures 8A-8BThe example is similar, except that the parameter “schema.TRIGGER_SOCIAL” presented in Tables 6 and 7 is used by the comparator to determine whether the node’s avatar trigger has met the conditions. This example illustrates the use of external avatars or user representation metadata and is not limited to the triggers proposed in Tables 6 and 7.

[0137] Figure 10 This is a flowchart illustrating an example process for handling avatar triggers according to some embodiments. For some embodiments, example process 1000 may include obtaining scene description data of a three-dimensional (3D) scene at 1002. For some embodiments of example process 1000, the scene description data may include behavioral information at 1004, including trigger information describing at least one trigger condition and action information describing actions to be performed on scene elements in the 3D scene. For some embodiments of example process 1000, at least one trigger condition corresponds to an avatar. For some embodiments, example process 1000 may further include performing an action on a scene element at 1006 in response to determining that at least one trigger condition has occurred.

[0138] Figure 11 This is a flowchart illustrating an example process for processing avatar triggers according to some embodiments. For some embodiments, example process 1100 may include obtaining 1102 scene description data of a three-dimensional (3D) scene. For some embodiments of example process 1100, the scene description data may include 1104 behavioral information, including trigger information describing at least one trigger condition, and action information describing actions to be performed on scene elements in the 3D scene. For some embodiments of example process 1100, at least one trigger condition corresponds to an avatar. For some embodiments of example process 1100, at least one trigger condition corresponds to an entire avatar trigger specified at the scene level of the 3D scene. For some embodiments, example process 1100 may further include performing 1106 an action on a scene element in response to determining that at least one trigger condition has occurred.

[0139] Figure 12This is a flowchart illustrating an example process for handling avatar triggers according to some embodiments. For some embodiments, example process 1200 may include obtaining 1202 scene description data of a three-dimensional (3D) scene. For some embodiments of example process 1200, the scene description data may include 1204 behavioral information, including trigger information describing at least one trigger condition, and action information describing actions to be performed on scene elements in the 3D scene. For some embodiments of example process 1200, at least one trigger condition corresponds to an avatar. For some embodiments of example process 1200, at least one trigger condition corresponds to a sub-avatar trigger specified at the scene level of the 3D scene, and for some embodiments, example process 1200 may further include performing 1206 an action on a scene element in response to determining that at least one trigger condition has occurred.

[0140] While the methods and systems according to some embodiments are generally discussed in the context of extended reality (XR), some embodiments can be applied to any XR context, such as, for example, virtual reality (VR) / mixed reality (MR) / augmented reality (AR) contexts. Furthermore, although the term "head-mounted display (HMD)" is used herein according to some embodiments, for some embodiments, some embodiments can be applied to, for example, wearable devices with XR, VR, AR, and / or MR capabilities (which may or may not be attached to the head).

[0141] Example methods according to some embodiments may include: obtaining scene description data of a three-dimensional (3D) scene, wherein the scene description data includes behavioral information, the behavioral information including: trigger information describing at least one trigger condition, and action information describing an action to be performed on a scene element in the 3D scene, wherein the at least one trigger condition corresponds to an avatar in the 3D scene, and wherein the at least one trigger condition tests a specific characteristic of the avatar; and performing an action on the scene element in response to determining that at least one trigger condition has occurred.

[0142] Example apparatuses according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0143] Example methods according to some embodiments may include: obtaining scene description data of a three-dimensional (3D) scene, wherein the scene description data includes behavioral information, the behavioral information including: trigger information describing at least one trigger condition, and action information describing an action to be performed on a scene element in the 3D scene, wherein the at least one trigger condition corresponds to an avatar in the 3D scene; and performing an action on the scene element in response to determining that at least one trigger condition has occurred.

[0144] For some embodiments of the example method, at least one triggering condition corresponds to a generic avatar trigger.

[0145] For some embodiments of the example method, at least one triggering condition corresponds to a combination of a general avatar triggering condition and a specific avatar triggering condition.

[0146] For some embodiments of the example method, at least one triggering condition corresponds to a specific avatar triggering condition.

[0147] For some embodiments of the example method, at least one triggering condition corresponds to a general avatar trigger that points to one or more specific avatar triggering conditions.

[0148] For some embodiments of the example method, specific avatar triggering conditions are selected from the group including: social actions, avatar permission, age restrictions, media device conditions, avatar abilities, and avatar disabilities.

[0149] For some embodiments of the example method, the triggering information may include: generic avatar trigger, comparator field, and node information.

[0150] For some embodiments of the example method, at least one triggering condition is based on a generic avatar trigger, a comparator field, and node information.

[0151] For some embodiments of the example method, the trigger information may include at least one metadata field.

[0152] For some embodiments of the example method, at least one triggering condition corresponds to at least one metadata field.

[0153] For some embodiments of the example method, at least one metadata field corresponds to an avatar.

[0154] For some embodiments of the example method, at least one metadata field corresponds to a second avatar in the 3D scene.

[0155] For some embodiments of the example method, at least one metadata field corresponds to a scene element that is separate from the avatar.

[0156] For some embodiments of the example method, at least one triggering condition corresponds to the avatar's social action.

[0157] For some embodiments of the example method, at least one triggering condition corresponds to a social action of a second avatar in the 3D scene.

[0158] For some embodiments of the example method, at least one triggering condition corresponds to an age restriction associated with the avatar.

[0159] For some embodiments of the example method, at least one triggering condition corresponds to a content restriction associated with the avatar.

[0160] For some embodiments of the example method, at least one triggering condition corresponds to an ability associated with the avatar.

[0161] For some embodiments of the example method, at least one triggering condition corresponds to a disability associated with the avatar.

[0162] For some embodiments of the example method, the scene description data is compatible with MPEG-I scene description (SD).

[0163] For some implementations of the example methods, the scene description data is compatible with the glTF configuration.

[0164] Some embodiments of the example method may further include: determining that the node associated with at least one triggering condition does not support MPEG node avatars; and transmitting the node configuration error to another device.

[0165] For some embodiments of the example method, at least one triggering condition corresponds to the entire avatar trigger specified at the scene level in the 3D scene.

[0166] For some embodiments of the example method, at least one triggering condition corresponds to a sub-avatar trigger specified at the scene level of the 3D scene.

[0167] Example apparatuses according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0168] This disclosure describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described in detail, and are generally described in a manner that may sound restrictive, at least for the purpose of illustrating the various characteristics. However, this is for the purpose of clarity and does not limit the disclosure or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. Furthermore, the aspects described can also be combined and interchanged with those described in earlier filings.

[0169] The aspects described and contemplated in this disclosure can be implemented in many different forms. While some embodiments are specifically illustrated, other embodiments are contemplated, and the discussion of particular embodiments does not limit the breadth of implementation. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to the transmission of generated or encoded bitstreams. These and other aspects can be implemented as methods, apparatus, computer-readable storage media having instructions stored thereon for encoding or decoding video data according to any of the methods, and / or computer-readable storage media having a bitstream generated according to any of the methods stored thereon.

[0170] In this disclosure, the terms “reconstruction” and “decoding” are used interchangeably, as are the terms “pixel” and “sample”, and the terms “image,” “picture,” and “frame.” Generally, but not necessarily, the term “reconstruction” is used on the encoder side, while “decoding” is used on the decoder side.

[0171] The terms HDR (High Dynamic Range) and SDR (Standard Dynamic Range) generally convey specific values ​​of dynamic range to those skilled in the art. However, they are also intended for use in additional embodiments, where reference to HDR is understood to mean "higher dynamic range," and reference to SDR is understood to mean "lower dynamic range." Such additional embodiments are not bound by any specific value of dynamic range that may often be associated with the terms "high dynamic range" and "standard dynamic range."

[0172] Various methods are described herein, and each of these methods includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined. Furthermore, in various embodiments, terms such as "first," "second," etc., may be used to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding." Unless specifically required, the use of such terms does not imply a reordering of the modified operations. Therefore, in this example, the first decoding does not need to be performed before the second decoding and can occur, for example, before, during, or within a time period overlapping with the second decoding.

[0173] For example, various numerical values ​​may be used in this disclosure. Specific values ​​are for illustrative purposes only, and the aspects described are not limited to these specific values.

[0174] The embodiments described herein can be executed by computer software implemented by a processor or other hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. As a non-limiting example, the processor can be of any type suitable for the technical environment and can encompass one or more of microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0175] Various implementations involve decoding. As used herein, “decoding” can encompass all or part of a process performed on a received encoded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also includes, or alternatively includes, processes performed by a decoder of the various implementations described herein, such as extracting an image from a tiled (packed) image, determining an upsampling filter to use, and then upsampling the image, as well as flipping the image back to its intended orientation.

[0176] As a further example, in one embodiment, "decoding" refers only to entropy decoding; in another embodiment, "decoding" refers only to differential decoding; and in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process will be clear based on the specific context of the description.

[0177] Various implementations involve encoding. In a manner similar to the discussion above regarding “decoding,” the term “encoding,” as used herein, can encompass all or part of a process performed, for example, on an input video sequence to produce an encoded bitstream. In various embodiments, such a process includes one or more processes typically performed by an encoder, such as segmentation, differential coding, transform, quantization, and entropy coding. In various embodiments, such a process also includes, or alternatively includes, processes performed by encoders of the various implementations described herein.

[0178] As a further example, in one embodiment, "encoding" refers only to entropy encoding; in another embodiment, "encoding" refers only to differential encoding; and in yet another embodiment, "encoding" refers to a combination of differential and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process will be clear based on the specific context of the description.

[0179] Various embodiments involve rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, usually with constraints on computational complexity. Rate-distortion optimization is generally expressed as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist to address the rate-distortion optimization problem. For example, these approaches can be based on extensive testing of all encoding options, including all considered modes or encoding parameter values, and a complete evaluation of their encoding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to save encoding complexity, particularly by calculating approximate distortion based on predicting or predicting the residual signal rather than the reconstructed signal. A hybrid of these two approaches can also be used, such as by using approximate distortion only for some of the possible encoding options and full distortion for others. Other methods evaluate only a subset of the possible encoding options. More generally, many methods employ any of a variety of techniques to perform optimization, but optimization is not necessarily a complete evaluation of both encoding costs and associated distortion.

[0180] When the accompanying drawings are presented as flowcharts, it should be understood that block diagrams of the corresponding devices are also provided. Similarly, when the accompanying drawings are presented as block diagrams, it should be understood that flowcharts of the corresponding methods / processes are also provided.

[0181] The implementations and aspects described herein can be implemented, for example, in a method or process, apparatus, software program, data stream, or signal. Even if discussed only in the context of a single form of implementation (e.g., discussed only as a method), the implementation of the discussed features can also be implemented in other forms (e.g., apparatus or program). Apparatus can be implemented, for example, in suitable hardware, software, and firmware. Methods can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate the transfer of information between end users.

[0182] References to “an embodiment” or “an embodiment” or “an implementation” or “implementation”, and other variations thereof, mean that a particular feature, structure, characteristic, etc., described in connection with the embodiment is included in at least one embodiment. Therefore, the phrases “in an embodiment” or “in an embodiment” or “in an implementation” or “in an implementation” appearing throughout this disclosure, and any other variations thereof, do not necessarily refer to the same embodiment.

[0183] Additionally, this disclosure may relate to "determining" fragments of various information. Determining information may include one or more of, for example, estimation information, calculation information, prediction information, or information retrieved from memory.

[0184] Furthermore, this disclosure may relate to “accessing” fragments of various information. Accessing information may include one or more of the following, such as receiving information, retrieving information (e.g., retrieving information from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0185] Additionally, this disclosure may relate to “receiving” fragments of various information. As with “access,” “receiving” is intended to be a broad term. Receiving information may include, for example, one or more of the following: accessing information or retrieving information (e.g., retrieving information from memory). Furthermore, “receiving” is generally referred to in one or more ways during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0186] To understand this, for example, in the cases of “A / B,” “A and / or B,” and “at least one of A and B,” the use of any of the following “ / ,” “and / or,” and “at least one of…” is intended to cover selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C,” such wording is intended to cover selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). This can be extended to as many items as possible listed.

[0187] Furthermore, among other things, as used herein, the term "signaling" also refers to instructing the corresponding decoder to do something. For example, in some embodiments, the encoder signals a specific one of several parameters for selecting region-based filter parameters for artifact removal filtering. Thus, in embodiments, the same parameter is used on both the encoder and decoder sides. Therefore, for example, the encoder can transmit (explicitly signal) the specific parameter to the decoder so that the decoder can use the same specific parameter. Conversely, if the decoder already has the specific parameter as well as other parameters, signaling can be used without transmission (implicitly signaling) to allow only the decoder to know and select the specific parameter. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in a variety of ways. For example, in various embodiments, information is signaled to the corresponding decoder using one or more syntax elements, flags, etc. Although the verb form of the term "signaling" has been referred to above, the word "signal" can also be used as a noun herein.

[0188] The implementation can generate various signals, which are formatted to carry information, such as information that can be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, the signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding the data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is well known that signals can be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.

[0189] We have described several embodiments. Features of these embodiments may be provided individually or in any combination across various claim classes and types. Furthermore, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination across various claim classes and types: • Includes a bitstream or signal of one or more of the described syntax elements or their variants.

[0190] • Includes bitstreams or signals that convey the syntax of information generated according to any embodiment of the described embodiments.

[0191] • Create and / or transmit and / or receive and / or decode bit streams or signals including one or more of the described syntax elements or their variants.

[0192] • Create and / or transmit and / or receive and / or decode according to any embodiment of the described embodiments.

[0193] • Methods, processes, apparatus, media for storing instructions, media for storing data, or signals according to any of the embodiments described.

[0194] Note that the various hardware elements in one or more of the described embodiments are referred to as “modules” that perform (i.e., execute, implement, and so on) the various functions described herein in connection with the respective modules. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more memory devices) that a person skilled in the art would consider suitable for a given implementation. Each described module may also include executable instructions for performing one or more functions described as being performed by the respective module, and note that those instructions may take the form of hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and / or such instructions, or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and / or such instructions, and may be stored in any suitable one or more non-transitory computer-readable media (such as commonly referred to as RAM, ROM, etc.).

[0195] Although the features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein can be implemented as computer programs, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROMs and digital universal discs (DVDs)). A processor associated with the software can be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.

Claims

1. A method comprising: Obtain scene description data for a three-dimensional (3D) scene. The scene description data includes behavioral information, which includes: Trigger information describing at least one trigger condition, and Action information describing the actions to be performed on scene elements in a 3D scene. Among them, at least one triggering condition corresponds to an avatar in the 3D scene, and Among them, at least one trigger condition tests a specific characteristic of the avatar; and In response to determining that at least one triggering condition has occurred, an action is performed on the scene element.

2. An apparatus comprising: processor; as well as A memory storing instructions that, when executed by a processor, are operable to cause the device to perform the method of claim 1.

3. A method comprising: Obtain scene description data for a three-dimensional (3D) scene. The scene description data includes behavioral information, which includes: Trigger information describing at least one trigger condition, and Action information describing the actions to be performed on scene elements in a 3D scene. Among them, at least one trigger condition corresponds to an avatar in the 3D scene; and In response to determining that at least one triggering condition has occurred, an action is performed on the scene element.

4. The method according to claim 3, wherein, At least one trigger condition corresponds to a universal avatar trigger.

5. The method according to claim 3, wherein, At least one trigger condition corresponds to a combination of general avatar trigger and specific avatar trigger conditions.

6. The method according to claim 3, wherein, At least one trigger condition corresponds to a specific avatar trigger condition.

7. The method according to claim 3, wherein, At least one trigger condition corresponds to a general avatar trigger that points to one or more specific avatar trigger conditions.

8. The method according to any one of claims 5-7, wherein, Specific avatar trigger conditions are selected from groups including the following: social actions, avatar permission, age restrictions, media device conditions, avatar abilities, and avatar disabilities.

9. The method according to any one of claims 3-8, wherein, Trigger information includes: Universal Avatar Trigger, Comparator field, and Node information.

10. The method according to claim 9, wherein, At least one trigger condition is based on the generic avatar trigger, the comparator field, and node information.

11. The method according to any one of claims 3-10, wherein, The trigger information includes at least one metadata field.

12. The method according to claim 11, wherein, At least one trigger condition corresponds to at least one metadata field.

13. The method according to claim 12, wherein, At least one metadata field corresponds to the avatar.

14. The method according to claim 12, wherein, At least one metadata field corresponds to a second avatar in the 3D scene.

15. The method according to claim 12, wherein, At least one metadata field corresponds to a scene element that is separate from the avatar.

16. The method according to any one of claims 3-15, wherein, At least one trigger condition corresponds to the avatar's social action.

17. The method according to any one of claims 3-15, wherein, At least one trigger condition corresponds to a social action of the second avatar in the 3D scene.

18. The method according to any one of claims 3-15, wherein, At least one trigger condition corresponds to an age restriction associated with the avatar.

19. The method according to any one of claims 3-15, wherein, At least one trigger condition corresponds to a content restriction associated with the avatar.

20. The method according to any one of claims 3-15, wherein, At least one trigger condition corresponds to an ability associated with the avatar.

21. The method according to any one of claims 3-15, wherein, At least one trigger condition corresponds to a disability associated with the avatar.

22. The method according to any one of claims 3-21, wherein, Scene description data is compatible with MPEG-I Scene Description (SD).

23. The method according to any one of claims 3-21, wherein, Scene description data is compatible with glTF configuration.

24. The method according to any one of claims 3-23, further comprising: The node associated with at least one triggering condition is determined to be unsupported by MPEG node avatars; as well as The node configuration error was transmitted to another device.

25. The method according to any one of claims 3-23, wherein, At least one trigger condition corresponds to the entire avatar trigger specified at the scene level in the 3D scene.

26. The method according to any one of claims 3-23, wherein, At least one trigger condition corresponds to a child avatar trigger specified at the scene level in the 3D scene.

27. An apparatus comprising: processor; as well as A memory storing instructions that, when executed by a processor, are operable to cause the apparatus to perform the method of any one of claims 3 to 26.