Event-based updates in scene descriptions

By introducing trigger conditions and action information into the XR scene description data, and updating the scene description using the JSON patch mechanism, the problem of insufficient user interaction in the prior art is solved, and a user-specific immersive XR experience is realized.

CN120345259APending Publication Date: 2025-07-18INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380084633.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-15
Filing Date
2023-11-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing Extended Reality (XR) scenario description framework, while ensuring the availability of timing media and virtual content during application rendering, does not provide a description of how users interact with scene objects at runtime, resulting in the inability to implement a user-specific immersive XR experience.

Method used

By adding trigger conditions and action information to the scene description data, including updating actions and setting actions, using the JavaScript Object Notation (JSON) patch mechanism, the scene description data is updated in real time to respond to trigger conditions, and user interactivity is achieved.

Benefits of technology

It realizes a user-specific immersive XR experience, allowing users to interact with scene objects at runtime, and enhances the interactivity and immersion of XR applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345259A_ABST
    Figure CN120345259A_ABST
Patent Text Reader

Abstract

Some embodiments of a method may include obtaining scene description data for a three-dimensional (3D) scene, wherein the scene description data includes scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, and action information describing at least one action to be performed on a first scene element of the plurality of scene elements, wherein the action information comprises at least a list of updated actions; and updating action information, wherein the updating action information comprises a patch parameter; and in response to determining that (i) a trigger condition of the at least one trigger condition has occurred and (ii) the action information indicates that an update action is to be performed, performing the update action to apply the patch parameter as an update to the scene description data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to European Patent Application Serial No. EP22306885, titled "EVENT BASED UPDATE IN SCENE DESCRIPTION", filed on December 15, 2022, which is incorporated herein by reference in its entirety. Background Art

[0003] This section is intended to introduce the reader to aspects of the field that may be relevant to aspects of the principles described and / or claimed below. This discussion is considered to be helpful in providing background information to the reader to facilitate a better understanding of the aspects of the principles. Accordingly, it should be understood that these statements are to be read in this light and not as an admission of prior art.

[0004] Extended Reality (XR) is a technology enabling interactive experiences where the real - world environment and / or video content are enhanced by virtual content, which can be defined across multiple sensory modalities (including vision, hearing, touch, etc.). During the runtime of an application, virtual content (such as 3D content or audio / video files) is rendered in real - time in a manner consistent with the user context (environment, viewpoint, device, etc.). Scene graphs (such as, for example, the scene graph proposed by Khronos / glTF and its extensions defined in the MPEG scene description format or Apple / USDZ) are a possible way to represent the content to be rendered. They combine, on the one hand, a declarative description of the scene structure that links real - world environment objects and virtual objects and, on the other hand, a binary representation of the virtual content.

[0005] Although such a scene description framework ensures that timed media and the corresponding associated virtual content are available at any time during the rendering of an application, such a framework does not provide a description of how a user can interact with scene objects at runtime to obtain an immersive XR experience. Therefore, user - specific XR experiences for using immersive media are not supported. Summary of the Invention

[0006] A first example method according to some embodiments may include: obtaining scene description data of a three-dimensional (3D) scene, where the scene description data includes: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, and action information describing at least one action to be performed on a first scene element among the plurality of scene elements, where the action information includes: a list of at least update actions, and update action information, where the update action information includes: patch parameters; and in response to determining: (i) that a trigger condition among the at least one trigger condition has occurred, and (ii) that the action information indicates that an update action is to be performed, performing the update action to apply the patch parameters as an update to the scene description data.

[0007] For some embodiments of the first example method, determining that a trigger condition has occurred includes a logical determination that the trigger condition has occurred.

[0008] For some embodiments of the first example method, determining that a trigger condition has occurred includes determining that the trigger condition has produced a true result.

[0009] For some embodiments of the first example method, the patch parameters indicate one or more operations, including one or more of adding, removing, replacing, moving, and copying, and the one or more operations are applied to JavaScript Object Notation (JSON) elements of the scene description data.

[0010] For some embodiments of the first example method, the update action information further includes a shared field that indicates a state shared with a device connected to the 3D scene.

[0011] Some embodiments of the first example method may further include: in response to determining that the shared field indicates a true result, performing the update action on at least one device sharing the 3D scene.

[0012] Some embodiments of the first example method may further include: in response to determining that the shared field indicates a true result, performing the update action on all devices sharing the 3D scene.

[0013] For some embodiments of the first example method, the update action information further includes a location description field and a placedItems field, the location description field identifies the source of the location information, the placedItems field is an array, and the placedIems array indicates a patch operation that references a node to be placed at a given location.

[0014] For some embodiments of the first example method, the update action includes setting the positions of at least the scene elements within the scene, and at least the scene elements of the scene are referenced in at least the path operations according to the placedItems field, and setting the positions of at least the scene elements uses position information from the identified source.

[0015] For some embodiments of the first example method, the update action information further includes: an array of childOps that indicates patch operations for nodes whose references are to be set as children of an existing node in the scene description data; and an array of parents that indicates the nodes to which each new child node referenced in the patch operations from the childOps array is to be added.

[0016] For some embodiments of the first example method, the scene description data further includes behavior information.

[0017] For some embodiments of the first example method, each of the one or more operations includes a set of operation parameters, and each set of operation parameters includes: an operation identifier, a path to a JavaScript Object Notation (JSON) element of the scene description data, and a value to be applied.

[0018] For some embodiments of the first example method, the path is a JSON pointer that references a first scene element.

[0019] For some embodiments of the first example method, each of the one or more operations includes a set of operation parameters, and each set of operation parameters includes: an operation identifier, a path to a JavaScript Object Notation (JSON) element, and a source JSON element.

[0020] For some embodiments of the first example method, the path is a JSON pointer that references a first scene element.

[0021] For some embodiments of the first example method, the patch parameter indicates a JavaScript Object Notation (JSON) patch to be applied to the scene description data.

[0022] For some embodiments of the first example method, the patch parameter includes a location pointer to a JavaScript Object Notation (JSON) file that includes the JSON patch to be applied to the scene description data.

[0023] A first example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any one of the methods listed above.

[0024] A second example method according to some embodiments may include: obtaining scene description data of a 3D scene, where the scene description data includes: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, and action information describing at least one action to be performed on a first scene element among the plurality of scene elements, where the action information includes: a list of at least set actions, and set action information, where the set action information includes: a path that identifies the first scene element; and a value that indicates a new value of the first scene element; and in response to determining: (i) that a trigger condition of at least one trigger condition has occurred, and (ii) that the action information indicates that a set action is to be performed, performing the set action to set the value associated with the first scene element to the new value of the first scene element.

[0025] For some embodiments of the second example method, determining that a trigger condition has occurred includes a logical determination that the trigger condition has occurred.

[0026] For some embodiments of the second example method, determining that a trigger condition has occurred includes determining that the trigger condition has produced a true result.

[0027] For some embodiments of the second example method, the path identifies the position of the first scene element in the scene tree.

[0028] For some embodiments of the second example method, setting the value associated with the first scene element sets the value associated with the first scene element to the new value.

[0029] For some embodiments of the second example method, the set action information further includes a shared field that indicates a state shared with a device connected to the 3D scene.

[0030] Some embodiments of the second example method may further include: in response to determining that the shared field indicates a true result, performing the set action among the at least one action on at least one device sharing the 3D scene.

[0031] Some embodiments of the second example method may further include: in response to determining that the shared field indicates a true result, performing the set action among the at least one action on all devices sharing the 3D scene.

[0032] For some embodiments of the second example method, the set action information further includes a location description field, and the location description field identifies the source of the location information.

[0033] For some embodiments of the second example method, setting the action includes setting the position of the first scene element within the scene, and setting the position of the first scene element uses the location information from the identified source.

[0034] For some embodiments of the second example method, the path is a JavaScript Object Notation (JSON) pointer that references a first scene element.

[0035] For some embodiments of the second example method, the path includes path information that references a first scene element, and the path information corresponds to Internet Engineering Task Force (IETF) standards.

[0036] For some embodiments of the second example method, the scene description data further includes behavior information.

[0037] A second example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0038] A third example method according to some embodiments may include: obtaining scene description data for a 3D scene, where the scene description data includes: scene element information that describes each of a plurality of scene elements in the scene, trigger information that describes at least one trigger condition, and action information that describes at least one action to be performed on a first scene element among the plurality of scene elements, where the action information includes: a list of at least update actions, and update action information, where the update action information includes: a scene description patch; and in response to execution of a trigger to update an action, performing a first action among the at least one action on the scene description data.

[0039] A third example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0040] A fourth example method according to some embodiments may include: obtaining scene description data for a 3D scene, where the scene description data includes: scene element information that describes each of a plurality of scene elements in the scene, trigger information that describes at least one trigger condition, and action information that describes at least one action to be performed on a first scene element among the plurality of scene elements, where the action information includes: a list of at least set actions, and set action information, where the set action information includes: a path that identifies the first scene element; and a value that indicates a new value for the first scene element; and in response to execution of a trigger to set an action, performing a first action among the at least one action on the first scene element.

[0041] A fourth example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0042] A fifth example apparatus according to some embodiments may include at least one processor configured to perform any of the methods listed above.

[0043] A sixth example apparatus according to some embodiments may include a computer-readable medium storing instructions for causing one or more processors to perform any of the methods listed above.

[0044] A seventh example apparatus according to some embodiments may include at least one processor and at least one non-transitory computer-readable medium storing instructions for causing at least one processor to perform any of the methods listed above.

[0045] An example signal according to some embodiments may include: a scene description file generated according to any of the methods listed above.

[0046] A further example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform the methods listed above. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1A is a schematic side view illustrating an example waveguide display that may be used with an extended reality (XR) application according to some embodiments.

[0048] Figure 1B is a schematic side view illustrating an example alternative display type that may be used with an extended reality application according to some embodiments.

[0049] Figure 1C is a schematic side view illustrating an example alternative display type that may be used with an extended reality application according to some embodiments.

[0050] Figure 1D is a system diagram illustrating a set of example interfaces of a system according to some embodiments.

[0051] Figure 1E is a system diagram illustrating a set of example interfaces of a scene description (stored as an item in glTF.json), three video tracks, one audio track, and one JavaScript Object Notation (JSON) patch update track in an ISOBMFF file according to some embodiments.

[0052] Figure 2 It is a system diagram showing a set of example interfaces of an MPEG-I node hierarchy that illustrates elements supporting scene interactivity according to some embodiments.

[0053] Figure 3 It is a block diagram showing an example of the logical relationship between trigger information (describing triggers 1 to n), action information (describing actions 1 to m), and behavior information (describing the relationship between triggers and actions) according to some embodiments, where the triggers and actions can refer to one or more nodes in a scene description (such as a hierarchical scene graph).

[0054] Figure 4 It is a schematic plan view showing an example relationship of extended reality scene description objects according to some embodiments.

[0055] Figure 5 It is a data syntax diagram showing an example syntax of a data stream encoding an extended reality (XR) scene description according to some embodiments.

[0056] Figure 6 It is a data structure diagram showing an example JSON patch operation.

[0057] Figure 7 It is a flowchart showing an example timing sample patch to be applied to a scene description file.

[0058] Figure 8 It is a flowchart showing another example timing sample patch to be applied to a scene description file.

[0059] Figure 9 It is a flowchart showing an example process of event-based update according to some embodiments.

[0060] Figure 10 It is a data structure diagram showing an example JSON patch operation according to some embodiments.

[0061] Figure 11 It is a data structure diagram showing an example glTF file excerpt according to some embodiments.

[0062] Figure 12 It is a data structure diagram showing an example SET action structure according to some embodiments.

[0063] Figure 13 It is a flowchart showing an example SET action process according to some embodiments.

[0064] Figure 14A - 14B It forms a data structure diagram showing an example UPDATE action structure according to some embodiments.

[0065] Figure 15 is a flowchart illustrating an example update operation process according to some embodiments.

[0066] Figure 16 is a flowchart illustrating an example process for performing a setup operation according to some embodiments.

[0067] Figure 17 is a flowchart illustrating an example process for performing an update operation according to some embodiments.

[0068] The entities, connections, arrangements, and the like depicted and described in connection with the various figures are presented by way of example and not by way of limitation. Accordingly, any and all statements or other indications of what is "depicted" in a particular figure, what a particular element or entity "is" or "has" in a particular figure, and any and all similar statements (which may be understood in isolation and out of context as being absolute and thus limiting) may only be properly understood as being constructively preceded by clauses such as "In at least one embodiment,...". For the sake of brevity and clarity of presentation, this implicit introductory clause is not repetitively recited in the detailed description. Detailed Description

[0069] Extended Reality (XR) Display Device

[0070] Figure 1A is a schematic side view of an example waveguide display that can be used with an extended reality (XR) application according to some embodiments. An image is projected by an image generator 102. The image generator 102 can project the image using one or more of a variety of techniques. For example, the image generator 102 can be a laser beam scanning (LBS) projector, a liquid crystal display (LCD), a light emitting diode (LED) display (including an organic LED (OLED) or a micro LED (μLED) display), a digital light processor (DLP), a liquid crystal on silicon (LCoS) display, or other types of image generators or light engines.

[0071] Light representing the image 112 generated by the image generator 102 is coupled into the waveguide 104 through a diffraction in-coupler 106. The in-coupler 106 diffracts the light representing the image 112 into one or more diffraction orders. For example, one of the light rays 108, which represents a portion of the bottom of the image, is diffracted by the in-coupler 106, and one of the diffraction orders 110 (e.g., the second order) is at an angle that can propagate through the waveguide 104 by total internal reflection. The image generator 102 displays the image according to the instructions of a control module 124, which operates to render image data, video data, point cloud data, or other displayable data.

[0072] At least a portion of the light 110 that has been coupled into waveguide 104 by diffractive in-coupler 106 is coupled out of the waveguide by diffractive out-coupler 114. At least some of the light coupled out of waveguide 104 replicates the angle of incidence of the light coupled into the waveguide. For example, in the illustration, the out-coupled light rays 116a, 116b, and 116c replicate the angle of the in-coupled light ray 108. Because the light leaving the out-coupler replicates the direction of the light entering the in-coupler, the waveguide substantially replicates the original image 112. The user's eye 118 can focus on the replicated image.

[0073] In Figure 1A the example of, out-coupler 114 out-couples only a portion of the light, where each reflection allows a single input beam (such as beam 108) to generate multiple parallel output beams (such as beams 116a, 116b, and 116c). In this way, even if the user's eye is not perfectly aligned with the center of the out-coupler, at least some of the light from each part of the image may reach the user's eye. For example, if the eye 118 is moved downward, even if beams 116a and 116b do not enter the eye, beam 116c may enter the eye, so the user can still perceive the bottom of the image 112 despite the movement in position. Thus, out-coupler 114 operates in part as an exit pupil expander in the vertical direction. The waveguide may also include one or more additional exit pupil expanders ( Figure 1A not shown in ) to expand the exit pupil in the horizontal direction.

[0074] In some embodiments, waveguide 104 is at least partially transparent to light originating outside the waveguide display. For example, at least some of the light 120 from a real-world object (such as object 122) passes through waveguide 104, allowing the user to see the real-world object while using the waveguide display. Since the light 120 from the real-world object also passes through diffractive grating 114, there will be multiple diffraction orders and thus multiple images. To minimize the visibility of the multiple images, it is desirable for the zero-order diffraction (no deviation by 114) to have a large diffraction efficiency for light 120 and the zero-order to be large, while the higher diffraction orders are lower in energy. Thus, in addition to expanding and out-coupling the virtual image, out-coupler 114 is preferably configured to let the zero-order of the real image pass through. In such an embodiment, the image displayed by the waveguide display may appear to be superimposed on the real world.

[0075] Figure 1Bis a schematic side view illustrating an example alternative display type that can be used with extended reality applications. In XR head-mounted display device 130, control module 132 controls display 134 (which can be an LCD) to display an image. The head-mounted display includes a partially reflective surface 136 that reflects (and in some embodiments, both reflects and focuses) the image displayed on the LCD to make the image visible to the user. The partially reflective surface 136 also allows at least some external light to pass through, thus allowing the user to see their surrounding environment.

[0076] Figure 1C is a schematic side view illustrating an example alternative display type that can be used with extended reality applications. In XR head-mounted display device 140, control module 142 controls display 144 (which can be an LCD) to display an image. The image is focused by one or more lenses of display optics 146 to make the image visible to the user. In Figure 1C the example, external light does not directly reach the user's eyes. However, in some such embodiments, external camera 148 can be used to capture an image of the external environment and display such an image on display 144 together with any virtual content that can also be displayed.

[0077] The embodiments described herein are not limited to any particular type or structure of XR display device.

[0078] Figure 1D is a system diagram illustrating a set of example interfaces of a system according to some embodiments. A system such as Figure 1D can be used to implement an extended reality display device and its control electronics. System 150 can be implemented as a device including various components described below and configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 150 can be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 150 are distributed across multiple ICs and / or discrete components. In various embodiments, system 150 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input ports and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described in this document.

[0079] System 150 includes at least one processor 152 configured to execute instructions loaded therein for implementing aspects such as those described in this document. The processor 152 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 150 includes at least one memory 154 (e.g., volatile memory devices and / or non-volatile memory devices). System 150 may include a storage device 158, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, the storage device 158 may include internal storage devices, attached storage devices (including detachable and non-detachable storage devices), and / or network-accessible storage devices.

[0080] System 150 includes an encoder / decoder module 156 configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 156 may include its own processor and memory. The encoder / decoder module 156 represents (one or more) modules that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module 156 may be implemented as a separate element of System 150 or may be incorporated into the processor 152 as a combination of hardware and software known to those skilled in the art.

[0081] The program code to be loaded onto the processor 152 or the encoder / decoder 156 to execute the aspects described in this document may be stored in the storage device 158 and subsequently loaded onto the memory 154 for execution by the processor 152. According to various embodiments, one or more of the processor 152, the memory 154, the storage device 158, and the encoder / decoder module 156 may store one or more of the various items during the execution of the processes described in this document. Such stored items may include but are not limited to input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from processing equations, formulas, operations, and operation logic.

[0082] In some embodiments, the memory internal to the processor 152 and / or the encoder / decoder module 156 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be the processor 152 or the encoder / decoder module 152) is used for one or more of these functions. The external memory can be the memory 154 and / or the storage device 158, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard developed by the Joint Video Exploration Team JVET).

[0083] Input to the elements of the system 150 can be provided through various input devices as indicated in block 172. Such input devices include, but are not limited to: (i) an RF (radio frequency) section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) component (COMP) input terminals (or a set of COMP input terminals); (iii) universal serial bus (USB) input terminals; and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 1C Other examples not shown include composite video.

[0084] In various embodiments, the input device of block 172 has corresponding input processing elements associated therewith as known in the art. For example, the RF section can be associated with elements suitable for: (i) selecting a desired frequency (also known as selecting a signal or band-limiting a signal band to one band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower band to select a signal band that can be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) de-multiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements for performing these functions, such as frequency selectors, signal selectors, band limiters, channel selectors, filters, down-converters, demodulators, error correctors, and de-multiplexers. The RF section can include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or down-converting to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive an RF signal transmitted through a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and again filtering to a desired frequency band. Various embodiments re-arrange the order of the (and other) elements described above, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, such as, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0085] Additionally, the USB and / or HDMI terminals can include corresponding interface processors for connecting system 150 to other electronic devices across the USB and / or HDMI connections. It is to be understood that various aspects of input processing (e.g., Reed-Solomon error correction) can be implemented, as needed, for example, within a separate input processing IC or within processor 152. Similarly, various aspects of USB or HDMI interface processing can be implemented, as needed, within a separate interface IC or within processor 152. The streams of demodulation, error correction, and de-multiplexing are provided to various processing elements, including, for example, processor 152 and encoder / decoder 156, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.

[0086] The various elements of system 150 can be provided within an integrated housing in which the various elements can be interconnected using a suitable connection arrangement 174 (e.g., an internal bus known in the art, including an inter-integrated circuit (I2C) bus, wiring, and a printed circuit board) and data can be transmitted therebetween.

[0087] System 150 includes a communication interface 160 that enables communication with other devices via a communication channel 162. The communication interface 160 can include, but is not limited to, a transceiver configured to transmit and receive data over the communication channel 162. The communication interface 160 can include, but is not limited to, a modem or a network card, and the communication channel 162 can be implemented, for example, within a wired and / or wireless medium.

[0088] In various embodiments, a wireless network such as a Wi-Fi network (e.g., IEEE 802.11, where IEEE refers to the Institute of Electrical and Electronics Engineers) is used to stream or otherwise provide data to system 150. The Wi-Fi signals of these embodiments are received via the communication channel 162 and the communication interface 160 suitable for Wi-Fi communication. The communication channel 162 of these embodiments is typically connected to an access point or a router that provides access to an external network (including the Internet) to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streamed data to system 150, and the set-top box delivers the data via an HDMI connection of the input block 172. Still other embodiments use an RF connection of the input block 172 to provide streamed data to system 150. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0089] System 150 can provide output signals to various output devices, including a display 176, speakers 178, and other peripheral devices 180. The display 176 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 176 can be used for a television, a tablet computer, a laptop computer, a cellular phone (mobile phone), or other devices. The display 176 can also be integrated with other components (e.g., as in a smart phone) or be separate (e.g., an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 180 include one or more of a standalone digital video disc (or digital versatile disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 180 that provide functions based on the output of system 150. For example, a disc player performs the function of playing the output of system 150.

[0090] In various embodiments, control signals are transmitted between system 150 and display 176, speaker 178, or other peripheral devices 180 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 1000 via dedicated connections through corresponding interfaces 164, 166, and 168. Alternatively, the output devices can be connected to system 150 using communication channel 162 via communication interface 160. In an electronic device such as a television, for example, display 176 and speaker 178 can be integrated with other components of system 150 in a single unit. In various embodiments, display interface 164 includes a display driver such as, for example, a timing controller (T Con) chip.

[0091] For example, if the RF portion of input 172 is part of a separate set-top box, display 176 and speaker 178 can alternatively be separate from one or more of the other components. In various embodiments in which display 176 and speaker 178 are external components, output signals can be provided via dedicated output connections including, for example, an HDMI port, a USB port, or a COMP output.

[0092] System 150 can include one or more sensor devices 168. Examples of sensor devices that can be used include one or more GPS sensors, gyroscopic sensors, accelerometers, light sensors, cameras, depth cameras, microphones, and / or magnetometers. Such sensors can be used to determine information such as the location and orientation of a user. In the case where system 150 is used as a control module for an extended reality display such as control modules 124, 132, the location and orientation of the user can be used when determining how to render image data so that the user perceives the correct part of a virtual object or virtual scene from the correct perspective. In the case of a head-mounted display device, for the purpose of rendering virtual content, the location and orientation of the device itself can be used to determine the location and orientation of the user. In the case of other display devices such as a phone, a tablet computer, a computer monitor, or a television, other inputs can be used to determine the location and orientation of the user for the purpose of rendering content. For example, a user can use a touch screen, a keypad or keyboard, a trackball, a joystick, or other inputs to select and / or adjust a desired viewpoint and / or viewing direction. In the case where a display device has sensors such as an accelerometer and / or a gyroscope, the viewpoint and orientation for the purpose of rendering content can be selected and / or adjusted based on the movement of the display device.

[0093] The embodiments can be executed by computer software implemented by the processor 152, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. As a non-limiting example, the memory 154 can be of any type suitable for the technical environment and can be implemented using any appropriate data storage technology (such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory). As a non-limiting example, the processor 152 can be of any type suitable for the technical environment and can include one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture).

[0094] Scene description framework for XR

[0095] This principle generally relates to the field of rendering for extended reality scene description and extended reality rendering. When rendering on an end-user device such as a mobile device or a head-mounted display (HMD), this document is also understood in the context of formatting and playing extended reality applications.

[0096] Extended reality (XR) is a technology enabling interactive experiences where the real-world environment and / or video content are enhanced with virtual content, which can be defined across multiple sensory modalities including vision, audition, touch, etc. During the runtime of an application, the virtual content (e.g., 3D content or audio / video files) is rendered in real time in a manner consistent with the user context (environment, viewpoint, device, etc.). A scene graph (such as, for example, the scene graph proposed by Khronos / glTF and its extensions defined in the MPEG scene description format or Apple / USDZ) is a possible way to represent the content to be rendered. They combine, on the one hand, a declarative description of the scene structure linking real-world environment objects and virtual objects, and on the other hand, a binary representation of the virtual content.

[0097] Although such an MPEG scene description framework ensures that the timed media and the corresponding associated virtual content are available at any time during the rendering of the application, there is no description of how a user can interact with the scene objects at runtime to obtain an immersive XR experience.

[0098] There is a lack of an XR system that can obtain an XR scene description, which includes metadata describing how a user can interact with the scene objects at runtime and how these interactions can be updated during the runtime of an XR application.

[0099] In an XR application, the scene description is used to combine an explicit and easily parsable description of the scene structure with some binary representation of the media content.

[0100] In a time-based media stream, the scene description itself can evolve over time to provide relevant virtual content for each sequence of the media stream. For example, for advertising purposes, a virtual bottle can be displayed during a video sequence where people are drinking.

[0101] This behavior can be achieved by relying on the framework defined in the scene description of an MPEG media document, which is Information technology–Coded representation of immersive media–Part14:SceneDescription for MPEG media, ISO / IEC DIS23090-14:2021(E). A scene update mechanism based on the JavaScript Object Notation (JSON) patch protocol defined in IETF RFC 6902 can be used to synchronize the virtual content with the MPEG media stream.

[0102] Figure 1E FIG. is a system diagram illustrating a set of example interfaces of a scene description (stored as an item in glTF.json), three video tracks, one audio track, and a JSON patch update track in an ISOBMFF (International Organization for Standardization (ISO) Base Media File Format) file according to some embodiments. The example ISOBMFF file structure 182 shows that the glTF.json scene description item 184 is connected to the glTF storage buffer binary 186 and the JSON patch update track 196. The example glTF.json scene description item 184 is also connected to three example video tracks 188, 190, 192 and the audio track 194. There are three update patches 198 in the JSON update track 196.

[0103] Although the MPEG-I scene description framework ensures that the timed media and the corresponding relevant virtual content are available at any time, it does not provide a description of how users can interact with the scene objects at runtime to obtain an immersive XR experience. Therefore, user-specific XR experiences are not supported for using immersive media.

[0104] Example embodiments described herein can be used to provide a scene description including virtual objects or light sources, but even if available, they do not necessarily display or render the virtual objects or light sources. In some embodiments, one or more of the following aspects can be considered when determining whether to display a virtual object or light source.

[0105] When determining whether to display a virtual object or a light source, spatial aspects can be considered. For example, if the user environment is not suitable (e.g., the user is too far away from the location of the rendered timed media), or if the user is not looking in the correct direction, or if the virtual object is supposed to be displayed on a user-specific area (e.g., above his undetected left hand), then the virtual object or the light source may not be displayed.

[0106] When determining whether to display a virtual object or a light source, temporal aspects can be considered. For example, if the user is not ready or wants to trigger the display of the object himself (e.g., using a specific gesture), then the virtual object or the light source may not be displayed until an appropriate trigger is detected.

[0107] In some embodiments, it is specified in the scene description which objects or light sources the user is allowed to manipulate or with which objects or light sources to interact via potential haptic feedback.

[0108] Runtime interactivity

[0109] Figure 2 is a system diagram showing a set of example interfaces of an MPEG-I node hierarchy that supports scene interactivity according to some embodiments. The example node hierarchy 200 shows some example items that may be in the scene. According to this principle, in addition to the node tree, behavior metadata items (examples of what are referred to herein as "behaviors") are added to the scene description. In an example embodiment, the scene description that evolves over time is enhanced by adding information that identifies the behaviors. These behaviors may be associated with predefined virtual objects on which runtime interactivity is allowed for a user-specific XR experience.

[0110] In some embodiments, these behaviors evolve over time. In such an embodiment, the behaviors can be updated through the already existing scene description update mechanism.

[0111] In an example embodiment, a behavior is characterized by one or more of the following attributes:

[0112] · One or more triggers that define the conditions to be met for activation.

[0113] · Trigger control parameters that define the logical operations between the defined triggers.

[0114] · Actions implemented in response to the activation of a trigger.

[0115] · Action control parameters that define the order of execution of the defined actions.

[0116] · A priority number that enables the selection of the highest priority behavior in case several behaviors occur simultaneously on the same virtual object.

[0117] · Optional interruption actions that specify how to terminate the behavior when it is no longer defined in the newly received scene update. For example, if the related object has been removed, or if the behavior is no longer relevant to the current media (e.g., audio or video) sequence, the behavior is no longer defined.

[0118] By adding these behaviors, time-related user interactivity in immersive content for XR experiences can be defined.

[0119] When a second scene description is received, some of the behaviors in the first scene description may be "in progress", e.g., they are triggered and their actions are running. The second scene description can be provided as updated metadata (e.g., metadata describing the differences between the first and second scene descriptions). The second scene description includes a node tree that describes objects that may be the same or different from the objects in the first scene description. The objects in the node tree of the first scene description may no longer exist in the second description. If the objects related to the running actions of the in-progress behaviors are missing in the second scene description, these in-progress behaviors are no longer applicable. Similarly, if the in-progress behaviors are not defined in the second description, the in-progress behaviors are no longer applicable. The interruption action field describes how to correctly interrupt the running actions of the in-progress behaviors.

[0120] Figure 3 FIG. is a block diagram showing an example of the logical relationship between trigger information (describing triggers 1 to n), action information (describing actions 1 to m), and behavior information (describing the relationship between triggers and actions) according to some embodiments, where the triggers and actions can refer to one or more nodes in a scene description (such as a hierarchical scene graph). For example, trigger 1 (308) directly references node 1 (310) and node 2 (312) (where node 31 (314) is a child node of node 1), and trigger 1 directly references node 1. The example set 300 of logical relationships can include behavior information 302, trigger information 304, and action information 306.

[0121] In an XR application, the scene description is used to combine a clear and easily parsable description of the scene structure with some binary representation of the media content. The above sections describe the action mechanism of the scene description. These behaviors are related to predefined virtual objects, and interactivity is allowed when running on these virtual objects for user-specific XR experiences. Figure 3 Illustrates the structure of an example behavior mechanism.

[0122] Figure 4It is a schematic plan view illustrating an example relationship of extended reality scene description objects according to some embodiments. In this example, the scene graph 400 includes a description of a real object 412, such as "flat horizontal surface" (which can be a table or a floor or a plate), and a description of a virtual object 414, such as an animation of a walking character. The scene graph node 414 is associated with a media content item 416, which is an encoding of data for rendering and displaying the walking character (such as a textured animated 3D mesh). The scene graph 400 also includes a node 410, which is a description of the spatial relationship between the real object described in node 412 and the virtual object described in node 414. In this example, node 410 describes the spatial relationship that enables the character to walk on the flat surface. When the XR application is started, the media content item 416 is loaded, rendered, and buffered for display when triggered. When a flat surface is detected by a sensor (or a camera for some embodiments) in the real environment, the application displays the buffered media content item as described in node 410. The timing is managed by the application according to the features detected in the real environment and the timing of the animation. The nodes of the scene graph may also not include a description and only serve as the parent node of child nodes.

[0123] XR applications are diverse and can be applied to different contexts and real or virtual environments. For example, in an industrial XR application, when a reference object (part B of the engine) is detected by a camera mounted on a head-mounted display device in the real environment, a virtual 3D content item (e.g., part A of the engine) is displayed. The 3D content item is positioned in the real world at a position and scale defined relative to the detected reference object.

[0124] For example, in an XR application for interior design, when a given image from a catalog is detected in the input camera view, a 3D model of furniture is displayed. The 3D content is positioned in the real world at a position and scale defined relative to the detected reference image. In another application, when the user enters an area near a church (real or virtual rendered in an extended real environment), some audio files may start playing. In another example, when the user sees a given can of soda in the real environment, an advertising jingle file may be played. In an outdoor game application, various virtual characters may appear according to the semantics of the scene observed by the user. For example, a bird character is suitable for a tree, so if the sensor of the XR device detects a real object described by the semantic label "tree", birds flying around the tree can be added. In an accompanying application implemented by smart glasses, when a car is detected within the field of view of the user's camera, a car noise can be emitted in the user's headphones to warn him of potential danger; in addition, the sound can be spatialized so that it arrives from the direction where the car is detected.

[0125] XR applications can also enhance video content rather than the real environment. The video is displayed on a rendering device, and when a timing event is detected in the video, the virtual objects described in the node tree are overlaid. In such an example context, the node tree only includes virtual object descriptions.

[0126] Referring to the scope description of an example embodiment of the MPEG-I scene description framework that uses the Khronos glTF (Graphics Language Transmission Format or GL Transmission Format) extension mechanism, which supports additional scene description features such as a node tree. However, the principles described herein are not limited to a specific scene description framework.

[0127] According to some example embodiments, the glTF scene description is extended to support interactivity. The interactivity extension is applied at the glTF scene level and is called MPEG_scene_interactivity.

[0128] The corresponding semantics are provided in Table 1.

[0129]

[0130] Table 1: Semantics of the example MPEG_Scene_Interactivity extension

[0131] In Table 1 and other semantic tables described herein, the "Usage" column indicates "M" for "mandatory" features and "O" for "optional" features. However, such features can be "mandatory" or "optional" only according to a specific proposed syntax. Features marked as "mandatory" are not necessarily required features for implementing an application. For example, in some embodiments, there are features marked as "mandatory" to meet the expectations of a specific type of parsing and rendering software; however, in other embodiments, the feature can be optional, or the feature can be completely omitted, where the corresponding functionality is implemented using default values or not implemented at all without departing from the scope of the present disclosure.

[0132] Figure 5It is a data syntax diagram that illustrates an example syntax of a data stream encoding an extended reality (XR) scene description according to some embodiments. This structure exists in a container that organizes the stream into elements of independent syntax. Structure 500 may include a header portion 502, which is a set of data common to each syntax element of the stream. For example, the header portion includes metadata about the syntax elements that describes the nature and role of each syntax element. The structure also includes a payload that includes a syntax element 504 and a syntax element 506. The syntax element 504 includes data representing a media content item described in a node of a scene graph related to a virtual element. Images, meshes, and other raw data may have been compressed according to a compression method. Element 506 is part of the payload of the data stream and includes data encoding the scene description.

[0133] In an XR application, the scene description can be used to combine a clear and easily parsable description of the scene structure with some binary representation of the media content. In a time-based media stream, the scene description itself can evolve over time to provide relevant virtual content for each sequence of the media stream. For example, for advertising purposes, a virtual bottle can be displayed on a table during a video sequence in which people are sitting around the table. Such behavior can be achieved by using the framework of INFORMATIONTECHNOLOGY–CODED REPRESENTATION OF IMMERSIVE MEDIA–PART 14:SCENE DESCRIPTIONFOR MPEG MEDIA,ISO / IEC DIS23090-14:2021(E)(2021)(“MPEG-I scene description”). Although the MPEG-I scene description framework allows timed media and the corresponding relevant virtual content to be available at any time, it does not describe how to update the scene based on runtime interactivity.

[0134] In the MPEG-I scene description document, a scene update mechanism is proposed. A special track in a media content file (e.g., an ISOBMFF (International Organization for Standardization (ISO) Base Media File Format) file) can provide timed patches to be applied to the original scene description. Such patches can be JSON (JavaScript Object Notation) patches, as available at datatracker <dot>ietf <dot>As described in the JavaScript Object Notation (JSON) Patch ("JSON Patch") of IETF RFC 6902 obtained from org / doc / html / rfc6902. This mechanism handles predefined scenario evolutions, but does not allow the description of event-based updates (e.g., after user actions or any event that may occur in the scenario object at any time).

[0135] Figure 6 is a data structure that illustrates example JSON Patch operations. As discussed in the JSON Patch document, a patch contains an array of operations (add, remove, replace, move, and copy). Each operation specifies a type ("op"); the JSON element to address ("path"); and (if needed) the value to apply ("value") or the source JSON element ("from"). Figure 6 Code listing 600 with example JSON Patches is provided, which includes two example JSON Patch operations. The first example operation is an "add" operation that adds a new node at the end of the glTF (Graphics Language Transmission Format) node array. The parameters of the new node according to the first example operation are from a JSON string ("value") representing the new node and have a "mesh" value of zero, which references the first element of the glTF mesh array, a "name" of "Sphere Yellow", and a "translation" value of [-3, 0, 0]. The second example operation is a "copy" operation that copies the existing node 0 (thus the source JSON element ("from")) to the end of the glTF node array.

[0136] Figure 7 is a flowchart that illustrates an example timed sample patch to be applied to a scenario description file. Non-timed items (which in this case are the glTF JSON document 728) are loaded as version 1 of the scenario description (SD(V1)). According to example 700, a series of JSON patches 716, 718, 720, 722, 724, 726 are applied to successive versions 702, 704, 706, 708, 710, 712, 714 of the scenario description. The dashed arrows indicate that the corresponding patches point to the previous scenario description, while the solid arrows indicate the resulting updated scenario description output as the new version. In other words, for from Figure 7 An example includes a JSON Patch 722, which is a JSON Patch removal operation summarized as "DelB" or "Delete B", being applied to a scenario description version 4 ("SD(V4)") 708 to remove the element "B" from the set of elements {a, b, c, A, B, C}, resulting in a scenario description version 5 ("SD(V5)") 710 having a reduced set of elements {a, b, c, A, C}.

[0137] More specifically, for Figure 7 the additional example shown in , the patch is a JSON Patch of the "application / json-patch+json" media type. In this example, version 1 (702) of the scenario has the elements "a", "b", and "c". The first patch 716 adds the element "A" so that version 2 (704) of the scenario elements is "a", "b", "c", and "A". The second patch 718 adds the element "B" so that version 3 (706) of the scenario elements is "a", "b", "c", "A", and "B". The third patch 720 adds the element "C" so that version 4 (708) of the scenario elements is "a", "b", "c", "A", "B", and "C". The fourth patch 722 deletes the element "B" leaving version 5 (710) of the scenario with the elements "a", "b", "c", "A", and "C". The fifth patch 724 deletes the element "a" leaving version 6 (712) of the scenario with the elements "b", "c", "A", and "C". The sixth patch 726 deletes the elements "a" and "b" leaving version 7 (714) of the scenario with the elements "c", "A", and "C". The element "a" has already been deleted, so attempting to delete the element "a" again has no effect.

[0138] Figure 8 It is a flowchart illustrating additional example timing sample patches to be applied to a scene description file. In the document EXPLORATION EXPERIMENTS ON DYNAMIC SCENE UPDATE, ISO / IEC JTC 1 / SC29 / WG 3 N0315 (August 2021) ("MPEG-I Scene Description Update"), an event-based scene update mechanism was proposed. When a predefined timed scene update is in progress, as described above, events that trigger additional updates to the scene description may occur (e.g., user input, a collision between two objects, or an object becoming visible). A new scene description version (SD version 3A) is created via patch A 812, which modifies the predefined timed scene update stream of patches 802, 804, 806, 808, 810. Then several scenarios were proposed, such as: (1) Scenario 1 (820): Apply the patch and switch to a new timed sample track (SD version 3A, SD version 5A,...); (2) Scenario 2 (822): Apply the patch and switch back to the previous version of the scene description and continue the same predefined timed scene update process; or (3) Scenario 3 (824): Apply the patch and skip one or more versions in the same track. Figure 8 A set 800 of examples of these scenarios is shown, which are labeled as Scenario 1, 2, and 3. For some embodiments, the second set of timed scene updates may have a series of patches 814, 816, 818.

[0139] The described event-based scene update mechanism is still closely related to the predefined scene evolution and does not specify how to describe the events that trigger updates in the scene description document. In addition, the mechanism does not handle the case where the same event that creates a new node may be fired multiple times. When a new action occurs, the number of times the same patch has been applied is unknown.

[0140] For example, creating a new node as a child of an existing node can be done using 2 patch operations. For the new node, a new JSON element is added at the end of the glTF node array (op = "add" and path = " / nodes / -"). Then, the child array of the parent node is modified by adding the index of the new node (op = "add" and path = " / nodes / 0 / children / -"). The challenge is that there is no way to know the index of the new node (which will be used as the value parameter) except for the first execution of the patch. This challenge is why additional parameters such as placedItems, chilOps, or parents may be used in the update operation.

[0141] Figure 9 The figure shows a flowchart of an example process for event-based updates according to some embodiments. In this example 900, the glTF scene (Vx) contains a description of the event-based update mechanism, where the same patch is applied each time an event occurs (such as a user clicks on a touchscreen). Some elements of the glTF scene are modified, but the event-based update description is not modified. For example, each time the user clicks on the touchscreen, a new object (a cube) is added to the scene. Version 0 of the glTF scene is V x0 902, as shown at the top of Figure 9 For the first event, the loop counter i is initialized 904 to 1, and the first JSON patch is applied to generate version 1 of the scene, i.e., V x1 . For each further event, the loop counter i is incremented, and the i-th JSON patch is applied 908 to generate version i of the scene, i.e., V xi 906.

[0142] European Patent Application No. 22305024.6 ("the '024 application") discusses enhancing the description of a scene evolving over time by adding "behavior" data. This enhancement is being integrated into a draft revision of the MPEG-I scene description document. These behaviors are associated with predefined virtual objects, and interactivity is allowed when running on these virtual objects for a user-specific XR experience. These behaviors also evolve over time and can be updated through the existing scene description update mechanism.

[0143] For some embodiments, a behavior is a metadata item that can include:

[0144] · A trigger, which defines the conditions for its activation to be satisfied

[0145] · Trigger control parameters, which define the logical operations between triggers

[0146] · An action to be processed when the trigger is activated

[0147] · Action control parameters, which define the order of execution of the defined actions

[0148] · A priority number, which enables the selection of the highest-priority behavior in the case where several behaviors occur simultaneously on the same virtual object

[0149] · An optional interrupt action, used to specify how to terminate the behavior when the behavior is no longer defined in the newly received scene update

[0150] For example, if the relevant object does not belong to the new scene, or if the behavior is no longer relevant to the current media (e.g., audio or video) sequence, the behavior is no longer defined.

[0151] The semantics of the actions are shown in Table 2.

[0152]

[0153]

[0154] Table 2: Action Semantics

[0155] The framework of Table 2 can be extended to allow other action types. This application discusses new action types for handling event-based scene updates according to some example embodiments. These new action types can be added to the framework of Table 2.

[0156] In addition, the scene update mechanism described above according to the MPEG-I scene description update only considers the private use of scenes with only local updates. This application also discusses a new parameter that allows sharing of updates with all clients that may be connected to the same scene. Figure 8

[0157] This application uses the MPEG-I scene description framework and interactivity framework described in the '024 application. According to some embodiments, the scene description can be enhanced by specializing the action semantics based on the interactivity framework of the '024 application to add two new action categories: set actions and update actions.

[0158] Set actions update the glTF elements in the scene description document. This element can be addressed using a JSON pointer, which is available in the datatracker <dot>ietf <dot>The document obtained from org / doc / html / rfc6901, the JAVASCRIPT OBJECT NOTIFICATION (JSON) POINTER, is described in IETF RFC 6901 ("JSON Pointer"). This action can replace actions that only update a single element of the glTF tree (e.g., the SET_MATERIAL action).

[0159] Although multiple set actions can be used to perform advanced updates in a scene, in some implementations, such an update process using multiple set actions can become cumbersome for the creation or removal of new nodes. The update action performs more complex updates based on the JSON patches used in the document MPEG-I scene description update. The update action contains a scene update patch, which contains an array of operations (add, remove, replace, move, or copy). A single replace operation can be performed by a set action. Each patch operation specifies its type ("op"), the JSON element to address ("path"), and (if required or applicable) the value to apply ("value") or the source JSON element ("from"). Figure 10 The example given in replaces (or sets) the scale value of node 10 to [2, 2, 2].

[0160] For both set and update actions, a location parameter (placeDescription) is specified to obtain the location from a user input device or function (such as, for example, the hand position or finger position on a touch screen). This location can be used to calculate the new value of a set action or the location of a new node (or node tree) added by an update action. In addition, for both set and update actions, a parameter (shared) is specified that allows the update to be applied to all devices that are connected and sharing the same 3D scene in the same session.

[0161] Figure 10 is a data structure that illustrates an example JSON patch operation according to some embodiments. For this example code list 1000, the operation type ("op") is a "replace" operation. The path ("path") parameter is a JSON pointer string. The path parameter points to the location of the node. The value ("value") parameter is the new value. In this example, the "scale" of node "10" is replaced with the value [2, 2, 2].

[0162] Table 3 below gives the semantics of the set action and the definitions of the parameters of the new action.

[0163]

[0164]

[0165] Table 3: Action Semantics with Set Actions

[0166] In the action structure shown in Table 3, the SET_MATERIAL action type is removed and the set action type is added. Elements of the set action structure are also added as shown in Table 3.

[0167] The path parameter indicates an element in the scene description glTF tree (e.g., a scene element). The scene description glTF tree is a JSON tree, and the JSON pointer semantics described in the JSON pointer document can be used with the path parameter.

[0168] The pathType parameter depends on the target element. For example, the pathType parameter can be the type of a leaf element: string, float, byte, short, boolean, or array.

[0169] If the element is a subtree (e.g., " / node / 1"), the value parameter is a JSON string to describe the glTF subtree.

[0170] To handle the use case where multiple connected clients share the same 3D scene, the shared parameter is introduced. When the shared parameter is set to true, the action is executed in all clients, so the scene rendering is synchronized everywhere. The action can be executed by forwarding the activation state, such that all clients trigger the action execution as if the trigger had occurred locally for each in the client. The action can also be executed by forwarding the scene update information to be applied by each client.

[0171] In the case where the set action manipulates a glTF node to be placed at a specific location (e.g., a location within the scene) related to user input, the placeDescription parameter is used. In that case, the path is the node root or the location element of the node. If the path is the node root (e.g., " / nodes / 1"), the translation and rotation values of the node are set according to the placeDescription value. If the path is the location element of the node (e.g., " / node / 2 / translation" or " / node / 2 / rotation"), only that element is set.

[0172] For some embodiments, the placeDescription value is a string parameter that contains a description of the user input, e.g., as available at www <dot>khronos <dot>org / registry / OpenXR / specs / 1.0 / html / xrspec <dot>The " / user / hand / left / grip" or " / user / hand / left / input / trackpad / touch" specified in THE OPENXR SPECIFICATION obtained from html#semantic - path - interaction - profiles is used for the interaction profile path.

[0173] This path is used by the 3D engine to provide position information, such as the translation and orientation of user elements or the 2D position of the user's finger touching the touchscreen. This information can be directly used to set the positioning elements of the glTF file, or additional operations (e.g., 2D - to - 3D conversion, or conversion from 3D XR space to another space) may be required.

[0174] Since placeDescription can refer to an absolute 3D position, in some embodiments, it should be positioned relative to the parent node (if any). Additional parameters can be added to specify how the placement should be handled: directly to the new position without any transition, linearly transition to the new position...

[0175] Figure 11 is a data structure that illustrates an example glTF file excerpt according to some embodiments. Figure 11 The example code list 1100 of the example glTF file excerpt in shows the example structure of the parent node and the child node. The parent node has a "name" parameter with the value "Cube". The "scale" parameter of the parent node is an array with the value [5, 0.05000000074505806, 5].

[0176] The child node has a "name" parameter with the value "Light". The "rotation" parameter of the child node is an array with the value [0.16907575726509094, 0.7558803558349609, - 0.27217137813568115, 0.570947527885437]. The "translation" parameter of the child node is an array with the value [4.076245307922363, 5.903861999511719, - 1.0054539442062378].

[0177] Figure 12 is a data structure that illustrates an example setting action structure according to some embodiments. For this example code list 1200, the proximity trigger activation setting action with type "TRIGGER_TYPE_PROXIMITY" and value 1 is shown in Figure 12 and has a value of 1 in. Figure 12 In it, the set action has type "ACTION_TYPE_SET" and value 7. When the user / camera approaches node 0 ("[0]") (a distance between 10m (distanceLowerLimit) and 20m (distanceUpperLimit)), the translation value of node 1 is set to the position pointed to by the user (the position of the user's finger...), because the "path" parameter is " / nodes / 1 / translation". To address the translation element of node 1, the path value is set to: " / nodes / 1 / translation". For Figure 12 the example shown in

[0178] Figure 13 is a flowchart illustrating an example set action process according to some embodiments. In the example set action process 1300, if a trigger is detected 1304 as fired, the flow advances to the start action box 1306. Otherwise, the flow returns to rendering 1302 the scene.

[0179] At the start action box 1306, the parameters are checked. Does the path exist in the glTF description file 1308? Do the pathType and value parameters match the glTF element pointed to? If the path does not exist or is invalid, the flow advances back to rendering the scene.

[0180] If the placeDescription parameter exists 1310, the flowchart advances to asking if the path is compatible 1312. If the path is not compatible, the flowchart advances to return to rendering the scene. If the path is compatible, the flowchart advances to obtain the place value 1314, and then the flow advances to the set path box 1316.

[0181] If the placeDescription parameter does not exist, the flow advances to the set path box 1316.

[0182] At the set path box, the glTF element identified by the path parameter is set with the value parameter, or is set with the location information retrieved from the placeDescription parameter ("obtain the place value" step). In the latter case, user input (click, touch) or a user body element (e.g., the left hand or finger) can be monitored to detect a trigger.

[0183] Table 4 below gives the semantics of the update action and the definitions of the parameters of the new action.

[0184]

[0185]

[0186] Table 4: Action semantics with update action #

[0187] The patch parameter describes the patch to be applied to the gtTF scene. Following the JSON patch semantics of the JSON Patch document, the patch parameter is a string that describes one or more operations (add, delete, replace, move, or copy) to be applied to a JSON element, addressed by a JSON pointer. Each operation contains a set of parameters: an operation identifier ("op"), a path to the JSON element ("path"), and a value ("value") or source JSON element ("from") depending on the operation.

[0188] The shared and placeDescription parameters are the same as in the set action, except that the placeDescription parameter requires additional parameters to identify the nodes: the placedItems array contains the indices of the patch operations in the patch array. These operations identify the nodes to be positioned (if the path of these operations is the node root (e.g., " / nodes / 1") or the position element of the node (e.g., " / node / 2 / translation")). If one or more glTF nodes are created using one or more "add" or "copy" patch operations, these nodes can be added as children of existing nodes.

[0189] One or more scene elements can be placed at a given location. The patch operations in the placedItems list are the operations that handle scene elements that can be moved (e.g., a node representing a 3D object can be moved). For some embodiments, the update action information can further include a position description field and a placedItems field. The position description field can identify the source of the position information, while the placedItems field can identify the patch operations that reference the nodes to be placed at the given location.

[0190] Two array parameters identify the child nodes and their respective parent nodes: the childOps and parents arrays. The childOps array gives the indices of the operations in the patch array that create the child nodes. The parent nodes of these child nodes are given in the parents array. The parents array contains only the indices of nodes that already exist in the original glTF document.

[0191] For some embodiments, the processing model for managing these parameters by a rendering application is as follows:

[0192] · The processor executes the patch operations, which results in the first modification of the glTF file.

[0193] · The processor obtains the index of the new node.

[0194] · The processor modifies the child array of the parent node, which results in the second modification of the glTF file.

[0195] Figure 14A - 14B A data structure is formed that illustrates an example update action structure according to some embodiments. For some embodiments, Figure 14A - 14B Code lists 1400 and 1450 are shown, which form an example excerpt of a (complete) glTF scene description file, and the descriptions of all objects included in the scene are not shown here. For this example, the glTF description includes a series of glTF nodes with at least 13 items (index 0 to 12). Figure 14A - 14B An example of this new update action is given. For Figure 14A - 14B the example, a proximity trigger activation update action is shown with the type of "TRIGGER_TYPE_PROXIMITY" and a value of 1. In Figure 14A , the update action has the type "ACTION_TYPE_UPDATE" and a value of 8. In this example, the patch is an array of 4 operations. The first operation (index 0 of the patch array) adds a new node (a cube) at the end of the glTF node array. The second operation (index 1 of the patch array) changes the "scale" element of the first node (index 10 in the glTF node array). The third operation (index 2 of the patch array) adds a new "translation" element to the second node (index 12 in the glTF node array). The last operation (index 3 of the patch array) changes the "translation" element of the third node (index 11 in the glTF node array).

[0196] The node referenced by the first operation (index 0 of the patch array) is the last node in the glTF node array (" / nodes / -"). In this case, the last node added to the end of the glTF node array is index 13.

[0197] The childOps array contains the index "[0]". Thus, index 0 of the patch array creates a child node. Additionally, the placedItems array contains the index "[0]". Thus, the operation at index 0 of the patch array (the first "add" operation) uses the placeDescription string as the value parameter. In this case, placeDescription has the string value " / input / aim / pose". Thus, the node is placed at the location pointed to by the user (which is the location of the user's finger). The parents array contains the index "1", so the node created by the first patch operation (index 0) is created as a child node of node 1. Thus, a cube is created as a new child node of node 1.

[0198] The second operation (index [1] of the patch array) replaces the scale element of node 10. Thus, the scale of node 10 is replaced with the value [2, 2, 2].

[0199] The third operation (index [2] of the patch array) adds a translation to node 12. Thus, a new "translation" element with the value [-40, -2.1, -10] is added to node 12.

[0200] The fourth operation (index [3] of the patch array) replaces the translation of node 11. Thus, the translation of node 11 is replaced with the value [20, -2, -10].

[0201] Figure 15 is a flowchart illustrating an example update action process according to some embodiments. In the example update action process 1500, if a trigger is detected 1504 as fired, the flow advances to the start action box 1506. Otherwise, the flow returns to rendering 1502 the scene.

[0202] At the start action box 1506, the parameters are checked for validity. The loop index i is initialized to 0. The next operation in the patch array at 1508 is executed for index i. If the childOps array contains the current index i in the list of indices, the node referenced in the operation is set 1510 as a child node of the node specified in the "parents" array. If the placedItems array contains the current index i in the list of indices, the location of the node referenced in the operation is set 1512 according to the placeDescription value.

[0203] If there are more operations to execute 1514, the flow returns to the top of the loop to increment index i. Otherwise, the flow returns to render the scene.

[0204] Figure 16 is a flowchart illustrating an example process for performing a set action according to some embodiments. For some embodiments, the example process 1600 may include obtaining 1602 scene description data of a 3D scene, where the scene description data includes: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, and action information describing at least one action to be performed on a first scene element among the plurality of scene elements, where the action information includes: a list of at least the set action, and set action information, where the set action information includes: a path that identifies the first scene element; and a value that indicates a new value for the first scene element. For some embodiments, the example process 1600 may further include: in response to determining 1604: (i) that a trigger condition of at least one trigger condition has occurred, and (ii) that the action information indicates that a set action is to be performed, performing the set action to set the value associated with the first scene element to the new value of the first scene element.

[0205] Figure 17 is a flowchart illustrating an example process for performing an update action according to some embodiments. For some embodiments, the example process 1700 may include obtaining 1702 scene description data of a 3D scene, where the scene description data includes: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, and action information describing at least one action to be performed on a first scene element among the plurality of scene elements, where the action information includes: a list of at least the update action, and update action information, where the update action information includes: patch parameters. For some embodiments, the example process 1700 may further include: in response to determining 1704: (i) that a trigger condition in at least one trigger condition has occurred, and (ii) that the action information indicates that an update action is to be performed, performing the update action to apply the patch parameters as an update to the scene description data.

[0206] Although the methods and systems according to some embodiments are discussed in the context of virtual reality (VR), some embodiments may also be applied to the mixed reality (MR) / augmented reality (AR) context. Additionally, although the term "head-mounted display (HMD)" is used herein according to some embodiments, for some embodiments, some embodiments may be applied to, for example, wearable devices with VR, AR, and / or MR capabilities (which may or may not be attached to the head).

[0207] Example methods according to some embodiments may include: obtaining scene description data of a three-dimensional (3D) scene, where the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, and action information describing at least one action to be performed on a first scene element among the plurality of scene elements, where the action information may include: a list of at least set actions, and set action information, where the set action information may include: a path that identifies the first scene element; and a value that indicates a new value of the first scene element; and in response to determining: (i) that a trigger condition of at least one trigger condition has occurred, and (ii) that the action information indicates that a set action is to be performed, performing a first action among the at least one action on the first scene element.

[0208] For some embodiments of the example method, determining that a trigger condition has occurred may include a logical determination that a trigger condition has occurred.

[0209] For some embodiments of the example method, determining that a trigger condition has occurred may include determining that a trigger condition has produced a true result.

[0210] For some embodiments of the example method, the path may identify the position of the first scene element in the scene tree.

[0211] For some embodiments of the example method, the first action may include setting the value associated with the first scene element to the new value of the first scene element.

[0212] For some embodiments of the example method, setting the value associated with the first scene element sets the value associated with the first scene element to a new value.

[0213] For some embodiments of the example method, the set action information may further include a shared field that indicates a state shared with a device connected to the 3D scene.

[0214] Some embodiments of the example method may further include: in response to determining that the shared field indicates a true result, performing a first action among the at least one action on at least one device sharing the 3D scene.

[0215] Some embodiments of the example method may further include: in response to determining that the shared field indicates a true result, performing a first action among the at least one action on all devices sharing the 3D scene.

[0216] For some embodiments of the example method, the set action information may further include a location description field, and the location description field may identify the source of the location information.

[0217] For some embodiments of the example method, the first action may include setting the position of a first scene element within a scene, and setting the position of the first scene element may use position information from an identified source.

[0218] For some embodiments of the example method, the path may be a JavaScript Object Notation (JSON) pointer that references the first scene element.

[0219] For some embodiments of the example method, the path may include path information that references the first scene element, and the path information may correspond to Internet Engineering Task Force (IETF) standards.

[0220] For some embodiments of the example method, the scene description data may further include behavior information.

[0221] An example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0222] An additional example method according to some embodiments may include: obtaining scene description data of a 3D scene, where the scene description data may include: scene element information that describes each of a plurality of scene elements in the scene, trigger information that describes at least one trigger condition, and action information that describes at least one action to be performed on a first scene element among the plurality of scene elements, where the action information may include: a list of at least update actions, and update action information, where the update action information may include: patch parameters; and in response to determining: (i) that a trigger condition among the at least one trigger condition has occurred, and (ii) that the action information indicates that an update action is to be performed, performing a first action among the at least one action on the scene description data.

[0223] For some embodiments of the additional example method, determining that a trigger condition has occurred may include a logical determination that the trigger condition has occurred.

[0224] For some embodiments of the additional example method, determining that a trigger condition has occurred may include determining that the trigger condition has produced a true result.

[0225] For some embodiments of the additional example method, the first action may include applying an update to the scene description data.

[0226] For some embodiments of the additional example method, the patch parameters may indicate one or more operations, including one of a plurality of add, remove, replace, move, and copy, and the one or more operations may be applied to JavaScript Object Notation (JSON) elements of the scene description data.

[0227] For some embodiments of the additional example method, each of one or more operations may include a set of operation parameters, and each set of operation parameters may include: an operation identifier, a path to a JavaScript Object Notation (JSON) element of scene description data, and a value to be applied.

[0228] For some embodiments of the additional example method, the path may be a JSON pointer referencing a first scene element.

[0229] For some embodiments of the additional example method, each of one or more operations may include a set of operation parameters, and each set of operation parameters may include: an operation identifier, a path to a JavaScript Object Notation (JSON) element, and a source JSON element.

[0230] For some embodiments of the additional example method, the path may be a JSON pointer referencing a first scene element.

[0231] For some embodiments of the additional example method, the patch parameter may indicate a JavaScript Object Notation (JSON) patch to be applied to the scene description data.

[0232] For some embodiments of the additional example method, the patch parameter may include a location pointer to a JavaScript Object Notation (JSON) file that includes the JSON patch to be applied to the scene description data.

[0233] For some embodiments of the additional example method, the update action information may further include a shared field that indicates a state shared with a device connected to the 3D scene.

[0234] Some embodiments of the additional example method may further include: in response to determining that the shared field indicates a true result, performing a first action of at least one action on at least one device sharing the 3D scene.

[0235] Some embodiments of the additional example method may further include: in response to determining that the shared field indicates a true result, performing a first action of at least one action on all devices sharing the 3D scene.

[0236] For some embodiments of the additional example method, the update action information may further include a location description field and a placedItems field, the location description field may identify a source of location information, and the placedItems field may identify a patch operation to be used.

[0237] For some embodiments of the additional example method, the first action may include setting the position of a first scene element within the scene, and setting the position of the first scene element may use position information from an identified source.

[0238] For some embodiments of the additional example method, updating the action information may further include: a childOps array that indicates patch operations for nodes that reference nodes to be set as children of an existing node in the scene description data; and a parents array that indicates the nodes where each new child node referenced in the patch operations from the childOps array is to be added.

[0239] For some embodiments of the additional example method, the scene description data may further include behavior information.

[0240] An additional example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0241] Another example method according to some embodiments may include: obtaining scene description data for a 3D scene, where the scene description data may include: scene element information that describes each of a plurality of scene elements in the scene, trigger information that describes at least one trigger condition, and action information that describes at least one action to be performed on a first scene element among the plurality of scene elements, where the action information may include: a list of at least set actions, and set action information, where the set action information may include: a path that identifies the first scene element; and a value that indicates a new value for the first scene element; and performing a first action among the at least one action on the first scene element in response to execution of the trigger set action.

[0242] An additional example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform the methods listed above.

[0243] Further example methods according to some embodiments may include: obtaining scene description data of a 3D scene, where the scene description data may include: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, and action information describing at least one action to be performed on a first scene element among the plurality of scene elements, where the action information may include: a list of at least update actions, and update action information, where the update action information may include: a scene description patch; and in response to the execution of a trigger update action, performing a first action among the at least one action on the scene description data.

[0244] Further example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform the methods listed above.

[0245] A first example method according to some embodiments may include: obtaining scene description data of a 3D scene, where the scene description data includes: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, and action information describing at least one action to be performed on a first scene element among the plurality of scene elements, where the action information includes: a list of at least update actions, and update action information, where the update action information includes: patch parameters; and in response to determining: (i) that a trigger condition among the at least one trigger condition has occurred, and (ii) that the action information indicates that an update action is to be performed, performing the update action to apply the patch parameters as an update to the scene description data.

[0246] For some embodiments of the first example method, determining that a trigger condition has occurred includes a logical determination that the trigger condition has occurred.

[0247] For some embodiments of the first example method, determining that a trigger condition has occurred includes determining that the trigger condition has produced a true result.

[0248] For some embodiments of the first example method, the patch parameters indicate one or more operations including one or more of adding, removing, replacing, moving, and copying, and the one or more operations are applied to JavaScript Object Notation (JSON) elements of the scene description data.

[0249] For some embodiments of the first example method, the update action information further includes a shared field that indicates a state shared with a device connected to the 3D scene.

[0250] Some embodiments of the first example method may further include: performing an update action on at least one device sharing the 3D scene in response to determining that the shared field indicates a true result.

[0251] Some embodiments of the first example method may further include: performing an update action on all devices sharing the 3D scene in response to determining that the shared field indicates a true result.

[0252] For some embodiments of the first example method, the update action information further includes a location description field and a placedItems field, the location description field identifying the source of the location information, the placedItems field being an array, and the placedIems array indicating a patch operation that references nodes to be placed at a given location.

[0253] For some embodiments of the first example method, the update action includes setting the position of at least a scene element within the scene, and at least a scene element of the scene is referenced in at least a path operation, and setting the position of at least the scene element uses the location information from the identified source.

[0254] For some embodiments of the first example method, the update action information further includes: a childOps array that indicates a patch operation that references nodes to be set as children of an existing node in the scene description data; and a parents array that indicates the nodes where each new child node referenced in the patch operation from the childOps array is to be added.

[0255] For some embodiments of the first example method, the scene description data further includes behavior information.

[0256] For some embodiments of the first example method, each of the one or more operations includes a set of operation parameters, and each set of operation parameters includes: an operation identifier, a path to a JavaScript Object Notation (JSON) element of the scene description data, and a value to be applied.

[0257] For some embodiments of the first example method, the path is a JSON pointer that references a first scene element.

[0258] For some embodiments of the first example method, each of the one or more operations includes a set of operation parameters, and each set of operation parameters includes: an operation identifier, a path to a JavaScript Object Notation (JSON) element, and a source JSON element.

[0259] For some embodiments of the first example method, the path is a JSON pointer that references a first scene element.

[0260] For some embodiments of the first example method, the patch parameter indicates a JavaScript Object Notation (JSON) patch to be applied to the scene description data.

[0261] For some embodiments of the first example method, the patch parameter includes a location pointer to a JavaScript Object Notation (JSON) file that includes the JSON patch to be applied to the scene description data.

[0262] A first example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any of the methods listed above.

[0263] A second example method according to some embodiments may include: obtaining scene description data of a 3D scene, where the scene description data includes: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, and action information describing at least one action to be performed on a first scene element among the plurality of scene elements, where the action information includes: a list of at least set actions, and set action information, where the set action information includes: a path that identifies the first scene element; and a value that indicates a new value of the first scene element; and in response to determining: (i) that a trigger condition of the at least one trigger condition has occurred, and (ii) that the action information indicates that a set action is to be performed, performing the set action to set the value associated with the first scene element to the new value of the first scene element.

[0264] For some embodiments of the second example method, determining that a trigger condition has occurred includes a logical determination that the trigger condition has occurred.

[0265] For some embodiments of the second example method, determining that a trigger condition has occurred includes determining that the trigger condition has produced a true result.

[0266] For some embodiments of the second example method, the path identifies the location of the first scene element in the scene tree.

[0267] For some embodiments of the second example method, setting the value associated with the first scene element sets the value associated with the first scene element to the new value.

[0268] For some embodiments of the second example method, the set action information further includes a shared field that indicates a state shared with a device connected to the 3D scene.

[0269] Some embodiments of the second example method may further include: performing a setup action among at least one action on at least one device sharing the 3D scene in response to determining that the shared field indicates a true result.

[0270] Some embodiments of the second example method may further include: performing a setup action among at least one action on all devices sharing the 3D scene in response to determining that the shared field indicates a true result.

[0271] For some embodiments of the second example method, the setup action information further includes a location description field, and the location description field identifies the source of the location information.

[0272] For some embodiments of the second example method, the setup action includes setting the position of a first scene element within the scene, and setting the position of the first scene element uses the location information from the identified source.

[0273] For some embodiments of the second example method, the path is a JavaScript Object Notation (JSON) pointer referencing the first scene element.

[0274] For some embodiments of the second example method, the path includes path information referencing the first scene element, and the path information corresponds to Internet Engineering Task Force (IETF) standards.

[0275] For some embodiments of the second example method, the scene description data further includes behavior information.

[0276] A second example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any one of the methods listed above.

[0277] A third example method according to some embodiments may include: obtaining scene description data of a 3D scene, where the scene description data includes: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, and action information describing at least one action to be performed on a first scene element among the plurality of scene elements, where the action information includes: a list of at least update actions, and update action information, where the update action information includes: a scene description patch; and in response to triggering the execution of an update action, performing a first action among at least one action on the scene description data.

[0278] A third example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any one of the methods listed above.

[0279] A fourth example method according to some embodiments may include: obtaining scene description data of a 3D scene, where the scene description data includes: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, and action information describing at least one action to be performed on a first scene element among the plurality of scene elements, where the action information includes: a list of at least setting actions, and setting action information, where the setting action information includes: a path that identifies the first scene element; and a value that indicates a new value of the first scene element; and in response to the execution of the trigger setting action, performing a first action among the at least one action on the first scene element.

[0280] A fourth example apparatus according to some embodiments may include: a processor; and a non-transitory computer-readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform any one of the methods listed above.

[0281] A fifth example apparatus according to some embodiments may include at least one processor configured to perform any one of the methods listed above.

[0282] A sixth example apparatus according to some embodiments may include a computer-readable medium storing instructions for causing one or more processors to perform any one of the methods listed above.

[0283] A seventh example apparatus according to some embodiments may include at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform any one of the methods listed above.

[0284] An example signal according to some embodiments may include: a scene description file generated according to any one of the methods listed above.

[0285] This disclosure describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are specifically described, and at least for the purpose of showing the respective characteristics, they are generally described in a way that may sound restrictive. However, this is for the purpose of clarity of description and does not limit the disclosure or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide further aspects. In addition, the aspects can also be combined and interchanged with the aspects described in earlier submissions.

[0286] Aspects described and contemplated in this disclosure can be implemented in many different forms. While some embodiments are specifically illustrated, other embodiments are contemplated, and the discussion of specific embodiments does not limit the breadth of the implementation. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting the generated or encoded bitstream. These and other aspects can be implemented as a method, an apparatus, a computer-readable storage medium storing instructions for encoding or decoding video data according to any of the methods, and / or a computer-readable storage medium storing a bitstream generated according to any of the methods.

[0287] Various methods are described herein, and each of the methods includes one or more steps or actions for implementing the method. Unless a particular order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, terms such as "first", "second", etc. can be used in various embodiments to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding". Unless specifically required, the use of such terms does not imply an ordering of the modified operations. Thus, in this example, the first decoding does not need to be performed before the second decoding, and can occur, for example, before, during, or within a time period overlapping with the second decoding.

[0288] For example, various numerical values can be used in this disclosure. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.

[0289] The embodiments described herein can be executed by computer software implemented by a processor or other hardware or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. As a non-limiting example, the processor can be of any type suitable for the technical environment and can encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0290] When the drawings are presented as flowcharts, it should be understood that they also provide a block diagram of the corresponding apparatus. Similarly, when the drawings are presented as block diagrams, it should be understood that they also provide a flowchart of the corresponding method / process.

[0291] The implementations and aspects described herein can be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., only as a method), the implementation of the features discussed can be implemented in other forms (e.g., an apparatus or a program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. A method can be implemented in, for example, a processor, which generally refers to a processing device that includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as, for example, a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate the transmission of information between end users.

[0292] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation" and other variations thereof mean that the particular features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" and any other variations thereof that occur throughout this disclosure are not necessarily all referring to the same embodiment.

[0293] Additionally, the present disclosure can relate to "determining" pieces of various information. Determining information can include, for example, one or more of the following: estimating information, calculating information, predicting information, or retrieving information from a memory.

[0294] Furthermore, the present disclosure can relate to "accessing" pieces of various information. Accessing information can include, for example, one or more of the following: receiving information, retrieving information (e.g., retrieving information from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0295] Additionally, the present disclosure can relate to "receiving" pieces of various information. Like "accessing", receiving is intended to be a broad term. Receiving information can include, for example, one or more of the following: accessing information or retrieving information (e.g., retrieving information from a memory). Additionally, "receiving" is generally involved in one way or another during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0296] It should be understood that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", any one of the following " / ", "and / or", and "at least one of..." is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of the first-listed option and the second-listed option (A and B), or the selection of the first-listed option and the third-listed option (A and C), or the selection of the second-listed option and the third-listed option (B and C), or the selection of all three options (A and B and C). This can be extended to as many items as are listed.

[0297] In addition, among other things, as used herein, the word "signal" also refers to indicating something to a corresponding decoder. For example, in some embodiments, the encoder signals a particular one of a plurality of parameters for region-based filter parameter selection for artifact filtering. In this way, in an embodiment, the same parameters are used at both the encoder side and the decoder side. Thus, for example, the encoder can transmit (explicitly signal) a particular parameter to the decoder such that the decoder can use the same particular parameter. Conversely, if the decoder already has a particular parameter and other parameters, signaling can be used without transmission (implicitly signal) to allow only the decoder to know and select the particular parameter. Bit savings are achieved in various embodiments by avoiding the transmission of any actual functionality. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, information is signaled to a corresponding decoder using one or more syntax elements, flags, and the like. Although the foregoing relates to the verb form of the word "signal", the word "signal" can also be used as a noun herein.

[0298] Implementations can generate a variety of signals that are formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry the bitstream of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. Signals can be stored on a processor-readable medium.

[0299] We have described multiple embodiments. The features of these embodiments can be provided individually or in any combination across various claim categories and types. Additionally, embodiments can include one or more of the following features, devices, or aspects, either individually or in any combination, across various claim categories and types:

[0300] · A bitstream or signal that includes one or more of the described syntax elements or variants thereof.

[0301] · A bitstream or signal that includes a syntax for conveying information generated according to any of the described embodiments.

[0302] · Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements or variants thereof.

[0303] · Creating and / or transmitting and / or receiving and / or decoding according to any of the described embodiments.

[0304] · A method, process, apparatus, medium storing instructions, medium storing data, or signal according to any of the described embodiments.

[0305] Note that various hardware elements of one or more of the described embodiments are referred to as "modules" that perform (i.e., execute, carry out, and the like) the various functions described herein in connection with the corresponding modules. As used herein, a module includes hardware that is considered suitable for a given implementation (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more storage devices). Each described module can also include executable instructions for performing one or more functions described as being performed by the corresponding module, and note that those instructions can take the form of hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and / or the like, or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and / or the like, and can be stored in any suitable one or more non-transitory computer-readable media (such as those commonly referred to as RAM, ROM, etc.).

[0306] Although the features and elements are described above in specific combinations, each feature or element can be used alone or in any combination with other features and elements. In addition, the methods described herein can be implemented as a computer program, software, or firmware, which is incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor storage devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROM disks and digital versatile disks (DVDs)). A processor associated with the software can be used to implement a radio frequency transceiver that is used in a WTRU, UE, terminal, base station, RNC, or any host computer.

[0307] Note that various hardware elements in the described embodiments are referred to as "modules" that perform (i.e., execute, carry out, and the like) the various functions described herein in connection with the corresponding modules. As used herein, a module includes hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more storage devices) that is considered suitable by those skilled in the relevant art for a given implementation. Each described module may also include executable instructions for performing one or more functions described as being performed by the corresponding module, and note that those instructions may take the form of hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and / or the like, or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, and / or the like, and may be stored in any suitable one or more non-transitory computer-readable media (such as those commonly referred to as RAM, ROM, etc.).

[0308] Although the features and elements are described above in specific combinations, those of ordinary skill in the art will understand that each feature or element can be used alone or in any combination with other features and elements. In addition, the methods described herein can be implemented as a computer program, software, or firmware, which is incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor storage devices, magnetic media (such as internal hard disks and removable disks), magneto-optical media, and optical media (such as CD-ROM disks and digital versatile disks (DVDs)). A processor associated with the software can be used to implement a radio frequency transceiver that is used in a WTRU, UE, terminal, base station, RNC, or any host computer.< / dot> < / dot> < / dot> < / dot> < / dot> ​< / dot> < / dot>

Claims

1. A method, comprising: Obtaining scene description data of a three-dimensional (3D) scene, wherein the scene description data includes: Scene element information describing each of a plurality of scene elements in the scene, Trigger information describing at least one trigger condition, and Action information describing at least one action to be performed on a first scene element among the plurality of scene elements, wherein the action information includes: A list of at least update actions, and Update action information, wherein the update action information includes: A patch parameter; and In response to determining: (i) that a trigger condition among at least one trigger condition has occurred, and (ii) that the action information indicates that an update action is to be performed, performing the update action to apply the patch parameter as an update to the scene description data.

2. The method according to claim 1, wherein determining that the trigger condition has occurred includes a logical determination that the trigger condition has occurred.

3. The method according to any one of claims 1-2, wherein determining that the trigger condition has occurred includes determining that the trigger condition has produced a true result.

4. The method according to any one of claims 1-3, wherein the patch parameter indicates one or more operations, including one or more of adding, removing, replacing, moving, and copying, and wherein the one or more operations are applied to JavaScript Object Notation (JSON) elements of the scene description data.

5. The method according to any one of claims 1-4, wherein the update action information further includes a shared field, and the shared field indicates a state shared with a device connected to the 3D scene.

6. The method according to claim 5, further comprising: In response to determining that the shared field indicates a true result, performing the update action on at least one device sharing the 3D scene.

7. The method according to claim 5, further comprising: In response to determining that the shared field indicates a true result, performing the update action on all devices sharing the 3D scene.

8. The method according to any one of claims 1-7, wherein the update action information further includes a position description field and a placedItems field, wherein the position description field identifies the source of the position information, wherein the placedItems field is an array, and wherein the placedIems array indicates a patch operation that references a node to be placed at a given position.

9. The method according to claim 8, wherein the update action includes setting the position of at least a scene element within the scene, and wherein at least a scene element of the scene is referenced in at least a path operation according to the placedItems field, and wherein setting the position of at least the scene element uses the position information from the identified source.

10. The method according to any one of claims 1-9, wherein the update action information further includes: A childOps array, the childOps array indicating a patch operation that references a node to be set as a child node of an existing node in the scene description data; And The parents array, where the parents array indicates the nodes to which each new child node referenced in the patch operation from the childOps array is to be added.

11. The method according to any one of claims 1 - 10, wherein the scenario description data further includes behavior information.

12. The method according to any one of claims 1 - 11, where each of the one or more operations includes a set of operation parameters, and where each set of operation parameters includes: An operation identifier, A path to a JavaScript Object Notation (JSON) element of the scenario description data, and A value to be applied.

13. The method according to claim 12, wherein the path is a JSON pointer referencing a first scenario element.

14. The method according to any one of claims 1 - 13, where each of the one or more operations includes a set of operation parameters, and where each set of operation parameters includes: An operation identifier, A path to a JavaScript Object Notation (JSON) element, and A source JSON element.

15. The method according to claim 14, wherein the path is a JSON pointer referencing a first scenario element.

16. The method according to any one of claims 1 - 15, wherein the patch parameter indicates a JavaScript Object Notation (JSON) patch to be applied to the scenario description data.

17. The method according to any one of claims 1 - 16, wherein the patch parameter includes a location pointer to a JavaScript Object Notation (JSON) file that includes the JSON patch to be applied to the scenario description data.

18. An apparatus, comprising: A processor; And A non - transitory computer - readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform the method according to any one of claims 1 to 17.

19. A method, comprising: Obtaining scenario description data of a three - dimensional (3D) scenario, where the scenario description data includes: Scenario element information describing each of a plurality of scenario elements in the scenario, Trigger information describing at least one trigger condition, and Action information describing at least one action to be performed on a first scenario element among the plurality of scenario elements, where the action information includes: A list of at least set actions, and Set action information, where the set action information includes: A path that identifies the first scenario element; and A value that indicates a new value of the first scenario element; and In response to determining: (i) that the trigger condition of at least one trigger condition has occurred, and (ii) that the action information indicates that a set action is to be performed, performing the set action to set the value associated with the first scenario element to the new value of the first scenario element.

20. The method according to claim 19, wherein determining that the trigger condition has occurred includes a logical determination that the trigger condition has occurred.

21. The method according to any one of claims 19 - 20, wherein determining that the trigger condition has occurred includes determining that the trigger condition has produced a true result.

22. The method according to any one of claims 19 - 21, wherein the path identifies the position of the first scene element in the scene tree.

23. The method according to any one of claims 19 - 22, wherein setting the value associated with the first scene element sets the value associated with the first scene element to a new value.

24. The method according to any one of claims 19 - 23, wherein setting the action information further includes a shared field that indicates the state shared with the device connected to the 3D scene.

25. The method according to claim 24, further comprising: performing the set action among the at least one action on at least one device sharing the 3D scene in response to determining that the shared field indicates a true result.

26. The method according to claim 24, further comprising: performing the set action among the at least one action on all devices sharing the 3D scene in response to determining that the shared field indicates a true result.

27. The method according to any one of claims 19 - 26, wherein setting the action information further includes a location description field, and wherein the location description field identifies the source of the location information.

28. The method according to claim 27, wherein setting the action includes setting the position of the first scene element within the scene, and wherein setting the position of the first scene element uses the location information from the identified source.

29. The method according to any one of claims 19 - 28, wherein the path is a JavaScript Object Notation (JSON) pointer referencing the first scene element.

30. The method according to any one of claims 19 - 29, wherein the path includes path information referencing the first scene element, and wherein the path information corresponds to the Internet Engineering Task Force (IETF) standard.

31. The method according to any one of claims 19 - 30, wherein the scene description data further includes behavior information.

32. An apparatus, comprising: a processor; and a non - transitory computer - readable medium storing instructions that, when executed by the processor, are operable to cause the apparatus to perform the method according to any one of claims 19 to 31.

33. A method, comprising: obtaining scene description data of a three - dimensional (3D) scene, wherein the scene description data includes: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, and action information describing at least one action to be performed on a first scene element among the plurality of scene elements, wherein the action information includes: a list of at least update actions, and update action information, wherein the update action information includes: a scene description patch; and performing the first action among the at least one action on the scene description data in response to the execution of the trigger update action.

34. An apparatus, comprising: a processor; and A non-transitory computer-readable medium storing instructions which, when executed by a processor, are operable to cause the apparatus to perform the method of claim 33.

35. A method comprising: obtaining scene description data of a three-dimensional (3D) scene, wherein the scene description data includes: scene element information describing each of a plurality of scene elements in the scene, trigger information describing at least one trigger condition, and action information describing at least one action to be performed on a first scene element among the plurality of scene elements, wherein the action information includes: a list of at least set actions, and set action information, wherein setting the action information includes: a path that identifies the first scene element; and a value that indicates a new value of the first scene element; and performing a first action among the at least one action on the first scene element in response to execution of the trigger setting action.

36. An apparatus comprising: a processor; and a non-transitory computer-readable medium storing instructions which, when executed by the processor, are operable to cause the apparatus to perform the method of claim 35.

37. An apparatus comprising at least one processor configured to perform the method of any one of claims 1-17, 19-31, 33, and 35.

38. An apparatus comprising a computer-readable medium storing instructions for causing one or more processors to perform the method of any one of claims 1-17, 19-31, 33, and 35.

39. An apparatus comprising at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform the method of any one of claims 1-17, 19-31, 33, and 35.

40. A signal comprising a scene description file generated according to any one of claims 1-17, 19-31, 33, and 35.