Reality Node Extension in Scene Description

By generating a scene graph that identifies real and virtual objects, the method enhances AR rendering efficiency and resource management, addressing the challenge of object differentiation in AR scene descriptions.

JP2025520319APending Publication Date: 2025-07-03INTERDIGITALCE PATENT HLDG SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024571835
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-17
Filing Date
2023-06-12
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing augmented reality (AR) scene description technologies fail to distinguish between real-world and virtual objects, leading to inefficiencies in rendering and resource utilization, particularly in managing collisions and lighting effects.

Method used

A scene graph is generated linking nodes to indicate whether they correspond to real or virtual objects, with attributes like 'Real Node' and 'Virtual Node' providing information to the renderer for accurate rendering and resource management.

Benefits of technology

Enables efficient rendering of AR scenes by distinguishing between real and virtual objects, optimizing resource use and improving collision detection and lighting consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520319000001_ABST
    Figure 2025520319000001_ABST
Patent Text Reader

Abstract

A method, device, and data stream are provided for generating, transmitting, and decoding a scene description of an augmented reality scene. According to this principle, a scene graph links nodes and includes information indicating which nodes correspond to objects in a first list of 3D models corresponding to objects in the real environment of the scene and / or which nodes correspond to objects in a second list of 3D models corresponding to virtual objects of the scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This principle generally relates to the fields of augmented reality scene description and augmented reality rendering. In particular, this principle relates to the description of the management of objects in the real environment in the description of 3D scenes. This document is also understood in the context of formatting and playing augmented reality applications when rendered on an end-user device such as a mobile device or a head-mounted display (HMD) such as see-through glasses.

Background Art

[0002] This section is intended to introduce the reader to various technical aspects that may be related to the various aspects of the principle described and / or claimed below. This discussion is thought to be useful in providing the reader with background information to facilitate a better understanding of the various aspects of the principle. Therefore, it should be understood that these descriptions should be read from this perspective and should not be read as an admission of prior art.

[0003] Extended reality (XR) is a technology that enables an interactive experience in which the real-world environment and / or video content is enhanced by virtual content that can be defined across multiple types of senses including vision, hearing, touch, etc. During the execution of an application, virtual content (e.g., 3D content or audio / video files) is rendered in real time in a manner that is consistent with the user's context (environment, viewpoint, device, etc.). A scene graph (such as those proposed by Khronos / glTF, and its extensions defined, for example, by the MPEG scene description format or Apple / USDZ) is a possible way to represent the content to be rendered. These combine, on the one hand, a declarative description of the scene structure that links real-world environment objects and virtual objects, and on the other hand, a binary representation of the virtual content. The binary representation of real-world objects can also be included in the nodes of the scene description. Such a representation of real-world objects may be obtained, for example, by scanning the real-world environment. The scene description framework ensures that time-based media and the corresponding associated virtual content are available at any time during the rendering of the application. The scene description can also carry scene-level data that describes how scene objects behave and interact during the execution of an immersive XR experience. The management of the representation of real-world objects (real-world objects and / or virtual objects, the user being a real-world object) is a technical challenge because they can be used for different processing in rendering. There is no solution to the scene description that indicates which objects in the scene contain the representation of virtual objects and which objects contain the representation of real-world environment objects. SUMMARY OF THE INVENTION

[0004] The following presents a simplified overview of the present principle to provide a basic understanding of some aspects of the present principle. This overview is not a comprehensive overview of the present principle. It is not intended to identify important or critical elements of the present principle. The following overview merely presents some aspects of the present principle in a simplified form as a prelude to the more detailed explanation provided below.

[0005] The present principle relates to a method for generating an augmented reality scene description. The method includes obtaining a first list of 3D models corresponding to objects in the real environment of the scene and a second list of 3D models corresponding to virtual objects in the scene. Next, a scene graph is generated, and the scene graph links nodes and includes information indicating which nodes correspond to objects in the first list and which nodes correspond to objects in the second list. The augmented reality scene description is encoded into a data stream using node-level information and scene-level information.

[0006] In one embodiment, the information is an array of indices of nodes corresponding to objects in the first list stored at the scene level. In another embodiment, the information is an array of indices of nodes corresponding to objects in the second list stored at the scene level. In another embodiment, the nodes of the scene graph include information indicating whether the node corresponds to an object in the first list or an object in the second list. All of these embodiments can be combined to duplicate or authenticate the information.

[0007] The present principle also relates to a device for implementing the above method. The present principle also relates to a method and a device for decoding and processing a scene description generated according to the above method. The present principle also relates to a data stream carrying data representing a scene description generated according to the above method.

Brief Description of the Drawings

[0008] This disclosure will be better understood when the following description is read, and other specific features and advantages will become apparent. In the following description, reference is made to the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

[0009] This principle will be described in more detail below with reference to the accompanying drawings in which examples of embodiments of the principle are shown. However, the principle can be embodied in many alternative forms and should not be construed as limited to the embodiments set forth herein. Accordingly, while there is room for various modifications and alternative forms, specific examples thereof are shown by way of example in the drawings and will be described in detail herein. However, it is to be understood that there is no intention to limit the principle to the particular forms disclosed, and on the contrary, the disclosure is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the principle as defined by the claims.

[0010] The terms used in this specification are for the purpose of describing particular embodiments only and are not intended to limit the principles. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. As used herein, the terms "comprises", "comprising", "includes", and / or "including" specify the presence of the stated feature, integer, step, operation, element, and / or component, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Further, when an element is referred to as "responding to" or "connected to" another element, it can respond directly to the other element, be connected to the other element, or there may be intervening elements. In contrast, when an element is referred to as "responding directly to" or "directly connected to" another element, there are no intervening elements. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and may be abbreviated as " / ".

[0011] In this specification, terms such as first, second, etc. may be used to describe various elements, but it will be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element can be referred to as a second element, and similarly, a second element can be referred to as a first element without departing from the teachings of the principles.

[0012] Some of the figures include arrows on the communication path to indicate the main direction of communication, but it should be understood that communication can occur in the direction opposite to the drawn arrows.

[0013] Several examples are described with respect to block diagrams and operational flowcharts that represent portions of a circuit element, module, or code, where each block contains one or more executable instructions for performing a specified logical function. It should also be noted that in other implementations, the functions described in the blocks may occur in the order described. For example, two blocks shown in sequence may actually be executed substantially simultaneously, or depending on the functions involved, the blocks may be executed in reverse order.

[0014] As used herein, "according to an example" or "in an example" means that a particular feature, structure, or characteristic described in connection with the example may be included in at least one implementation of the principle. The appearances of the phrases "according to an example" or "in an example" in various places in this specification are not necessarily all referring to the same example, and in separate or alternative examples, they are not necessarily mutually exclusive of other examples.

[0015] The reference signs appearing in the claims are for illustrative purposes only and shall not have a limiting effect on the scope of the claims. Although not explicitly stated, the examples and variations of this example may be used in any combination or partial combination.

[0016] Figure 1 shows an exemplary graph 10 for augmented reality scene description. In this example, the scene graph includes a description of a real object, such as a "horizontal surface" (which could be a table or a road), and a description of a virtual object 12, such as an animation of a car. The scene description is organized as an array of nodes 10. Nodes can be linked to child nodes to form a scene structure 11. A node can carry a description of a real object (e.g., a semantic description) or a description of a virtual object. In the example of Figure 1, node 101 describes a virtual camera located within the 3D volume of an XR application. Node 102 describes a virtual car and includes an index of the car's representation, e.g., an index within an array of 3D meshes. According to this principle, a node can also include an index of the representation of this real object. This representation may be, for example, obtained by scanning the real environment, or, for example, a model generated by a content creator. This representation must correspond to the real object described by this node and thus be incorporated into the 3D scene. In such a scene description, since both real nodes and virtual nodes include 3D representations, there is no longer a means to distinguish between real nodes and virtual nodes. According to this principle, information about the nature of the nodes in the scene description is provided to the renderer.

[0017] The scene description may include a number of arrays containing descriptions of various aspects of the scene, such as an array containing several animation instructions, an array containing the meshes of several virtual or real objects, or an array containing a description of materials. Node 103 is a child of node 102 and includes a description of one of the car's wheels. Similarly, it includes an index to the 3D mesh of the wheel. Since the scale, location, and orientation of the object are described in the scene node, the same 3D mesh may be used for several objects within the 3D scene. The scene graph 10 also includes nodes that are descriptions of the spatial relationships between real objects and virtual objects.

[0018] XR applications are diverse and can be applied to different contexts and real or virtual environments. For example, in an industrial XR application, an item of virtual 3D content (e.g., part A of an engine) is displayed when a reference object (part B of the engine) in the real environment is detected by a camera mounted on a head-mounted display device. The 3D content item is placed in the real world at a position and scale defined relative to the detected reference object. Similarly, an action on a virtual object may be executed when a collision, i.e., a collision between two real objects, between two virtual objects, or between one real object and one virtual object, is detected.

[0019] For example, in an XR application for interior design, the color of the displayed virtual furniture parts changes when the user touches a virtual control object or when the user touches a real table. In another application, the playback of an audio file may start when two moving virtual objects being displayed collide. In another example, a short advertising sound file may be played when the user grabs a given soda can in the real environment. However, detecting a collision (or contact) between real or virtual objects and rendering a realistic reaction of the virtual object requires the use of a physics engine with huge memory and processing resources. Therefore, in order to optimize the use of the physics engine, it is important to have a scene description that accurately describes different types of behaviors linked to the collision. Such an XR scene description format is provided herein in accordance with the present principle.

[0020] The XR application can also extend video content rather than the real environment. The video is displayed on the rendering device, and the virtual objects described in the node tree are overlaid when time-based events are detected in the video. In such a context, the node tree only contains the description of the virtual objects.

[0021] Real assets, which are representations of real objects, are very useful for improving the XR experience. For example, scanning a room makes it possible to manage the collision between virtual objects and real objects. This can also be used for rendering purposes, for example, to prevent real objects from being rendered in the case of see-through XR devices, or to ensure consistent lighting between real objects and virtual objects, including, for example, shadow management. This principle provides a solution that indicates the nature of any node in the scene description.

[0022] Figure 4 shows an example of the rendering of an augmented reality scene. In the example of Figure 5, object 51 is a real table, object 52 is a real speaker, and object 53 is a virtual teddy bear. The scene description generated according to this principle enables the rendering system to estimate the shadows on the real objects in order to generate non-contradictory shadows for the teddy bear. In such an example, it is important for the rendering system to distinguish between real objects and virtual objects in the augmented reality scene.

[0023] According to this principle, the scene description includes information about nodes that contain representations of real objects in the scene in order to obtain a rendering close to reality. In the first embodiment, the scene description includes an array of real nodes at the scene level. For example, within the scope of the MPEG-I scene description framework that uses the Khronos glTF extension mechanism to support additional scene description features, the "MPEG_real_nodes" extension is defined at the scene level. The semantics of MPEG_real_nodes at the scene level are defined in the following table.

[0024] [Table 1]

[0025] An AR scene with an actual scan of the environment enables the management of collisions between the real world and the virtual world. In an AR experience, the meshes and textures of real objects are not visible (the real world is already visible through the AR device), and only the mesh collision detection function is used. In other types of applications, such as remote XR experiences, the representation of real objects (scanned points of a cloud or 3D model) can be rendered and displayed together with virtual objects.

[0026] According to this principle, the "Real Node" attribute provides information to the renderer to manage the visibility of real objects. In an AR scene, lighting can be very important. Lighting enables the consistent management of the shadows of virtual objects. To calculate the output of the lighting shader, various cases must be considered. The output of the lighting shader provides the color of the pixels in the final composite image, and its calculation varies depending on whether the pixel belongs to a real object that is not in shadow, a real object in the shadow cast by a real surface, a real object in the shadow cast by a virtual surface, a virtual object that is not in shadow, or a virtual object that is in shadow.

[0027] To distinguish cases, a binary flag is supplied to the shader and set to 0 for vertices corresponding to the 3D model of a real object and to 1 for vertices corresponding to additional virtual objects. Two depth maps are also generated from the estimated position of each light, one for real objects (i.e., texture meshes) and the other for virtual objects. Since the "Real Node" attribute provides the necessary information to the renderer, no shadow is generated for real objects as the shadow already exists.

[0028] Since the concept of "Real Node" is related to nodes in the scene graph, this concept is also related to light sources. In an AR framework, real lights are realized through virtual lights that illuminate virtual objects for consistent lighting rendering. Each light is associated with at least one node. Lights inherit the "Real Node" attribute from the nodes. This "Real Node" attribute can be used for possible control of the lights.

[0029] On the rendering side, during execution, the application iterates over each defined node at the scene level. Information included in the scene description (collisions, virtual lights) and information obtained from the application are used differently by the renderer. For the sake of explanation, the details of the present invention are given within the scope of the MPEG-I scene description framework using the Khronos glTF extension mechanism, supporting additional scene description features. However, this principle may be applied to other existing or future descriptions of XR scenes. As illustrated in FIG. 1, each glTF node can include an array called children that contains the indices of its child nodes. Thus, each node is one element of the node hierarchy, and together they define the structure of the scene as a scene graph. An example of the possible syntax for such a real node description in the MPEG-I scene description is shown below.

[0030]

Table 2-1

[0031]

Table 2-2

[0032] In the second embodiment, the scene description includes a list of virtual objects at the scene level. Thus, by parsing this list, the renderer is informed that nodes having pointers to the 3D models listed are virtual objects, and nodes having pointers to 3D models not listed are real objects. In this second embodiment, the array attribute is called "VirtualNodes". The second embodiment may be combined with the first embodiment. In fact, since the scene graph may include nodes that are not directly related to objects having 3D models, it may be beneficial for the renderer to have both a list of real objects and a list of virtual objects, enabling the renderer to distinguish object nodes from other types of nodes.

[0033] In the third embodiment, information regarding the nature of a node is indicated by an attribute at the node level. For example, the attribute "isRealObject" at the node level may be set to 1 when the associated 3D model corresponds to an object in the real environment of the application, and may be set to 0 when the associated 3D model corresponds to a virtual object. In a variant, the "isVirtual" attribute at the node level is set to 1 when the object is virtual and set to 0 when the object is real.

[0034] The third embodiment can also be combined with the first or second embodiment. In fact, information regarding the real or virtual nature of a node may be indicated at the scene level and at the node level. The renderer can use this replicated information for post - processing steps.

[0035] FIG. 2 shows an exemplary architecture of an XR processing engine 30 that can be configured to implement the present principle. Devices with the architecture of FIG. 2 are linked to other devices via their buses 31 and / or via the I / O interface 36.

[0036] Device 30 includes the following elements linked together by data and address bus 31. - A microprocessor 32 (or CPU), for example a DSP (or Digital Signal Processor), - A ROM (or Read Only Memory) 33, - A RAM (or Random Access Memory) 34, - A storage device interface 35 - An I / O interface 36 for receiving data transmitted from an application, and - A power source, such as a battery (not shown in FIG. 2).

[0037] According to one example, the power supply is external to the device. In each of the memories mentioned, the word "register" as used herein can correspond to a small capacity area (a few bits) or a very large area (e.g., an entire program or a large amount of received or decoded data). The ROM 33 includes at least programs and parameters. The ROM 33 can store algorithms and instructions for executing techniques according to this principle. When the switch is turned on, the CPU 32 uploads the program in the RAM and executes the corresponding instructions.

[0038] The RAM 34 includes, within the register, a program executed by the CPU 32 and uploaded after turning on the switch of the device 30, input data within the register, intermediate data of different states of the method within the register, and other variables used for the execution of the method within the register.

[0039] The implementations described herein can be realized, for example, in a method or process, an apparatus, a computer program product, a data stream, or a signal. Even when considered only in the context of a single form of implementation (e.g., considered only as a method or a device), the implementation of the features considered can also be implemented in other forms (e.g., a program). The apparatus can be implemented, for example, with appropriate hardware, software, and firmware. This method can be executed, for example, in an apparatus such as a processor that generally refers to a processing device, including a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors include, for example, communication devices such as computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the communication of information between end users.

[0040] Device 30 is linked, for example, via bus 31 to a set of sensors 37 and a set of rendering devices 38. The sensors 37 may be, for example, cameras, microphones, temperature sensors, inertial measurement units, GPS, humidity measurement sensors, IR or UV sensors, or wind sensors. The rendering devices 38 may be, for example, displays, speakers, vibrators, heat sources, fans, etc.

[0041] According to an example, device 30 is configured to implement a method according to the present principle and belongs to a set including the following. - Mobile devices, - Communication devices, - Gaming devices, - Tablets (or tablet computers), - Laptops, - Still cameras, - Video cameras.

[0042] FIG. 3 shows an example of an embodiment of the syntax of a data stream encoding an extended reality scene description according to the present principle. FIG. 3 shows an exemplary structure 4 of an XR scene description. The structure is within a container that organizes the stream in independent elements of the syntax. The structure may include a header portion 41 that is a set of data common to all syntax elements of the stream. For example, the header portion includes some metadata regarding the syntax elements and describes the nature and role of each of them. The structure also includes a payload including a syntax element 42 and a syntax element 43. The syntax element 42 includes data representing a media content item described at a node of a scene graph associated with a virtual element. Images, meshes, and other raw data may be compressed according to a compression method. The syntax element 43 is part of the payload of the data stream and includes data encoding a scene description as described by the present principle.

[0043] The implementations described in this specification may be realized, for example, in a method or process, an apparatus, a computer program product, a data stream, or a signal. Even when considered in the context of a single form of implementation (e.g., only considered as a method or a device), the implementation of the features considered can also be implemented in other forms (e.g., a program). The apparatus can be implemented, for example, with appropriate hardware, software, and firmware. This method can be executed, for example, in an apparatus such as a processor that generally refers to a processing device, including a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor can also include, for example, a communication device such as a smartphone, a tablet, a computer, a mobile phone, a portable / personal digital assistant (``PDA''), and other devices that facilitate the communication of information between the end user.

[0044] The implementations of the various processes and features described in this specification can be embodied in various different devices or applications, particularly, for example, devices or applications associated with data encoding, data decoding, view generation, texture processing, and other processing of images and related texture information and / or depth information. Examples of such devices include an encoder, a decoder, a post-processor that processes the output from the decoder, a pre-processor that provides input to the encoder, a video coder, a video decoder, a video codec, a web server, a set-top box, a laptop, a personal computer, a mobile phone, a PDA, and other communication devices. As should be clear, the devices can be mobile and can be installed in a mobile vehicle.

[0045] Additionally, the method may be implemented by instructions executed by a processor, and such instructions (and / or data values resulting from the implementation) may be stored, for example, on a processor-readable medium such as an integrated circuit, a software carrier, or other storage devices such as a hard disk, a compact diskette (CD), an optical disk (e.g., a digital versatile disc, often referred to as a DVD), a random access memory (RAM), or a read-only memory (ROM). The instructions may form an application program tangibly embodied on the processor-readable medium. The instructions may be, for example, hardware, firmware, software, or a combination. The instructions may be found, for example, in an operating system, a separate application, or a combination of the two. Thus, a processor may be characterized as both a device configured to execute a process and a device including a processor-readable medium (such as a storage device) having instructions for executing the process. Further, the processor-readable medium can store data values resulting from the implementation in addition to, or instead of, the instructions.

[0046] As will be apparent to those skilled in the art, the implementation forms can generate various signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for executing a method or data generated by one of the described implementation modes. For example, the signal can be formatted to carry, as data, rules for writing or reading the syntax of the described embodiments, or actual syntax values written by the described embodiments. Such signals can be formatted, for example, as electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or as baseband signals. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog information or digital information. The signal can be transmitted by various different wired or wireless links, as is known. The signal can be stored in a processor-readable medium.

[0047] Numerous implementation forms have been described. Nevertheless, it will be understood that various modifications can be made. For example, elements of different implementation forms can be combined, supplemented, modified, or deleted to yield other implementation forms. In addition, those skilled in the art can substitute other structures and processes for those disclosed, and it will be understood that the resulting implementation forms will perform at least substantially the same functions in at least substantially the same manner to achieve at least substantially the same results as the disclosed implementation forms. Accordingly, these and other implementation forms are contemplated by this application.

Claims

1. A method for generating a scene description of an augmented reality scene, the method comprising: - obtaining a first list of 3D models corresponding to objects in the real environment of the augmented reality scene and a second list of 3D models corresponding to virtual objects in the augmented reality scene; - generating the scene description including a scene graph that links nodes and includes information indicating which node corresponds to an object in the first list and which node corresponds to an object in the second list; - encoding the augmented reality scene description in a data stream.

2. The method according to claim 1, wherein the information is an array of indexes of the nodes corresponding to the objects in the first list stored at the scene level.

3. The method according to claim 1 or 2, wherein the information is an array of indexes of the nodes corresponding to the objects in the second list stored at the scene level.

4. The method according to any one of claims 1 to 3, wherein each node of the scene graph includes information indicating whether the node corresponds to an object in the first list or an object in the second list.

5. A device comprising a memory associated with a processor, the processor being configured to: - obtain a first list of 3D models corresponding to objects in the real environment of the augmented reality scene and a second list of 3D models corresponding to virtual objects in the augmented reality scene; - generate a scene description including a scene graph that links nodes and includes information indicating which node corresponds to an object in the first list and which node corresponds to an object in the second list; and - encode the augmented reality scene description in a data stream.

6. The device according to claim 5, wherein the information is an array of indexes of the nodes corresponding to the objects in the first list stored at the scene level.

7. The device according to claim 5 or 6, wherein the information is an array of indexes of the nodes corresponding to the objects in the second list stored at the scene level.

8. The device according to any one of claims 5 to 7, wherein each node of the scene graph includes information indicating whether the node corresponds to an object in the first list or an object in the second list.

9. A method for rendering an augmented reality scene, the method comprising: - obtaining a scene description of the augmented reality scene from a data stream; - decoding, from the scene description, a scene graph that links nodes and includes information indicating which nodes correspond to objects in the real environment of the augmented reality scene and which nodes correspond to virtual objects in the augmented reality scene; - processing the nodes of the scene graph according to the information.

10. The method according to claim 9, wherein the information is an array of indices of the nodes corresponding to the objects in the first list stored at the scene level.

11. The method according to claim 9 or 10, wherein the information is an array of indices of the nodes corresponding to the objects in the second list stored at the scene level.

12. The method according to any one of claims 9 to 11, wherein each node of the scene graph includes information indicating whether the node corresponds to an object in the first list or an object in the second list.

13. A device comprising a memory associated with a processor, the processor being configured to: - obtain a scene description of an augmented reality scene from a data stream; - decode, from the scene description, a scene graph that links nodes and includes information indicating which nodes correspond to objects in the real environment of the augmented reality scene and which nodes correspond to virtual objects in the augmented reality scene; and - process the nodes of the scene graph according to the information.

14. The device according to claim 13, wherein the information is an array of indices of the nodes corresponding to the objects in the first list stored at the scene level.

15. The device according to claim 13 or 14, wherein the information is an array of indices of the nodes corresponding to the objects in the second list stored at the scene level.

16. The device according to any one of claims 13 to 15, wherein each node of the scene graph includes information indicating whether the node corresponds to an object in the first list or an object in the second list. **Claim 17** A data stream including a scene description of an augmented reality scene, the scene description including a scene graph that links nodes and includes information indicating which nodes correspond to objects in the real environment of the augmented reality scene and which nodes correspond to virtual objects in the augmented reality scene. **Claim 18** The data stream according to claim 17, wherein the information is an array of indices of the nodes corresponding to the objects in the first list stored at the scene level. **Claim 19** The data stream according to claim 17 or 18, wherein the information is an array of indices of the nodes corresponding to the objects in the second list stored at the scene level. **Claim 20** The data stream according to any one of claims 17 to 19, wherein each node of the scene graph includes information indicating whether the node corresponds to an object in the first list or an object in the second list.