Method for updating time evolution scans

By segmenting AR/XR scenes into mesh objects and using the 'canRepair' flag, the problem of repairing elements after partial scanning is solved, simplifying the interaction and rescanning steps, and improving the consistency and efficiency of the AR/XR experience.

CN121548799APending Publication Date: 2026-02-17INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480047302.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-19
Filing Date
2024-07-04
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

In augmented reality and extended reality, existing technologies lack effective methods to repair real-world environmental elements after partial scanning has been moved. This results in complex and time-consuming rescanning and semantic segmentation steps, affecting the coherence and efficiency of the AR/XR experience.

Method used

By segmenting the extended reality scene into mesh objects and assigning a first state (fully visible), a second state (partially visible), or a third state (unknown) to each object, a node tree is generated, and the 'canRepair' flag is inserted into the scene description file to allow for the repair of locally scanned meshes at runtime.

Benefits of technology

It simplifies interaction with segmented grids, reduces the need for rescanning, improves the coherence and efficiency of the AR/XR experience, and supports flexible updates and sharing of scene descriptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121548799A_ABST
    Figure CN121548799A_ABST
Patent Text Reader

Abstract

A method and apparatus for generating a scene description of an evolved augmented reality scene and updating the evolved augmented reality scene upon detection of a change. A scene description associated with the scene scan is generated and stored in a file. The nodes have a state indicating whether they can be repaired, and optionally have a fallback grid of the relevant object. At runtime, the AR / XR application loads the first (previous) scan, and if the mesh of the object does not coincide with the real object, the mesh is repaired by the user through a graphical interface or automatically by using a picture. When the moving object affects other grids, the other grids are also repaired.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present principles relate generally to the field of augmented or extended reality applications. This document is also understood in the context of generating a scene description for an evolving extended reality scene and updating the evolving extended reality scene when changes are detected, for example for rendering the evolving extended reality scene on an end-user device such as a mobile device or a head-mounted display (HMD). BACKGROUND

[0002] This section is intended to introduce the reader to various aspects of art that can be related to various aspects of the present principles that are described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present principles. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art.

[0003] In augmented reality (AR) or extended reality (XR) experiences, virtual content is seamlessly inserted into the user’s real environment captured by using a camera or a video see-through device. Different tasks require a description of elements of the user’s real environment, such as positioning of virtual objects attached to AR anchors, consistent management of collisions between virtual and real objects, or consistent rendering of virtual and real objects, including occlusion and lighting / shading aspects. Some existing AR frameworks are able to capture, compute, store and load real environment data for AR experiences. Real environment data is computed from raw data of embedded sensors. An AR device can have several embedded sensors to scan the real environment, such as color camera(s) and / or light detection and ranging (LiDAR). The raw data generated can be point clouds, depth maps and / or pictures. When acquiring these data, an inertial measurement unit (IMU) is also needed to estimate the current pose of the AR device (i.e. position and orientation in the 3D space of the real environment). Based on these sensor raw data, a representation of the real environment is computed and the resulting real environment data can have various formats. For coherent collision processing and lighting, a mesh representation can be enough. However, for the definition of advanced anchoring and / or interaction, a semantic representation can be needed (e.g. “desk”, “laptop”, “screen”, “floor”, “ceiling”, “wall”). Mesh segmentation is needed for detecting and managing individual real objects. Therefore, a description of the real environment requires a scanning step and a semantic segmentation step. These steps are complex and time consuming.

[0004] When loading a scan of a real environment, some elements of the real environment can have been moved since the scan was performed. Re-scanning the objects that have moved (local scan) and starting a new segmentation is a performance and time-expensive approach. Moreover, the new scan can be performed from a different position and orientation than the first scan and a complex matching step is needed in addition to the semantic segmentation. There is a lack of a technique to allow a user to locally and optionally repair a first scan described in an AR / XR scene description file. SUMMARY

[0005] The following presents a simplified summary of the present principles to provide a basic understanding of some aspects of the present principles. This summary is not an extensive overview of the present principles. It is not intended to identify key or critical elements of the present principles. The following summary merely presents some aspects of the present principles in a simplified form as a prelude to the more detailed description provided below.

[0006] The present principles relate to a method for generating a scene description of an extended reality scene. The method comprises segmenting a scan of the extended reality scene into mesh objects. Then, for each object, a first state and a third state are assigned to the object when the object is fully visible in the scan. Otherwise, a second state is assigned to the object when the object is partially visible in the scan. The method checks whether the object is complete and, if so, replaces the partial mesh with a complete mesh of the object and assigns the third state to the object. The scene description is then generated as a tree of nodes and the third state is assigned to the nodes corresponding to the objects having the third state.

[0007] The present principles also relate to a device comprising a memory associated with a processor configured for implementing the above method. BRIEF DESCRIPTION OF DRAWINGS

[0008] The present disclosure will be better understood when read in conjunction with the appended figures, to which reference will now be made: - Figure 1 Figuratively illustrates the steps required for generating a description of an AR scene representing a real environment; - Figure 2 Illustrates a scene description formatted as a tree; - Figure 3 shows an example architecture of a device 30 that can be configured to implement the method for generating a scene description of an evolving extended reality scene and updating the evolving extended reality scene when a change is detected according to the present principles; - Figure 4 shows an example of an embodiment of the syntax of a stream when data is transmitted through a packet-based transmission protocol; - Figure 5Figuratively illustrates the initial steps of allowing insertion of a "canRepair" flag in a scene description file and construction of the scene description file according to the present principles; Figure 6 Figuratively illustrates a run-time processing model according to the present principles. DETAILED DESCRIPTION

[0009] The present principles will be described more fully hereinafter with reference to the accompanying drawings, in which example embodiments of the present principles are shown. The present principles may, however, be embodied in many alternate forms and should not be construed as limited to the examples set forth herein. Accordingly, while the present principles are susceptible to various modifications and alternative forms, specific examples thereof are shown by way of example in the drawings and will be described herein in detail. It should be understood that there is no intent to limit the present principles to the particular examples disclosed, but on the contrary, the present disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present principles as defined by the claims.

[0010] The terminology used herein is for the purpose of describing particular examples only and is not intended to be limiting of the present principles. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises", "comprising", "includes" and / or "including" when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Additionally, it is to be understood that when a element is referred to as being "responsive" or "connected" to another element, it can be directly responsive or connected to the other element, or indirectly responsive or connected to the other element through one or more other elements. Conversely, when an element is referred to as being "responsive to" or "connected to" one or more other elements, it can be directly responsive to or connected to the other element(s), or indirectly responsive to or connected to the other element(s) through one or more other elements. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items and can be abbreviated as " / ".

[0011] It will be understood that, although the terms first, second, etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the teachings of the present principles.

[0012] Although some of the diagrams include arrows on communication paths to show a primary direction of communication, it is to be understood that communication can occur in the opposite direction to the depicted arrows.

[0013] ​Some examples are described with respect to block and operational flow diagrams in which each block represents a circuit element, module, or portion of code that includes one or more executable instructions for implementing the specified logic function(s). It should also be noted that in other implementations, the function(s) noted in the blocks can occur out of the order noted in the figure. For example, two blocks shown in succession can in fact be executed substantially concurrently or the blocks can sometimes be executed in reverse order, depending on the functionality involved.

[0014] Reference herein to "in accordance with an example" or "in an example" means that a particular feature, structure, or characteristic described in connection with an example can be included in at least one implementation of the present principles. The appearances of the phrase "in accordance with an example" or "in an example" in various places in the specification are not necessarily all referring to the same example, nor are they necessarily mutually exclusive or alternative examples to one another.

[0015] Reference signs appearing in the claims are merely for illustration and should not be limiting to the scope of the claims. Although not explicitly described, the present examples and variants can be taken in any combination or sub-combination.

[0016] AR devices can have several embedded sensors to scan the real environment, such as color camera(s) and Light Detection and Ranging (LiDAR). The raw data generated can be point clouds, depth maps, and / or pictures. When acquiring these data, an Inertial Measurement Unit (IMU) is also needed to estimate the current pose of the AR device (i.e. position and orientation in the 3D space of the real environment). Based on these sensor raw data, a representation of the real environment is computed, and the resulting real environment data can have various formats. For coherent collision handling and lighting, a mesh representation can be enough. However, for the definition of advanced anchors and / or interactions, a semantic representation can be needed (e.g. "desk", "laptop", "screen", "floor", "ceiling", "wall"). Mesh segmentation is needed for detecting and managing individual real objects. Therefore, the description of a real environment requires a scanning step and a semantic segmentation step. These steps are complex and time consuming.

[0017] Figure 1 The required steps for generating an AR scene description representative of a real environment are illustrated. For example, the Microsoft Mixed Reality Framework has been developed for the HoloLens 2 device. It consists of a spatial computation module and a scene understanding module, the spatial computation module generating a mesh representation of the real environment as Figure 1the real environment depicted in FIG. 11, the scene understanding module detects and labels planar surfaces for placement of virtual content according to the OpenXR-based Mixed Reality Toolkit (MRTK) version 2.7. Apple’s ARKit uses the LiDAR scanner to create a mesh representation of the user’s real environment on a fourth-generation iPad Pro running iPad OS 13.4 or later. The mesh is further segmented, and multiple anchors called ARMeshAnchor are assigned to the resulting set of segmented meshes. As Figure 1 As illustrated in FIG. 12, semantic labeling of real objects that ARKit can identify is performed, such as ceiling, door, floor, seat, table, wall, and window labels. The Meta / Oculus framework has been developed for Meta Quest 2 and Meta Quest Pro devices. The scene understanding system provides a scene model, a representation of the user’s real environment. Currently, floor, ceiling, wall, desk, sofa, door frame, and window frame labels are supported. The scene model is generated from a scene capture system stream that lets the user walk around and capture their scene.

[0018] In each case, after the scan of the scene is captured, a segmentation operation can be initiated. Semantic segmentation of meshes has been a research topic explored for many years. Frameworks are characterized by the number of labels available in their database. The higher this number, the more difficult it will be to perform the segmentation operation in real time. Deep learning has achieved significant success in 3D segmentation, giving an overview of different frameworks. For example, ScanNet provides an online benchmark for 3D semantic segmentation evaluation on a test set, which is a reference for mesh segmentation frameworks. The computation of real environment data can be done locally in the AR device or remotely in a computing server. The process of completing the segmented mesh can be replaced, for example, with a 3D model of the segmented object in the scene, or with a complete model, provided that the complete model has been scanned (e.g., a chair in the room).

[0019] Thus, when incomplete scan information does not identify an object, but a complete model of the object can be obtained from a database of known objects, this complete model is used to complete the scan. This implies that if an object is rotated between sessions in the real world and a previously un-scanned side of the object is actually displayed, the object can still be identified because a complete model of the object can be obtained from a database of objects storing the correct features from all sides. According to the present principles, this process involves three steps. First, the scene region that is consistent with a previously stored partial scene model is effectively identified, i.e. a region that does not require any modification of the model. Second, the regions that can be repaired are effectively identified and these repairs are made automatically or with user assistance. Third, any region remaining after the first and second steps is reverted to partial re-scanning (with segmentation / object identification). This repair action is possible when the volume representation of the bounding box of a given object is less than a threshold, e.g. 90% or 8% or 1% of the volume of the bounding box of the 3D scene.

[0020] A scene is composed of real assets (e.g. a scan of a room) and virtual assets. Real assets are useful to increase the AR / XR experience. For example, a scan of a room allows to manage collisions between virtual objects and real objects. This can also be used for rendering purposes, e.g. not rendering real objects in the case of see-through XR devices, or ensuring coherent lighting between real and virtual objects, including shadow management. A scene can advantageously be represented by a graph composed of nodes corresponding to real and virtual objects. There are solutions to indicate the nature of each node of the scene description, e.g. by the flag “MPEG_node_nature” in the MPEG standard. A segmented scene is composed of a set of real nodes that improve the AR / XR experience. A scan of a scene can be stored and reloaded later for new AR / XR experiences, avoiding new scanning operations.

[0021] When a first scan of a real environment is loaded, some elements of the real environment can have been moved since the first scan was performed. According to the present principles, when the segmented mesh of the real environment is available from the first scan and the semantic segmentation process, the user is displayed the possibility to locally modify the scan and, optionally, provided with a handle to revert the mesh. According to the present principles, two main steps are performed: processing the segmented mesh to set the flag “canRepair” and inserting the flag “canRepair” in the scene description file. The flag “canRepair” is used by the rendering device at runtime. This flag will be used to initiate repair actions when the scene that was scanned has locally evolved. Thus, the interaction with the segmented mesh is enhanced and simplified. According to the present principles, the model is updated on the server side when possible.

[0022] Figure 2A scene description formatted as a tree is illustrated. After a first scan and the associated semantic segmentation process, the result of the segmentation is checked according to the present principles. The nodes of the scene description tree are labeled according to three states: complete (the segmentation of the object corresponds to a complete model), partial (only a part of the object is extracted) and unknown (the object is not identified; nothing will be done). A flag related to the previous state of the node in the scene description file (e.g. glTF file) is added to follow the changes. The use of a scene description file allows to distribute an AR / XR scene from a single server to several users. It also allows to reload the models of a scanned scene without having to perform the scan and the semantic segmentation again.

[0023] After the segmentation process, a state "complete", "partial" or "unknown" is associated with the nodes of the scene description tree, as illustrated in Figure 2 The state "complete" means that the object has been completely scanned and identified. The state "partial" represents that a region of the object has not been completely scanned but the object is identified. For example, a chair under a table has not been completely scanned but the label "chair" can be associated with this object mesh. This operation can be performed automatically by the segmentation framework or driven by the user. When the state "complete" is associated with a node, a repair action is possible, presenting the corresponding display. At this step, different factors like the size of the object are also taken into account.

[0024] For example, a repair action is possible when the volume of the bounding box of a given object represents less than a threshold, for example 10% or 50% or 2% of the volume of the bounding box of the 3D scene.

[0025] Figure 5 The initial steps to allow the insertion of a "canRepair" flag in the scene description file and to build the scene description file are illustrated schematically. In step 51, the scene is scanned and in step 52, the semantic segmentation is performed. Then in step 53, the segmentation state is associated with the nodes. The state "complete" and the state "partial" are related to a threshold (rate of identification). They are determined by comparison with existing models in a database. If the determined state is "partial", a complete check step 54 is performed. If this check step is successful, the object associated with the node is replaced by a complete object in step 55. When the determined state is "complete" or when the complete check is successful, in step 56, the canRepair flag is associated with the object. Another segmented object at step 52 is then considered until each object is considered.

[0026] Using a scene description file allows export to a server to share the file. As part of the scene description, parameters can be added to the standard as shown in the following example based on MPEG-1 scene description. MPEG-I scene description is a framework to support additional scene description features using Khronos glTF extension mechanism. The semantics of MPEG_node_nature are provided in the following table. .

[0027] According to the present principles, the semantics of MPEG_node_nature are extended as follows: where the usage "M" stands for "mandatory" and "O" stands for "optional".

[0028] As described herein above, mobile objects can expose holes in the mesh. A fallback mesh can have been provided to solve this problem. In the following example scene description, meshes 0 and 4 are affected. In the "meshFallback" line, the value -1 indicates no fallback. In this example, only mesh 0 has a fallback.

[0029] Figure 6 A runtime processing model according to the present principles is illustrated diagrammatically. At the beginning of the AR / XR experience, the application loads the scene description file at step 60. If at step 61, a split scan of the real scene with an extension including the flag "canRepair" is read in the scene description, then the user can apply a repair action 62 when the scan and the scene are not perfectly aligned.

[0030] When the AR / XR application loads the first (previous) scan, if the mesh of the object is not consistent with the real object, a graphical interface can be displayed to allow the user to move the mesh to superimpose it to the real object. If there is an uplink, the update of the pose of the mesh can be sent to the server, in which case only the TRS (translation / rotation / scale values) are transmitted. This operation can have an impact on other parts of the mesh. In this case, a fallback operation is implemented. For example, when a chair is placed on a floor, there will be holes in the mesh of the floor at the contact points with the legs of the chair. Thus, if the mesh of the chair is displaced, the holes will be visible. In the case of the floor, the mesh can be replaced by a 3D model of a plane. For other cases, such as re-illumination of the real scene with virtual light, it can be necessary to add textures to the repaired mesh. When the user identifies a discrepancy between the loaded model and the real scene, he can initiate the process of a repair action. In this embodiment, the action is performed based on visual observation. In another embodiment, an automatic comparison mode is considered by matching pictures captured from known viewpoints with the textured mesh. In a third embodiment, an automatic comparison of the mesh of the first scan with the mesh of the second (new) scan is performed.

[0031] In an embodiment, the repair process is used when the object is no longer part of the scene. In this case, the mesh is deleted. In another embodiment, the fallback can be to use a geometric primitive (sphere, cube, disc...) as a model of the real object.

[0032] Figure 3 An example architecture of a device 30, which can be configured to implement the method for generating an evolving extended reality scene and updating the evolving extended reality scene when a change is detected according to the present principles, is shown. The device can implement the method described with respect to Figure 5 and Figure 6 The encoder and / or the decoder can each be a device according to the architecture of Figure 3 , linked together, for example, via their buses 31 and / or via I / O interfaces 36. The device 30 comprises the following elements linked together by data and address buses 31 : - a microprocessor 32 (or CPU), which is, for example, a DSP (or Digital Signal Processor); - a ROM (or Read Only Memory) 33; - a RAM (or Random Access Memory) 34; - a storage interface 35; - an I / O interface 36 for receiving data to be transmitted from an application; and - a power supply, for example a battery.

[0033] According to the example, the power supply is external to the device. In each of the mentioned memories, the term "register" used in the specification can correspond to a small (a few bits) area or a very large area (e.g., the entire program or a large amount of received or decoded data). ROM 33 includes at least the program and parameters. ROM 33 can store algorithms and instructions to execute techniques according to these principles. When powered on, CPU 32 uploads the program to RAM and executes the corresponding instructions.

[0034] RAM 34 includes in registers the program executed by CPU 32 and uploaded after device 30 is turned on, input data in registers, intermediate data of different states of the method in registers, and other variables in registers used to execute the method.

[0035] The embodiments described herein can be implemented, for example, in methods or processes, apparatus, computer program products, data streams, or signals. Even if discussed only in the context of a single embodiment (e.g., discussed only as a method or apparatus), embodiments of the discussed features can be implemented in other forms (e.g., programs). Apparatus can be implemented, for example, in suitable hardware, software, and firmware. Methods can be implemented, for example, in apparatus, such as, for example, a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, computers, cellular phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0036] According to the example, device 30 belongs to a set that includes the following items: - mobile device; - Communication equipment; - Gaming equipment; - Tablet (or tablet computer); - Laptop; - Still image camera; - Video camera; - Encoding chip; - Servers (such as broadcast servers, video-on-demand servers, or web servers).

[0037] Figure 4 An example of an implementation of the syntax for streams is shown when data is transmitted via a packet-based transport protocol. Figure 4Example structure 4 is shown for a stream that encodes a scene description of an extended reality scene according to this principle. This structure is contained within a container that organizes the stream into individual syntax elements. The structure may include a header section 41, which is a dataset common to each syntax element of the stream. For example, the header section includes metadata about the syntax elements, describing the properties and roles of each syntax element. The structure includes a payload, which includes syntax element 42 and at least one syntax element 43. Syntax element 42 includes the scene description itself, such as an organized tree of nodes. Syntax element 43 is part of the payload of the data stream and may include data referenced by the nodes of the scene description. There may be syntax elements 43 for different types of data, such as one for nodes, one for meshes, one for textures, and so on.

[0038] The embodiments described herein can be implemented, for example, in methods or processes, apparatus, computer program products, data streams, or signals. Even if discussed only in the context of a single embodiment (e.g., discussed only as a method or apparatus), embodiments of the discussed features can be implemented in other forms (e.g., programs). Apparatus can be implemented, for example, in suitable hardware, software, and firmware. Methods can be implemented, for example, in apparatus, such as, for example, a processor, which generally refers to a processing device, including, for example, a computer, microprocessor, integrated circuit, or programmable logic device. Processors also include communication devices, such as, for example, smartphones, tablets, computers, mobile phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate information communication between end users.

[0039] The various processes and features described herein can be implemented in a wide variety of different equipment or applications, particularly those associated with data encoding, data decoding, view generation, texture processing, and other processing of image and associated texture and / or depth information. Examples of such equipment include encoders, decoders, post-processors that process the output from the decoder, pre-processors that provide input to the encoder, video encoders, video decoders, video codecs, network servers, set-top boxes, laptops, personal computers, cellular phones, PDAs, and other communication devices. It should be understood that the equipment can be mobile and even mounted in mobile vehicles.

[0040] Furthermore, the method can be implemented by instructions executed by a processor, and such instructions (and / or data values ​​generated by the implementation) can be stored on a processor-readable medium, such as an integrated circuit, software carrier, or other storage device, such as a hard disk, floppy disk (“CD”), optical disk (such as a DVD, commonly referred to as a digital multifunction disk or digital video disk), random access memory (“RAM”), or read-only memory (“ROM”). These instructions can form an application tangibly embodied on the processor-readable medium. The instructions can be in the form of, for example, hardware, firmware, software, or a combination thereof. The instructions can be found, for example, in an operating system, a standalone application, or a combination of both. Thus, a processor can be characterized as, for example, a device configured to implement the process and a device comprising a processor-readable medium (such as a storage device) having instructions for implementing the process. In addition to or instead of the instructions, the processor-readable medium can store data values ​​generated by the implementation.

[0041] As will be apparent to those skilled in the art, implementations can generate a wide variety of signals that are formatted to carry, for example, information that can be stored or transmitted. This information may include, for example, instructions for performing a method, or data generated by one of the described implementations. For example, a signal may be formatted to carry as data rules the syntax for writing or reading the described embodiments, or as data actual syntax values ​​written by the described embodiments. Such a signal may be formatted as, for example, electromagnetic waves (e.g., using the radio frequency portion of the spectrum) or baseband signals. Formatting may include, for example, encoding a data stream and modulating a carrier wave using the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is well known that signals can be transmitted over a wide variety of wired or wireless links. The signal may be stored on a processor-readable medium.

[0042] Many embodiments have been described. However, it should be understood that various modifications can be made. For example, elements of different embodiments can be combined, supplemented, modified, or removed to produce other embodiments. Furthermore, those skilled in the art will understand that other structures and processes can replace those disclosed, and the resulting embodiments will perform at least substantially the same functions in at least substantially the same manner to achieve at least substantially the same results as the disclosed embodiments. Therefore, these and other embodiments are contemplated in this application.

Claims

1. A method for generating a scene description of an extended reality scene, the method comprising: - Obtain a scan of an extended reality scene segmented within a grid object; - For each object, • When the object is fully visible in the scan, assign the object a flag indicating that the object can be repaired; • When the scan of the object is partially visible in the scan, check whether the object is complete, and if so, replace the partial mesh with the complete mesh of the object, and assign a flag indicating that the object can be repaired to the object; as well as - Generate a scene description as a node tree, and assign the flag to the node corresponding to the object with the flag.

2. The method of claim 1, wherein semantic segmentation of the scan is performed to determine whether the object is fully visible.

3. The method according to claim 2, wherein, Based on the semantic segmentation, the determination of whether an object is fully visible is performed by comparing it with a model in the database.

4. The method according to any one of claims 1 to 3, wherein, When the scan of the object is fully visible, the node is assigned a first state (complete); when the scan of the object is partially visible, the node is assigned a second state (partial); and in other cases, the node is assigned a third state (unknown).

5. The method according to any one of claims 1 to 4, wherein, The nodes of the node tree include information indicating whether the node corresponds to a real object.

6. The method according to any one of claims 1 to 5, wherein, The scene description is encoded in the data stream.

7. An apparatus for generating a scene description of an extended reality scene, the apparatus comprising a processor configured to: - A scan of the extended reality scene segmented within the grid object; - For each object, • When the object is fully visible in the scan, assign the object a flag indicating that the object can be repaired; • When the object is partially visible in the scan, check whether the object is complete, and if so, replace the partial mesh with the complete mesh of the object, and assign a flag indicating that the object can be repaired to the object; as well as - Generate a scene description as a node tree, and assign the flag to the node corresponding to the object with the flag.

8. The device according to claim 7, wherein, The processor is configured to perform semantic segmentation on the scan to determine whether an object is fully visible.

9. The device according to claim 8, wherein, Based on the semantic segmentation, the determination of whether an object is fully visible is performed by comparing it with a model in the database.

10. The device according to any one of claims 7 to 9, wherein, When the scan of the object is fully visible, the node is assigned a first state (complete); when the scan of the object is partially visible, the node is assigned a second state (partial); and in other cases, the node is assigned a third state (unknown).

11. The device according to any one of claims 7 to 10, wherein, The nodes of the node tree include information indicating whether the node corresponds to a real object.

12. The device according to any one of claims 7 to 11, wherein, The scene description is encoded in the data stream.