Object regions for proximity triggers in augmented reality scene descriptions

By incorporating geometric primitives in AR scene descriptions, the method improves interaction accuracy and realism in AR environments, enabling complex triggering scenarios and enhanced user experiences.

JP2026509795APending Publication Date: 2026-03-25INTERDIGITALCE PATENT HLDG SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-20
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing augmented reality (AR) scene descriptions lack efficient methods for defining and triggering interactions between virtual objects based on proximity and collision, particularly in complex scenarios involving articulated bodies and varied geometric representations.

Method used

The introduction of additional geometric primitives, such as spheres, boxes, and cylinders, to virtual objects in the scene graph, which define 3D regions for proximity and collision triggers, enabling more accurate and complex interactions.

Benefits of technology

Enhances the realism and interactivity of AR experiences by allowing precise triggering of actions based on object proximity and collision, supporting social behaviors, privacy considerations, and improved haptic feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509795000001_ABST
    Figure 2026509795000001_ABST
Patent Text Reader

Abstract

Methods, devices, and data streams are provided for generating, transmitting, and decoding scene descriptions of augmented reality scenes. According to this principle, the scene description graph contains information indicating how proximity triggers should be evaluated to link nodes and trigger actions on virtual objects. The proximity conditions are based on a set of geometric primitives plus boundary values ​​that modify the associated 3D region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This principle generally relates to the fields of augmented reality scene description and augmented reality scene rendering. In particular, this principle relates to the description of proximity triggers between regions of objects in scene description. This document is also understood in the context of the format and playback of augmented reality applications when rendered on an end-user device such as a mobile device or a head-mounted display (HMD) such as see-through glasses.

Background Art

[0002] This section is intended to introduce the reader to various aspects of the art, which may be related to various aspects of the present disclosure described and / or claimed below. This discussion is thought to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the principle. Accordingly, it should be understood that these descriptions are to be read from this perspective and should not be read as an admission of prior art.

[0003] Augmented reality (XR) is a technology that enables interactive experiences in which real-world environments and / or video content are enhanced by virtual content that can be defined across multiple sensory modalities, including sight, hearing, and touch. During the runtime of an application, virtual content (e.g., 3D content or audio / video files) is rendered in real time in a manner consistent with the user context (environment, viewpoint, device, etc.). Scene graphs (e.g., those proposed by Khronos / glTF, and their extensions defined in the MPEG scene description format or Apple / USDZ) are possible ways of representing the content being rendered. These combine, on the one hand, a declarative description of a scene structure that links nodes describing virtual objects and actions on these virtual objects, and on the other hand, a binary representation of the virtual content. Binary representations of real objects may be included in the nodes of the scene description. Such representations of real objects may be obtained, for example, by scanning the real environment. The scene description framework ensures that timed media and their corresponding associated virtual content are available at any point during the rendering of the application. Scene descriptions can also carry data at the scene level, describing how scene objects behave and interact at runtime for an immersive XR experience. Some actions on virtual objects are triggered by proximity triggers. A general description of proximity triggers is a technical challenge because they involve a wide range of interactions between virtual objects, e.g., objects controlled by a human user. [Overview of the project]

[0004] Below is a simplified overview of the Principle to provide a basic understanding of some aspects of it. This overview is not a comprehensive summary of the Principle. It is not intended to identify the main or important elements of the Principle. The following overview merely presents some aspects of the Principle in a simplified form as a prelude to the more detailed explanation provided below.

[0005] This principle relates to a method for encoding an augmented reality (XR) scene description into a data stream. The method involves linking nodes to obtain a scene graph belonging to the XR scene description, where the nodes describe objects in the augmented reality (XR) scene and include triggers. Then, for all nodes in the scene graph that include triggers, which are collision or proximity triggers, a description of at least one primitive is added to the trigger. The primitive comprises the primitive type, a description of the 3D region depending on the primitive type, and boundary values ​​indicating the extent of the 3D region. The modified XR scene description is encoded into a data stream. According to this principle, primitives belong to at least two groups of primitives, including spheres, boxes, round boxes, box frames, tori, capped tori, links, infinite cylinders, capped cylinders, rounded cylinders, cones, infinite cones, capped cones, rounded cones, planes, hexagonal prisms, triangular prisms, capsules, lines, solid angles, truncated spheres, truncated hollow spheres, ellipsoids, rhombuses, octahedrons, pyramids, triangles, and quadrilaterals.

[0006] This principle relates to a device comprising a processor and memory associated with the processor, configured to carry out the method described above.

[0007] This principle relates to a method for rendering an augmented reality (XR) scene. This method involves linking nodes and obtaining a scene graph belonging to a scene description. Nodes describe objects in the XR scene and are associated with triggers. When a trigger is a collision trigger or proximity trigger, it includes a description of at least one primitive, including the primitive type, a description of a 3D region depending on the primitive type, and boundary values ​​indicating the extent of the 3D region. At least one trigger is activated when a virtual object enters the region of the XR scene within the extent of the 3D region. At least one action is associated with at least one trigger. At least one action is executed when at least one trigger is activated. Actions belong to a non-exhaustive group of actions, including haptic feedback, physics simulation, and audio feedback.

[0008] This principle relates to a device comprising a processor and memory associated with the processor, configured to carry out the method described above. [Brief explanation of the drawing]

[0009] The following description will help to better understand this disclosure and reveal other specific features and advantages, and this description refers to the attached drawings. [Figure 1] An example graph illustrating the description of an augmented reality scene based on this principle is shown. [Figure 2] This demonstrates an example of human-user interaction with virtual objects via a renderer-side mechanical device, i.e., when an XR experience is being performed. [Figure 3] An exemplary architecture of an XR processing engine that may be configured to implement the methods described in relation to this principle is shown. [Figure 4] An example of an embodiment of the syntax for a data stream encoding an augmented reality scene description based on this principle is shown. [Figure 5] This demonstrates the triggering of proximity triggers based on the volume of the sphere. [Modes for carrying out the invention]

[0010] The principle is described more fully below with reference to the accompanying drawings illustrating examples of the principle. However, the principle may be embodied in many alternative forms and should not be construed as being limited to the examples expressed herein. Thus, the principle is open to various modifications and alternative forms, specific examples of which are shown as examples in the drawings and described in detail herein. However, there is no intention to limit the principle to the specific forms disclosed, but rather this disclosure should be understood to encompass all modifications, equivalents, and alternatives that fall within the spirit and scope of the principle as defined in the claims.

[0011] The terminology used herein is intended solely to illustrate specific examples and is not intended to limit the principles. Where used herein, the singular forms “a,” “an,” and “the” are intended to include the plural form unless otherwise specified in the context. Where used herein, the terms “equip,” “equip,” “contain,” and / or “contain” identify the presence of the described feature, integer, step, action, element, and / or component, but do not exclude the presence or addition of one or more other features, integers, steps, actions, elements, components, and / or groups thereof. Furthermore, where an element is referred to as “corresponding to” or “connected to” another element, it may directly correspond to, be able to connect to, or interpose with the other element. In contrast, where an element is referred to as “directly corresponding to” or “directly connected to” another element, there is no interposing element. Where used herein, the terms “and / or” include any combination of one or more of the enumerated items relating to each other and may be abbreviated as “ / .”

[0012] In this specification, terms such as "first" and "second" may be used to describe various elements, but it should be understood that these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, without deviating from the teachings of this principle, the first element may be called the second element, and similarly, the second element may be called the first element.

[0013] Some diagrams include arrows on the communication path to indicate the primary direction of communication, but please understand that communication may occur in the opposite direction to the depicted arrow.

[0014] Some examples are illustrated with block diagrams and operation flowcharts, where each block represents a circuit element, module, or portion of code containing one or more executable instructions to implement a specified logical function. Note that in other implementations, the functions described in a block may occur in a different order than those listed. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or blocks may sometimes be executed in reverse order depending on the functions they involve.

[0015] Any reference in this specification to “by example” or “in one example” means that a particular feature, structure, or characteristic described in relation to the example may be included in at least one implementation of the principle. The occurrence of the phrase “by example” or “in one example” in various places in this specification does not necessarily refer to the same example, nor do separate or alternative examples necessarily exclude each other.

[0016] Reference numerals appearing in the claims are illustrative only and shall not limit the scope of the claims. Although not expressly described, these examples and modifications may be used in any combination or partial combination.

[0017] Figure 1 shows an exemplary scene graph 10 of an augmented reality scene description. In this example, the scene graph may include descriptions of real objects, e.g., a "plane horizontal plane" (which could be a table or a road), and descriptions of virtual objects 12, e.g., an animation of a car. The scene graph is organized as an array of nodes 10. Nodes can be linked to child nodes to form a scene structure 11. Nodes can carry descriptions of real objects (e.g., semantic descriptions) or descriptions of virtual objects. In the example in Figure 1, node 101 describes a virtual camera located within a 3D volume of the XR application. Node 102 describes a virtual car and includes an index of the car's representation, e.g., an index in an array of 3D meshes. This representation may be, for example, obtained by scanning a real environment, or it may be, for example, a model generated by a content creator. This representation corresponds to the object described in this node and is therefore localized within the 3D scene.

[0018] A scene description can include multiple arrays containing descriptions of various aspects of the scene, such as an array containing instructions for several animations, an array containing meshes for several virtual or real objects, or an array containing material descriptions. It can also include descriptions of actions to apply to these virtual objects when events occur when playing the XR application, for example, when an object controlled by a human user touches and presses or grabs a virtual object, or more generally, when a virtual object collides with or approaches another virtual object or the boundary of the scene. Node 103 is a child of node 102 and contains a description of one wheel of the car. Similarly, it contains an index to the wheel's 3D mesh. Since the scale, position, and orientation of objects are described in the scene nodes, the same 3D mesh can be used for several objects in the 3D scene. In the example in Figure 1, the same mesh can be used for the four wheels of the virtual car. Scene graph 10 also includes nodes that describe the spatial relationships between virtual objects.

[0019] Digital human representation is an active area of ​​research, initially proposed for the visualization and playback of video content. In recent years, research has been driven by animated and manipulable digital human representations. This can be achieved through skeletal skinning techniques or learning-based methods. Advances in virtual environment development and animation techniques make human-machine interaction more natural and intuitive. Figure 2 shows an example of human-user interaction with a virtual object 21 via a renderer-side mechanical device 22, i.e., when an XR experience is being performed. This interaction is activated (i.e., an associated action is triggered) when a user-controlled virtual object 23 approaches an area of ​​interest containing the virtual object 21. The triggered action implements different functions such as grasping for object manipulation, haptic feedback, sound effects, or collision detection. Other types of triggers, e.g., timer-based triggers or triggers based on combinations of user actions, can coexist within the scene description. The renderer obtains a scene description generated according to this principle. When a trigger is activated, the associated action is executed.

[0020] Figure 5 illustrates the triggering of proximity triggers based on the volume of a sphere. While proximity triggers based on the volume of a sphere are easy to compute, they are limited in practice. In reality, virtual objects can have different areas of interest with different geometric representations that consequently affect user interaction. The volume of the sphere of the virtual object 21 in Figure 2 is described by 3D coordinates with a center 51 and radius 52. This description defines the outer boundary sphere 53. When the volume of the sphere of a virtual object, for example, object 23 in Figure 2, is initially at position 54 and intersects with the outer boundary sphere 53 at position 55, the associated proximity trigger is activated and the corresponding action is applied to both objects. The use of this single primitive is limited when dealing with articulated bodies and is limited when using simple objects such as the rabbit character 21 in Figure 2.

[0021] According to this principle, additional geometric primitives are fixed to virtual objects and / or parts of virtual objects. These additional primitives constitute valuable information that facilitates the rendering engine in generating accurate interactions between virtual objects. These primitives enable the leveraging of scenarios common in the real world and the world of social technologies, such as social behavior, time-based animation, and privacy issues. According to this principle, geometric primitives are defined and added to triggers on nodes in the scene graph describing virtual objects (even the root node encompassing the entire scene) with the intention of activating triggers (for example, collision or proximity triggers where an action is performed when two primitives intersect, and this action can represent any activity such as collision, proximity, social behavior, privacy, ability / characteristics, acoustics, or tactile sensation).

[0022] According to this principle, the representation of the interaction space of virtual objects with new primitives is intended to be compatible with existing scene description formats. Here, the format proposed according to this principle follows the glTF format and is compatible with current MPEG efforts to extend glTF using MPEG extensions. However, its meaning and usage are general and can be encoded in any other format (e.g., XML, USD, ...).

[0023] The volumes to be contained may be of different natures. According to this principle, there are at least two primitives that define volumes that may have different roles. A primitive is defined by its type (e.g., a box (=cube in 3D), a square (in 2D), a cylinder, a capsule, a sphere (in 3D), or a signed distance field). A primitive is also defined by a set of parameters that depend on its type. The values ​​of the primitive's parameters, and therefore its default values, depend on the primitive's type and the virtual object to which the neighboring primitive is applied. For example,

[0024] [Table 1]

[0025] This list of possible primitives is not exhaustive.

[0026] A primitive is a geometric item that defines a region in the 3D space of an XR scene that surrounds a given virtual object (or part of a virtual object). Different types of regions are possible according to the action or behavior to which the proximity triggers associated with them are linked. For example, a "social region" corresponds to a given social distance of any object from this given object. It can be general (for example, the default distance can be set to 1m or 4 feet) or user-defined. On the renderer side, when interacting with another object, a trigger is activated if an intersection is detected between the region of the first object and the region of the given object. Thus, there is a difference between the region defined by the primitive and the region that triggers the proximity action. In other words, there is a distance value (called the boundary) described in the parameters of the relevant trigger in the scene description that indicates at what distance from the centroid of the primitive the trigger is activated, i.e., when this social region becomes active. A "contact region" is a direct collision detection trigger. For this type of region, the boundary is set to zero. This means that a proximity trigger is activated when a first primitive intersects with a given primitive. Such areas can trigger haptic feedback, physical simulation, audio feedback, or any other event. Also, if necessary, if two contact areas overlap, for example, when an avatar's "vital space" intersects with the "vital space" of a social event, it can affect the avatar's attributes, which can trigger actions in sports activities and allow all users / objects to perform actions that were previously impossible, such as flying, jumping, or speaking. A "legal space" can be used, for example, in a legal permission scenario. The permissible displacement of a virtual object may be restricted for legal reasons (age, restrictions, access rights, etc.). Such a space represents the limited space in which this avatar can move. An "experience space" corresponds to the space in which an object is permitted to move, taking obstacles into consideration, and corresponds to the generated space.In a related manner, the “parent region” may be configured to protect children and young adults by restricting their interaction with permitted content. For any type of region, default boundary values ​​may be defined or set in a table associated with the scene description.

[0027] To demonstrate a possible syntax for primitives for proximity triggers in scene descriptions, an extension to the glTF node "MPEG_node_interactivity" element is proposed below. Since the MPEG interactivity glTF extension enables triggering in collision and proximity situations, the proposed extension contributes to an extension of the new primitive description for proximity triggers in the node "MPEG_node_interactivity". The general node implementation form can also be applied to avatar representations, allowing for the addition of interactivity constraints to avatars and their respective elements.

[0028] The glTF node element "MPEG_node_interactivity" property is extended to define an additional geometric primitive for "TRIGGER_PROXIMITY" to define the region of interactivity. This element extension describes an additional sub-primitive of the node characterized as an interactivity element. This element provides the client application with additional information about the area of ​​interaction between the scene and the objects within it. This can be applied as, for example, a box, cylinder, sphere, or distance function object. For example, the surrounding region of an avatar could be defined as a cubic area for trigger activity, for example, to allow only interactive object trigger actions that consequently affect the avatar. The avatar's node body element can be individually set with different primitives to interact with scene objects within this primitive boundary region, for example, the avatar could use its arms to grab or grasp an object.

[0029] Several definitions need to be set for the scene description.

[0030] For the MPEG_node_interactivity_primitive extension description:

[0031] [Table 2]

[0032] Semantic description or primitive property:

[0033] [Table 3]

[0034] The semantics of primitive types:

[0035] [Table 4]

[0036] The semantics for each individual domain are provided in the table below, where "M" indicates "required".

[0037] [Table 5]

[0038] The semantic descriptions of each primitive available in SDF are as follows:

[0039] [Table 6]

[0040] The following glTF is an example (not exhaustive) of instantiations of "MPEG_node_interactivity_primitives" in clients that support "MPEG_node_interactivity," otherwise it reverts to the standard "MPEG_node_interactivity."

[0041] Example of extending the "MPEG_node_interactivity" property:

[0042] [Table 7]

[0043] Example of extending the "MPEG_node_interactivity" extension property:

[0044] [Table 8]

[0045] [Table 9]

[0046] [Table 10]

[0047] [Table 11]

[0048] Figure 3 shows an exemplary architecture of an XR processing engine 30 that can be configured to implement this principle. Devices according to the architecture of Figure 2 are linked to other devices via their bus 31 and / or via the I / O interface 36.

[0049] Device 30 comprises the following elements, which are linked to each other by a data and address bus 31: - For example, a DSP (or digital signal processor), processor 32 (or CPU), -ROM (or read-only memory) 33, -RAM (or random access memory) 34, -Storage interface 35, - I / O interface 36 that receives data sent from the application, and - Power source (not shown in Figure 2), e.g., battery.

[0050] For example, the power supply is external to the device. In each of the above-mentioned memories, the word "register" as used herein may correspond to a small area (a few bits) or a very large area (e.g., an entire program or a large amount of received or decoded data). ROM33 contains at least a program and parameters. ROM33 may store algorithms and instructions for performing the technology according to this principle. When switched on, CPU32 uploads the program into RAM and executes the corresponding instructions.

[0051] RAM34 contains registers for a program that is executed by CPU32 and uploaded after device30 is switched on, input data, intermediate data for different states of the method, and other variables used to execute the method.

[0052] The implementations described herein may be implemented, for example, in methods or processes, apparatus, computer program products, data streams, or signals. Even when considered only in the context of a single implementation (for example, only as a method or device), the implementation of the features considered may also be implemented in other forms (for example, programs). Apparatus may be implemented, for example, in appropriate hardware, software, and firmware. These methods may be implemented, for example, in apparatus, and may be implemented in processing devices in general, such as processors, including computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices, such as computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the communication of information between end users.

[0053] Device 30 is linked, for example, via a bus 31 to a set of sensors 37 and a set of rendering devices 38. The sensors 37 may be, for example, a camera, microphone, temperature sensor, inertial measuring device, GPS, humidity measuring sensor, IR or UV light sensor, or wind sensor. The rendering devices 38 may be, for example, a display, speaker, vibrator, heat sensor, fan, etc.

[0054] For example, device 30 is configured to implement the method according to the present principle and belongs to a set that includes: - Mobile devices, -Communication devices, - Game devices, - A tablet (or tablet computer), -Laptop, -Still camera and, - Video camera.

[0055] Figure 4 shows an example of an embodiment of the syntax for a data stream encoding an augmented reality scene description according to this principle. Figure 3 shows an exemplary structure 4 of an XR scene description. The structure resides within a container that organizes the stream into individual syntax elements. This structure may include a header section 41, which is a set of data common to all syntax elements of the stream. For example, the header section includes some metadata about the syntax elements, describing their respective properties and roles. The structure also includes a payload containing elements of syntax 42 and elements of syntax 43. The syntax elements 42 comprise data representing media content items described in the scene graph nodes related to virtual elements. Images, meshes, and other raw data may be compressed according to a compression method. The elements of syntax 43 are part of the data stream payload and contain data encoding the scene description as described according to this principle.

[0056] The implementations described herein may be implemented, for example, in methods or processes, apparatus, computer program products, data streams, or signals. Even when considered only in the context of a single implementation (for example, only as a method or device), the implementation of the features considered may also be implemented in other forms (for example, programs). Apparatus may be implemented, for example, in appropriate hardware, software, and firmware. These methods may be implemented, for example, in apparatus, and may be implemented in processing devices in general, such as processors, including, for example, computers, microprocessors, integrated circuits, or programmable logic devices. Processors also include communication devices such as, for example, smartphones, tablets, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate the transmission of information between end users.

[0057] The various processes and features described herein may be embodied in a variety of different devices or applications, specifically, for example, in devices or applications associated with data encoding, data decoding, view generation, texture processing, and other image processing, as well as related texture information and / or depth information. Examples of such devices include encoders, decoders, post-processors that process the output from decoders, pre-processors that supply inputs to encoders, video coders, video decoders, video codecs, web servers, set-top boxes, laptops, personal computers, mobile phones, PDAs, and other communication devices. As should be obvious, the devices may be portable and may even be mounted on mobile vehicles.

[0058] Furthermore, the method may be implemented by instructions executed by the processor, and such instructions (and / or data values ​​generated by the implementation) may be stored in a processor-readable medium such as an integrated circuit or a software carrier, or in other storage devices such as a hard disk, a compact diskette ("CD"), an optical disc (such as a DVD, often referred to as a digital multipurpose disc or digital video disc), random access memory ("RAM"), or read-only memory ("ROM"). The instructions may form an application program that is tangibly embodied in the processor-readable medium. The instructions may be in hardware, firmware, software, or a combination of the two. The instructions may be found in an operating system, a separate application, or a combination of the two. Thus, a processor may be characterized as both, for example, a device configured to perform processing and a device including a processor-readable medium (such as a storage device) having instructions for performing processing. Furthermore, the processor-readable medium may store data values ​​generated by the implementation in addition to, or instead of, the instructions.

[0059] As will be apparent to those skilled in the art, implementations can generate a variety of signals formatted to carry information, which can, for example, be stored or transmitted. The information may include, for example, instructions to perform a method, or data generated by one of the implementations described. For example, a signal may be formatted as data to convey rules for writing or reading the syntax of the embodiment described, or as data to convey actual syntax values ​​described from the embodiment described. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a wide variety of different wired or wireless links, as is known. The signal may be stored in a processor-readable medium.

[0060] Several implementations have been described. Nevertheless, it will be understood that various modifications are possible. For example, elements of different implementations may be combined, supplemented, modified, or deleted to produce other implementations. Furthermore, those skilled in the art will understand that other structures and processes may be substituted for the disclosed structures and processes, and that the resulting implementations may perform at least substantially the same functions(s) in at least substantially the same manner(s) as the disclosed implementations to achieve at least substantially the same results(s). Accordingly, these and other implementations are conceived in this application.

[0061] Appendix A glTF schema Semantics MPEG_node_interactivity_primitive object

[0062] [Table 12]

[0063] MPEG_node_interactivity_primitive.objects object

[0064] [Table 13-1]

[0065] [Table 13-2] (Continuation of the table above)

[0066] MPEG_node_interactivity_primitive.objects.square_region object

[0067] [Table 14]

[0068] MPEG_node_interactivity_primitive.objects.box_region object

[0069] [Table 15]

[0070] MPEG_node_interactivity_primitive.objects.capsule_region object

[0071] [Table 16]

[0072] MPEG_node_interactivity_primitive.objects.cylinder_region object

[0073] [Table 17]

[0074] MPEG_node_interactivity_primitive.objects.sphere_region object

[0075] [Table 18]

Claims

1. A method for encoding an augmented reality (XR) scene description into a data stream, wherein the method is - Linking nodes and obtaining a scene graph belonging to the XR scene description, wherein the nodes describe objects in the XR scene and are associated with triggers, - For a node in the scene graph associated with a trigger that is a collision trigger or a proximity trigger, add a description of at least one primitive to the trigger, wherein the primitive is: • Primitive types, - Description of the 3D region according to the type of primitive, - Includes boundary values ​​indicating the range of the 3D region, - Encoding the XR scene description into the data stream, Methods that include...

2. The method according to claim 1, wherein the primitive is a geometric item that defines a region of the XR scene surrounding a virtual object or a part of a virtual object.

3. The method according to claim 1 or 2, wherein the boundary value indicates the distance from the centroid of the at least one primitive at which the trigger is activated.

4. The method according to any one of claims 1 to 3, wherein the default boundary value is defined for the type of region associated with the primitive.

5. The method according to any one of claims 1 to 4, wherein the primitive belongs to at least two groups of primitives, including spheres, boxes, round boxes, box frames, tori, capped tori, links, infinite cylinders, capped cylinders, rounded cylinders, cones, infinite cones, capped cones, rounded cones, planes, hexagonal prisms, triangular prisms, capsules, lines, solid angles, truncated spheres, truncated hollow spheres, ellipsoids, rhombuses, octahedrons, pyramids, triangles, and quadrilaterals.

6. A method for rendering an augmented reality (XR) scene, wherein the method is: - Linking nodes and obtaining a scene graph belonging to the XR scene description, wherein each node describes an object in the XR scene and is associated with a trigger, and the trigger, when it is a collision trigger or proximity trigger, includes a description of at least one primitive, and the primitive is • Primitive types, - Description of the 3D region according to the type of primitive, - Includes boundary values ​​indicating the range of the 3D region, - The at least one trigger is activated when the virtual object enters the region of the XR scene within the range of the 3D region. A method wherein at least one action is associated with the at least one trigger, the at least one action is performed when the at least one trigger is activated, and the at least one action belongs to a group of actions including haptic feedback, physical simulation, and audio feedback.

7. A device for encoding an augmented reality (XR) scene description into a data stream, comprising a processor and a memory associated with the processor, wherein the processor - Link nodes and obtain the scene graph belonging to the XR scene description, the nodes describe the objects of the XR scene and are associated with triggers, - For a node in the scene graph that includes a trigger which is a proximity trigger, add a description of at least one primitive to the node, and the primitive is, • Primitive types, - Description of the 3D region according to the type of primitive, - Includes boundary values ​​indicating the range of the 3D region, - A device configured to encode the XR scene description into the data stream.

8. The device according to claim 7, wherein the primitive is a geometric item that defines a region of the 3D scene surrounding a virtual object or a part of a virtual object.

9. The device according to claim 7 or 8, wherein the boundary value indicates the distance from the centroid of the at least one primitive at which the trigger is activated.

10. The device according to any one of claims 7 to 9, wherein the default boundary value is defined for the type of region associated with the primitive.

11. The device according to any one of claims 7 to 10, wherein the primitive belongs to at least two groups of primitives, including spheres, boxes, round boxes, box frames, tori, capped tori, links, infinite cylinders, capped cylinders, rounded cylinders, cones, infinite cones, capped cones, rounded cones, planes, hexagonal prisms, triangular prisms, capsules, lines, solid angles, truncated spheres, truncated hollow spheres, ellipsoids, rhombuses, octahedrons, pyramids, triangles, and quadrilaterals.

12. A device for rendering an augmented reality (XR) scene, the device comprising a processor and memory associated with the processor, the processor is - Link nodes and obtain the scene graph belonging to the XR scene description, where each node describes an object in the XR scene and is associated with a trigger, and the trigger, when it is a collision trigger or proximity trigger, includes a description of at least one primitive, and the primitive is, • Primitive types, - Description of the 3D region according to the type of primitive, - Includes boundary values ​​indicating the range of the 3D region, - The at least one trigger is activated when the virtual object enters the region of the XR scene within the range of the 3D region. A device configured such that at least one action is associated with the at least one trigger, the at least one action is performed when the at least one trigger is activated, and the at least one action belongs to a group of actions including haptic feedback, physical simulation, and audio feedback.