Immersive reality data processing for temporal persistence
By integrating a previous scene version into the current one using temporal occultation, immersive reality systems achieve efficient, real-time display with reduced computational load and improved user interaction.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2026-03-12
AI Technical Summary
Existing immersive reality systems struggle to display both the current and previous states of a virtual object in real-time without overwhelming computing resources, and fail to distinguish between the two states, leading to interaction errors.
Implement a method of integrating a previous version of the scene into the current version, using temporal occultation to partially obscure the previous version with the current version, allowing for a near-real-time display with reduced computational load.
Enhances user interaction by providing a persistent view of scene evolution, reducing computing resources and minimizing interaction errors by clearly distinguishing between current and past states.
Smart Images

Figure EP2025074965_12032026_PF_FP_ABST
Abstract
Description
Description Title: Immersive reality data processing for temporal afterimage technical field
[0001] This description relates to data processing in the context of virtual reality or augmented reality, or immersive reality, particularly in three-dimensional representation. Previous technique
[0002] When a real-world environment is captured by at least one sensor and rendered in a virtual environment in real time, it is important to update the displayed data to reflect the current state of the captured scene. In some cases, displaying previous states can be useful, particularly when the sensor does not capture the entire environment and is moving, for example.
[0003] A known solution for displaying the current state and previous states in the same space is to use fusion algorithms such as image fusion or mesh fusion (especially in the case of volumetric data) which allow obtaining an object typically grouping current and past information.
[0004] Such algorithms are resource-intensive and most often executed offline, and therefore not in real time, for example on CPUs in linear processing mode (CPU stands for "central processing unit"). Furthermore, they do not (as they currently stand) allow for the distinction between "fresh" and old data.
[0005] Typically, immersive realities result from the three-dimensional reproduction of an environment in such a way that the user has the impression of moving within that environment. In particular, the concept of environmental reproduction was introduced with the emergence of virtual reality, made possible by virtual reality systems such as VR headsets. These immersive reality systems now encompass systems ranging from virtual reality to mixed reality, including augmented reality and augmented virtuality. The virtual environment reproduced by these immersive reality systems consists of universes generated from scratch (also called "virtually generated environments") or existing locations (also called "real environments"), possibly remote, in which the user has the impression of moving.
[0006] Figure 1 illustrates an immersive reality device, such as a headset that a user can wear in front of their eyes. In the example shown, the headset includes at least one SCR screen in front of the user's eyes and at least one sensor (two sensors in the example shown, C1 and C2).
[0007] Thus, each sensor C1, C2 can capture a physical quantity such as, for example, depth (via time of flight), temperature, or other parameters. Significant elements of this data capture are located in a scene of a real environment (labeled SCEN in the example of Figure 3), and from this, for example, a point cloud Gk of points pk corresponding to this scene can be deduced, each point pk of the cloud being represented by: - The x coordinates P , y P , z P from point pk in the scene, and - By a color or grey level attribute (or size or other) of this point in the point cloud, depending on the intensity of the signal captured by the sensor(s) C1, C2.
[0008] Sensors C1, C2 thus capture a SCEN scene at a given time t=k, and a point cloud Gk results from this capture at time k. The point cloud Gk can be representative of a quantity (depth, and / or temperature, and / or positions, or other) of one or more real objects that are in the captured scene.
[0009] For example, by way of illustration (and not limitation), each virtual environment that is reproduced can be considered to consist of one or more virtual objects corresponding respectively to one or more real objects in the scene. Thus, an object in the scene can be reproduced in its position and shape at the current time k of its evolution in the real environment. In particular, when a virtual environment corresponds to the reproduction of a real environment in real time, the virtual object reproduced at the current time t corresponds to the real object captured at time k, with t > k, but t very close to k (t ~ k). Thus, typically, the position and shape of the virtual object reproduced in the virtual environment at the current time correspond to the position and shape of the real object captured in the real environment at the capture time t ~ k.For the sake of simplicity, we will consider the reproduction time t, during real-time reproduction, to correspond to the capture time k. Strictly speaking, there is a transmission delay between the sensor and the reproduction system, but the delays, particularly those related to the transmission and / or processing of the captured data, are generally negligible and will be treated as such hereafter.
[0010] Currently, simply reproducing the virtual object in a current state k does not allow a user to anticipate how the virtual object's state will evolve (movement, deformation, addition of context, etc.) at future times. Because of this limitation, the user may make errors when interacting with the virtual object.
[0011] To mitigate this type of error, the user can request the display of at least one previous state of the virtual object. However, analyzing state evolution can be challenging because the system only reproduces a single state of the virtual object: either the current state in real-time mode or one of the previous states upon request. Consequently, improving interaction with the virtual object remains very limited at present.
[0012] One solution would be to display multiple states of the same virtual object, including its current state within the same virtual environment, using image or mesh fusion algorithms (as previously introduced). This solution has several drawbacks, including: - The use of systems with high computing power, otherwise the merger will fail and only the virtual object in its current state is reproduced. - The inability to distinguish between the virtual object in its current state and the virtual object in a past state, since the merging algorithm provides a merged virtual object that is the virtual object reproduced in place of the virtual object in its current state. Consequently, when the user initiates an interaction with the merged virtual object, the interaction operates on both the virtual object in its current state and its past state. This can lead to an interaction error if the user only intended to interact with the virtual object's current state, for example. Summary
[0013] This disclosure improves the situation.
[0014] A method for generating a scene in immersive reality is proposed, comprising integrating, in a current version of the scene, at least one previous version of the scene.
[0015] Thus, in immersive reality, a user is offered a persistence of the scene, allowing them to observe its evolution over time relative to one or more moments prior to the current moment. Such a feature enhances the user experience in immersive reality, particularly by enabling users to perceive their own movements relative to the scene, or even the movement of objects within the scene.
[0016] To give temporality to this representation of afterimage, the previous and current versions of the scene are superimposed as illustrated in Figure 4 (left part) and, to mark this evolution over time, it is preferred to at least partially obscure the previous version with the current version in the respective parts which coincide spatially.
[0017] The term "to at least partially obscure" means an occultation which can be total (where all the pixels of the previous version are replaced by those of the current version), or a "partial" occultation by applying a higher or lower degree of transparency to the pixels of the current version so as not to completely hide those of the previous version.
[0018] Furthermore, the shapes of the earlier version (reference Gk-1 in Figure 4) may not completely correspond with those of the current version (reference Gk) so that only part of the current version may at least partially obscure only part of the earlier version.
[0019] Thus, according to one possible definition of the aforementioned integration, it will be understood that this integration is carried out in such a way that at least a part of the current version (Gk), having the same position as at least a part of the previous version (Gk-i), at least partially obscures said part of the previous version (Gk-i).
[0020] When this "same position" is located in a screen space representing immersive reality, the advantages of such an achievement include a reduction in computing resources and a limitation of the overload of the rendering of immersive reality.
[0021] In an embodiment where the process involves obtaining data acquired by at least one sensor, at successive times, the aforementioned integration includes the current version of the scene represented by data acquired at a current time, and a previous version of the scene represented by data acquired at a time preceding the current time.
[0022] In such a design, the process may include storing in memory the data acquired in correspondence of successive moments.
[0023] The aforementioned acquired data can define a point cloud, for each acquisition time, and include, for each point in the point cloud, at least: - 3D coordinates, and - a color attribute.
[0024] In such a design, by implementing an immersive reality device comprising at least one screen with a surface, the process can then include a display on the screen: - of a projection onto the surface, when it exists, of each point of a point cloud at a given instant, and - of a projection onto the surface, when it exists and is not already occupied by the projection of a point from the cloud of the given instant, of each point from at least one cloud of points from an instant preceding the given instant.
[0025] The aforementioned "given" instant can be the current instant (so that a point from a previous instant is not displayed if its position on the screen is already occupied by a point from the current instant), or, algorithmically, it can be a point from a previous instant (so that a point from an instant even earlier than this previous instant is not displayed if its position on the screen is already occupied).
[0026] Such an achievement, implementing a computer processing called "temporal occultation" in the detailed description that follows, offers a high speed of processing allowing a near real-time display of the current version of the scene and its previous version(s).
[0027] Thus, when dealing with point clouds, for example, the current scene and the previous scene can be blended at common positions. For instance, if the current scene contains a bright yellow object, and the previous scene contains a previous object with a degree of transparency (pale yellow, for example), then the common positions blending the two objects might be "medium yellow" (intermediate between bright yellow and pale yellow), blurring the boundaries of the current object. Consequently, interaction with, or even grasping, the current object is less precise, with a risk of interaction / grasping errors. Temporal occlusion, which masks the previous object with the current object at common positions, reduces or even eliminates this risk of error.
[0028] Thus, for example, it is possible to limit, if necessary, the aforementioned persistence and temporal occultation to only objects moving in the scene.
[0029] Thus, in one embodiment, when the previous and current versions of the scene include a moving object, said occlusion is applied to the moving object (at a minimum).
[0030] Typically, in such an implementation, the color attribute of the display projection of a point in a given point cloud can be modulated according to the age of the data acquisition time defining the given point cloud.
[0031] More generally, the earlier version of the scene can be integrated into immersive reality with a modulation of a color attribute that depends on the age of the earlier version.
[0032] For example, the aforementioned modulation may include a transparency feature, to a degree of transparency that is a function of seniority.
[0033] Alternatively or in addition, the modulation may involve a discoloration, to a degree of discoloration which is a function of seniority.
[0034] In another approach, a computer program is proposed that includes instructions for implementing the process as defined above, when executed by a processor. This could be, for example, a graphics processing unit (GPU). Alternatively, a non-transient, computer-readable recording medium is proposed on which such a program is recorded.
[0035] In another aspect, a generator for an immersive reality scene is proposed, comprising an integrator, within a current version of the scene, of at least one previous version of the scene. The integrator is configured such that at least a part of the current version, having the same position as at least a part of the previous version, at least partially obscures said part of the previous version.
[0036] According to another aspect, an immersive realization device is proposed comprising at least one display screen and a generator of the aforementioned type, to display the current version of the scene with at least one previous version of the scene.
[0037] Such a device may further include at least one scene data sensor, and a motion sensor for the data sensor to compensate for relative movements of the scene relative to the data sensor between a previous scene data acquisition time and a current scene data acquisition time. Brief description of the drawings
[0038] Other features, details, and advantages will become apparent upon reading the detailed description below and analyzing the attached drawings, on which: Fig. 1
[0039] [Fig. 1] illustrates an example of an implementation context in immersive reality. Fig. 2
[0040] [Fig. 2] shows an example of a sequence of steps in a process according to one embodiment. Fig. 3
[0041] [Fig. 3] shows a device for implementing the process according to one embodiment. Fig. 4
[0042] [Fig. 4] illustrates an example of temporal occultation according to one embodiment. Description of the implementation methods
[0043] Referring to Figure 2, during a first step S1, sensors C1 and C2 capture a scene from the real environment at a current time t=k. An input interface IN of an immersive reality DIS device, illustrated as an example in Figure 3, then forms a data stream of a point cloud Gk, as described previously with reference to Figure 1. More specifically, this point cloud Gk is defined by: - its points pk, with coordinates x P , y P , z P (with p=1, ..., N if the cloud contains N points in total), each point pk having a color or grayscale attribute, for example, denoted "COLpk" below; and - for a given acquisition time t=k.
[0044] In the next step, S2, this data (coordinates, attributes, of each point in the point cloud Gk) is stored in MEM memory with reference to the time k when this data was acquired. This step S2 is performed repeatedly so as to keep in memory the current version of the point cloud Gk and n-1 previous versions of point clouds Gk-1, Gk-2, ..., Gk-n, of the scene. Thus, it will be understood that these versions are identical to the point cloud Gk for the same scene where the objects remain unchanged, and in particular, stationary for the same orientation axes of sensors C1, C2. On the other hand, if, for example, an object in the scene is moving while the orientation axes of sensors C1, C2 remain fixed, then the previous versions of the point cloud Gk can characterize an evolution over time of this object in the scene. This property is used below to represent such an evolution.
[0045] In step S3, a DISPL display control signal feeding an interface of the SCR screen (and which can be generated, for example, by a graphics processing unit, or GPU) is constructed as follows, at the current time t=k (quasi-real time). Consider the plane coinciding with the surface of the SCR screen and comprising a set of pixels (x,y). If there exists a point pk in the currently acquired point cloud Gk, and such that its projection proj(pk) onto the aforementioned plane can be displayed on the SCR screen, then this projection proj(pk) is displayed on the SCR screen by activating the corresponding screen pixel (x,y) with its color or grayscale attribute COLpk. This display is conventional and simply corresponds to the display of the last acquired point cloud Gk on the SCR screen, for example, of an immersive reality headset, as described previously with reference to Figure 1.In contrast, in this S3 step, a display of previous versions (Gk-1, Gk-2, ..., Gk-n) of the point cloud is added to the standard display as follows: for i=1, ..., n, if there exists a point pk-i of the point cloud Gk-i acquired at time ki, and such that its projection proj(pk-i) onto the plane can be displayed on the SCR screen without coinciding with the aforementioned pixel (x,y) (already occupied), then this projection proj(pk-i) is displayed on the SCR screen, with its color or grayscale attribute COLpk-i modulated by a function Att(i) that depends on the acquisition time of this point cloud Gk-i. This modulation can, for example, consist of applying a "decolorization" of the point (i.e., making the point display more "black and white") depending on the value of the acquisition age, represented here by the parameter i. Alternatively, this modulation could consist of making the point more "transparent" depending on the seniority of acquisition i (i(For example, assigning a color closer to the background of the scene at that point). Other examples of modulation are possible, of course.
[0046] Thus, if the projection proj(pk-i) of a point in an older cloud Gk-i is already occupied by the projection proj(pk-j) of a point in a more recent cloud version Gk-j (with j less than i), then the projection of the older point proj(pk-i) is not displayed. It can be understood that older points are "occulted" by more recently acquired points. Such an implementation (according to the principle known as "temporal occultation") requires very few resources. A standard GPU (but programmed according to the algorithm illustrated in Figure 2) can be used to implement the processing described above.
[0047] Referring now to Figure 4, if the user UT remains stationary so that the axis of sensors C1, C2 remains fixed, while an object in the scene is moving or deforming, the variation in time t of the object captured at different successive times t=k-2, t=k-1, t=k (on the t axis illustrated on the left part of Figure 4), is represented at the end of the processing of Figure 2 as illustrated as an example on the right part of Figure 4: the different versions of the object in time (or more generally of the scene) are superimposed, with a degree of transparency or "image coldness" which increases according to the age of the version displayed, the points acquired earlier being always occulted by the points more recently acquired according to the principle of temporal occultation presented above.
[0048] This principle of "concealing" an object, in the general sense (elements of a scene or a given object within a scene), is thus applied when one object is placed in front of another, as illustrated on the right-hand side of figure 4. In computer science, it is possible to use the same principle, so that here, a newer object hides an older object.
[0049] By using the aforementioned type of system axis modification, it is possible to place all the data captured from a previous scene into the current scene without merging this previous data with that of the current scene. The "occultation" operation then allows the most "fresh" data of an object (in the broadest sense) to be displayed in a single pixel (x,y) of the SCR screen area dedicated to that object, placing it "in front" of the oldest data.
[0050] One of the advantages of this solution is that it avoids significant computation, preserves a unique object for each capture over time, and therefore allows for differentiation (for example, by adjusting its color or transparency) across its various versions. Occlusion can be performed on multiple objects in parallel (i.e., for each screen pixel (x,y) by a GPU).
[0051] To avoid overloading the immersive reality rendering for the user, a reasonable number of previous versions to be displayed can be optimized. This number, n-1, could, for example, be between 2 and 5. Thus, the processing power of a typical immersive reality headset's GPU is more than sufficient to handle this display of previous versions.
[0052] In the implementation example shown on the right side of Figure 4, the orientation of the axes of sensors C1 and C2 is fixed (the user's head is stationary). However, the user may be moving, and it can be advantageous for the user to see the relative displacements of the scene with respect to their own movements: the processing shown in Figure 2 allows the user to see the afterimage of their movements relative to the scene in a virtual environment.
[0053] However, in a variant, and in order to visualize only the movements / deformations of objects in the scene, without the influence of any possible movements or head movements of the user UT, a user motion sensor (for example, a gyroscope) can be used to compensate for orientation deviations of the axes of sensors C1 and C2 based on the user's movements. In such a variant, only the movements / deformations of objects in the scene are reproduced in the immersive environment that the user can view. More specifically, in this implementation, the time-dependent variations in the coordinates of the points in the point clouds Gk, Gk-1, ..., successively obtained, are then due solely to the movements / deformations of objects in the scene.
[0054] Figure 3 schematically illustrates a device for implementing the above process. Such a DIS device typically includes a processing circuit incorporating: - a PROC processor to implement the processing described above, for example with reference to Figure 2, - a MEM memory to store at least instructions of a computer program for implementing the above process when executed by the PROC processor, as well as, in particular, point cloud data Gk, Gk-1, Gk-2, ..., according to a current version Gk and one or more earlier versions Gk-1, Gk-2, ..., corresponding to current times k, and previous k-1, k-2, - an IN input interface to receive data captured by sensor(s) C1, C2, and possibly the GYR motion sensor, and - an OUT output interface to deliver a VE immersive reality display control signal to the SCR screen, allowing the user UT to view, for example, a current version of a scene object OBJ(k) represented by a point cloud Gk, and one or more previous versions of this object OBJ(k-1), represented by one or more previously acquired point clouds Gk-1. Industrial application
[0055] Among the possible applications of immersive reality, the rendering of real-world scenes in virtual reality can generally be implemented in real time, after capturing spatially positionable data. Typically, a possible application is in games or video learning that implement a concept based on display over time, for example, to perfect a movement, such as a sporting gesture (tennis, or others).
[0056] Another example of an application is the remote rendering of an environment captured in real time. The implementation described above then allows for the addition of a prior context (for example, the sensors being mobile while the environment is static), which enables a more complete remote rendering.
Claims
Demands
1. 1. Method of generating a scene in immersive reality, comprising an integration, in a current version of the scene, of at least one earlier version of the scene, the integration being carried out in such a way that at least a part of the current version (Gk), having the same position as at least a part of the earlier version (Gk-i), at least partially obscures said part of the earlier version (Gk-i).
2. 2. A method according to claim 1, comprising obtaining data acquired by at least one sensor, at successive times, and in which the integration comprises the current version of the scene represented by data acquired at a current time, and a previous version of the scene represented by data acquired at a time preceding the current time.
3. 3. Method according to claim 2, comprising storing in memory the data acquired in correspondence of successive instants.
4. 4. A method according to any one of claims 2 and 3, wherein the acquired data define a point cloud (Gk) for each acquisition time (k), and comprise, for each point in the point cloud (Gk), at least: - 3D coordinates (x P , y P , z P ), And - a color attribute (COLpk).
5. 5. A method according to claim 4, implementing an immersive reality device comprising at least one screen (SCR) having a surface, the method comprising a display on the screen (SCR): - of a projection onto the surface, when it exists, of each point of a point cloud (Gk) at a given instant (k), and - of a projection onto the surface, when it exists and is not already occupied by the projection of a point from the cloud (Gk) at the given instant (k), of each point from at least one point cloud (Gk-i) of an instant (ki) preceding the given instant (k).
6. 6. Method according to claim 5, wherein the color attribute of the display of the projection of a point of a given point cloud (Gk-i) is modulated according to an age (i) of the instant of acquisition of the data defining the given point cloud (Gk-i).
7. 7. A method according to any one of the preceding claims, wherein the earlier version of the scene is integrated into the immersive reality with a modulation (Att(i)) of a color attribute (COLpk-i) that depends on an age (i) of the earlier version.
8. 8. A method according to any one of claims 6 and 7, wherein the modulation includes making transparent, to a degree of transparency that is a function of seniority (i).
9. 9. A method according to any one of claims 6 and 7, wherein the modulation includes a discoloration, to a degree of discoloration that is a function of age (i).
10. 10. A method according to any one of the preceding claims, wherein the earlier and current versions of the scene include a moving object, and said occlusion is applied to the moving object (Gk, Gk-1, Gk-2).
11. 11. Computer program comprising instructions for carrying out the method according to any one of the preceding claims, when executed by a processor (PROC).
12. 12. Generator of an immersive reality scene, comprising an integrator, in a current version of the scene, of at least one earlier version of the scene, the integrator being configured such that at least a part of the current version (Gk), having the same position as at least a part of the earlier version (Gk-i), at least partially obscures said part of the earlier version (Gk-i).
13. 13. Immersive realization device comprising at least one display screen (SCR) and a generator according to claim 12 for displaying the current version of the scene with at least one previous version of the scene.
14. 14. Device according to claim 13, further comprising at least one scene data sensor (C1, C2), and a motion sensor (GYR) of the data sensor to compensate for relative movements of the scene with respect to the data sensor between an earlier scene data acquisition time (ki) and a current scene data acquisition time (k).
Citation Information
Patent Citations
Railway comprehensive inspection method
CN114440767A
Control device for alternate reality system, alternate reality system, control method for alternate reality system, program, and recording medium
WO2014027681A1