Immersive reality data processing for temporal afterimage
The method integrates previous scene versions into the current version using temporal occultation, addressing the challenge of real-time display and interaction errors by modulating older data, thus enhancing immersive reality systems.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2026-03-13
AI Technical Summary
Existing immersive reality systems struggle with displaying both current and previous states of a virtual object in real-time without excessive computational resources, and fail to distinguish between these states, leading to interaction errors.
Implement a method that integrates previous versions of a scene into the current version using temporal occultation, where older data is modulated by transparency or discoloration based on age, allowing near-real-time display with minimal computational resources.
Enables efficient display of current and previous scene states with minimal computational overhead, enhancing user interaction by distinguishing between different states of virtual objects.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Immersive reality data processing for temporal afterimage technical field
[0001] The present description relates to data processing in the context of virtual reality or augmented reality, or immersive reality, in particular in three-dimensional representation. Previous technique
[0002] When a real environment is captured by at least one sensor and rendered in a virtual environment in real time, it is important to update the displayed data to reflect the current state of the captured scene. In some cases, displaying previous states can be useful, particularly when the sensor does not capture the entire environment and is moving, for example.
[0003] A known solution for displaying the current state and previous states in the same space consists in using, in particular, fusion algorithms such as image fusion or mesh fusion (especially in the case of volumetric data) which make it possible to obtain an object typically grouping current and past information.
[0004] Such algorithms are resource-intensive and most often executed offline, and therefore not in real time, for example on CPUs in linear processing mode (CPU stands for "central processing unit"). Furthermore, they do not (as they currently stand) allow for distinguishing "fresh" data from old data.
[0005] Typically, in the case of immersive realities, these result from the three-dimensional reproduction of an environment in such a way that the user has the impression of moving within that environment. In particular, the notion of reproducing environments was introduced with the emergence of virtual reality, made possible notably by virtual reality systems, such as virtual reality headsets. These immersive reality systems now include systems ranging from virtual reality to mixed reality, including augmented reality and augmented virtuality. The virtual environment reproduced by these immersive reality systems consists of universes generated ab nihilo (also called "virtually generated environments") or existing locations (also called "real environments"), possibly remote, in which the user has the impression of moving.
[0006] Figure 1 illustrates an immersive reality device, such as a headset that a user can wear in front of their eyes. In the illustrated example, the headset includes at least one SCR screen in front of the user's eyes and at least one sensor (two sensors in the illustrated example, Cl and C2).
[0007] Thus, each sensor Cl, C2 can capture a physical quantity such as, for example, depth (via time of flight), temperature, or other parameters. Significant elements of this capture are located in a scene of a real environment (bearing the reference SCEN in the example of [Fig. 3]), and it can be deduced, for example, a cloud Gk of points pk corresponding to this scene, each point pk of the cloud being represented by: - The coordinates xp, yp, zp of the point pk in the scene, and - By a color or grey level attribute (or size or other) of this point in the point cloud, depending on the intensity of the signal captured by the sensor(s) Cl, C2.
[0008] The sensors Cl, C2 thus capture a SCEN scene at a given time t=k, and a point cloud Gk results from this capture at time k. The point cloud Gk can be representative of a quantity (depth, and / or temperature, and / or positions, or other) of one or more real objects that are in the captured scene.
[0009] For example, by way of illustration (and not limitation), each virtual environment that is reproduced can be considered to consist of one or more virtual objects corresponding respectively to one or more real objects in the scene. Thus, an object in the scene can be reproduced in its position and shape at the current instant k of its evolution in the real environment. In particular, when a virtual environment corresponds to the reproduction of a real environment in real time, the virtual object reproduced at the current instant t corresponds to the real object captured at instant k, with t > k, but t very close to k (t ~ k). Thus, typically, the position and shape of the virtual object reproduced in the virtual environment at the current instant correspond to the position and shape of the real object captured in the real environment at the capture instant t ~ k.For the sake of simplicity, we will consider the reproduction time t, during real-time reproduction, to correspond to the capture time k. Strictly speaking, there is a transmission delay between the sensor and the reproduction system, but the delays, particularly those related to the transmission and / or processing of the captured data, are generally negligible and will be treated as such hereafter.
[0010] Currently, simply reproducing the virtual object in a current state k does not allow a user to anticipate the evolution of the virtual object's state (movement, deformation, addition of context, etc.) at future times. Because of this limitation, the user may make errors in interacting with the virtual object.
[0011] To limit this type of error, the user can request the display of at least one previous state of the virtual object. However, analyzing state evolution can be difficult because the system only reproduces a single state of the virtual object: either the current state in real-time mode or one of the previous states upon request. Consequently, improving interaction with the virtual object remains very limited at present.
[0012] One solution would be to display several states of the same virtual object, including the current state of the virtual object in the same virtual environment, using image or mesh merging algorithms as previously introduced. This solution has several drawbacks, including: - The use of systems with high computing power is necessary; otherwise, the merging process fails and only the virtual object in its current state is reproduced. - The inability to distinguish between the virtual object in its current state and the virtual object in a past state, since the merging algorithm provides a merged virtual object that is the virtual object reproduced in place of the virtual object in its current state. Consequently, when the user initiates an interaction with the merged virtual object, the interaction operates on both the virtual object in its current state and its past state. This can lead to an interaction error if the user only intended to interact with the virtual object's current state, for example. Summary
[0013] This disclosure improves the situation.
[0014] A method for generating a scene in immersive reality is proposed, comprising an integration, in a current version of the scene, of at least one earlier version of the scene.
[0015] Thus, a user in immersive reality is offered a persistence of the scene, allowing them to observe an evolution of the scene over time relative to one or more moments prior to the current moment. Such a feature improves the user experience in immersive reality, in particular by allowing them to be aware of their own movements relative to the scene, or even the movement of objects within the scene.
[0016] In an embodiment where the method involves obtaining data acquired by at least one sensor, at successive times, the aforementioned integration includes the current version of the scene represented by data acquired at a current time, and a previous version of the scene represented by data acquired at a time preceding the current time.
[0017] In such an embodiment, the method may include storing in memory the data acquired in correspondence of successive instants.
[0018] The aforementioned acquired data can define a point cloud, for each acquisition time, and include, for each point in the point cloud, at least: - 3D coordinates, and - a color attribute.
[0019] In such an embodiment, by implementing an immersive reality device comprising at least one screen having a surface, the method can then include a display on the screen: - of a projection onto the surface, when it exists, of each point of a point cloud at a given instant, and - of a projection onto the surface, when it exists and is not already occupied by the projection of a point from the cloud at the given instant, of each point from at least one point cloud of a moment preceding the given moment.
[0020] The aforementioned "given" instant can be the current instant (so that a point from a previous instant is not displayed if its position on the screen is already occupied by a point from the current instant), or, algorithmically, it can be a point from a previous instant (so that a point from an instant even earlier than this previous instant is not displayed if its position on the screen is already occupied).
[0021] Such an implementation, implementing a computer processing called "temporal occultation" in the detailed description that follows, offers high processing speed allowing a near real-time display of the current version of the scene and its previous version(s).
[0022] Typically, in such an embodiment, the color attribute of the display of the projection of a point in a given point cloud can be modulated according to the age of the instant of acquisition of the data defining the given point cloud.
[0023] More generally, the earlier version of the scene can be integrated into immersive reality with a modulation of a color attribute that depends on the age of the earlier version.
[0024] For example, the aforementioned modulation may include a transparency feature, to a degree of transparency that is a function of seniority.
[0025] Alternatively or in addition, the modulation may include a discoloration, to a degree of discoloration which is a function of seniority.
[0026] According to another aspect, a computer program is proposed comprising instructions for implementing the method as defined above, when executed by a processor. This could be, for example, a graphics processing unit (GPU). According to another aspect, a non-transient, computer-readable recording medium is proposed on which such a program is recorded.
[0027] According to another aspect, a generator of an immersive reality scene is proposed, comprising an integrator, in a current version of the scene, of at least one earlier version of the scene.
[0028] According to another aspect, an immersive realization device is proposed comprising at least one display screen and a generator of the aforementioned type, to display the current version of the scene with at least one previous version of the scene.
[0029] Such a device may further include at least one scene data sensor, and a data sensor movement sensor to compensate for relative movements of the scene with respect to the data sensor between a previous scene data acquisition time and a current scene data acquisition time. Brief description of the drawings
[0030] Other features, details and advantages will become apparent from reading the detailed description below and from analyzing the accompanying drawings, in which: Fig. 1
[0031] [Fig.1] illustrates an example of an implementation context in immersive reality. Fig. 2
[0032] [Fig.2] shows an example of a succession of steps of a process according to an embodiment. Fig. 3
[0033] [Fig.3] shows a device for implementing the process according to one embodiment. Fig. 4
[0034] [Fig.4] illustrates an example of temporal occultation according to one embodiment. Description of the implementation methods
[0035] With reference to [Fig. 2], during a first step SI, sensors Cl, C2 capture a scene from the real environment at a current time t=k. An input interface IN of an immersive reality DIS device illustrated as an example in [Fig. 3] then forms a data stream of a point cloud Gk, as described previously with reference to [Fig. 1]. More specifically, this point cloud Gk is defined by: - its points pk, with coordinates xp, yp, zp (with p=l, ..., N if the cloud has N points in total), each point pk having a color or grayscale attribute, for example, denoted "COLpk" below; and - for a given acquisition time t=k.
[0036] In the next step S2, this data (coordinates, attribute, of each point of the point cloud Gk) is stored in MEM memory with reference to the acquisition time k of this data. This step S2 is performed repeatedly to store in memory the current version of the point cloud Gk and n-1 previous versions of point clouds Gk-1, Gk-2, ..., Gk-n, of the scene. Thus, it will be understood that these versions are identical to the point cloud Gk for the same scene where the objects remain unchanged, and in particular, stationary for the same orientation axes of sensors Cl, C2. On the other hand, if, for example, an object in the scene is moving while the orientation axes of sensors Cl, C2 remain fixed, then the previous versions of the point cloud Gk can characterize an evolution over time of this object in the scene. This property is used below to represent such an evolution.
[0037] In step S3, a DISPL display control signal feeding an interface of the SCR screen (and which can be generated, for example, by a graphics processing unit or "GPU") is constructed as follows, at the current time t=k (quasi-real time). Consider the plane coinciding with the surface of the SCR screen and comprising a set of pixels (x,y). If there exists a point pk in the currently acquired point cloud Gk, and such that its projection proj(pk) onto the aforementioned plane can be displayed on the SCR screen, then this projection proj(pk) is displayed on the SCR screen by activating the corresponding screen pixel (x,y) with its color or grayscale attribute COLpk. This display is conventional and simply corresponds to the display of the last acquired point cloud Gk on the SCR screen, for example, of an immersive reality headset as described previously with reference to [Fig. 1].In contrast, in this S3 step, a display of previous versions (Gk-1, Gk-2, ..., Gk-n) of the point cloud is added to the standard display as follows: for i = 1, ..., n, if there exists a point pk in the point cloud Gk-i acquired at time ki, and such that its projection proj(pk i) onto the plane can be displayed on the SCR screen without coinciding with the aforementioned pixel (x,y) (already occupied), then this projection proj(pk i) is displayed on the SCR screen, with its color or grayscale attribute COLpk i modulated by a function Att(i) that depends on the acquisition time of this point cloud Gk-i. This modulation can, for example, consist of applying a "decolorization" to the point (i.e., making the point display more "black and white") depending on the acquisition age value represented here by the parameter i. Alternatively, this modulation could consist of making the point more "transparent" depending on the seniority of acquisition i (i(For example, assigning a color closer to the background of the scene at that point). Other examples of modulation are possible, of course.
[0038] Thus, if the projection proj(pk i) of a point in an older cloud Gk-i is already occupied by the projection proj(pk j) of a point in a more recent cloud version Gk-j (with j less than i), then the projection of the older point proj(pk i) is not displayed. It will thus be understood that older points are "occulted" by more recently acquired points. Such an implementation (according to the principle known as "temporal occultation") requires very few resources. A standard GPU (but programmed according to the algorithm illustrated in [Fig. 2]) can be used to implement the processing described above.
[0039] Now with reference to Figure 4, if the user UT remains stationary so that the axis of the sensors Cl, C2 remains fixed, while an object in the scene is moving or deforming, the variation in time t of the object captured at different successive times t=k-2, t=kl, t=k (on the t axis illustrated on the left part of [Fig.4]), is represented at the end of the processing of [Fig.2] as illustrated by way of example on the right part of [Fig.4]: the different versions of the object in time (or more generally of the scene) are superimposed, with a degree of transparency or "image coldness" which increases according to the age of the version displayed, the points acquired earlier being always occulted by the points more recently acquired according to the principle of temporal occultation presented above.
[0040] This principle of "occulting" an object, in the general sense (elements of a scene or a given object in a scene), is thus applied when one object is placed in front of another, as illustrated on the right-hand side of [Fig. 4]. In computer science, it is possible to use the same principle, so that here, a newer object hides an older object.
[0041] By using the aforementioned type of system axis modification, it is possible to place all the data captured from a previous scene into the current scene without merging this previous data with that of the current scene. The "occultation" operation then makes it possible to display the "freshest" data of an object (in the broadest sense) in one pixel (x,y) of the SCR screen area dedicated to that object, placing it "in front" of the oldest data.
[0042] One of the advantages of the solution is that it avoids significant computation, preserves a unique object for each capture over time, and therefore allows it to be distinguished (for example, by manipulating its color or transparency) according to its different versions over time. Occlusion can be performed on several objects in parallel (i.e., for each pixel (x,y) of the screen by a GPU).
[0043] However, to avoid overloading the immersive reality rendering for the user, a reasonable number of previous versions to be displayed can be optimized. This number n-1 can, for example, be between 2 and 5. Thus, the processing power of a typical immersive reality headset's GPU is more than sufficient to support this display of previous versions.
[0044] In the embodiment illustrated on the right-hand side of [Fig. 4], the orientation of the axes of sensors C1, C2 is fixed (the user's head is stationary). However, The user may be in motion and it may be advantageous for the user to see the relative displacements of the scene with respect to their own movements: the processing of [Fig.2] allows the user to be presented with a persistence of their movements relative to the scene in a virtual environment.
[0045] However, in one embodiment, and in order to visualize only the movements / deformations of objects in the scene, without the influence of any possible displacements or head movements of the user UT, a user motion sensor (for example, a gyroscope) can be used to compensate for orientation deviations of the axes of sensors C1, C2 as a function of the user's movements. In such an embodiment, only the movements / deformations of objects in the scene are reproduced in the immersive environment that the user can visualize. More specifically, in this embodiment, the time-dependent variations in the coordinates of the points in the point clouds Gk, Gk-1, ..., successively obtained, are then due only to movements / deformations of objects in the scene.
[0046] Figure 3 schematically illustrates an embodiment of a device for implementing the above process. Such a DIS device comprises a processing circuit typically integrating: - a PROC processor to implement the processing described above, for example with reference to [Fig.2], - a MEM memory to store at least instructions from a computer program for implementing the above process when executed by the PROC processor, as well as, in particular, point cloud data Gk, Gk-1, Gk-2, ..., according to a current version Gk and one or more previous versions Gk-1, Gk-2, ..., corresponding to current times k, and previous times k-1, k-2, ..., - an IN input interface to receive data captured by the sensor(s) Cl, C2, and possibly the GYR motion sensor, and - an output interface OUT to deliver a VE immersive reality display control signal to the SCR screen, allowing the user UT to view, for example, a current version of a scene object OBJ(k) represented by a point cloud Gk, and one or more previous versions of this object OBJ(kl), ..., represented by one or more previously acquired point clouds Gk-1,.... Industrial application
[0047] Among the possible applications in immersive reality, the rendering of real-world scenes in virtual reality, in general, can be implemented in real time, after capturing spatially positionable data. Typically, a possible application is in games or video learning implementing a concept based on the display as a function of time, for example to perfect a movement, for example a sporting gesture (tennis, or others).
[0048] Another example of an application is the remote rendering of an environment captured in real time. The embodiment described above then makes it possible to add a prior context (the sensors being mobile, for example, while the environment is static), which allows for a more complete remote rendering.
Claims
Demands
1. 1. Method for generating a scene in immersive reality, comprising an integration, in a current version of the scene, of at least one earlier version of the scene.
2. 2. A method according to claim 1, comprising obtaining data acquired by at least one sensor, at successive times, and in which the integration comprises the current version of the scene represented by data acquired at a current time, and a previous version of the scene represented by data acquired at a time preceding the current time.
3. 3. Method according to claim 2, comprising storing in memory the data acquired in correspondence of successive instants.
4. 4. A method according to any one of claims 2 and 3, wherein the acquired data define a point cloud (Gk), for each acquisition instant (k), and include, for each point of the point cloud (Gk), at least: - 3D coordinates (xp, yp, zp), and - a color attribute (COLpk).
5. 5. A method according to claim 4, implementing an immersive reality device comprising at least one screen (SCR) having a surface, the method comprising a display on the screen (SCR): - of a projection onto the surface, where it exists, of each point of a point cloud (Gk) of a given instant (k), and - of a projection onto the surface, where it exists and not already occupied by the projection of a point of the cloud (Gk) of the given instant (k), of each point of at least one point cloud (Gk-i) of an instant (ki) preceding the given instant (k).
6. 6. Method according to claim 5, wherein the color attribute of the display of the projection of a point of a given point cloud (Gk-i) is modulated according to an age (i) of the instant of acquisition of the data defining the given point cloud (Gk-i).
7. 7. A method according to any one of the preceding claims, wherein the earlier version of the scene is integrated into the immersive reality with a modulation (Att(i)) of a color attribute (COLpk i) which depends on an age (i) of the previous version.
8. 8. A method according to any one of claims 6 and 7, wherein the modulation includes making transparent, to a degree of transparency that is a function of seniority (i).
9. 9. A method according to any one of claims 6 and 7, wherein the modulation includes a discoloration, to a degree of discoloration that is a function of age (i).
10. 10. Computer program comprising instructions for carrying out the method according to any one of the preceding claims, when executed by a processor (PROC).
11. 11. Generator of an immersive reality scene, comprising an integrator, in a current version of the scene, of at least one earlier version of the scene.
12. 12. Immersive realization device comprising at least one display screen (SCR) and a generator according to claim 11 for displaying the current version of the scene with at least one previous version of the scene.
13. 13. Device according to claim 12, further comprising at least one scene data sensor (Cl, C2), and a motion sensor (GYR) of the data sensor to compensate for relative movements of the scene with respect to the data sensor between an earlier scene data acquisition time (ki) and a current scene data acquisition time (k).
Citation Information
Patent Citations
Railway comprehensive inspection method
CN114440767A
Control device for alternate reality system, alternate reality system, control method for alternate reality system, program, and recording medium
WO2014027681A1