Method for the dynamic observation of stereoscopic images of a 3D video, head-mounted display and 3D display device implementing such a method

The method dynamically adjusts stereoscopic 2D image pairs to align with viewer perspective changes, enhancing immersion by modifying pixel positions and depth information, addressing the static viewpoint issue in existing 3D display technologies.

WO2025224387A1PCT designated stage Publication Date: 2025-10-30FOGALE OPTIQUE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/FR2024/050544
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-25
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing 3D video display technologies in head-mounted and 3D display devices provide a fixed view of the scene, lacking an immersive experience due to a static viewpoint, which does not align with the observer's changing perspective.

Method used

A method for dynamically modifying stereoscopic 2D image pairs by adjusting pixel positions and adding or removing pixels based on depth information and viewpoint changes, allowing for a dynamic change of perspective in real-time.

Benefits of technology

Enhances the immersive experience by providing a dynamic change of viewpoint, aligning the viewer's perspective with their actual movements, thereby improving the perceived depth and realism of the 3D scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FR2024050544_30102025_PF_FP_ABST
    Figure FR2024050544_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for the dynamic observation of stereoscopic images of a 3D video of a scene, the 3D video comprising a set of pairs of stereoscopic 2D images, referred to as one or more pairs of source images, capable of providing 3D video rendering by being displayed in a head-mounted display or on a 3D display device. The method comprises the steps of obtaining a change of viewpoint with respect to the imaged scene, modifying each of the images of at least one pair of source images so as to render at least one pair of stereoscopic 2D images of the scene, referred to as at least one rendered pair of images. The at least one rendered pair of images is intended to provide 3D rendering by being displayed in a head-mounted display or on a 3D display device and provides a change of viewpoint with respect to the 3D video of the scene. The images of the at least one pair of source images are modified by moving and / or spreading and / or contracting pixels or one or more groups of pixels of the at least one pair of source images of the scene and / or by adding and / or removing pixels or groups of pixels, to and / or from the images of the at least one pair of source images, as a function of a depth map of the imaged scene and of the obtained change of viewpoint.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DESCRIPTION

[0002] Method for dynamically observing stereoscopic images from a 3D video, head-mounted display and 3D display device implementing such a method

[0003] technical field

[0004] The present invention relates to stereoscopic 3D videos intended to be displayed in a head-mounted display or in a 3D display device in order to reproduce a 3D rendering when displayed.

[0005] The invention relates to the dynamic observation of 3D videos displayed in a head-mounted display or in a 3D display device.

[0006] Prior art

[0007] In the prior art, 3D videos are known that allow, from 2D images, the introduction of a perspective effect into the displayed 3D videos, such as would be perceived if the imaged scene were observed directly by the viewer. The 3D videos according to the invention are not computer-generated images or three-dimensional computer graphics, but three-dimensional images in which perspective or depth has been generated or introduced by computer-aided design (CAD). The synthesis of 3D images can be performed directly from a model of a real scene. The production of 3D images is costly in terms of time and the computing resources required for processing and storing these large images. The term 3D images can be understood as stereoscopic images obtained, for example, by 3D surface modeling of space or from two images with different disparities, typically two images obtained from two distinct viewpoints.The workload is tenfold when dealing with a stream of 3D images, particularly a stream of 3D images derived from a 3D surface model of space. Furthermore, the processing time is incompatible with the dynamic visualization of 2D images acquired in real time. The time required to generate synthetic 3D images is incompatible with the dynamic visualization of 2D images acquired in real time. The production of synthetic 3D images is costly in terms of time, computing resources required for processing, and image storage. The invention relates to stereoscopic 3D videos that provide a 3D rendering when displayed on a suitable display medium, by exploiting the binocular vision of the observer's brain to reconstruct depth and provide the perception of depth.

[0008] Stereoscopic vision is a known phenomenon in the current state of technology. Binocular human vision, or more broadly animal vision, based on two images that are compared by the brain, is the most widespread example. The biological processing of the resulting images is extremely efficient, as it provides a real-time sense of depth in the observed scene, allowing, for example, movement while knowing the relative distance of observed objects.

[0009] The state of the art includes viewing stereoscopic 3D videos, which are inspired by or mimic the principle of stereoscopic vision. In practice, a stereoscopic image comprises two distinct images of the same scene acquired from two different positions in space. Human stereoscopic vision allows us to reconstruct a scene containing depth information from a 3D video composed of a set of stereoscopic images. The principle of stereoscopic viewing is to enhance the immersive nature of image observation by adding a depth effect through the projection of a single composite image, or monocular 3D video. This involves combining two 2D images of the same scene acquired from different viewpoints onto the same 3D display screen or projecting 2D images of the same scene acquired from different viewpoints onto two separate display devices, each viewed by one of the observer's eyes.

[0010] We are also familiar with head-mounted displays, 3D glasses, and 3D displays. These display devices allow users to view 3D videos. However, these display devices have the drawback of offering a fixed view of the scene, which remains the same as that seen by the camera with which the video was acquired and which moves along with the display device. The viewpoint from which the scene is observed is fixed during the dynamic observation of the scene. The observed scene gives the observer the impression of moving with the display device, which does not provide an immersive experience when watching the video. One objective of the invention is to propose a method for processing 3D videos, a head-mounted display, and a 3D display device:

[0011] - enabling the correction of problems with state-of-the-art processes, and / or

[0012] - rendering a change of viewpoint of a video scene to a user of a head-mounted display or 3D display device, and / or

[0013] - enabling the enhancement and / or improvement of the immersive experience for a user of a head-mounted display or 3D display device showing stereoscopic 3D videos of a scene, and / or

[0014] - providing an alternative to head-mounted displays and 3D display devices showing synthetic 3D images.

[0015] Presentation of the invention

[0016] To this end, a method for dynamically observing stereoscopic images from a 3D video of a scene is proposed. The 3D video comprises a set of stereoscopic 2D image pairs, called source image pair(s), capable of providing a 3D rendering to the video when displayed in a head-mounted display or on a 3D display device.

[0017] According to the invention, source images can be understood as any stored or available image, for example a stack of images, particularly from a 3D video. As detailed elsewhere, the source images can originate from or be transmitted by external means, preferably in real time, and preferably even more so be processed by the method according to the invention, also in real time.

[0018] Preferably, the images of the source image pair were acquired by one or more imaging systems.

[0019] The images from at least one pair of source images are 2D images that can include depth information.

[0020] The process includes the step of obtaining a change of viewpoint relative to the imaged scene.

[0021] It can be understood as "image scene": an image of a pair of source images, that is to say unmodified or before modification.

[0022] The process further includes the step of modifying each of the images, preferably each image, preferably each of the two images, of at least one pair of source images, preferably of each pair of source images, to reproduce at least one pair of stereoscopic 2D images of the scene, referred to as at least one pair of reproduced images.

[0023] Preferably, the two images of at least one pair of restored images correspond to the two images modified, by implementation of the process, of at least one pair of source images.

[0024] Preferably, the number of restored image pairs corresponds to the number of modified source image pairs.

[0025] At least one pair of rendered images is intended or capable of providing 3D rendering by display in a head-mounted display or on a 3D display device.

[0026] At least one pair of rendered images provides or is intended to provide a change of viewpoint relative to the 3D video of the scene, preferably by display in a head-mounted display or on a 3D display device.

[0027] The images of at least one pair of source images are modified, or the step of modifying each of the images of at least one pair of source images is implemented, by moving and / or spreading and / or shrinking pixels or groups of pixels of at least one pair of source images of the scene and / or by adding and / or deleting pixels or groups of pixels, in said images of at least one pair of source images, according to:

[0028] - a depth map of the imaged scene,

[0029] - the change in perspective achieved.

[0030] Preferably, according to the invention, only certain pairs, that is, at least some of the source image pairs, from among the set of source image pairs, are modified. Indeed, in some cases, it may be that no change of viewpoint occurs for, or only occurs for, some of the images in the set of source image pairs of the 3D video. Such a case can occur, for example, if the 3D video viewer remains stationary for a period of time while watching the 3D video and, consequently, no change of viewpoint occurs during that period.

[0031] Preferably, the image modification step of at least one pair of source images of the scene is carried out using a processing unit.

[0032] Preferably, the change of viewpoint corresponds to the transition from one viewpoint, called the previous viewpoint, to a different viewpoint, called the new viewpoint. Preferably, the images in the image pair obtained by implementing the method, or the images modified according to the method, correspond to the images of the scene as they would be observed from the new viewpoint.

[0033] It can be understood as a point of view, a position, or coordinates of space.

[0034] Preferably, each data processing step of the process according to the invention, including for example any operation performed from or on data, for example a calculation, a determination or a comparison, is implemented by a processing unit or any device or system capable of processing data.

[0035] Preferably, the two images in the source image pair correspond to two distinct viewpoints of the scene.

[0036] Preferably, a pixel corresponds to or is associated with a point in the imaged scene and a group of pixels corresponds to or is associated with a part of an object in the imaged scene.

[0037] Preferably, the displacement of a portion of the pixels or group(s) of pixels within each image of the source image pair allows for, aims to reproduce, or consists of reproducing a parallax effect induced by the change in viewpoint. Preferably, the displacement of a portion of the pixels or group(s) of pixels is understood to mean: a modification of the position and / or a translation and / or a spreading and / or a contraction of a portion of the pixels or group(s) of pixels within each image of the source image pair, preferably in a plane or in a reference frame specific to each of the two images of at least one source image pair, or in a reference frame common to the two images of at least one source image pair.

[0038] Preferably, it can be understood as contraction: a settling or a reduction or a compression.

[0039] The term "plane" or "reference frame" can be understood as: a plan of or a reference frame common to the images considered, or a plan of or a reference frame distinct from the images considered, preferably a plan or a reference frame specific to each of the images considered.

[0040] Preferably, each pixel or group of pixels is moved and / or translated and / or stretched and / or contracted relative to a reference depth, preferably relative to the reference depth. Preferably, the reference depth is associated with one or more objects in the scene or a group of objects. Depth information can be understood as data representing, corresponding to, or providing information about the depth associated with a pixel, corresponding to a point in the imaged scene, or with a group of pixels, corresponding to an object in the imaged scene, of an image of the scene.

[0041] Depth information can be contained in a map or an image. Depth information can be contained in a depth map or an inverse depth map. Also, according to the invention, the depth map can be substituted for, or is equivalent to, an inverse depth map.

[0042] Depth can be understood as the distance between a point or object in the imaged scene and the optical sensor of the imaging device from which the image of the scene was acquired.

[0043] The term "depth information" can be understood as a depth field. The depth field may contain depth information for all or some of the points or objects in the scene, preferably a depth field of the entire imaged scene. The depth field may be contained within a channel of all or some of the pixels of one or both images of the source image pair.

[0044] At least some of the pixels in one or more images of the at least one pair of source images of the scene may include, but not necessarily, depth information and / or sharpness information.

[0045] Preferably, the modification step consists of:

[0046] - to modify, identically, each of the two images of a given pair of source images in order to produce at least one pair of images comprising two modified images, or

[0047] - to modify, distinctly, each of the two images of a pair of source images considered in order to reproduce at least one pair of images comprising two modified images.

[0048] Preferably, the image modification step of at least one pair of source images is carried out, in addition, relative to, or as a function of, an area, preferably a single area of ​​each of the two images of at least one pair of source images or an area comprising several parts of each of the two images of at least one pair of source images, of the images of at least one pair of source images, corresponding to an area called the attention area.

[0049] It can be understood by area of ​​attention: one or more pixels or a group of pixels of, preferably common to, each of the images of the pair of source images.

[0050] Preferably, the area of ​​attention is common to both images of at least one pair of source images, or corresponds to or is associated with the same object in the scene or the same area of ​​the scene within each of the two images of at least one pair of source images.

[0051] Preferably, the image modification step of at least one pair of source images is carried out, in addition, relative to, or as a function of, a depth associated with the area of ​​attention.

[0052] Preferably, the area of ​​focus should remain unchanged when modifying the images of at least one pair of source images, or

[0053] - is modified when modifying the images of at least one pair of source images.

[0054] The area of ​​attention can be modified in accordance with or identically to, or differently from, the modification of the images of at least one pair of source images.

[0055] Preferably, the modification of at least one image of the scene is performed on each image of at least one pair of source images, excluding the area of ​​attention. In other words, preferably, the displacement of some pixels or groups of pixels in each image of at least one pair of source images and / or the addition and / or deletion of pixels or groups of pixels in each image of at least one pair of source images is performed, preferably only, outside the area of ​​attention.

[0056] Preferably, the pixel or pixels of the attention area, preferably only the pixel or pixels of the attention area, remain unchanged and / or immobile during the modification step.

[0057] Preferably, the area of ​​attention remains unchanged, when modifying at least one pair of source images, by anchoring the area of ​​attention relative to the plane(s) or reference frame(s) of one or each of the images of at least one pair of source images.

[0058] Preferably, the modification of the images of at least one pair of source images is proportional to the change in viewpoint.

[0059] According to the invention, modification of the images of at least one image can be understood as: the displacement and / or spreading and / or contraction of pixels or group(s) of pixels within the images of at least one pair of source images and / or the quantity or number of pixels or groups of pixels added and / or removed from the images of at least one pair of source images.

[0060] Preferably:

[0061] - the modification of a part, called the first part, of the pixels or groups of images from at least one pair of source images is inversely proportional to the depth associated with the pixels or groups of pixels in the first part, and / or

[0062] - the modification of a part, called the second part, distinct from the first part, of the pixels or groups of pixels of the images of at least one pair of source images is proportional to the depth associated with the pixels or groups of pixels of the second part, and

[0063] - preferably, a third part, corresponding to the area of ​​attention, distinct from the first and second parts, of the pixels or groups of pixels of the images from at least one pair of source images:

[0064] • remains unchanged, or

[0065] • is modified in accordance, in whole or in part, with the first part, or in accordance, in whole or in part, with the second part, or differently from the first and second parts.

[0066] According to a first alternative, the first part can consist of all the pixels or groups of pixels in the images of at least one pair of source images, preferably excluding the pixel(s) or group(s) of pixels in the area of ​​attention. According to the first alternative, the modification of all the pixels in at least one pair of source images, preferably excluding the pixel(s) or group(s) of pixels in the area of ​​attention, is inversely proportional to the depth associated with the pixels or groups of pixels in the first part. According to a second alternative, the second part can consist of all the pixels or groups of pixels in the images of at least one pair of source images, preferably excluding the pixel(s) or group(s) of pixels in the area of ​​attention.According to the second alternative, the modification of all pixels of at least one pair of source images, preferably with the exception of the pixel(s) or pixel groups of the attention area, is inversely proportional to the depth associated with the pixels or pixel groups of the second part.

[0067] Preferably: for objects in the scene, called distant objects, whose depth is greater than the depth of the attention area, the modification step includes shifting the distant objects in the scene (or the pixel or group of pixels corresponding to each distant object) in the same direction and sense as the direction and sense of the viewpoint change, and / or

[0068] - for objects in the scene, called near objects, whose depth is less than the depth of the attention area, the modification step includes a shift of the near objects (or the pixel or group of pixels corresponding to each near object) in the same direction as the direction of the change of viewpoint and in the opposite direction to the direction of the change of viewpoint.

[0069] Preferably, the process includes, for each pixel or group(s) of pixels in at least one image, a calculation of a deviation or difference:

[0070] - depth between the area of ​​attention and each pixel or group(s) of pixels located outside the area of ​​attention, and / or

[0071] - between the inverse of the depth of the attention area (or the depth associated with the pixel(s) of the attention area) and the inverse of the depth of each pixel or group(s) of pixels located outside the attention area (or the depth associated with each pixel or group(s) of pixels located outside the attention area).

[0072] Preferably, the step of modifying at least one image is performed based on the calculated gap. Preferably, the calculation of the depth gap between the area of ​​attention and each pixel or group of pixels located outside the area of ​​attention can be defined as the calculation of a gap between the inverse of a depth associated with the area of ​​attention and the inverse of a depth associated with each pixel or group of pixels located outside the area of ​​attention.

[0073] Preferably, the process includes a calculation of at least one angle associated with or corresponding to the change of viewpoint.

[0074] Preferably, the image modification step of at least one pair of source images is performed based on at least one calculated angle.

[0075] The at least one angle corresponding to or associated with the change of viewpoint can be defined as an angle formed between the optical axis of the observer before the change of viewpoint and the optical axis of the observer after the change of viewpoint.

[0076] Preferably, the process includes a calculation of a variation in distance, associated with or corresponding to the change of viewpoint, relative to the imaged scene.

[0077] Preferably, the image modification step for at least one pair of source images is performed based on the calculated distance variation.

[0078] It can be understood as an angle associated with the change of viewpoint and / or by a variation of distance from the pictured scene: the change of viewpoint of the observer in relation to the scene as he would see it.

[0079] As an example, the change of viewpoint may include or correspond to a movement, preferably an angle and / or a variation in distance, of an observer relative to the 3D video of the imaged scene.

[0080] Preferably, the modification step is applied to a stream or stack of pairs of stored source images, transmitted, for example by a data transmission means and / or an imaging device, or acquired, for example by an imaging device.

[0081] Preferably, the stream or stack of stored source image pairs forms a 3D video. Preferably, the modification step is applied to a stream of source image pairs. Preferably, both the acquisition and modification steps are applied to a stream of source image pairs.

[0082] Preferably, the modification of the images of at least one pair of source images is a function of an interpupillary distance.

[0083] The image modification step of at least one pair of source images is implemented based on an interpupillary distance, preferably when it is observed, determined or calculated that a distance between the viewpoint of one of two images of a considered pair of source images and the other viewpoint of the two images of the considered pair of source images presents a distance that is too large relative to the interpupillary distance and / or a distance that is too small relative to the interpupillary distance.

[0084] Preferably, the modification of the images of at least one pair of source images is also a function of:

[0085] - a distance between two separate imaging devices, from which the two images of each pair of source images were respectively acquired, and / or

[0086] - of a position or arrangement or spatial coordinates of said two distinct imaging devices relative to each other and / or relative to the scene, and / or

[0087] - a difference in distance between said two separate imaging devices and the scene.

[0088] Preferably, the two separate imaging devices are arranged in two separate spatial positions and one of the images from each pair of source images comes from, or has been acquired or is being acquired by, one of the two imaging devices and the other of the two images from each pair of source images comes from, or has been acquired or is being acquired by, the other of the two imaging devices.

[0089] Preferably, this step is implemented when it is observed, determined, or calculated that the two images of at least one source image pair must, in addition to the modification implemented to reproduce the dynamic observation as the viewpoint changes, also be modified to improve the stereoscopic 3D rendering that the reproduced image pair is intended to provide. Depth can be understood as the distance between a point or object in the imaged scene and the optical sensor of the imaging device from which the scene image (e.g., an image from the source image pair) was acquired.

[0090] Taking into account, during the image modification step of at least one pair of source images, the position of the two distinct imaging devices in relation to each other and / or relative to the scene and / or preferably allows each pair of restored images to consist of an image, called the right image, as it would be perceived by the right eye of an observer of the scene and an image, called the left image, as it would be perceived by the left eye of an observer of the scene.

[0091] Preferably, the images of at least one pair of source images are further modified relative to each other by displacement and / or spreading and / or contraction of pixels or groups of pixels of at least one pair of source images of the scene and / or by addition and / or deletion of pixels or groups of pixels in said images of at least one pair of source images so that the difference in viewpoint between the two images of at least one pair of restored images corresponds to the pupillary distance.

[0092] Preferably, the process includes the step of displaying, in a virtual reality / augmented reality (VR / VA) headset, called a head-mounted display, or in a 3D display device, called the display step, at least one pair of rendered images or the images from at least one pair of source images.

[0093] The display stage may also include the display of one or more pairs of source images, for example when no change of viewpoint is obtained.

[0094] Preferably, the head-mounted display or 3D display device includes both separate imaging devices.

[0095] Preferably, the change of viewpoint corresponds to, is or is a function of a change in position of the head-mounted display or 3D display device in space and / or a movement relative to the scene as seen by the observer.

[0096] Preferably, the area of ​​attention corresponds, in or for each of the images of at least one pair of images, to the same object(s) in the imaged / displayed scene. Preferably, the area of ​​attention corresponds, in or for each of the images of at least one pair of images, to the pixel(s) or group of pixels corresponding to the same object(s) in the imaged / displayed scene.

[0097] Preferably:

[0098] - the change in viewpoint obtained is a function of a change in position and / or orientation and / or tilt, in space, of the head-mounted display or the 3D display device, or of a change in position and / or orientation and / or tilt, in space, of an observer relative to the 3D display device, and / or

[0099] - The attention zone corresponds to an area of ​​the displayed scene on which the observer's gaze is focused.

[0100] The viewpoint can correspond to the center of the line segment connecting the same part of each of the observer's eyes, for example to the center of the line connecting the fovea or the cornea or the iris of each of the observer's eyes.

[0101] Preferably, the viewpoint corresponds to the respective positions of each of the observer's eyes, for example, the respective positions of the fovea, cornea, or iris of each of the observer's eyes. In this case, the change in viewpoint, according to which the images of at least one pair of source images will be modified, will be distinct for each of the two images of at least one pair of source images.

[0102] The change in the observer's point of view can be defined as the observer's shift from the previous point of view to the new point of view.

[0103] The observer's point of view can be defined as:

[0104] - the position and / or orientation of the observer, preferably such as the position and / or orientation of the observer's head and / or eyes, or

[0105] - the position and / or orientation of each of the observer's eyes, preferably as the respective position and / or orientation of the foveae or corneas or irises of each of the observer's eyes.

[0106] Preferably, the area of ​​attention corresponds to the point or object, or to the pixels or group(s) of pixels, or to the area of ​​the images in at least one pair of source images on which the observer's gaze is focused or intended to be focused. Preferably, the area of ​​the images in at least one pair of source images on which the observer's gaze is focused corresponds to the same point, object, or group(s) of object(s) in the imaged scene.

[0107] Preferably, the area of ​​attention can be determined or calculated from information relating to the area of ​​the images of at least one pair of source or displayed images on which the observer's gaze is focused or intended to be focused.

[0108] Preferably, the area of ​​attention corresponds to a point, for example a pixel or group of pixels, in each of the images of at least one pair of source images located or intended to be located within the cone of the respective fields of vision of each of the observer's eyes. For example, the area of ​​attention may correspond to the point or area of ​​the images of at least one pair of source images of the scene having the shallowest depth within the cone of the observer's field of vision. Preferably, the cone of the observer's field of vision, preferably extending around an axis of revolution of the optical axis of one of the observer's eyes, is formed by a solid angle of approximately 60°, including a solid angle between 3° and 5° where visual acuity is maximal, relative to the optical axis of one of the observer's eyes.

[0109] Preferably:

[0110] - the change of viewpoint is obtained from position data, for example position and / or orientation and / or tilt data, or from position change data of the head-mounted display and / or movement data of the head-mounted display, for example a change or variation in position data, referred to as head-mounted display data, and / or

[0111] - the change of viewpoint is obtained from images of the observer and / or the observer's environment and / or the head-mounted display or 3D display device, and / or

[0112] - the area of ​​attention is obtained from images of the observer's eyes.

[0113] Preferably, at least one image of the observer's eyes comes from an imaging device, preferably from one or more imaging devices of the head-mounted display or 3D display device, capable of imaging the observer and / or the observer's eyes.

[0114] The observer's viewpoint and / or area of ​​attention can be determined using eye tracking. Data from the head-mounted display (HUD) or 3D display device can be measured or detected by the HUD or 3D display device. Changes in viewpoint can be determined or calculated from this data.

[0115] Images of the observer and / or the head-mounted display or 3D display device can be acquired by an external imaging device, preferably one that is not part of or integrated into the head-mounted display or 3D display device. Images of the observer's environment can be acquired by the external imaging device and / or by one or more imaging devices of the head-mounted display or 3D display device.

[0116] The external imaging device can be arranged to image the observer and / or the head-mounted display or the 3D display device.

[0117] Preferably, in this application, "position data" means: a position, preferably relative, and / or an orientation, preferably relative, and / or an inclination, preferably relative.

[0118] Preferably, the method includes linear or non-linear filtering of the data from the head-mounted display or 3D display device to limit the amplitude of the change in viewpoint obtained, as a function of time, to a value less than a limit value.

[0119] Preferably, the method includes non-linear clipping of the data from the head-mounted display or 3D display device to limit the amplitude of the viewpoint change obtained to a value below a limit value.

[0120] Preferably, the method includes a nonlinear dynamic compression of a gain applied to the data of the head-mounted display or 3D display device to attenuate or amplify a variation in amplitude or an amplitude of the change in viewpoint obtained, as a function of time, so that said amplitude does not exceed a threshold value of amplitude or variation in amplitude.

[0121] Preferably, the method includes a nonlinear dynamic gain amplification applied to the data from the head-mounted display or 3D display device to increase the amplitude variation of the resulting viewpoint change over time. The filtering, clipping, and / or dynamic gain compression steps limit and / or reduce the amplitude and / or the amplitude variation over time of the viewpoint change.

[0122] The change in position and / or orientation and / or tilt, in space, of the 3D headset or display device and / or of the observer can be defined or qualified as a change of real or effective point of view.

[0123] The effective change of viewpoint can be defined as the change of viewpoint measured or detected by the head-mounted display or 3D display device, or determined or calculated from the data of the head-mounted display or 3D display device.

[0124] The resulting change in viewpoint may be equal to, correspond to, or be proportional to the actual change in viewpoint. This is particularly the case, for example, when the magnitude and / or the change in magnitude of the change in viewpoint is less than the limit value(s) or the threshold value.

[0125] The resulting change of viewpoint may differ from the actual change of viewpoint. This is particularly the case, for example, when the magnitude and / or the variation in magnitude of the change of viewpoint exceeds the limit value(s) or the threshold value.

[0126] Preferably, the process includes an acquisition step, preferably implemented by the head-mounted display or the 3D display device:

[0127] - of scene images, said acquired scene images comprising all pairs of source images, and / or

[0128] - images of the observer, which may include images of the observer's eyes, and / or

[0129] - images of at least part of a headset environment or at least part of a 3D display device environment.

[0130] According to the invention, a data processing device is also proposed, comprising means arranged and / or programmed and / or configured to implement the method according to the invention. The invention also proposes a computer program comprising instructions which, when the program is executed by a computer, cause the computer to implement the method according to the invention.

[0131] According to the invention, a computer-readable medium is also proposed comprising instructions which, when executed by a computer, lead the computer to implement the process according to the invention.

[0132] According to the invention, a virtual reality / augmented reality (VR / VA) headset, referred to as a headset, is also proposed, comprising one or more means for displaying 3D videos, and:

[0133] - means arranged and / or programmed and / or configured to implement the process according to the invention, and / or

[0134] - means of communication arranged and / or programmed and / or configured to and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the process according to the invention.

[0135] Preferably, the display means of the head-mounted display are capable or arranged to provide an observer with a 3D rendering of the displayed 3D video, said displayed 3D video comprising or consisting of a set of stereoscopic 2D image pairs.

[0136] The head-mounted display is designed or equipped to display 3D videos in monocular and / or binocular vision. The term "head-mounted display" can also refer to 3D glasses.

[0137] Preferably, the head-mounted display includes at least one imaging device arranged to image the eyes of an observer wearing said head-mounted display.

[0138] Preferably, the headset should include:

[0139] - at least one imaging device, preferably different from the observer's eye imaging device, preferably comprising both separate imaging devices, arranged to image at least a part of an environment of said VR / VA headset and / or at least a part of an environment of the observer, and / or

[0140] - means arranged to detect a movement and / or a position and / or an orientation and / or an inclination of said VR / VA head-mounted display.

[0141] According to the invention, a 3D display device is also proposed, arranged to display 3D videos, or comprising one or more means for displaying 3D videos, including:

[0142] - means arranged and / or programmed and / or configured to implement the process according to the invention, and / or

[0143] - means of communication arranged and / or programmed and / or configured to and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the process according to the invention.

[0144] Preferably, the 3D display device includes at least one imaging device arranged to image the eyes of an observer.

[0145] Preferably, the 3D display system includes:

[0146] - at least one imaging device, preferably different from the observer's eye imaging device, preferably comprising both distinct imaging devices, arranged to image at least a part of an environment of said 3D display device and / or the observer and / or an environment of the observer, and / or

[0147] - means arranged to detect a movement and / or a position and / or an orientation and / or an inclination of the observer relative to said 3D display device.

[0148] By way of non-limiting examples, the 3D display device can be a screen, for example LCD or plasma, or a video projector or a tablet or a smartphone.

[0149] Preferably, at least a portion of the environment of the head-mounted display or 3D display device includes at least a portion of the scene facing the head-mounted display, that is, the scene located in front of or opposite the observer, preferably the entire scene. Preferably, the scene facing the head-mounted display is the imaged scene, in other words, the scene contained or represented in each of the images in the set of source image pairs.

[0150] Preferably, at least part of the observer's environment includes all or part of the lateral and / or rear environment of the head-mounted display or 3D display device, i.e., located laterally and / or behind the observer.

[0151] At least one or each of the source image pairs may be acquired or originate or be obtained, but not necessarily, by the imaging devices of the head-mounted display or the 3D display device.

[0152] Preferably, the head-mounted display or 3D display device includes both separate imaging devices.

[0153] Preferably, the head-mounted display or 3D display device is arranged to measure and / or detect the change in viewpoint or to determine and / or calculate the change in viewpoint from the data in the head-mounted display or 3D display device.

[0154] Preferably, the head-mounted display and the 3D display device according to the invention are arranged to provide, enable, or deliver 3D rendering to the pairs of reconstructed stereoscopic 2D images. Where appropriate, the head-mounted display may include, or the 3D display device may be paired with, or used with:

[0155] - synchronous shuttering devices, active (or synchronous) glasses in the case of the 3D display device,

[0156] - means of polarization or polarizing glasses, for example synchronous, polarizing glasses or glasses with red / blue filters.

[0157] Preferably, according to the invention, 3D rendering to stereoscopic 2D images means any display method or process that provides a 3D rendering from a pair of displayed 2D images.

[0158] A person skilled in the art will be able to consider the different display methods that can provide a 3D rendering from two 2D images to be displayed.

[0159] Preferably, the two images of each pair of rendered images, whose stream of pairs of rendered images forms or constitutes the 3D video, can be displayed simultaneously in the binocular head-mounted display, with one image displayed being seen by the right eye and one image displayed being seen by the left eye.

[0160] Preferably, the two images of each pair of rendered images, whose stream of rendered image pairs forms or constitutes the 3D video, can be displayed simultaneously, by spatial multiplexing, in the monocular head-mounted display or the 3D display device, for example as a single composite or rasterized image comprising a portion of each of the two images of each pair of rendered images. In this case, only a portion of each of the two images of each pair of rendered images is displayed in a single image.

[0161] Preferably, the two images of each pair of rendered images can be displayed simultaneously in the monocular head-mounted display or 3D display device by superimposing the two images of each pair of rendered images, each of the two images of each pair of rendered images displayed having a distinct polarization.

[0162] Preferably, the two images of the scene from each pair of rendered images can be displayed consecutively or alternately or iteratively, for example by time-division multiplexing, in a synchronous monocular head-mounted display or on a 3D display device, preferably matched by the wearing of synchronous glasses by the user.

[0163] Preferably, the head-mounted display is a binocular head-mounted display in which each image of the scene from each pair of rendered images is displayed so as to be seen by only one of the eyes in order to allow stereoscopic vision of the scene.

[0164] Preferably, the headset is worn by the user during the implementation of the process.

[0165] Preferably, the method according to the invention is suitable, more preferably is particularly adapted, even more preferably is designed, and most advantageously is specially designed, for implementation by the VR / VA head-mounted display or the 3D display device according to the invention. Furthermore, any feature of the method according to the invention is directly transferable to the VR / VA head-mounted display or the 3D display device according to the invention, and vice versa. Description of the figures

[0166] Other advantages and features of the invention will become apparent from the detailed description of implementations and embodiments, which are by no means limiting, and the following accompanying drawings: FIGURE 1 is an example of a 3D video with spatial immersion viewed in prior art virtual reality glasses; FIGURE 2 represents a flowchart of an embodiment of the method according to the invention; FIGURES 3A and 3B are schematic representations of two examples of changes in viewpoint of a scene by an observer; FIGURES 4A and 4B are schematic representations illustrating a pair of stereoscopic 2D source images of a scene; FIGURES 4C and 4D are schematic representations illustrating a pair of stereoscopic 2D images of the reconstructed scene; FIGURES 5 to 7 illustrate a schematic representation of a perspective top view of the scene displayed in FIGURES 4A and 4B.Figure 8 is a schematic representation of a depth map of the scene illustrated in Figures 4A and 4B.

[0167] Description of the implementation methods

[0168] The embodiments described below are not exhaustive; variants of the invention may include, in particular, a selection of the described features, isolated from the other described features (even if this selection is isolated within a sentence containing these other features), provided that this selection of features is sufficient to confer a technical advantage or to differentiate the invention from the prior art. This selection includes at least one feature, preferably a functional feature without structural details, or with only a portion of the structural details if this portion alone is sufficient to confer a technical advantage or to differentiate the invention from the prior art.

[0169] With reference to FIGURES 1 to 7, an embodiment of the dynamic 3D video observation method according to the invention is presented, referred to as the method in the remainder of the description.

[0170] The person skilled in the art will be able to understand the concepts of 3D images, stereoscopy, monocular vision and binocular vision as well as their implementation.

[0171] Figure 1 illustrates an example of monocular 3D video viewed through state-of-the-art virtual reality glasses, designed to provide dynamic viewing for an enhanced sense of immersion. Figure 1 shows four images of a static scene from the 3D video in which the observer's feet are visible facing a beach with the sun in the background. The images illustrate how the scene is rendered by the virtual reality glasses during a lateral panning motion from left to right and vice versa. The image on the left corresponds to the image rendered when the observer, wearing the virtual reality glasses, has their head turned to the left, and the image on the right corresponds to the image rendered when the observer has their head turned to the right.

[0172] Figure 1 shows that state-of-the-art dynamic observation devices and methods rotate the images of the scene to enhance the observer's experience. Furthermore, vanishing points (such as the foreground and the sun) maintain their orientation relative to the observer when the observer turns their head. From a perceptual standpoint, it is as if the scene were rotating with the observer. In practice, the observer continues to see the scene from the same perspective as they would from a fixed camera. However, the rotational movements, visible on the outer frame of the image, aim to create the effect of movement relative to the scene. This effect, however, is disruptive and detracts from the immersive experience.

[0173] The dynamic 3D video observation method according to the invention aims to correct these drawbacks.

[0174] Figure 2 illustrates a flowchart of the process according to the invention.

[0175] The 3D video of a scene 1 comprises a set of n pairs of stereoscopic 2D images, called source image pairs, denoted (IS1, IS2)n. A given source image pair is denoted (IS1, IS2)i, where, for the considered pair, denoted i, of source images: IS1 is one of the two images of the considered pair i, and IS2i is the other of the two images of the considered pair i. The 3D video comprises a set, a stream, or a stack of n source image pairs (IS1, IS2). IS1 is an image acquired by a first imaging device, and IS12 is an image acquired by a second imaging device. Each IS1 image of each source image pair i is acquired by the first imaging device, and each IS12 image of each source image pair i is acquired by the second imaging device. An imaging device according to the invention is a camera or any device comprising, among other things, a lens associated with a photographic sensor.

[0176] The two images II, 12 of each source image pair are two distinct images of the same scene acquired, preferably simultaneously, by two separate cameras, in two distinct positions in space. The two images II, 12 of each source image pair correspond to the observation of the scene from a different viewpoint separated by a distance corresponding to the spacing between the two cameras II, 12. Each stereoscopic 2D image pair (II, 12) of the 3D video is therefore capable of providing 3D rendering to the video by displaying the stream of source image pairs (II, I2)n in a head-mounted display or on a 3D display device.

[0177] Any stereoscopic 2D source image pair consisting of two distinct images of the same scene acquired from two different positions in space is capable of providing, and will provide, a 3D rendering by displaying the source image pair in a head-mounted display or on a 3D display device. According to the invention, "capable of providing a 3D rendering" does not refer solely to the case where the two images were acquired with a horizontal spacing equal to the interpupillary distance. This scenario is a special case. The method according to the invention is intended to be implemented using a virtual reality / augmented reality (VR / VA) head-mounted display, referred to as a head-mounted display, or any 3D display device capable of providing a 3D rendering. The invention also proposes a VR / VA head-mounted display and a 3D display device.The VR / VA headset and the 3D display device each include one or more means for displaying 3D videos, in monocular and / or binocular mode, consisting of a set n of stereoscopic 2D image pairs. The VR headset and the display device each further include:

[0178] - means arranged and / or programmed and / or configured to implement the process according to the invention, and / or

[0179] - means of communication arranged and / or programmed and / or configured to and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the process according to the invention, and / or

[0180] - at least one imaging device.

[0181] The method according to the invention includes the step of obtaining a change of viewpoint of the observer. The change of viewpoint is, according to the non-limiting embodiment, a function of a change in the position and / or orientation and / or inclination of the observer, and in particular of the observer's head, in space.

[0182] According to one improvement, the change of viewpoint is a function of a change in the relative position and / or orientation and / or tilt of the observer. The step of obtaining the change of viewpoint results in obtaining the new viewpoint, or the new position, from which the observer would see scene 1 if they were observing it directly.

[0183] For this purpose, the head-mounted display includes means arranged to detect movement and / or position and / or orientation and / or tilt of the head-mounted display. By way of non-limiting examples, the means arranged to detect movement and / or position and / or tilt and / or orientation may be a gyroscope or an accelerometer. The means arranged to detect movement and / or position and / or orientation and / or tilt may be an inertial measurement unit (IMU). For this purpose, the 3D display device, and possibly the head-mounted display in addition to the IMU, include means arranged to detect movement and / or position and / or orientation and / or tilt of the observer relative to said 3D display device or the head-mounted display. In this case, the change of viewpoint is obtained from images of the observer and / or the observer's environment and / or the 3D display device or the head-mounted display.The change of viewpoint can be achieved, alternatively or in combination:

[0184] - from data originating from observer images, for example from images from at least one imaging system of the 3D display device, or from at least one imaging system of the head-mounted display, configured to image an observer's eyes, and / or

[0185] - from data from the observer's environment, for example from images from at least one imaging device of the 3D display device arranged to image at least part of the environment, or from at least one imaging device of the head-mounted display arranged to image at least part of the environment.

[0186] Advantageously, the change of viewpoint obtained is calculated and / or determined from data from means arranged to detect a movement and / or a position and / or an inclination and / or an orientation of the head-mounted display and / or the observer relative to said 3D display device and / or the head-mounted display.

[0187] Preferably:

[0188] - the change of viewpoint includes a lateral (du, dv) (along the U axis) and / or vertical (dv) (along the V axis) displacement of the observer, that is, in the observer's frontal plane, preferably of the observer's head; this displacement (du, dv) may, in some cases, be a displacement (du, dv) in, or relative to, a frame of reference of the images of the source image pair or in, or relative to, the frame of reference of one of the images of the source image pair, and / or

[0189] - The change of viewpoint includes a lateral and / or vertical displacement (du, dv) of the head-mounted display or a lateral and / or vertical displacement (du, dv) of the observer, i.e., in the observer's frontal plane, relative to the 3D display device, and / or - The change of viewpoint includes a longitudinal displacement (dz) (along the Z-axis) of the observer, i.e., in the observer's sagittal plane, preferably of the observer's head; this displacement (dz) may in some cases be perpendicular to the image reference frame of the source image pair or to the reference frame of one of the images of the source image pair, and / or

[0190] - the change of viewpoint includes a displacement (dz), perpendicular to the plane formed by the axes (du, dv), of the visioheadset or a displacement (dz), perpendicular relative to the plane formed by the axes (du, dv), that is to say along the axis connecting the observer to the 3D display device, of the 3D display device, or of the means of displaying the 3D device.

[0191] With reference to Figures 3A and 3B, two illustrative examples of a change in the observer's viewpoint are presented. The left-hand images in Figures 3A and 3B illustrate the observer's viewpoint before the change, and the right-hand images in Figures 3A and 3B illustrate the change in viewpoint as the observer would see it if they were directly observing Scene 1. The observer's field of view 11 is also shown. In Figure 3A, the change in viewpoint corresponds to a displacement dv parallel to the plane of Scene 1 and / or the 3D display device, which would correspond to a horizontal displacement of the observer if they were actually observing Scene 1 as depicted. In this case, there is no change in distance between the observer and Scene 1 and / or between the observer and the 3D display device.In FIGURE 3B, the change of viewpoint corresponds to a displacement dv parallel to the plane of scene 1 and / or the 3D display device and to a displacement dz perpendicular to the plane of scene 1 and / or the 3D display device. The displacement dz would correspond to a horizontal displacement moving the observer who was actually observing scene 1 as depicted closer to or further away from the scene.

[0192] Figure 3A shows the observer's optical axis 6 before the change of viewpoint and the observer's optical axis 7 after the change of viewpoint. Figure 3A illustrates a change of viewpoint during which the observer keeps their gaze focused on the same area, object, or group of objects in scene 1, referred to as the attention zone 12. According to the invention, the method includes filtering the data from the head-mounted display and the 3D display device to limit the amplitude of the resulting change of viewpoint, particularly as a function of time, to a value below a limit. For example, the limit value could be the absolute value or the magnitude of the displacement vector determined from the data from the head-mounted display or the 3D display device. Preferably, the filtering is linear.

[0193] For this purpose, an effective viewpoint change can be defined as the viewpoint change measured or detected by the head-mounted display or 3D display device, or determined or calculated from the data of the head-mounted display or 3D display device. In other words, the effective viewpoint change corresponds to the actual viewpoint change experienced by the head-mounted display and by the user, or by the observer relative to the 3D display device or head-mounted display. The resulting viewpoint change may, in some cases, correspond to or be equal to the effective viewpoint change. However, in most cases, the resulting viewpoint change will be a function of, calculated from, or determined from the effective viewpoint change, but will not be equal to the effective viewpoint change.

[0194] The primary objective of the filtering step is to eliminate changes in viewpoint with a large amplitude. In other words, the filtering step aims to prevent abrupt variations or gradients in viewpoint changes, for example, when the observer wearing the headset turns abruptly or pivots sharply 90°. Indeed, since the modification made to the images of the source image pairs (IS1, IS2) is a function of the viewpoint change, beyond a certain amplitude of viewpoint change, the objects in the images of the rendered image pairs (IR1, IR2) would be too strongly distorted and / or spread out, potentially blurring or rendering the modified images (IR1, IR2) unobservable. Indeed, although it is planned to add and / or remove certain pixels, groups of pixels or objects from the source image pairs, the invention is not intended to generate data or information not present in the source image pairs.In other words, the filtering step aims to truncate or disregard excessively large movements within a short timeframe, ensuring that the resulting changes in viewpoint do not deviate too much from the initial position and orientation of the observer's head, whether wearing the headset or viewing the 3D display. The filtering step may also include filtering or smoothing small amplitude changes, such as minor physiological head movements. This prevents the observer from experiencing frequent, subtle changes in the source image pairs (IS1, IS2), which could induce motion sickness in the observer wearing the headset or viewing the 3D display.In other words, the filtering step can be defined as involving the application of a bandpass filter, specifically a high-pass filter, or subtracting an estimate of the average position of the headset, or of the observer relative to the 3D display device, from the actually detected viewpoint changes. This therefore smooths out the actually detected viewpoint change.

[0195] Viewpoint changes, which can be described as high-amplitude viewpoint changes, are changes or variations in amplitude occurring at low frequencies. High amplitude can be defined as an amplitude greater than 1 cm. Low-frequency amplitude changes or variations can be defined as having a frequency less than 1 Hz. High-frequency viewpoint changes can generally be observed as small-amplitude movements of the head-mounted display, typically less than 0.1 cm. High-frequency viewpoint changes can be defined as having a frequency greater than 10 Hz. The concepts of "low" and "high" amplitude changes or variations are relative; frequency and amplitude are not synonymous.

[0196] Depending on the embodiment, the method includes nonlinear clipping of the data from the head-mounted display or 3D display device to limit the amplitude of the resulting viewpoint change to a value below a limit. Clipping can be defined as a cut applied to the actual viewpoint changes, or to the data relating to the actual viewpoint changes, so as to limit or maintain the amplitude of the resulting viewpoint change used to modify the source image pairs (IS1, IS2) below a maximum amplitude. Indeed, since the modification made to the source image pairs (IS1, IS2) is a function of the viewpoint change, beyond a certain amplitude of the resulting viewpoint change, the objects contained in the rendered image pairs (IR1, IR2) would be too strongly distorted and / or spread out, potentially rendering the modified images (IR1, IR2) blurry or unobservable.This step aims to limit the large effective viewpoint change values ​​to keep them below predefined maximum values.

[0197] The method may include nonlinear dynamic gain compression applied to the data from the head-mounted display or 3D display device to attenuate a variation in amplitude or the amplitude of the resulting viewpoint change over time, such that said amplitude does not exceed a threshold value. This nonlinear dynamic gain compression of the effective viewpoint change is complementary to clipping the effective viewpoint change. Nonlinear dynamic gain compression can be defined as a modulation of a gain applied to the effective viewpoint changes, or to the data relating to the effective viewpoint changes, so as to limit the amplitude of the resulting viewpoint change used to modify the source image pairs (IS1, IS2).Indeed, besides cutting off the strong sudden amplitude variations of the effective viewpoint change, it may be wise to limit the amplitude of the viewpoint change to be restored to the limitation of the modification of the source image pairs (IR1, IR2) below a predefined value beyond which the restored image pairs (IR1, IR2) are no longer usable.

[0198] Finally, the method includes a nonlinear dynamic gain amplification applied to the data of the head-mounted display or 3D display device to attenuate or amplify a time-dependent variation in the amplitude of the resulting viewpoint change. Preferably, this dynamic gain amplification step can advantageously be implemented following the nonlinear dynamic gain compression step applied to the data of the head-mounted display or 3D display device. The nonlinear dynamic gain amplification of the data of the head-mounted display or 3D display device can advantageously be implemented during a change, particularly a reversal, in the direction of movement of the head-mounted display or the user relative to the 3D display device. Nonlinear dynamic gain amplification can be defined as a modulation of the gain applied to the actual viewpoint changes, or to the data relating to the actual viewpoint changes.In other words, nonlinear dynamic gain amplification applied to the data from the head-mounted display or 3D display is advantageously implemented when previous movements (or the amplitude of previous movements) of the head-mounted display or the user relative to the 3D display were such that nonlinear dynamic gain compression was required. Also, in such a case, when the movement of the head-mounted display or the user relative to the 3D display reverses, the gain applied to the actual viewpoint changes can be increased to restore a rapid perception of movement to the user.

[0199] The process further includes the modification step of modifying each of the two images of at least a portion of the set of n source image pairs (IS1, IS2)n to reproduce at least one pair of stereoscopic 2D images of the scene (IR1, IR2), referred to as at least one pair of reproduced images (IR1, IR2). The pairs of reproduced images (IR1, IR2) are intended to provide 3D rendering by display in a head-mounted display or on a 3D display device and offer a change of viewpoint relative to the 3D video of the scene. Thus, each pair of reproduced images (IR1, IR2) provides a change of viewpoint relative to the corresponding pair of source images (IS1, IS2) (whose two images (IS1, IS2) have been modified according to the process to reproduce the two images (IR1, IR2)).

[0200] The method further includes a step of displaying, in the head-mounted display or 3D display device, at least one pair of rendered images (IR1, IR2) when a change of viewpoint is obtained, or at least one pair of source images (IS1, IS2) when no change of viewpoint is obtained. Indeed, it is possible that at certain moments in the 3D video, there may be no change of viewpoint for some of the images in the set of n pairs of images in the 3D video (for example, if the observer remains stationary in the reference position for a period of time while viewing the video).

[0201] Preferably, the source image pairs (IS1, IS2) and the rendered image pairs (IR1, IR2) can be displayed:

[0202] - simultaneously, in monocular or binocular vision, or consecutively or alternately or iteratively, in monocular vision.

[0203] It can be understood as "consecutively displaying pairs of images or the stream of pairs of images": a successive, alternating, sequential, or looping display of the pairs of images in the scene. It can also be understood as "consecutively displaying pairs of images or the stream of pairs of images in the scene": a time-division multiplexing display of the two images in the pairs of images.

[0204] Simultaneous display of image pairs can be understood as the superimposition of the two images of the pairs, with each image of a displayed pair exhibiting a distinct polarization. Simultaneous display of image pairs can also be understood as the spatial multiplexing of the two images of the pairs, that is, the display of a single composite or rasterized image comprising a portion of each of the two images of the displayed pair.

[0205] The two images of the pairs / couples of images, or the stream of pairs / couples of images forming or constituting the 3D video, can be displayed simultaneously in binocular vision, with one displayed image being seen by the right eye and one displayed image being seen by the left eye.

[0206] For the same reason, the modification step of the process may only be implemented for a subset of the n source image pairs (IS1, IS2)n. Indeed, if no change in viewpoint is achieved, the modification step is not implemented, and the source image pairs (IS1, IS2) are therefore displayed without any change in viewpoint.

[0207] Advantageously, as soon as a change of viewpoint is obtained, the source image pairs (IS1, IS2) are modified according to the process so as to reproduce the change of viewpoint of the scene by displaying, in the head-mounted display or the 3D display device, the pairs of reproduced images (IR1, IR2) offering an observation of the scene from the new viewpoint of the observer.

[0208] The modification step includes moving and / or spreading and / or shrinking pixels or groups of pixels in or within the images of at least one pair of source images (IS1, IS2) and / or adding and / or deleting pixels or groups of pixels in or within the images of at least one pair of source images (IS1, IS2).

[0209] The modification step is implemented based on:

[0210] • a depth map of the imaged scene,

[0211] • of the change of viewpoint obtained. For a simple implementation of the invention, the modification of the images of at least one pair of source images (IS1, IS2) can be proportional to the change of viewpoint.

[0212] According to the invention, the modification is achieved, among other things, by parallax effect.

[0213] For example, the modification is implemented using a matrix transformation operator, also called a transform operator. Advantageously, the transform operator is arranged and / or configured to apply a parallax effect to the modified images during the modification step.

[0214] The two images thus modified constitute the two images of the pair of restored images (IR1, IR2) providing the change of viewpoint.

[0215] In particular, the modification step includes a displacement and / or contraction, or may include or consist of moving, for example by translation, by means of any function, operator, matrix or grid, called a transformation operator, coordinates or positions of all or part of the pixels or groups of pixels within the grid or pixel matrix of the image to be modified.

[0216] In some cases, the modification step may include spreading and / or shrinking and / or weighting the displacement, using the transformation operator, of the coordinates or positions of some of the pixels or groups of pixels within the grid or pixel matrix of the image to be modified. In particular, spreading and / or shrinking and / or weighting the displacement may be applied to pixels or groups of pixels whose value or local variance of a parameter or channel, by way of non-limiting examples: sharpness, depth, or any other parameter known to those skilled in the art, such as, by way of non-limiting examples, the parameters R, G, B (for R (Red), VI (Green, or G (Green)), V2 (Green, or G (Green)), B (Blue)), a dilation parameter, and / or an edge parameter.Furthermore, in these cases, the modification step may include the deletion and / or generation of pixels, using the transformation operator, particularly pixels near the edges of the images to be modified. The pixels to be generated may be created, for example, by transposition, stretching, and / or shrinking, from adjacent or neighboring pixels or groups of pixels.

[0217] The transformation operator described above, used for the modification step, may include, be supplemented by, or implement all the features described below, according to which the modification step is implemented. Furthermore, the modification step may include pixel interpolation in order, for example, to add missing pixels (particularly during significant changes in viewpoint). Those skilled in the art are familiar with pixel interpolation techniques, such as linear interpolation based on the respective distances to neighboring pixels used as reference points for the interpolation, or interpolation using spline functions.

[0218] Advantageously, with reference to Figures 5 to 7, the method includes calculating at least one angle corresponding to the change of viewpoint. This at least one angle corresponds to the angle formed between the observer's optical axis 6 before the change of viewpoint and the optical axis 7 after the change of viewpoint. The modification of the source image pair 21, 22 is performed according to this calculated at least one angle.

[0219] The process may include calculating at least one angle, called angle alpha (a), whose axis of rotation is axis U. In other words, angle a defines a rotation around axis U. In this embodiment, axis U is a vertical axis. In this embodiment, axis U is an axis parallel to or contained within scene 1 as seen by the observer.

[0220] The process may include calculating at least one angle, called a Beta angle (3), whose axis of rotation is the V axis. In other words, angle 3 defines a rotation around the V axis. The V axis is perpendicular to the U axis. The axis is, according to the embodiment, a horizontal axis. The V axis is, according to the embodiment, an axis parallel to or contained within scene 1 as seen by the observer.

[0221] The method may include calculating at least one angle, called angle Gamma(y), whose axis of rotation is the Z-axis. In other words, angle y defines a rotation around the Z-axis. The Z-axis is, according to the embodiment, a horizontal axis. The Z-axis is, according to the embodiment, an axis perpendicular to or included in scene 1 as seen by the observer.

[0222] Thus, the process may include the calculation of an angle a, an angle 3 and / or an angle y.

[0223] The angles α, γ, and γ can be defined relative to a frame of reference of scene 1 as seen by the observer, or to a frame of reference linked to or corresponding to scene 1 as seen by the observer, and / or to a frame of reference of the user, the head-mounted display, and / or the 3D display device. In most cases, each of the two images of a given pair of source images (IS1, IS2) is modified identically during the modification step. In some specific cases, the two images of a given pair of source images (IS1, IS2) are modified separately during the modification step.

[0224] Indeed, in most cases, the resulting change in viewpoint will reflect an identical displacement, translation, and / or rotation by each of the observer's two eyes. In this case, the modification applied to each of the two images of a given source image pair (IS1, IS2) is identical. However, in certain specific cases, the change in viewpoint of one of the observer's eyes is different from the other, for example, in the case of rotation around the optical axis of one of the observer's eyes.

[0225] In a particular case, the process may also include the calculation of an angle, called angle Gamma (y) according to the embodiment, corresponding to a change of viewpoint, or to a component of the change of viewpoint, by rotation (by precession) around the optical axis of the observer (the attention zone 12 being, preferably, intersected by the optical axis of the observer).

[0226] In the specific case where the change of viewpoint includes, or a component of the change of viewpoint includes, a rotation by angle y, the change of viewpoint of one of the observer's eyes is different from the other. In this case, the modification applied, as a function of angle y, to the two images of a considered source image pair i (IS1, IS2)i may be different, or the modification applied, as a function of angle y, may be made to only one of the images of the considered source image pair i (IS1, IS2)i.

[0227] Furthermore, and possibly in combination with the translations of dv or dv, or of +dv, or with the angles α, θ, and γ, the change of viewpoint may also include a translation, called a longitudinal translation, denoted dz. Such a change of viewpoint may be equivalent to, or associated with, a movement of a person directly observing scene 1 along or around its optical axis, preferably with the observer's gaze fixed on the attention zone 12, in the direction connecting the observer to scene 1 (the observer would move closer to scene 1) or in the opposite direction to the direction connecting the observer to scene 1 (the observer would move away from scene 1). In this case, the modification step may include or consist of, for example, a homothety.

[0228] According to an advantageous embodiment, the image modification step of at least one pair of source images (IS1, IS2) is further performed relative to, or as a function of, an area 12 of the images of at least one pair of source images (IS1, IS2), referred to as the attention area 12, during the modification of the images of at least one pair of source images (IS1, IS2). The attention area 12 corresponds to the area of ​​the images ((IR1, IR2) or (IS1, IS2) as the case may be) or of the objects in scene 1 on which the observer's gaze is focused. Attention zone 12 is a unique area of ​​the scene image, specifically of the images ((IR1, IR2) or (IS1, IS2) or of the objects in scene 1, that is, of the two images ((IR1, IR2) or (IS1, IS2) of the 3D video at a given moment. Attention zone 12 corresponds, in particular, to one or more objects in the scene, specifically to one or more objects in the scene having the same depth in the displayed scene images.

[0229] Modifying the images of at least one pair of source images (IS1, IS2) based on attention zone 12 enhances the viewing experience for the 3D video observer. Furthermore, modifying the images of at least one pair of source images (IS1, IS2) based on attention zone 12 increases the immersive experience for the 3D video observer. Additionally, modifying the images of at least one pair of source images (IS1, IS2) based on attention zone 12 reduces the resources required to process the source image stream (IS1, IS2).

[0230] Preferably, attention zone 12 is obtained from images of the observer's eyes.

[0231] According to the non-limiting embodiment, the attention area 12 remains unchanged during the image modification step of at least one pair of source images (IS1, IS2).

[0232] Also, the displacement, spreading, and / or contraction of pixels is achieved through parallax relative to the attention zone 12, which remains unchanged during the modification step. In other words, the transformation operator is configured to shift the coordinates or positions of pixels, through parallax, relative to the pixels of the attention zone 12. If necessary, when considering the attention zone 12 during the modification step, the modification step may include the deletion and / or generation of pixels, using the transformation operator, particularly pixels in the vicinity of the attention zone. The pixels to be generated can be created, for example, by transposition, spreading, and / or contraction of adjacent or neighboring pixels or groups of pixels. However, pixels located at the edge of the attention zone 12 are only slightly displaced by the parallax effect implemented relative to the attention zone 12.

[0233] Taking into account attention zone 12 during the modification step preferably involves modifying the images of the source image pairs (IS1, IS2) as soon as a change in attention zone 12 is detected / obtained, but this approach offers several advantages. Indeed, an observer watching a video focuses their attention on a different area or objects within scene 1 over time. Therefore, modifying the images as a whole can detract from immersion because the object in scene 1 on which the observer focuses their gaze while watching the 3D video is altered.

[0234] Advantageously, objects are moved, preferably translated, during the modification step proportionally to a depth gap AP corresponding to the difference between the depth of an object in the scene 1 under consideration, located outside the attention zone 12, and the depth of the attention zone 12.

[0235] The attention zone 12 can be deduced from the observer's gaze data. The user's gaze data can be obtained by eye tracking. The method may therefore include the step of acquiring the attention zone 12 and / or the step of determining the attention zone 12 from this data. In this case, the user's gaze data can be acquired by one or more imaging systems of the head-mounted display or the 3D display device arranged to image the eyes of an observer wearing said head-mounted display. A person skilled in the art is familiar with eye movement recording techniques and also with the data that can be derived from them, such as, for example, the foveal scan. The method may include the step of acquiring the observer's gaze data, for example, by means of one or more optical systems which may include one or more cameras.With reference to FIGURES 4A, 4B, 4C and 4D, an advantageous embodiment of the process is illustrated in which the modification step is implemented according to the depth map of the imaged scene, the change of viewpoint obtained and the area of ​​attention 12. On FIGURE 4A is illustrated the left image ISli of a pair, denoted i, of source images (IS1, IS2)i considered, on FIGURE 4B is illustrated the right image IS2i of the pair i of source images (IS1, IS2)i considered. Figure 4C illustrates the left image IRli of the pair, denoted i, of restored images (IR1, IR2)i, associated with the pair i of source images (IS1, IS2)i considered, and Figure 4D illustrates the left image IRli of the pair i of restored images (IR1, IR2)i, associated with the pair i of source images (IS1, IS2)i considered.

[0236] According to this non-limiting embodiment, the observer maintains their gaze focused on the same attention zone 12 during the change of viewpoint. This embodiment allows for consideration of the scenario in which the modification of the pair of source images (IS1, IS2)i is performed according to the attention zone 12 on which the observer's gaze remains focused during the change of viewpoint.

[0237] A person skilled in the art will be able to adapt and transpose the description of this non-limiting embodiment to the case where the observer changes his area of ​​attention 12 over time, for example between the change in the pair i of source images (IS1, IS2)i considered (at time t) and the change in the pair i+1 of source images (IS1, IS2)i (at time t+1).

[0238] In this embodiment, the modification of the pair of source images (IS1, IS2)i considered is also a function of the attention zone 12. According to this embodiment, the attention zone 12 could remain unchanged, during the modification of the pair of source images (IS1, IS2)i considered, by anchoring the attention zone 12 relative to the observer or to a reference frame other than or external to the observer.

[0239] To facilitate description, Figures 5 to 7 show a top-view perspective of scene 1, imaged (IS1, IS2)i, as represented in Figures 3A to 3D; that is, scene 1 with the concept of depth. The optical axis 6 of the observer before the change of viewpoint and the optical axis 7 of the observer after the change of viewpoint are illustrated. In the specific case presented, the given change of viewpoint of the observer corresponds to a displacement parallel to the image plane of the considered pair of source images (IS1, IS2)i displayed, which corresponds to a horizontal displacement of the observer if they were actually observing scene 1 as imaged.

[0240] The attention zone 12, on which the observer's gaze remains focused during the change of viewpoint, is the tree 8, which is located behind the individual 9 but in front of the background landscape 10, for FIGURE 6. The attention zone 12, on which the observer's gaze remains focused during the change of viewpoint, is the background 10, that is, the background landscape 10, for FIGURE 5. The attention zone 12, on which the observer's gaze remains focused during the change of viewpoint, is the foreground 9, that is, the individual 9, for FIGURE 7. The modification of the images of the source image pair (IS1, IS2) illustrated in FIGURES 4A and 4B corresponds to the embodiment presented in FIGURE 6, that is, when the observer's gaze, that is to say, attention zone 12, remains focused on tree 8 during the change of viewpoint.

[0241] Advantageously, the modification of the source image pair (IS1, IS2) involves, for each pixel or group of pixels in each image of the source image pair (IS1, IS2), calculating an AP deviation between the inverse of a depth associated with attention zone 12 and the inverse of a depth associated with each pixel or group of pixels in each image of the source image pair (IS1, IS2) located outside attention zone 12. This AP deviation is calculated from the depth map data. The modification of the source image pair (IS1, IS2) is performed based on the calculated AP depth deviation.

[0242] Advantageously, according to the non-limiting embodiment described, the translation of a given object in an image of the source image pair (IS1, IS2), outside the object of the attention zone 12, will be proportional to the angle a.

[0243] Figures 5 to 7 illustrate the modification of scene 1 as seen by an observer following a change of viewpoint.

[0244] The translation will be performed along the axis connecting the observer's optical axis 6 before the change of viewpoint and the optical axis 7 after the change of viewpoint. Since the objects are translated within the plane of the image to be modified, a trigonometric relationship can be established between the calculated depth difference AP, the angle a, the angle p and / or the angle y, and the distance, denoted m, by which the object must be moved to reproduce the change of viewpoint.

[0245] For example, with reference to FIGURE 6 and FIGURES 4A to 4D, for individual 9 in the foreground of the imaged scene 1, the depth difference AP is equal to PI-PO, or to the absolute value of Pl-PO, where PO corresponds to the depth of individual 9 in scene 1 and PI corresponds to the depth of the attention zone 8, 12 in scene 1. In the case of the background 10 of the image, the depth difference AP is equal to P1-P2, or to the absolute value of P1-P2, where P2 corresponds to the depth of the background 10 in scene 1.

[0246] Thus, by way of non-limiting example, the relationship can be established.

[0247] Furthermore, the method may include calculating a variation in the observer's distance Ad relative to the imaged scene 1. The modification of the images of at least one pair of source images (IS1, IS2) will advantageously be carried out based, in addition, on the calculated variation in distance Ad. The variation in distance can be associated with the change in viewpoint. In practice, in most cases this variation in distance will be negligible compared to the depths of the objects in the imaged scene 1. However, in cases where this variation is not negligible compared to the depths of the objects in the imaged scene 1, it would be possible, as a non-limiting example, to use the relation + Ad) the modification to be made to the images of at least one pair of source images (IS1, IS2).

[0248] It is emphasized that the value "m" and the relationships described above correspond to the physical changes in scene 1 (i.e., the actual or scale changes in scene 1 brought about by a change of viewpoint) as they would be perceived by an observer of scene 1.

[0249] It is specified that the relationships described in this request are indicative only. A person skilled in the art will be able to propose any matrix operator based on simple mathematical tools allowing the calculation of the displacement (spreading and / or contraction and / or shifting) of each pixel or group of pixels in the images of a pair of source images (IS1, IS2) (such as trigonometry and simple geometry (Thales' theorem, etc.)). The following description will describe a specific, non-limiting example of implementing the IS1, IS2 image modification step. This description concerns the modification of the IS1, IS2 images of scene 1 (that is, the modification of pixels within the images (or in the image grid itself) resulting from a change of viewpoint) as they would be displayed to the observer.In this embodiment, the displacement of the pixel(s), or groups of pixels, is described as being inversely proportional to the depth. The modification (the displacement), denoted dpi, of a portion of the pixels located at a given depth Z can be written as: dpi = F*((1 / ZR) - (1 / Z))*TT, equation 1, where F is the so-called focal length of the sensor, or more precisely the distance between the objective optical center and the sensor, used for the acquisition of the image IS1 or IS1 of a pair of source images considered (IS1,IS2)i, TT the component, called transverse, of the viewpoint translation (TT being equal to or corresponding to du or dv or du+dv) and ZR is the depth of, or associated with, the attention zone 12.

[0250] TT can be defined as a direction perpendicular to the central axis of the camera module, or a translation in a plane perpendicular to the optical axis of the observer, or more precisely parallel to the plane of the rendered image, and dpi is the displacement of pixels in the same direction as the TT translation.

[0251] As a further, non-limiting example, in the case of viewpoint changes by longitudinal translation dz, the modification to the coordinates (U1, V1) of a pixel (or group of pixels) in each of the IS1, IS2 images of a pair of source images (IS1, IS2) can be defined as: U2 = (Z / (Z- TL))*U1 and V2 = (Z / (Z-TL))*V1, where Z is the depth of the point or object in the scene under consideration (or associated with the pixel or group of pixels corresponding to the object in the scene under consideration), TL is the longitudinal translation (dz), and (U2, V2) are the coordinates of the pixel after displacement following implementation of the process. It is noted that when Z is large (the point or object in scene 1 has a large depth), the shift is small.Also, an object located at infinity (in practice, objects located at a depth greater than a threshold value, for example located 50 meters or more from the camera that acquired the source images, can be considered as being at infinity) will not be moved and the homothety ratio will be equal to 1.

[0252] In other words, the modification of the source images IS1, IS2 by contraction can correspond to a homothety with a ratio less than 1. Conversely, the spreading can correspond to a homothety with a ratio greater than 1.

[0253] As a further non-limiting example, in the case of changes of viewpoint by transverse translation of the or dv or of the+dv, the modification made to the coordinates (Ul, VI) of a pixel (or groups of pixels) considered of each of the images IS1, IS2 of a pair of source images (IS1, IS2) can be defined as: U2 = Ul + ((1 / (ZR)) - (1 / Z))*F*TTU and V2 = VI + ((1 / (ZR)) - (1 / Z))*F*TTV where TTU is the component in transverse (or perpendicular) translation along the U axis and TTV is the component in transverse (or perpendicular) translation along the V axis.

[0254] Preferably, in this case, the units of U, V, TTU, TTV, and F are homogeneous; they can be physical distances (physical distances or distances of a unit linked to the camera sensor for U, V, and focus distance F, and physical displacement distances for TTU and TTV and the depths Z and ZR). Alternatively, TTU, TTV, ZR, and Z are homogeneous (in physical distances or distances of a unit linked to the camera sensor), and U, V, and F are in the same unit (for example, pixels (F being pixels per radian, the radian being unitless)) which is different from the unit of TTU, TTV, ZR, and Z.

[0255] In practice, the change of viewpoint can include a perpendicular translation component TT and a longitudinal translation component TL. Also, for a change of viewpoint including several translational components, the modification made to the coordinates (Ul, VI) of a pixel (or groups of pixels) considered in each of the images IS1, IS2 of a pair of source images (IS1, IS2) can be defined as: U2 = (Ul + F*((1 / ZR) - (1 / Z))*TTU)*(1 / (1 - TL / Z)) and V2 = (VI + F*((1 / ZR) - (1 / Z))*TTV)*(1 / (1 - TL / Z)).

[0256] In cases where the change of viewpoint involves very large rotation angles (a, [3 and / or y]) that must be taken into account when modifying the source images IS1, IS2, the modification made to the coordinates (Ul, VI) of a pixel (or groups of pixels) considered in each of the images IS1, IS2 of a pair of source images (IS1, IS2) can be defined as:

[0257] - U2 = [(Ul + F*((1 / ZR) - (1 / Z))*TTU)*(1 / (1-TL / Z))]*COS(Y) - [(VI + F*((1 / ZR) - (l / Z))*TTV)*(l / (l-TL / Z)]*sin(Y), and

[0258] - V2 = [(Ul + F*((1 / ZR) - (l / Z))*TTU)*(l / (l-TMZ))]*sin(Y) + [(VI + F*((1 / ZR) - (1 / Z))*TTV)*( 1 / (1-TL / Z)]*C0S(Y) .

[0259] It should be noted that in the particular case of the change of viewpoint including a rotation component of angle Y, it is possible to apply a modification to one of the images of a considered source image pair i (IS1, IS2)i which is different from the modification of the other of the images of the considered source image pair i (IS1, IS2)i.

[0260] In most cases, the changes in viewpoint of the headset or the observer viewing the 3D display are sufficiently small to keep the entire attention zone 12 unchanged. Furthermore, the headset's data processing (filtering, clipping, and / or compression) allows, or at least contributes to, maintaining tolerable changes in viewpoint that keep the attention zone 12 unchanged.

[0261] In certain specific cases of significant viewpoint changes, particularly for viewpoint changes by translation dz and / or for viewpoint changes by significant rotation (especially around angles a, p, and y), the source image pairs (IS1, IS2), including the attention zone 12, may also be modified by spreading, dilation, or stretching and / or by contraction or compression. In the case of significant viewpoint changes by translation dz, the modification step may include a modification of the attention zone 12, for example, a homothety. In the case of viewpoint changes by significant rotation (o, p, y), the modification step may include, for example, stretching and / or compressing the attention zone along one or two axes.However, even in the rare cases requiring a change to attention area 12, the change made to attention area 12 may be different or limited compared to the change to the rest of the source images (IS1, IS2).

[0262] With reference to Figures 4A to 4D and Figure 6, all objects in scene 9, 10, other than the one (the tree 8, 12) on which the observer's gaze is focused, will be translated within the image plane in response to the change in viewpoint. Furthermore, some objects or parts of objects in the scene, corresponding to a few pixels on the edge of the attention zone 8, 12—that is, pixels located in the immediate vicinity of the attention zone 8, 12—may be obscured by the object on which the observer's gaze is focused. Additionally, some objects or parts of objects in the scene, corresponding to a few pixels on the edge of the attention zone 8, 12, which were obscured before the change in viewpoint, may be generated on the pair of reconstructed images (IR1, IR2) according to the process.

[0263] According to the particular embodiment illustrated in FIGURES 4A to 4D and FIGURE 6, corresponding to the case where the attention zone 12 is located between the background of the imaged scene 1 and the foreground of the imaged scene 1 (in other words, the attention zone 12 does not correspond to or is not an object with the greatest depth and does not correspond to or is not an object with the shallowest depth), the modification step is implemented, according to a non-limiting embodiment:

[0264] - for the objects (individual 9, tree and landscape 10 and 12) of scene 1 whose depth is different from the depth of the attention zone 8, 12, the objects are shifted inversely proportionally to the depth of an object in scene 1 considered, in inverse difference from the depth of field of the attention zone 8, 12, and

[0265] - for objects whose depth is not very different from the depth of the attention zone 8, 12, it is possible to shift these objects proportionally to the difference in depth between the object considered and the depth of the attention zone 8, 12, as a first approximation,

[0266] - for more distant objects, whose depth is greater than the depth of the attention zone 8, 12, these objects are shifted in the same direction and sense as the direction and sense of the change in viewpoint, and

[0267] - for objects that are less distant, whose depth is less than the depth of the attention zone 8, 12, these objects are shifted in the same direction as the direction of the change of viewpoint and in the opposite direction to the direction of the change of viewpoint.

[0268] According to the particular embodiment illustrated in FIGURE 5, corresponding to the case where the attention zone 12 is located in the background of the imaged scene 1 (in other words, the attention zone 12 corresponds to or is an object in the background and / or has the greatest depth), the objects are shifted in a manner inversely proportional to the depth of an object in the scene 1 under consideration.

[0269] According to the specific embodiment illustrated in FIGURE 7, corresponding to the case where the attention zone 12 is located in the foreground of the imaged scene 1 (in other words, where the attention zone 12 corresponds to or is a foreground object and / or one with the shallowest depth), the objects are shifted inversely proportionally to the depth of an object in the scene 1 under consideration. In this description, "proportional" (or inversely proportional) means a proportional relationship up to a constant (in particular, the depth value of the attention zone 12), that is, an affine relationship.

[0270] In almost all 3D videos, each pair of source images (IS1, IS2) comprises a left image IS1 and a right image IS2. However, in some cases, the two images in the source image pairs (IS1, IS2) do not allow for binocular display (with a right image for the right eye and a left image for the left eye); in other words, they do not constitute a right and a left image. Indeed, as described above, the method according to the invention is applicable to any type of 3D video, provided that it comprises two 2D image streams of the same scene capable of producing a 3D rendering.Therefore, in some cases, for example, when two videos are acquired simultaneously by two camera sensors on the same connected device (such as a smartphone or tablet) filming the same scene, the images in the source image pairs (IS1, IS2) may not consist of a left and right image, but rather, for example, an upper and a lower image. While it is possible to display this type of video in monocular 3D vision, in order to display the 3D video in binocular vision, it is necessary to consider, during the image modification step for at least one source image pair (IS1, IS2), the distance between the two separate cameras from which the source image pairs (IS1, IS2) were acquired, and the position of the two separate cameras relative to each other and / or relative to scene 1.

[0271] In this way, each pair of restored images (IR1, IR2) consists of an IR2 image, called the right image IR2, as it would be perceived by the right eye of an observer of scene 1 and an IR1 image, called the left image IR1, as it would be perceived by the left eye of an observer of scene 1.

[0272] The entire description relating to the modification step is transposable to the modification of the source image pairs (IS1, IS2) relative to the distance between the two separate cameras from which the source image pairs (IS1, IS2) were acquired, the position of the two separate cameras relative to each other and / or relative to scene 1. Thus, according to the non-limiting embodiment, the modification step can be implemented further depending on the distance between the two separate cameras from which the source image pairs (IS1, IS2) were acquired, the position of the two separate cameras relative to each other and / or relative to scene 1.

[0273] In some cases, the source image pairs (IS1, IS2) may not provide optimal 3D rendering. This can occur, for example, when the two separate imaging devices from which the source image pairs (IS1, IS2) were acquired are too far apart. In this case, according to a non-limiting embodiment, it is possible to implement the modification step based, in addition, on the interpupillary distance.

[0274] A person skilled in the art is familiar with interpupillary distance (IPD). It is possible to use an average IPD corresponding, for example, to the average IPD of users, or the IPD for which the headsets and / or synchronous, polarizing, or red / blue filter glasses are designed. As a non-limiting example, IPD can be considered to be between 45 and 75 mm.

[0275] Preferably, to provide the observer with optimal 3D rendering, it may be preferable to combine the consideration of interpupillary distance with the relative position data of the two separate cameras from which the source image pairs (IS1, IS2) were acquired. Advantageously, in this case, the step of modifying the source image pairs (IS1, IS2) is also a function of the interpupillary distance and the distance between the two separate cameras from which the source image pairs (IS1, IS2) were acquired, as well as the position of the two separate cameras relative to each other and / or relative to scene 1. In particular, this can be implemented by further modifying the IS1 and IS2 images of at least one source image pair (IS1, IS2), preferably relative to each other.The step of modifying the IS1, IS2 images of at least one pair of source images (IS1, IS2), relative to each other, is carried out by moving and / or spreading and / or contracting pixels or groups of pixels of at least one pair of source images (IS1, IS2) of scene 1, and / or by adding and / or deleting pixels or groups of pixels, in the images of at least one pair of source images, so that the difference in viewpoint between the two images IR1, IR2 of at least one pair of restored images (IR1, IR2) corresponds to the pupillary distance.

[0276] In this case, the modification applied to the images of a source image pair (IS1, IS2) may differ for each image in the pair. For example, the spacing between the observation points can be readjusted when the two images of the source image pair (IS1, IS2) come from two cameras spaced at a distance different from the interpupillary distance and / or when the two cameras are not aligned along a predominantly horizontal axis.

[0277] When the two images of the source image pair (IS1, IS2) come from two cameras spaced (for example along the V-axis, like the observer's eyes) by a distance different from the pupillary distance, we can then apply to the right image (IS2) a translation of value TTV = (DPU - IC), where IC is the distance between the two cameras and DPU is the pupillary distance.

[0278] We can then deal with the other changes of viewpoint (which can be written algebraically as a common transformation if desired, so that only one transformation is applied to each image).

[0279] When the two images of the source image pair (IS1, IS2) come from two that are not aligned along a predominantly horizontal axis (along the V axis), it is possible to modify the image (for example the left image (IS1)) of the source image pair (IS1, IS2) to be rectified by a pure y rotation.

[0280] When the two images of the source image pair (IS1, IS2) come from two cameras spaced at a different distance than the interpupillary distance, and when the two cameras are not aligned along a predominantly horizontal axis, a translational modification (TTU and / or TTV) should be applied to the right image of the source image pair (IS1, IS2), followed by a rotational modification by angle y to the left image of the source image pair (IS1, IS2). The translational value is TTV = DPU*(1 - cos(y)), TTU = DPU*sin(y), if the interpupillary distance DPU of the observer's eyes is initially oriented along the V-axis, rotated by an angle y, perpendicular to the observation axis.

[0281] In this application, it is assumed that the person skilled in the art is familiar with the concept of a depth chart and the methods for obtaining a depth chart.

[0282] By way of non-limiting example, the process may include the step of generating the depth map of the imaged scene from one or more images of a pair of source images (IS1, IS2) and / or from sharpness information from one or more images of a pair of source images (IS1, IS2) and / or from a local variance.

[0283] A skilled professional will know how to choose the most appropriate sharpness information for each situation. The sharpness information contained in a pixel or block of pixels can include the entropy, energy, and / or variance of an image parameter between neighboring pixels.

[0284] Preferably, any energy or entropy estimator commonly used in focus detection can be used to determine the variance. Those skilled in the art are familiar with a range of techniques for calculating local variance. As a non-limiting example, local variance can be the variance of pixel intensity, for instance, the average intensity of the R, G, and B channels, calculated over a given set of pixel matrices, adjusted according to the image resolution and / or the processing power / capacity of the processing unit. Then, for each pixel, group of pixels, or object in the depth map image to be generated, the scene image with the highest variance is selected.

[0285] The depth map can correspond to an image comprising a single depth channel. Depth is understood to be the physical distance between a point or physical object in scene 1, corresponding to a pixel or group of pixels in a given image, and the sensor from which the image was acquired. As an illustration, a simplified schematic depth map associated with the scene 1 image shown in Figures 4 to 7 is presented in Figure 8. In this depth map, the white area 4 corresponds to an object located, for example, five meters from the camera's optical sensor, and the gray area 5 at the background of the image corresponds to objects located beyond, for example, fifty meters from the optical sensor.A distance of five meters corresponds, for example, to the minimum focusing distance of the camera lens and a distance of fifty meters corresponds, for example, to the maximum focusing distance of the camera lens.

[0286] A disparity map of the imaged scene 1 can also be determined. The depth map can be calculated and / or determined from the disparity map. Thus, any feature relating to the depth map can be transposed to the disparity map, and vice versa. The disparity map can be defined as a map or image containing depth information. In other words, the disparity map can be an image or map comprising at least one depth channel of the imaged scene. Preferably, the depth information constitutes a depth field of the imaged scene.

[0287] Therefore, a group or block of pixels corresponding to one of the objects in the imaged scene 1 will have a different depth information than one or more other groups or blocks of pixels corresponding to the other object(s) in the imaged scene 1. Consequently, each image acquired by the same mobile imaging system can be enriched by a depth field; that is, each pixel or group or block of pixels (in the matrix or pixel table representing the scene image) corresponding to an object in the scene will have a distinct depth information that corresponds to the distance between the object and the optical sensor; the depth of the different objects in the imaged scene being, in most cases, different for at least some of the objects in the scene.

[0288] According to the invention, a VR / VA head-mounted display and a 3D display device are also proposed. The head-mounted display includes means for displaying images in monocular or binocular vision arranged to provide 3D rendering from two stereoscopic 2D images. The 3D display device includes means for displaying images arranged to provide 3D rendering from two stereoscopic 2D images.

[0289] During the optional display step, the pairs of reconstructed images (IR1, IR2) can be displayed binocularly in a binocular head-mounted display. One IR1 of the two images in a reconstructed image pair (IR1, IR2) is displayed monocularly, by a display means of the head-mounted display, for the left eye of the observer wearing the head-mounted display, and the other IR2 of the two images in a reconstructed image pair (IR1, IR2) is displayed, by a display means of the head-mounted display, monocularly for the right eye of the observer wearing the head-mounted display.

[0290] During the optional display step, the pairs of restored images (IR1, IR2) can be displayed in monocular vision in a head-mounted display or in a 3D display device. The two images of a pair of rendered images (IR1, IR2) are displayed on a display means of the 3D display device or the head-mounted display either simultaneously or alternately so that one IR1 of the two images of a pair of rendered images (IR1, IR2) or part of one IR1 of the two images of a pair of rendered images (IR1, IR2) is seen only by one eye of the observer and the other image IR2 of the two images of a pair of rendered images (IR1, IR2) or part of the other image IR2 of the two images of a pair of rendered images (IR1, IR2) is seen only by the other eye of the observer (where appropriate by means of additional viewing means (for example glasses)).As detailed previously, the person skilled in the art will be able to understand, among other things, the concepts of display by spatial and temporal multiplexing as well as the display of polarized images.

[0291] A person skilled in the art will be able to understand the concepts of stereoscopy and monocular and binocular vision, as well as their applications. In particular, a person skilled in the art will observe that the two images in a pair of reconstructed images correspond to two different viewpoints of scene 1, so that the observer wearing the headset or viewing the 3D display device is able to observe a perspective or three-dimensional view of the scene.

[0292] Advantageously, the head-mounted display and the 3D display device include means arranged and / or programmed and / or configured to implement the method according to the invention. In practice, the means arranged and / or programmed and / or configured to implement the method according to the invention constitute a processing unit. In this case, the stream of source image pairs (IS1, IS2)n and / or the stream of rendered image pairs (IR1, IR2)n of scene 1 can be loaded into a storage means of the head-mounted display or the 3D display device. The method will therefore be implemented autonomously by the head-mounted display and the 3D display device.

[0293] However, the head-mounted display and the 3D display device may include communication means arranged and / or programmed and / or configured for and / or capable of communicating with external or remote means. In this case, the stream of source image pairs (IS1, IS2)n and / or the stream of rendered image pairs (IR1, IR2)n of scene 1 may be transferred, preferably in real time, from the external means to the head-mounted display and from the 3D display device.

[0294] Alternatively, external means can be arranged and / or programmed and / or configured to implement the method according to the invention. In this case, the head-mounted display and the 3D display device may not implement the method according to the invention, or may not implement all the steps of the method according to the invention. In an extreme case, the head-mounted display and the 3D display device may not implement any of the steps of the method according to the invention and may only display the stream of source image pairs (IS1, IS2)n and / or the stream of rendered image pairs (IR1, IR2)n of scene 1.

[0295] According to an advantageous embodiment, the head-mounted display and the 3D display device include at least one imaging system arranged to image the observer and / or the eyes of an observer.

[0296] The processing unit can be arranged and / or programmed and / or configured to determine and / or calculate the change of viewpoint based on data from at least one imaging system arranged to image the observer and / or the eyes of an observer.

[0297] The head-mounted display, and respectively the 3D display device, include means arranged to detect movement and / or the relative position of the head-mounted display, and respectively of the 3D display device. By way of non-limiting examples, the means arranged to detect movement and / or position may be a gyroscope or an accelerometer for the head-mounted display, and respectively one or more imaging systems for the 3D display device.

[0298] Advantageously, the resulting change of viewpoint is a function of, or is calculated and / or determined from, data originating, for example, from means arranged to detect movement and / or the relative position of the head-mounted display (referred to as head-mounted display data) and from the imaging system(s) for the 3D display device. Of course, the invention is not limited to the examples just described, and numerous modifications can be made to these examples without departing from the scope of the invention.

[0299] Thus, in combinable variants of the previously described embodiments: the change of viewpoint includes a displacement (du, dv and / or dz) in a reference frame of one or both of the source image pairs (IS1, IS2) and / or in a reference frame of the user of the head-mounted display or the 3D imaging device, and / or the method may include the acquisition of all or part of the source image pairs (IS1, IS2) of scene 1 from the two separate imaging systems in two positions or two separate viewpoints, and / or the method may include the acquisition of images of the observer, and / or the method may include the acquisition of images of the observer's eyes, and / or the method may include the acquisition of images of at least a part of a head-mounted display environment or at least a part of a 3D display device environment, and / or

[0300] - in this description, the two separate imaging systems may be one, the imaging system(s) of the VR / VA headset or 3D display device, and / or

[0301] - in the present description, the observer's images and / or the observer's eye images and / or the images of at least a part of the environment of the head-mounted display or at least a part of the environment of the 3D display device may be acquired by one or more cameras of the head-mounted display or the 3D display device, and / or the two separate imaging systems, from which the images of the source image pairs (IS1, IS2) were acquired, are stationary and / or at a constant distance and position relative to each other during image acquisition, and / or the invention provides a computer program comprising instructions which, when the program is executed by a computer, cause the computer to implement the method according to any one of the described embodiments, and / or the invention provides a readable medium, in particular,by computer or by any device comprising a processing unit including instructions which, when executed by said computer or device, cause it to implement the process according to any one of the embodiments described.

[0302] Furthermore, the different features, forms, variants and embodiments of the invention can be associated with each other in various combinations insofar as they are not incompatible or mutually exclusive.

Claims

DEMANDS 1. A method for dynamically observing stereoscopic images from a 3D video of a scene, said 3D video comprising a set of stereoscopic 2D image pairs, referred to as source image pair(s), capable of rendering the video in 3D by display in a head-mounted display or on a 3D display device, said method comprising the steps of: - to obtain a change of perspective in relation to the imaged scene, - modify each of the images from at least one pair of source images to reproduce at least one pair of stereoscopic 2D images of the scene, referred to as at least one pair of reproduced images, referred to as at least one pair of reproduced images: • being designed to offer 3D rendering via display in a head-mounted display or on a 3D display device, • providing a change of viewpoint relative to the 3D video of the scene; the images of at least one pair of source images being modified by displacement and / or spreading and / or contraction of pixels or groups of pixels of at least one pair of source images of the scene and / or by addition and / or deletion of pixels or groups of pixels, in said images of at least one pair of source images, as a function of: • a depth map of the imaged scene, • the change in perspective achieved.

2. A method according to the preceding claim, wherein the modification step consists of: - to modify, identically, each of the two images of a given pair of source images in order to produce at least one pair of images comprising two modified images, or - to modify, distinctly, each of the two images of a pair of source images considered in order to reproduce at least one pair of images comprising two modified images.

3. Method according to claim 1 or 2, wherein the image modification step of at least one pair of source images is carried out, in addition, with respect to an area of ​​the images of at least one pair of source images, called the area of ​​attention.

4. A method according to the preceding claim, wherein the area of ​​attention: remains unchanged when modifying the images of at least one pair of source images, or - is modified when modifying the images of at least one pair of source images.

5. A method according to any one of the preceding claims, wherein the modification of the images of at least one pair of source images is proportional to the change in viewpoint.

6. A method according to any one of the preceding claims, wherein: - the modification of a part, called the first part, of the pixels or groups of images from at least one pair of source images is inversely proportional to the depth associated with the pixels or groups of pixels in the first part, and / or - the modification of a part, called the second part, distinct from the first part, of the pixels or groups of pixels of the images of at least one pair of source images is proportional to the depth associated with the pixels or groups of pixels of the second part.

7. A method according to the preceding claim, taken in combination with claim 3 or 4, comprising, for each pixel or group(s) of pixels of the images from at least one pair of source images, a calculation of a deviation: - depth between the area of ​​attention and each pixel or group(s) of pixels located outside the area of ​​attention, and / or - between the inverse of the depth of the attention area and the inverse of the depth of each pixel or group(s) of pixels located outside the attention area; the modification of the images of at least one pair of source images is carried out according to the calculated difference.

8. A method according to any one of the preceding claims, comprising a calculation of at least one angle associated with or corresponding to the change of viewpoint; the modification of the images of at least one pair of source images is carried out as a function of at least one calculated angle.

9. A method according to any one of the preceding claims, comprising a calculation of a variation in distance, associated with or corresponding to the change of viewpoint, relative to the imaged scene; the modification of the images of at least one pair of source images is carried out according to the calculated variation in distance.

10. A method according to any one of the preceding claims, wherein the modification step is applied to a stream or stack of stored source image pairs.

11. A method according to any one of the preceding claims, wherein the modification of the images of at least one pair of source images is a function of an interpupillary distance.

12. A method according to any one of the preceding claims, wherein the modification of the images of at least one pair of source images is a function of: - a distance between two separate imaging devices, from which the two images of each pair of source images were respectively acquired, and / or - the position of said two distinct imaging devices relative to each other and / or relative to the scene, and / or - a distance difference between said two separate imaging devices and the scene; said two separate imaging devices are arranged in two separate spatial positions and one of the images of each pair of source images comes from one of the two imaging devices and the other of the two images of each pair of source images comes from the other of the two imaging devices.

13. A method according to the preceding claim taken in combination with claim 11, wherein the images of at least one pair of source images are further modified relative to each other by displacement and / or spreading and / or contraction of pixels or group(s) of pixels of at least one pair of source images of the scene so that the difference in viewpoint between the two images of at least one pair of restored images corresponds to the pupillary distance.

14. A method according to any one of the preceding claims, comprising the step of displaying, in a virtual reality / augmented reality (VR / VA) headset, referred to as a visiohelmet, or in a 3D display device, at least one pair of rendered images or the images from at least one pair of source images.

15. A method according to the preceding claim, wherein: - the change in viewpoint obtained is a function of a change in position and / or orientation and / or tilt, in space, of the head-mounted display or the 3D display device or of a change in position and / or orientation and / or tilt, in space, of an observer relative to the 3D display device, and / or - The attention zone corresponds to an area of ​​the displayed scene on which the observer's gaze is focused.

16. A method according to the preceding claim, wherein: - the change of viewpoint is obtained from position data of the head-mounted display and / or movement data of the head-mounted display, referred to as head-mounted display data, and / or - the change of viewpoint is obtained from images of the observer and / or the observer's environment and / or the head-mounted display or 3D display device, and / or - the area of ​​attention is obtained from images of the observer's eyes.

17. Method according to claim 16, comprising filtering the data from the head-mounted display or 3D display device to limit the amplitude of the change in viewpoint obtained, as a function of time, to a value less than a limit value.

18. Method according to claim 16 or 17, comprising non-linear clipping of the data from the head-mounted display or 3D display device to limit the amplitude of the change of viewpoint obtained to a value less than a limit value.

19. A method according to any one of claims 16 to 18, comprising a nonlinear dynamic compression of a gain applied to the data of the head-mounted display or 3D display device to attenuate a variation in amplitude or amplitude of the change of viewpoint obtained, as a function of time, such that said variation in amplitude does not exceed a threshold value of amplitude or variation in amplitude.

20. A method according to any one of claims 16 to 18, comprising a nonlinear dynamic amplification of a gain applied to the data of the head-mounted display or 3D display device to increase a variation in the amplitude of the change of viewpoint obtained, as a function of time.

21. A method according to any one of the preceding claims, comprising an acquisition step: - of scene images, said acquired scene images comprising all pairs of source images, and / or - of the observer's images, and / or - images of at least part of a headset environment or at least part of a 3D display device environment.

22. Data processing device comprising means arranged and / or programmed and / or configured to implement the method according to any one of claims 1 to 21.

23. Computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method according to any one of claims 1 to 21.

24. Computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the method according to any one of claims 1 to 21.

25. Virtual reality / augmented reality (VR / AV) headset including one or more means for displaying 3D videos, and: - means arranged and / or programmed and / or configured to implement the process according to any one of claims 1 to 21, and / or - means of communication arranged and / or programmed and / or configured to and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the process according to any one of claims 1 to 21.

26. VR / VA head-mounted display according to the preceding claim, comprising at least one imaging device arranged to image the eyes of an observer wearing said VR / VA head-mounted display.

27. VR / VA head-mounted display according to claim 25 or 26, comprising: - at least one imaging device arranged to image at least part of an environment of said VR / VA headset and / or an environment of the observer, and / or - means arranged to detect a movement and / or a position and / or an orientation and / or an inclination of said VR / VA head-mounted display.

28. A 3D display device, arranged to display 3D videos, comprising: - means arranged and / or programmed and / or configured to implement the process according to any one of claims 1 to 21, and / or - means of communication arranged and / or programmed and / or configured to and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the process according to any one of claims 1 to 21.

29. 3D display device according to the preceding claim, comprising at least one imaging device arranged to image the eyes of an observer.

30. 3D display device according to claim 28 or 29, comprising: - at least one imaging device arranged to image at least a part of an environment of said 3D display device and / or of the observer and / or of an environment of the observer, and / or - means arranged to detect a movement and / or a position and / or an orientation and / or an inclination of the observer relative to said 3D display device.

Citation Information

Patent Citations

  • Device, method and computer program for 3D rendering

    EP3001680A1

  • Image display within a three-dimensional environment

    WO2022192306A1