Method for processing images of a scene in order to render a change in the viewpoint of the scene

The method modifies stereoscopic 2D images using depth maps and viewpoint changes to achieve real-time 3D rendering in head-mounted and 3D display devices, addressing the inefficiencies of traditional 3D image generation and enhancing immersion.

WO2025141251A1PCT designated stage expired Publication Date: 2025-07-03FOGALE OPTIQUE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/FR2023/052128
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing methods for generating 3D images are costly in terms of time and computing resources, and are incompatible with real-time dynamic visualization of 2D images, limiting the immersive experience in head-mounted and 3D display devices.

Method used

A method for modifying stereoscopic 2D images by moving or adding/deleting pixels based on a depth map and change of viewpoint, using a processing unit to create a 3D rendering effect in head-mounted or 3D display devices.

Benefits of technology

Enables real-time 3D rendering with reduced computational resources, enhancing the immersive experience without the need for traditional 3D images, thus overcoming the limitations of existing 3D image generation methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FR2023052128_03072025_PF_FP_ABST
    Figure FR2023052128_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for modifying images of a scene in order to render a change in the viewpoint of the scene, the method comprising the steps of: obtaining a change in the viewpoint with respect to the imaged scene, modifying a pair of stereoscopic 2D images of the scene, called the pair of source images, by moving or spreading pixels or one or more groups of pixels of at least one of the images of the pair of source images and / or by adding and / or deleting pixels or groups of pixels in at least one of the images of the pair of source images, as a function: of a depth map of the imaged scene, of the change of viewpoint obtained.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DESCRIPTION

[0002] Method of processing images of a scene to restore a change in point of view of the scene

[0003] Technical field

[0004] The present invention relates to stereoscopic 2D images intended to be displayed in a head-mounted display or in a 3D display device in order to render a 3D rendering when they are displayed.

[0005] The invention relates to the dynamic observation of stereoscopic 2D images displayed in a head-mounted display or in a 3D display device. The invention also relates to the immersive observation of stereoscopic 2D images displayed in a head-mounted display or in a 3D display device.

[0006] The invention does not relate to 3D images.

[0007] State of the prior art

[0008] 3D images are known in the prior art which make it possible to introduce into displayed 2D images a perspective effect such as it would be perceived if the imaged scene were observed directly by the observer. According to the invention, the 3D images are not stereoscopic images but three-dimensional images offering a relief representation of the objects. The invention relates to stereoscopic 2D images making it possible to provide a 3D rendering, when they are displayed in a suitable display medium, by exploiting the binocular vision of the observer which gives him a perception of depth. The synthesis of 3D images can be carried out directly from a model of a real scene. The production of 3D images is costly in terms of time and the computer resources required for processing and storing these large images.3D images can be understood to mean relief images obtained, for example, by 3D surface modeling of space or from two images with different disparities, typically two images obtained from two distinct viewpoints. The load is increased tenfold when it comes to a stream of 3D images, in particular a stream of 3D images resulting from 3D surface modeling of space. In addition, the time required for processing is incompatible with dynamic visualization of 2D image(s) acquired in real time. Stereoscopic vision is known in the state of the art. Binocular human vision, or more broadly animal vision, from two images that are compared by the brain, constitutes the most widespread example.The biological processing of the images obtained is extremely efficient, since it provides in real time a notion of depth of the observed scene, allowing for example to move knowing at what relative distance the observed objects are.

[0009] The observation of stereoscopic images is known in the state of the art, which is inspired by or mimics the principle of stereoscopic vision. The principle of stereoscopic observation consists of increasing the immersive nature of image observation by adding a depth effect by projecting a different image onto each eye of an observer. In practice, a stereoscopic image comprises two distinct images of the same scene acquired in two distinct positions in space, in particular by two cameras spaced by the pupillary distance. Human stereoscopic vision will make it possible, from a stereoscopic image, to reconstruct a scene including depth information. The time required to generate 3D images is incompatible with dynamic visualization of 2D image(s) acquired in real time. The production of 3D images is costly in terms of time and computing resources required for processing and storing these images.

[0010] To enhance the immersive experience, 3D head-mounted displays are also available. These headsets allow you to view images or films in 3D. However, these head-mounted displays require the display of 3D images and therefore have the disadvantages associated with 3D images as described above.

[0011] One aim of the invention is to propose an image processing method, a video headset and a 3D display device:

[0012] - enabling the resolution of problems with state-of-the-art processes, and / or

[0013] - rendering a change of viewpoint of a scene to a user of a head-mounted display or a 3D display device displaying 2D images, and / or

[0014] - rendering a parallax effect to a user of a head-mounted display or a 3D display device displaying 2D images, and / or - making it possible to increase the immersive experience for a user of a head-mounted display or a 3D display device displaying 2D images, and / or

[0015] - providing an alternative to head-mounted displays and 3D display devices that display 3D images.

[0016] Presentation of the invention

[0017] To this end, a method is proposed for modifying images of a scene to restore a change in point of view of the scene, said method comprising the steps of:

[0018] - obtain a change of point of view in relation to the imaged scene,

[0019] - modifying a pair of stereoscopic 2D images of the scene, called a pair of source images, by moving or spreading pixels or group(s) of pixels, for example at least part, only part or all of the pixels or group of pixels, of at least one of the images of the pair of source images, preferably of each of the two images of the pair of source images, and / or by adding and / or deleting pixels or groups of pixels in at least one of the images of the pair of source images, preferably of each of the two images of the pair of source images, as a function of:

[0020] • a depth map of the imaged scene,

[0021] • the change of point of view obtained.

[0022] Preferably, the step of modifying the at least one image of the scene is carried out by means of a processing unit.

[0023] Preferably, the change of point of view corresponds to the passage from a point of view, called the previous point of view, to a different point of view, called the new point of view. Preferably, the image(s) obtained by implementing the method or the image(s) modified according to the method correspond(s) to the image(s) observed from the new point of view.

[0024] It can be understood as a point of view, a position or coordinates of space.

[0025] Preferably, each data processing step of the method according to the invention, comprising for example any operation carried out from or on data, for example a calculation, a determination or a comparison, is implemented by a processing unit or any device or system capable of processing data.

[0026] Preferably, both images of the source image pair are intended to provide a 3D rendering when displayed on a head-mounted display or on a 3D display device.

[0027] Preferably, the two images of the source image pair correspond to two distinct viewpoints of the scene.

[0028] Preferably, at least one image of the pair of source images, more preferably both images of the pair of source images, are modified during the modification step. In some cases, only one of the images of the pair of source images is modified during the modification step.

[0029] In the remainder of the description, particularly for the sake of brevity, the modification of both images of the pair of source images will be described or referred to. The reference to the modification of both images of the pair of source images will also include the case in which only one of the images of the pair of source images is modified. In other words, in the remainder of the description, the modification of both images of the pair of source images will be described or referred to without excluding from the scope of the invention the case of the modification of only one of the images of the pair of source images.

[0030] Preferably, the modification of the pair of source images is carried out or implemented or performed on at least a portion of the images of the pair of source images, preferably on all of the images of the pair of source images. In other words, preferably, the displacement of pixels or group(s) of pixels of the images of the pair of source images and / or the addition and / or deletion of pixels or groups of pixels in the images of the pair of source images is carried out on at least a portion of the images of the pair of source images, preferably on all of the images of the pair of source images.

[0031] Preferably, a pixel corresponds or is associated with a point in the imaged scene and a group of pixels corresponds or is associated with an object in the imaged scene.

[0032] Preferably, the displacement of a portion of the pixels or group(s) of pixels within each image of the pair of source images makes it possible to or aims to reproduce a parallax effect induced by the change of point of view. Preferably, displacement of a portion of the pixels or group(s) of pixels is understood to mean: a modification of the position and / or a translation and / or a spreading of a portion of the pixels or group(s) of pixels within each image of the pair of source images, preferably in a plane or in a reference frame specific to the modified image or in a reference frame common to the two images of the pair of source images.

[0033] Preferably, each pixel or each group of pixels is moved and / or translated and / or spread according to a reference depth, more preferably relative to the reference depth. Preferably, the reference depth is associated with one or more objects in the scene or a group of objects.

[0034] Preferably, all pixels of the source image pair are modified during the modification step.

[0035] Depth information can be understood as data representing or corresponding to or providing information on the depth associated with a pixel, corresponding to a point in the imaged scene, or to a group of pixels, corresponding to an object in the imaged scene, of an image of the scene.

[0036] Preferably, the images of the source image pair were acquired by one or more imaging systems.

[0037] Preferably, at the end of the step of modifying the pair of source images, a modified pair of images is obtained, said modified pair of images comprising:

[0038] - two modified images of the scene, that is to say that the two images of the source image pair have been modified, or

[0039] - an image of the modified scene, that is to say one of the two images of the pair of source images has been modified, and one of the images of the pair of source images (unmodified).

[0040] Preferably, the modified image pair is a new stereoscopic 2D image pair of the scene, called the modified image pair.

[0041] The modified image pair may possibly constitute a source image pair to be modified in a new implementation or new iteration of the process.

[0042] Preferably, as for the pair of source images, the pair of modified images is intended to provide a 3D rendering when they are displayed on a head-mounted display or 3D display device. Preferably, the modification step is applied to a stream of pairs of source images. Preferably, the obtaining step and the modification step are applied to a stream of pairs of source images.

[0043] Depth can be understood as a distance between a point or object in the imaged scene and the optical sensor of the imaging system from which the image of the scene was acquired.

[0044] The depth field may comprise depth information of all or part of the points or objects in the scene, preferably a depth field of the imaged scene. The depth field may be contained in a channel of all or part of the pixels of one or more images of the source image pair.

[0045] Preferably, the method comprises: prior to the modification step, the step of displaying, in a video headset or on a 3D display device, the pair of source images, and / or the step of displaying, in a video headset or on a 3D display device, the pair of modified images.

[0046] The term "imaged scene" may be understood to mean any image of the scene, for example the pair of displayed source images, i.e. unmodified or before modification, and / or the pair of modified images and / or at least two images of the scene from which a depth map of the scene is generated and / or from which depth information of the scene is extracted.

[0047] Preferably, according to the invention, the images of the scene are 2D images which can include depth information.

[0048] Preferably, the step of displaying the modified pair of images makes it possible to restore the change of point of view, in other words the image of the scene as the observer would view it after having changed point of view.

[0049] The display step, in the head-mounted display or in the 3D display device, may comprise the display of a stream of pairs of images, source or modified, of the scene. Preferably, the method may be implemented for at least some of the images, more preferably for each image, of the stream of source images. It may be understood by plane or reference frame of the displayed images: a plane or reference frame common to the displayed images or a separate plane or reference frame, preferably a plane or reference frame specific to each of the displayed images.

[0050] Preferably, the modification, or the modification step, of the pair of source images is also a function of an area of ​​the scene, preferably an area of ​​the imaged scene, called the attention area.

[0051] Preferably, the change of viewpoint corresponds to, is or is a function of a change in position of the head-mounted display in space and / or of a displacement, preferably relative to each pixel or group(s) of pixels of one or both images of the pair of images displayed, of the attention zone.

[0052] Preferably, the attention zone corresponds, in or for each of the images of the pair of images, to the same object(s) of the imaged / displayed scene. Preferably, the attention zone corresponds, in or for each of the images of the pair of images, to the pixel(s) or group of pixels corresponding to the same object(s) of the imaged / displayed scene.

[0053] The modification of the at least one image of the scene can be operated or implemented or carried out for at least a part of one or each of the images of the pair of source images, preferably on the whole of one or each of the images of the pair of source images, with the exception of the attention zone. In other words, preferably, the displacement of a part of the pixels or group(s) of pixels of one or each of the images of the pair of source images and / or the addition and / or deletion of pixels or groups of pixels in one or each of the images of the pair of source images is carried out outside the attention zone.

[0054] It can be understood by attention zone: one or more pixels of one or each of the images of the pair of source images.

[0055] At least a portion of the pixels of one or each of the images of the source image pair, more preferably the pixel(s) of the attention zone, preferably only the pixel(s) of the attention zone, may remain unchanged and / or immobile during the modification step, in particular when the change in viewpoint of the observer comprises or corresponds to or is a variation in the distance between an observer and the imaged scene, preferably between an observer and one or each of the images of the source image pair. Preferably, the two images of the source image pair are modified, during the modification step, with respect to or relative to the attention zone and / or to the change in viewpoint obtained or detected.

[0056] The attention zone can remain unchanged, when modifying the pair of source images, by anchoring the attention zone relative to the plane(s) or reference frame(s) of one or each of the images in the pair of source images.

[0057] Preferably, the method comprises a step of generating the depth map of the imaged scene from depth information of the imaged scene and / or from sharpness information of at least one of the images of the pair of source images.

[0058] Preferably, the step of generating the depth map can be defined as enriching one of each image of the pair of source images with depth information, preferably with a depth field, associated with the imaged scene. Preferably, the enriched pair of source images comprises a depth field channel, preferably a single depth field channel.

[0059] By way of non-limiting example, the information contained in a pixel or block of pixels may include an entropy value, energy, a color channel, typically R, G, B (Red, Green, Blue), intensity detected by the photosites, a gain and / or depth information or a depth field.

[0060] Preferably, the information contained in a pixel or block of pixels is a value or information, called sharpness, relating to or representative of sharpness.

[0061] Generating the depth map of the imaged scene may be understood to mean extracting depth information from one or, preferably, two or more images of the scene. Preferably, generating the depth map of the imaged scene may comprise, or consist of, a step of: extracting depth information from the pair of source images, or from at least two images of the scene, and / or from sharpness information of at least one of the images of the pair of source images.

[0062] According to the present invention, the depth information or the depth field can be obtained by comparing images, for example by photogrammetry, or by a distance measuring device, for example a device for detecting and estimating distance by light, called "LIDAR" for "light detection and ranging".

[0063] Preferably, according to a first aspect of the invention, the depth information and / or the depth field is determined or calculated and / or the depth map is generated, produced, calculated or determined from at least two images of the scene originating from:

[0064] - from the same stationary imaging system and each being acquired with a different focusing distance, and / or

[0065] - of the same mobile imaging system having acquired, at least in part, the scene in two distinct positions, and / or

[0066] - at least two separate imaging systems, arranged in two separate positions so as to each acquire, at least in part, the scene.

[0067] The at least two images of the scene originating from at least two distinct imaging systems arranged in two distinct positions and / or the at least two images of the scene originating from the same mobile imaging system having acquired, at least in part, the scene in two distinct positions may be the two images of the source image pairs.

[0068] Preferably, the at least two images of the scene, from which the depth map is generated, comprise the two images of the pair of source images. Preferably, the at least two images of the scene, from which the depth map is generated, are the two images of the pair of source images.

[0069] In the present description, unless otherwise indicated, the term "at least two images of the scene" used alone corresponds to the at least two images of the scene from which the depth map of the scene is generated and / or from which the depth information of the scene is extracted.

[0070] Preferably, according to the first aspect of the invention, the depth information and / or the depth field is determined or calculated and / or the depth map is generated, produced, calculated or determined:

[0071] - by selection and / or identification of pixels or groups of pixels from the at least two images of the scene or by identification and / or selection of an image or image(s) from among the at least two images, - by combination, assembly or association of pixels or groups of pixels from the at least two images of the scene.

[0072] The determination of the depth information or the depth field and / or the depth map may comprise, prior to the selection and / or identification step, a step of comparing the at least two images of the scene or pixels or groups of pixels of the at least two images of the scene.

[0073] In this application, generating the depth map may be understood to mean producing, calculating or determining the depth map.

[0074] In the present application, it may be understood by determining the depth information or the depth field: calculating or deducing the depth information.

[0075] Preferably, according to a first alternative of the first aspect of the invention, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by:

[0076] - for each point of the depth map to be generated, identification of the image of the scene, among the at least two images of the scene, for which the sharpness is maximum, and

[0077] - for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene for which the sharpness is maximum.

[0078] Preferably, according to a second alternative of the first aspect of the invention, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by:

[0079] - for each point of the depth map to be generated, interpolation or extrapolation, from the respective focusing distances of the at least two images, of an optimal focusing distance for which each point of the depth map to be generated considered would present the maximum local sharpness, and,

[0080] - for each point of the depth map to be generated, association of said optimal focusing distance. It can be understood by "a point of the depth map": a pixel or a group of pixels of the depth map.

[0081] Preferably, each pixel considered or group of pixels considered correspond(s) to an area of ​​the scene or to an object of the scene.

[0082] Preferably, the pixel considered or the group of pixels considered corresponding) or common to the at least two images of the scene correspond(s) to the same zone of the scene or to the same object of the scene.

[0083] The local maximum or maxima may correspond to an area of ​​the scene or to an object in the scene.

[0084] Preferably the method comprises, according to a third alternative of the first aspect of the invention, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by:

[0085] - for each point of the depth map to be generated, identification of the image of the scene, among the at least two images of the scene, for which a local variance is the highest,

[0086] - for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene for which local variance is the highest.

[0087] Preferably the method comprises, according to a fourth alternative of the first aspect of the invention, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by:

[0088] - for each point of the depth map to be generated, interpolation or extrapolation, from the respective focusing distances of the at least two images, of an optimal focusing distance for which the depth map would present the highest local variance, and

[0089] - for each point of the depth map to be generated, or association of said optimal focusing distance.

[0090] Preferably, according to a second aspect of the invention, the depth information is determined and / or the depth map is generated, from an image of the scene, preferably from at least one of the images of the pair of source images, preferably from both images of the pair of source images, by means of a neural network, preferably an acyclic neural network, more preferably a convolutional neural network.

[0091] Preferably, the determination of the depth information is carried out and / or the generation of the depth map, more preferably according to the second aspect of the invention, is implemented from a single image of the scene, preferably from one of the two images of the pair of source images. Thus, preferably, a depth map is generated for each image of the scene.

[0092] Preferably, the determination of the depth information is carried out and / or the generation of the depth map, more preferably according to the second aspect of the invention, comprises a step of minimizing an objective function.

[0093] Preferably, the determination of the depth information is carried out and / or the generation of the depth map, more preferably according to the second aspect of the invention, is implemented from the R, G, B values ​​of one or more images of the scene.

[0094] The characteristics described in this description apply, unless otherwise indicated, to each of the aspects of the invention and to each of the alternatives according to the invention.

[0095] Preferably, the depth information and / or the sharpness information contained in, or extracted from, an image of the scene considered, preferably an image considered among the pair of source images, or, respectively, the depth information contained in the depth map is refined or improved by processing the image of the scene considered, or an image associated with the image of the scene considered, or, respectively, the depth map. At the end of the processing phase, an image of the scene whose depth and / or sharpness information or, respectively, a depth map whose depth information is refined or improved is obtained.

[0096] The processing phase comprises an iterative modification of the image of the scene considered or, respectively, of the depth map so as to minimize a function E comprising: a term D, called a difference term, determined by comparing the image of the scene considered during iterative processing, or respectively the depth map during iterative processing, convolved by a point spread function (PSF), dependent on the depth map, with the image of the scene considered, or respectively with the depth map, before convolution, the PSF describing the response of an imaging system from which the image of the scene considered, or respectively the depth map, is obtained, and a term A, called anomaly term, representative of defects or anomalies within the image of the scene considered during iterative processing, or respectively the depth map during iterative processing.

[0097] Preferably, the processing phase, or the refining or improving the information and / or sharpness, is implemented for one or each of the images of the pair of source images and / or for the at least two images from which the depth map of the scene is generated and / or from which the depth information of the scene is extracted.

[0098] At least a portion of the pixels of one or more images of the pair of source images of the scene may include, but not necessarily, depth information and / or sharpness information.

[0099] It can be understood by image associated with an image of the scene considered: a reconstructed image. At least part of the pixels of an image associated with the imaged scene can include depth and / or sharpness information.

[0100] Preferably, a reconstructed image is obtained from one or more images of the source image pair.

[0101] Preferably, the reconstructed image can be enriched by the depth information, contained in at least one part of the pixels of said reconstructed image, obtained by a dual pixel sensor, or by comparison of the at least two images of the scene, preferably acquired simultaneously, more preferably comprising the pair of source images, obtained in two positions or two distinct points of view. The reconstructed image can be obtained by photogrammetry.

[0102] The step of generating the depth map can be defined as determining and / or associating a distance between each of the points or each group(s) of point(s) of the observed or imaged scene and an optical sensor, preferably an optical sensor of an imaging system, for example a camera, having imaged the scene.

[0103] Preferably:

[0104] - the change of point of view includes a displacement (du, dv) in a reference frame of the images or to the reference frame of one of the images of the pair of source images, and / or

[0105] - the change of point of view includes a lateral and / or vertical movement (of the, dv) of the head-mounted display or a lateral and / or vertical movement (of the, dv) relative to the 3D display device, and / or

[0106] - the change of point of view includes a displacement (dz) perpendicular to the reference frame of the images or to the reference frame of one of the images in the pair of source images, and / or

[0107] - the change of point of view includes a displacement (dz), perpendicular to the plane formed by the axes (du, dv), of the head-mounted display or a displacement (dz), perpendicular to the plane formed by the axes (du, dv), that is to say along the axis connecting the observer to the 3D display device, relative to one or two images of the scene displayed.

[0108] Preferably, the change in the source image pair is proportional to the change in viewpoint.

[0109] Preferably, the modification of the source image pair is inversely proportional to the depth associated with the pixel or group of pixels considered.

[0110] Preferably, the method comprises, for each pixel or group(s) of pixels of one or each image of the pair of source images, a calculation, preferably from the depth map, of a difference between a depth, or an average depth, associated with the attention zone and a depth associated with each pixel or group(s) of pixels located outside the attention zone. Preferably, the modification of the pair of source images is carried out as a function of the calculated difference.

[0111] Preferably, the modification of the pair of source images is carried out relatively and / or proportionally to the calculated difference. Preferably, the method comprises a calculation of an angle corresponding to the change of point of view. Preferably, the modification of the pair of source images is carried out as a function of the calculated angle.

[0112] Preferably, the modification of the source image pair is carried out relatively and / or proportionally to the calculated angle.

[0113] The angle corresponding to the change of viewpoint can be defined as the angle formed between the optical axis of the observer before the change of viewpoint and the optical axis of the observer of the observer after the change of viewpoint.

[0114] Preferably:

[0115] - the change of point of view obtained is a function of, or linked to or corresponds to, a change of position and / or orientation, in space, of the head-mounted display or a change of position and / or orientation, in space, of an observer relative to the 3D display device, and, preferably, preferably to a change of position and / or orientation, in space, of the observer's head, and / or

[0116] - the attention zone corresponds to an area of ​​the displayed scene on which the observer's gaze is focused.

[0117] The change of viewpoint can be obtained in the absence of a change in position and / or orientation, in space, of the head-mounted display and / or the observer. The change of viewpoint obtained can be predetermined. The change of viewpoint can be chosen by the observer.

[0118] The point of view can correspond to the center of the straight line segment connecting the same part of each of the observer's eyes, for example to the center of the straight line connecting the foveas or the cornea or the iris of each of the observer's eyes.

[0119] The change of the observer's point of view can be defined as the observer's passage from the previous point of view to the new point of view.

[0120] The observer's point of view can be defined as the position and / or orientation of the observer, preferably as the position and / or orientation of the observer's head and / or eyes.

[0121] The area of ​​attention may be predetermined, selected, predefined, or defined. Preferably, the area of ​​attention may be predetermined, selected, predefined, or defined based on information relating to the area of ​​the pair of source or displayed images on which the observer's gaze is focused or intended to be focused.

[0122] Preferably, the area of ​​the images of the pair of source or displayed images on which the observer's gaze is focused or intended to be focused corresponds to the point, pixel or group(s) of pixels of the images of the pair of source or displayed images corresponding to the same point, object or group(s) of object(s) of the imaged scene.

[0123] Preferably, the area of ​​attention corresponds to the point, pixel or group of pixels or the area of ​​the images of the pair of source or displayed images on which the observer's gaze is focused or intended to be focused.

[0124] Preferably, the attention zone corresponds to a point, for example to a pixel or a group of pixels, of each of the images of the pair of source or displayed images located or intended to be located within the cone of the respective fields of vision of each of the observer's eyes. Preferably, the attention zone corresponds to the point of the images of the pair of source or displayed images of the scene having the shallowest depth within the cone of the observer's field of vision. Preferably, the cone of the observer's field of vision, preferably extending around an axis of revolution of the optical axis of an eye of the observer, formed by a solid angle of between 3 and 5° relative to the optical axis of an eye of the observer.

[0125] Preferably:

[0126] - the change of viewpoint is obtained from position data of the head-mounted display and / or movement data of the head-mounted display, called head-mounted display data, and / or

[0127] - the change of viewpoint is obtained from images of the observer and / or the observer's environment and / or the head-mounted display or the 3D display device, and / or

[0128] - the attention area is obtained, measured, detected, determined or calculated from images of the observer's eyes.

[0129] Preferably, the at least one image of the observer's eyes comes from an imaging system capable of imaging the observer and / or the observer's eyes.

[0130] Determining the observer's viewpoint and / or attention area may be achieved by eye tracking. Data from the head-mounted display or 3D display device may be measured or detected by the head-mounted display or 3D display device. Changes in viewpoint may be determined or calculated from the data from the head-mounted display or 3D display device.

[0131] The images of the viewer and / or the head-mounted display or 3D display device may be acquired by an external imaging system, preferably not belonging to or being part of the head-mounted display or 3D display device. The images of the viewer's environment may be acquired by the external imaging system and / or by an imaging system of the head-mounted display or 3D display device.

[0132] The external imaging system may be arranged to image the viewer and / or the head-mounted display or 3D display device.

[0133] Preferably, the method comprises linear or non-linear filtering of the data from the head-mounted display to limit an amplitude of the change in viewpoint obtained, as a function of time, to a value lower than a limit value.

[0134] Preferably, the method comprises non-linear clipping of the head-mounted display data to limit an amplitude of the resulting viewpoint change to a value less than a limit value.

[0135] Preferably, the method comprises non-linear dynamic compression of the head-mounted display data to attenuate an amplitude variation of the obtained viewpoint change, as a function of time, such that said amplitude variation does not exceed a variation amplitude threshold value.

[0136] The filtering and / or clipping and / or dynamic compression steps have the effect of limiting and / or reducing the amplitude and / or the variation in amplitude, as a function of time, of the change in point of view.

[0137] The change in position and / or orientation, in space, of the head-mounted display and / or the observer can be defined or qualified as a real or effective change of point of view.

[0138] The effective viewpoint change may be defined as the viewpoint change measured or detected by the head-mounted display or 3D display device, or determined or calculated from data from the head-mounted display or 3D display device.

[0139] The resulting viewpoint change may be equal to, correspond to, or be proportional to the actual viewpoint change. This is particularly the case, for example, when the amplitude and / or variation in amplitude of the viewpoint change is less than the limit value(s) or threshold value.

[0140] The resulting viewpoint change may differ from the actual viewpoint change. This is particularly the case, for example, when the amplitude and / or variation in amplitude of the viewpoint change is greater than the limit value(s) or threshold value.

[0141] Preferably, the method comprises a step of acquiring images of the scene, and in particular of the pair of source images, and / or images of the observer and / or of at least part of an environment of the head-mounted display or of at least part of an environment of the 3D display device.

[0142] The method may comprise a step of acquiring the pair of source images and / or the at least two images of the scene from the same stationary imaging system, preferably by means of the two separate imaging systems and / or from the same mobile imaging system.

[0143] According to another aspect of the invention, a device is proposed comprising means configured to implement all the steps of the method according to the invention.

[0144] According to the invention, a data processing device is also proposed comprising means arranged and / or programmed and / or configured to implement the method according to the invention.

[0145] According to the invention, there is also provided a computer program comprising instructions which, when the program is executed by a computer, cause the latter to implement the method according to the invention. According to the invention, there is also provided a computer-readable medium, for example a recording medium, comprising instructions which, when executed by a computer, cause the latter to implement the method according to the invention.

[0146] According to the invention, a computer-readable data carrier is also provided, on which the computer program according to the invention is recorded.

[0147] According to the invention, a virtual reality / augmented reality head-mounted display, known as a VR / VA head-mounted display, is also proposed, comprising means for displaying images in monocular or binocular vision and:

[0148] - means arranged and / or programmed and / or configured to implement the method according to the invention, and / or

[0149] - communication means arranged and / or programmed and / or configured for and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the method according to the invention.

[0150] Preferably, the VR / VA head-mounted display comprises at least one imaging system arranged to image the eyes of an observer wearing said VR / VA head-mounted display.

[0151] Preferably, the VR / VA head-mounted display includes:

[0152] - at least one imaging system arranged to image at least part of an environment of said, or of a space surrounding said, VR / VA head-mounted display, and / or of an environment of the observer, and / or

[0153] - means arranged to detect a movement and / or a relative position of said VR / VA head-mounted display.

[0154] According to the invention, there is also provided a 3D display device, arranged to display 3D images, comprising:

[0155] - means arranged and / or programmed and / or configured to implement the method according to the invention, and / or - communication means arranged and / or programmed and / or configured for and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the method according to the invention.

[0156] Preferably, the 3D display device comprises at least one imaging system arranged to image the eyes of an observer.

[0157] Preferably, the 3D display device comprises:

[0158] - at least one imaging system arranged to image at least part of an environment of said 3D display device and / or of the observer and / or of an environment of the observer, and / or

[0159] - means arranged to detect a movement and / or a position of an observer relative to said 3D display device.

[0160] By way of non-limiting examples, the 3D display device may be a screen, for example LCD or plasma, or a video projector.

[0161] The at least two images of the scene, from which a depth map of the scene is generated and / or from which depth information of the scene is extracted, and / or the pair of source images may be acquired or originate or be obtained, but not necessarily, by the imaging system of the VR / VA head-mounted display or the 3D display device.

[0162] The VR / VA head-mounted display or 3D display device may include multiple imaging systems.

[0163] The same or the two separate imaging systems from which the at least two images of the scene, from which a depth map of the scene is generated and / or from which depth information of the scene is extracted, are obtained or acquired may be the imaging system(s) of the VR / VA head-mounted display or by the 3D display device.

[0164] Preferably, the head-mounted display or the 3D display device is arranged to measure and / or detect the change of viewpoint or to determine and / or calculate the change of viewpoint from the data of the VR / VA head-mounted display or the 3D display device. The data of the head-mounted display or the 3D display device may be measured, detected, determined or calculated from: images from at least one imaging system arranged to image at least a portion of an environment of said VR / VA head-mounted display or at least a portion of an environment of said 3D display device, and / or

[0165] - means arranged to detect a movement and / or a relative position of said VR / VA head-mounted display or a relative position of the observer relative to the 3D display device.

[0166] Preferably, the head-mounted display and the 3D display device according to the invention are arranged to provide or enable or provide a 3D rendering from the pair of stereoscopic 2D images of the scene displayed. Where appropriate, the head-mounted display may comprise, or the 3D display device may be combined with or used with:

[0167] - synchronous shutter means, active (or synchronous) glasses in the case of the 3D display device,

[0168] - polarizing means or polarizing glasses.

[0169] Preferably, according to the invention, stereoscopic rendering is understood to mean any display method or process making it possible to provide a 3D rendering from a pair of displayed 2D images.

[0170] A person skilled in the art will be able to envisage the different display modes capable of providing a 3D rendering from a pair of 2D images to be displayed.

[0171] Preferably, both images of the pair of displayed images may be displayed simultaneously in the binocular head-mounted display, with one displayed image being viewed by the right eye and one displayed image being viewed by the left eye.

[0172] Preferably, the two images of the pair of displayed images can be displayed simultaneously, by spatial multiplexing, in the monocular head-mounted display or the 3D display device, for example in the form of a single composite or raster image comprising a portion of each of the two images of the pair of images. In this case, only a portion of each of the two displayed images is displayed in a single image.

[0173] Preferably, the two images of the pair of images displayed can be displayed simultaneously in the monocular head-mounted display or the 3D display device by superimposing the two images of the pair of images, each image of the pair of images displayed having a distinct polarization. Preferably, the two images of the scene displayed can be displayed consecutively or alternately or iteratively, for example by time multiplexing, in a synchronous monocular head-mounted display or on a 3D display device, preferably matched by the user wearing synchronous glasses. Preferably, the head-mounted display is a binocular head-mounted display in which each image of the scene is displayed so as to be seen by only one of the eyes so as to allow stereoscopic vision of the scene.

[0174] Preferably, the head-mounted display is worn by the user when implementing the method.

[0175] Preferably, the method according to the invention is suitable, more preferably is particularly suitable, more preferably is designed and particularly advantageously is specially designed, to be implemented by the VR / VA head-mounted display or by the 3D display device according to the invention. Also, any characteristic of the method according to the invention is directly transposable to the VR / VA head-mounted display or to the 3D display device according to the invention, and vice versa.

[0176] Description of figures

[0177] Other advantages and particularities of the invention will appear on reading the detailed description of implementations and embodiments which are in no way limiting, and the following appended drawings: FIGURE 1 represents a flowchart of an embodiment of the method according to the invention, FIGURES 2A and 2B are schematic representations illustrating a pair of stereoscopic 2D images of a scene displayed in a head-mounted display or in a 3D display device, FIGURE 3 illustrates, on the left, a schematic representation of a displayed image and, on the right, a schematic representation of the modified image, restoring a change of viewpoint of the observer horizontally with respect to the displayed image, obtained by implementing the method, FIGURES 4A and 4B are schematic representations of two examples of changes of viewpoint of a scene by an observer,FIGURE 5 is a schematic representation of a depth map of the scene shown in FIGURES 2A and 2B, FIGURE 6 is a schematic representation of an object projected through a lens onto an optical sensor, FIGURES 7 and 8 show a schematic representation of a top perspective view of the scene shown in FIGURES 2A and 2B.,

[0178] Description of the embodiments

[0179] The embodiments described below being in no way limiting, it will be possible in particular to consider variants of the invention comprising only a selection of the described characteristics, isolated from the other described characteristics (even if this selection is isolated within a sentence comprising these other characteristics), if this selection of characteristics is sufficient to confer a technical advantage or to differentiate the invention compared to the state of the prior art. This selection comprises at least one characteristic, preferably functional without structural details, or with only a part of the structural details if this part only is sufficient to confer a technical advantage or to differentiate the invention compared to the state of the prior art.

[0180] With reference to FIGURES 1 to 7, an embodiment of the image modification method according to the invention is presented, referred to as the method in the remainder of the description.

[0181] The person skilled in the art will be able to understand the concepts of 3D images, stereoscopy, monocular vision and binocular vision as well as their implementations.

[0182] FIGURE 1 illustrates a flowchart of the method according to the invention. The method comprises the step of obtaining n changes in viewpoint with respect to the imaged scene 21, 22. The change in viewpoint is obtained from a pair of stereoscopic 2D images of the scene 21, 22. This pair of stereoscopic 2D images of the scene 21, 22 is called a pair of source images.

[0183] The person skilled in the art will know that the pair 21, 22 or 31, 32 of stereoscopic 2D images is intended to provide a 3D rendering. In other words, the person skilled in the art will know that the two images of a pair 21, 22 or 31, 32 of stereoscopic 2D images correspond to the observation of the scene from a different point of view separated by the pupillary distance.

[0184] The method comprises the step of modifying at least one of the two images of the pair of source images of the scene 21, 22. The modification is carried out by moving a portion of the pixels or group(s) of pixels of at least one of the two images of the pair of source images 21, 22. In particular, the displacement of a portion of the pixels or group(s) of pixels of at least one of the two images of the pair of source images 21, 22 consists of a spreading of certain pixels or groups of pixels, preferably a spreading of the pixels not occluded by objects of the scene after a change of viewpoint. More preferably, the displacement of a portion of the pixels or group(s) of pixels of at least one of the two images of the pair of source images 21, 22 comprises a spreading of certain pixels or groups of pixels onto pixels occluded by objects of the scene after a change of viewpoint.

[0185] The modification of at least one of the two images of the pair of source images 21, 22 is a function of: a depth map of the imaged scene 21, 22, the change of point of view obtained.

[0186] A source image can be understood to mean any stored or available image, for example a stack of images or a video.

[0187] FIGURE 1 shows, in dotted lines, data that can be transmitted or provided, preferably in real time, by:

[0188] - external means, for example storage means, data processing means or image acquisition means, or by means of the video headset or the 3D display device according to the invention.

[0189] As non-limiting examples, the images of the scene, in particular the pair of source images or the at least two images of the scene from which the depth map of the scene is generated and / or from which the depth information of the scene is extracted, the depth information and / or the depth map, may be transmitted or provided.

[0190] The method according to the invention aims to modify at least one image of a pair of stereoscopic 2D images of a scene 21, 22 to restore, in a head-mounted display or in a 3D display device, to the observer a pair of stereoscopic 2D images corresponding to the observation of the scene from a different point of view. With reference to FIGURES 2A and 2B, and in the case of observation by a user of the pair of source images (unmodified) in a binocular head-mounted display, the pair of source images (unmodified) displayed 21, 22 is illustrated, the displayed image 21 is observed by the left eye of the observer in the binocular head-mounted display and the displayed image 22 is observed by the right eye of the observer in the binocular head-mounted display.With reference to FIGURES 3A and 3B, in the case of observation by a user of the pair of images in a binocular head-mounted display, the pair of modified images 31, 32 is illustrated, the modified image 31 is observed by the left eye of the observer in the binocular head-mounted display and the modified image 32 is observed by the right eye of the observer in the binocular head-mounted display. The implementation of the method allows the observer to observe the scene from a different point of view.

[0191] It can be heard by observer: the user of the head-mounted display, preferably the user wearing the head-mounted display, or the user of the 3D display device.

[0192] It is also provided to implement at least one of the steps of the method, advantageously to implement the method, for a stream of pairs of source images 21, 22. The method may be implemented for at least part of the images, preferably for each image, of the stream of source images.

[0193] The method may comprise a step of displaying, in the video headset or in the 3D display device, a stream of source image pairs 21, 22 before modification.

[0194] The method comprises a step of displaying, in the video headset or in the 3D display device, the pair of modified images 31, 32 and / or restored images 31, 32, preferably successively.

[0195] Preferably, the pair of images, sources 21, 22 or modified 31, 32, or the stream of pairs of images, sources 21, 22 or modified 31, 32, can be displayed:

[0196] - simultaneously, or

[0197] - consecutively or alternately or iteratively.

[0198] In the remainder of the description, it will be considered that the pair of source images 21, 22 will be the pair of displayed images 21, 22. Indeed, advantageously, the pair of source images 21, 22 will be displayed before the implementation of the modification step.

[0199] It can be understood by "consecutively displaying the pair of images or the stream of pairs of images": a successive or alternating or sequential or looped display of the pair(s) of images 21, 22, 31, 32 of the scene. It can be understood by "consecutively displaying the pair of images 21, 22, 31, 32, or the stream of pairs of images 21, 22, 31, 32": a time-division multiplexed display of the two images of the pair of images 21, 22, 31, 32 displayed.

[0200] By simultaneous display of the pair of images 21, 22, 31, 32, it can be understood the superposition of the two images of the pair of images, each image of the pair of images displayed having a distinct polarization. By simultaneous display of the pair of images 21, 22, 31, 32, it can be understood the spatial multiplexing of the two images of the pair of images, that is to say the display of a single composite or raster image comprising a part of each of the two images of the pair of images displayed.

[0201] Also, with reference to FIGURES 2A and 2B illustrating a schematic representation of the pair of stereoscopic 2D images of a scene displayed 21, 22 in a head-mounted display or a 3D display device. Those skilled in the art will observe that the two displayed images 21, 22 correspond to two different viewpoints of the scene so that the observer is able to observe a perspective or relief view of the scene. It will be considered in the present description that the observer of the scene displayed 21, 22 in the head-mounted display wears the head-mounted display or observes the scene in a 3D display device, where appropriate with suitable observation means (synchronous glasses, polarizing glasses or glasses equipped with red / blue filters).

[0202] The method also comprises the step of obtaining a change of point of view with respect to the displayed scene 21, 22. The change of point will preferably be a function of a change in position and / or orientation, in space, of the head-mounted display, and therefore of a change in position and / or orientation of the observer, and in particular of the observer's head, in space.

[0203] However, it is not excluded according to the invention that the change of point of view or the area of ​​attention is predetermined. The change of point of view, for example, corresponds to a model or a sequence of given change(s) which will be implemented in the step of obtaining change of point of view to propose or offer or submit to the observer a change of point of view without the head-mounted display or the observer changing position and / or orientation.

[0204] The change of viewpoint or the area of ​​attention may be provided or indicated by a third party or preferably by the observer himself. By way of non-limiting examples, the change of viewpoint or the area of ​​attention may be indicated by means of a touch zone (a smartphone or tablet screen for example), a movement of the observer's hand or limb, preferably detected by the head-mounted display or by the 3D display device, or any other means that a person skilled in the art will be able to envisage.

[0205] Alternatively, and by way of example, the observer may look at the image from a first position constituting the observer's point of view and may provide a second position different from his own, with the aim of obtaining the image as he would see it from the point of view associated with this second position.

[0206] As previously mentioned, for the purpose of conciseness and simplification of the description of the invention, the step of modifying one or more images of the pair of source images 21, 22 will be described as applying to both images of the pair of source images 21, 22. Indeed, in most practical cases, both images of the pair of source images 21, 22 will be modified. This corresponds to the common case where the observer, preferably the observer's head, will change position and / or orientation.

[0207] However, in certain cases, only one of the two images of the pair of source images 21, 22 will be modified. These may be cases in which the change of point of view occurs on only one of the user's eyes, and therefore on only one of the two source images 21, 22. By way of non-limiting example, this may be, for example, the case of a rotation of the user's head around the visual axis of one of the user's eyes. In this case, the change of point occurs on only one of the two images of the pair of displayed source images 21, 22.

[0208] According to the invention, a monocular or binocular virtual reality / augmented reality head-mounted display, known as a VR / VA head-mounted display, is also described.

[0209] According to the invention, a 3D display device is also described, arranged to display 3D images.

[0210] The VR / VA head-mounted display may be binocular and include means for displaying binocular images of the two images 21, 22, 31, 32 of the pair of images.

[0211] The VR / VA head-mounted display and the 3D display device may be monocular and comprise means for monocularly displaying the two images 21, 22, 31, 32 of the pair of images.

[0212] The VR / VA head-mounted display and the 3D display device further comprise means arranged and / or programmed and / or configured to implement the method according to the invention. Also, in this case, the entire method can be implemented by or in the head-mounted display and by the 3D display device.

[0213] Alternatively or in combination, the VR / VA head-mounted display and the 3D display device further comprise communication means arranged and / or programmed and / or configured for and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the method according to the invention. Also, in this case, only part of the steps of the method, by way of non-limiting example only the display of the two images 21, 22, 31, 32 can be implemented by or in the head-mounted display or by the 3D display device.

[0214] The VR / VA head-mounted display device and the 3D display device may comprise a processing unit or a processor arranged and / or programmed and / or configured to implement the method according to the invention. Alternatively or in combination, the external means may comprise a processing unit or a processor arranged and / or programmed and / or configured to process data from the VR / VA head-mounted display device or the 3D display device and / or to implement the method according to the invention.

[0215] The VR / VA head-mounted display and the 3D imaging device comprise at least one imaging system arranged to image the eyes of an observer wearing said VR / VA head-mounted display or of an observer of the 3D display device.

[0216] In this description, the term "imaging system" means a camera. The camera comprises an optical lens and an optical sensor.

[0217] The processing unit may be arranged and / or programmed and / or configured to determine and / or calculate the change of viewpoint and / or the attention zone 8 based on data from the at least one imaging system arranged to image the eyes of an observer.

[0218] The VR / VA head-mounted display and the 3D display device comprise at least one imaging system arranged to image at least a portion of an environment of, or a space surrounding, the VR / VA head-mounted display and / or the user and / or the 3D display device.

[0219] The processing unit may be arranged and / or programmed and / or configured to determine and / or measure and / or calculate the change of viewpoint based on data from the at least one imaging system arranged to image at least part of an environment of the VR / VA head-mounted display and / or the user and / or the 3D display device.

[0220] The VR / VA head-mounted display comprises means arranged to detect a movement and / or a relative position of the VR / VA head-mounted display. By way of non-limiting examples, the means arranged to detect a movement and / or a position of the VR / VA head-mounted display may be a gyroscope or an accelerometer.

[0221] Also, advantageously, the change of point of view obtained is a function of or is calculated and / or determined from data, coming for example from the means arranged to detect a movement and / or a relative position of the VR / VA head-mounted display, called head-mounted display data.

[0222] The effective viewpoint change can be defined as the viewpoint change measured or detected by the head-mounted display or 3D display device, or determined or calculated from the data from the head-mounted display or 3D display device. In other words, the effective viewpoint change corresponds to the actual viewpoint change made by the head-mounted display and by the user. The resulting viewpoint change may, in some cases, correspond to or be equal to the effective viewpoint change. However, in most cases, the resulting viewpoint change will be a function of, calculated from, or determined from the effective viewpoint change, but will not be equal to the effective viewpoint change.

[0223] The change of viewpoint is obtained from images of the observer and / or the observer's environment and / or the head-mounted display and / or the 3D display device. The change of viewpoint may be obtained, alternatively or in combination, from data originating from images of the observer, for example from the at least one imaging system of the head-mounted display or the 3D display device arranged to image the eyes of an observer, and / or from the observer's environment, for example from the at least one imaging system of the head-mounted display or the 3D display device arranged to image at least a portion of the environment.

[0224] Preferably, the attention zone 8 is obtained from images of the observer's eyes. With reference to FIGURES 2A, 2B, 3A and 3B, the pairs of modified images 31, 32 obtained or restored are illustrated, i.e. the two modified images of the scene, after implementation of the method. The two images 31 and 32 of the pair of modified images 31, 32 correspond respectively to the two source images 21 and 22 of the pair of source images 21, 22 (illustrated respectively in FIGURES 2A and 2B) modified or after modification or after implementation of the method.

[0225] According to the non-limiting embodiment presented, only a portion of the pixels or groups of pixels of the images 21, 22 of the pair of source images 21, 22 is moved and / or only a portion of the pixels or groups of pixels of the images 21, 22 of the pair of source images 21, 22 is spread, during the modification step, as a function of the depth information that the pixel considered or the group of pixels considered contains and of the change of point of view. This embodiment has the advantage of limiting the resources necessary for the modification step. In this case, and as illustrated in FIGURES 2A, 2B, 3A and 3B, it may be preferred not to modify the pixels and / or group(s) of pixels located in the background 10 of the imaged scene 21, 22 in order to provide a more faithful or realistic rendering.

[0226] However, in an advantageous embodiment making it possible to restore the most realistic rendering possible, all of the pixels or groups of pixels of the images 21, 22 of the pair of source images 21, 22 are moved and / or all of the pixels or groups of pixels of the images 21, 22 of the pair of source images 21, 22 are spread, during the modification step, as a function of the depth information that the pixel considered or the group of pixels considered contains and of the change of point of view.

[0227] According to this non-limiting embodiment, each pixel or each group of pixels will be moved proportionally to the depth of the point, object or group of objects in the imaged scene 21, 22. According to the invention, proportionally is understood to mean: proportionally or inversely proportionally. Furthermore, preferably, each pixel or each group of pixels will be moved according to a reference depth. The reference depth is advantageously associated with one or more objects in the scene or a group of objects. In practice, each pixel or each group of pixels will be moved relative to this reference depth. According to the implementation illustrated in FIGURE 7 of this embodiment, the background 10 of the scene (or for example objects with which the highest depth of the depth map is associated) has been chosen as the reference depth.In this case, all the objects located between the background 10 and the observer will be displaced inversely proportional to their depth (the objects in the foreground (character 9 in FIGURE 7 for example) will be displaced more than those in mid-depth (tree 8 in FIGURE 7 for example) which will be displaced more than the objects in the background 10. Also, according to the embodiment illustrated in FIGURE 7, the closer the pixel considered or a group of pixels considered is to the imaging system used to image the scene, the greater the displacement and / or the spread will be (relative to the pixels and group of pixels having a greater depth).

[0228] According to this non-limiting embodiment, each pixel or each group of pixels will also be moved proportionally to the angle a corresponding to the change of point of view.

[0229] The modification of the source image pair can be inversely proportional to the depth associated with the pixel or group of pixels considered.

[0230] In some cases, the displacement of pixel(s) or each group of pixels may be weighted based on sharpness or any other parameter known to those skilled in the art, such as, by way of non-limiting examples, the parameters R, G, B (for R (Red), VI (Green, or G (Green)), V2 (Green, or G (Green)), B (Blue)), a dilation parameter and / or an outline parameter.

[0231] According to another non-limiting embodiment, only a portion of the images of the pair of source images 21, 22 will be modified. This embodiment applies, in particular but not exclusively, to the case in which it is detected that the observer focuses his gaze on the same point, object or group of objects of the displayed scene during the change of viewpoint. The attention zone corresponds to zone 8 of the pair of displayed images 21, 22 on which the observer's gaze is focused.

[0232] Although the attention zone 8 may be data used according to the method, this attention zone 8 may be deduced from the data relating to the observer's gaze. The data relating to the user's gaze may be obtained by eye tracking without the acquisition step or even the step of determining the attention zone from this data necessarily being part of the method according to the invention. In this case, the data relating to the user's gaze may be acquired by one or more imaging systems of the head-mounted display arranged to image the eyes of an observer wearing said head-mounted display. A person skilled in the art knows the techniques for recording eye movement and also knows the data that may be derived therefrom, such as, for example, the foveal path.The method may include the step of acquiring data relating to the observer's gaze, for example, by means of one or more optical systems which may include one or more cameras.

[0233] With reference to FIGURES 4A and 4B, two examples of a change in viewpoint relative to the displayed scene 21, 22 are illustrated, for example in a monocular head-mounted display or a 3D display device. FIGURES 4A and 4B illustrate the change in viewpoint as the observer would see it if he were directly observing the scene. Also illustrated is the field of vision 11 of the observer. In FIGURE 3A, the change in viewpoint corresponds to a displacement dv parallel to the plane in which the pair of images is displayed 21, 22, which would correspond to a horizontal displacement of the observer if he were actually observing the scene as imaged. In FIGURE 3B, the change in viewpoint corresponds to a displacement dv parallel to the plane in which the pair of images is displayed and to a displacement dz perpendicular to the plane in which the pair of images is displayed 21, 22.The displacement dz would correspond to a horizontal displacement moving away or bringing closer the observer who was actually observing the scene as imaged.

[0234] The observer's optical axis 6 is shown before changing viewpoint and the observer's optical axis 7 after changing viewpoint. FIGURE 3A illustrates the case of a change of viewpoint during which the observer maintains his gaze focused on the same attention zone 8.

[0235] The change of viewpoint obtained may be effective, that is to say it corresponds to a change in position of the observer, that is to say to a movement of the observer's head, relative to the pair of images displayed 21, 22. However, the change of viewpoint may also be fictitious, that is to say the position of the observer's eyes relative to the pair of images displayed 21, 22 remains identical. When the change of viewpoint is fictitious, the new viewpoint therefore does not correspond to the position from which the observer looks at the image but to the position from which the observer would look at it.

[0236] The method may comprise the step of generating the depth map of the displayed image from depth information of one or more images of the pair of source images 21, 22 and / or from sharpness information of one or more images of the pair of source images 21, 22.

[0237] A person skilled in the art will be able to choose the most appropriate sharpness information depending on the situation. The sharpness information contained in a pixel or block of pixels may include the entropy, energy and / or variance of an image parameter between neighboring pixels.

[0238] According to the embodiment, the depth map corresponds to an image comprising a single depth channel. Depth is understood to mean the physical distance between a point or a physical object of the scene, corresponding to a pixel or a group of pixels of a considered image, and the sensor from which the considered image was acquired. By way of illustration, a simplified schematic depth map associated with the image of the scene illustrated in FIGURE 2A is presented in FIGURE 5. On this depth map, the white part 4 corresponds to an object located, for example, five meters from the optical sensor of the camera and the gray area 5 of the background of the image corresponds to objects located beyond, for example, fifty meters from the optical sensor.The distance of five meters corresponds, for example, to the minimum focusing distance of the camera lens and the distance of fifty meters corresponds, for example, to the maximum focusing distance of the camera lens.

[0239] According to a first aspect of the invention, the depth information and / or the depth map is determined from at least two images of the scene coming from:

[0240] - from the same stationary imaging system and each being acquired with a different focusing distance, or

[0241] - of the same mobile imaging system having acquired, at least in part, the scene in two distinct positions, or

[0242] - at least two separate imaging systems arranged in two separate positions so as to each acquire, at least in part, the scene. In the case where the at least two images of the scene come from the same stationary imaging system for each shot, the at least two images are each acquired with a different focusing distance. A shooting technique known in the prior art is called "focus bracketing". It consists of acquiring a set of images of the same scene from the same stationary optical system, each image is acquired with a different focusing distance. The image acquired with the shortest focusing distance, noted Ido, will give the image with the largest viewing angle while the image acquired with the highest focusing distance, noted Idmax, will give the smallest viewing angle.Thus, some objects in the scene will not be found on Idmax either because they are outside the image field or because they are hidden by objects in the scene.

[0243] In the case of the same mobile imaging system having acquired, at least in part, the scene in two distinct positions and in the case where two distinct imaging systems are arranged in two distinct positions, a depth map can be determined from at least two images of the scene.

[0244] Preferably, in the case of the at least two separate imaging systems, they are stationary relative to each other.

[0245] In the case of the same mobile imaging system and in two imaging systems arranged in two separate positions, a disparity map of the imaged scene 21, 22 can be determined. The depth map can be calculated and / or determined from the disparity map. Thus, any characteristic relating to the depth map can be transposed to the disparity map, and vice versa. The disparity map can be defined as a map or an image containing depth information. In other words, the disparity map can be an image or a map comprising at least one depth channel of the imaged scene. Preferably, the depth information constitutes a depth field of the imaged scene.

[0246] Also, a group or block of pixels corresponding to one of the objects of the imaged scene 21, 22 will have depth information different from one or more other groups or blocks of pixels corresponding to the other or other objects of the imaged scene 21, 22. Consequently, each of the at least two images acquired by the same mobile imaging system or by the two imaging systems arranged in two distinct positions may be enriched by a depth field, that is to say that each pixel or group or block of pixels (of the matrix or table of pixels representing the image of the scene) corresponding to an object of the scene will have distinct depth information which corresponds to the distance between the object and the optical sensor; the depth of the different objects of the imaged scene being, in most cases, different for at least some of the objects of the scene.

[0247] In the remainder of this description, the term “at least two images of the scene” used alone designates the at least two images from which the depth information and / or the depth map is determined.

[0248] Preferably, but not necessarily, the at least two images of the scene comprise the pair of source images 21, 22. Even more preferably, the depth information and / or the depth map is determined from the pair of source images 21, 22.

[0249] Preferably, the depth information and / or the depth map is generated by selecting and / or identifying pixels or groups of pixels from the at least two images of the scene or by identifying and / or selecting one or more images from the at least two images. Alternatively, or in combination, the depth information and / or the depth map is generated by combining, assembling or associating pixels or groups of pixels from the at least two images of the scene from which the depth information and / or the depth map is determined.

[0250] Thus, the depth map according to the first aspect of the invention is obtained by aggregating pixels or groups of pixels from different images of the scene. Prior to their aggregation, the group or groups of pixels of the depth map to be generated are selected from the at least two images.

[0251] According to a first alternative of the first aspect of the invention, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by: - ​​for each point of the depth map to be generated, identification of the image of the scene, among the at least two images of the scene, for which the sharpness is maximum, and

[0252] - for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene for which the sharpness is maximum.

[0253] According to a variant of the first alternative, the depth information can be determined and / or the depth map can be generated, from the at least two images of the scene, by:

[0254] - identification of a local maximum or maxima of sharpness or locals within each of the at least two images, and

[0255] - for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene presenting the maximum local sharpness.

[0256] According to a second alternative of the first aspect of the invention, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by:

[0257] - for each point of the depth map to be generated, interpolation or extrapolation, from the respective focusing distances of the at least two images, of an optimal focusing distance for which each point of the depth map to be generated considered would present the maximum local sharpness, and,

[0258] - for each point of the depth map to be generated, association of said optimal focusing distance.

[0259] According to a variant of the second alternative, the depth information can be determined and / or the depth map can be generated, from the at least two images of the scene, by:

[0260] - for each point or object or pixel or group(s) of pixels within the at least two images interpolation or extrapolation, from the respective focusing distances of the at least two images, of an optimal focusing distance for which each point or object or pixel or group(s) of pixels within each of the at least two images would present the maximum local sharpness, and for each point of the depth map to be generated, association of said optimal focusing distance.

[0261] In other words, the generation of the depth map can be described as consisting of combining groups of pixels, from several images of the scene, each presenting a maximum of sharpness and enriching each group of pixels with the depth information associated with the object of the scene to which they relate.

[0262] In other words, the first and second alternatives of the first aspect of the invention can be defined as consisting of, from the at least two images of the scene:

[0263] - for each of the at least two images of the scene, identify or calculate or determine or extrapolate or interpolate a local maximum or maxima of sharpness within each of the at least two images, and

[0264] - determining the depth information and / or the depth map to be generated from the local or local sharpness maximum or maxima of each of the at least two images, for example by comparing the local or local sharpness maximum or maxima of each of the at least two images, and the focusing distance corresponding to the local or local sharpness maximum or maxima.

[0265] According to a third alternative of the first aspect of the invention, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by:

[0266] - for each point of the depth map to be generated, identification of the image of the scene, among the at least two images of the scene, for which a local variance is the highest,

[0267] - for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene for which local variance is the highest.

[0268] According to a variant of the third alternative, the depth information can be determined and / or the depth map can be generated, from the at least two images of the scene, by:

[0269] - for each image of the scene among the at least two images, calculation of a local variance, - for each point of the depth map to be generated, selection of the image of the scene whose local variance is the highest, and

[0270] - for each point of the depth map to be generated, association of the focusing distance corresponding to the selected image.

[0271] According to a fourth alternative, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by:

[0272] - for each point of the depth map to be generated, interpolation or extrapolation, from the respective focusing distances of the at least two images, of an optimal focusing distance for which the depth map would present the highest local variance, and

[0273] - for each point of the depth map to be generated, or association of said optimal focusing distance.

[0274] According to a variant of the fourth alternative, the depth information can be determined and / or the depth map can be generated, from the at least two images of the scene, by:

[0275] - for each point or object or pixel or group(s) of pixels within the at least two images interpolation or extrapolation, from the respective focusing distances of the at least two images, of an optimal focusing distance for which each point or object or pixel or group(s) of pixels within each of the at least two images presents(would) the maximum local variance, and

[0276] - for each point of the depth map to be generated, association of said optimal focusing distance.

[0277] In other words, the third and fourth alternatives of the first aspect of the invention can be defined as consisting of, from the at least two images of the scene:

[0278] - for each pixel considered or group of pixels considered, preferably corresponding or common to the at least two images of the scene, of the at least two images of the scene, identify or calculate or determine or extrapolate or interpolate the image of the scene, among the at least two images of the scene, for which the local variance of the pixel considered or of the group of pixels considered is maximum, - for each pixel or group of pixels of the at least two images of the scene, determine the depth information and / or generate the depth map for which the local variance of a pixel considered or of a group of pixels considered is maximum and of the focusing distance corresponding to the image of the scene for which the pixel considered or the group of pixels considered has the maximum local variance.

[0279] According to an embodiment of the third and fourth alternatives, the depth map is generated by calculating, for each pixel or group of pixels or objects of each image of the scene among the at least two images, a local variance. Preferably, any energy, entropy estimator, conventionally used in focus detection, can be used to determine the variance. Those skilled in the art know a set of techniques for calculating the local variance. As a non-limiting example, the local variance can be the variance of the intensity of the pixels, for example the averaged intensity of the R, G, B channels, calculated on a set of given pixel matrix, to be adjusted according to the resolution of the image and / or the processing power / capacity of the processing unit. Then, for each pixel, group of pixels or object of the image of the depth map to be generated, the image of the scene whose variance is the highest is selected.As defined above, and by way of non-limiting example, this step may be defined as selecting the sharpest image for a pixel, group of pixels or object under consideration. The focusing distance corresponding to the selected image, or the interpolated or interpolated focusing distance associated with the selected image, is associated with the pixel, group of pixels or object of the depth map image to be generated. When selecting the scene image whose calculated or determined or interpolated or extrapolated variance is the highest, a comparison may be made between the scene images and an identification of common objects and / or a correlation between the pixels or objects (groups of pixels) of the scene images.

[0280] As a non-limiting example, and with reference to FIGURE 6, it is possible to establish a discrete relationship between each image among the at least two images of the scene and the focusing distance, denoted f, associated with each of the images. According to the geometric optical relationship ^ + ^7 = [Math 1], it is possible to link the focal length df of the objective of the optical system, to the physical distance di between the object of the scene and the optical system (also called objective), object corresponding to the pixel or group of pixels appearing as sharp in the selected image, and to the distance do between the optical sensor and the optical system of the selected image. Considering, for example, a set of n images each acquired with a distinct focusing distance di between dimin, for example 20 cm, and dimax, corresponding, for example, to an infinite distance or 50 meters.The number of images n is an integer between one and the number of the at least two images. Considering also that there is a linear relationship between the focusing distances of each of the images, the physical distance or depth, denoted P, at which the point of the scene or the object of the scene is located, corresponding to the pixel or group of pixels whose local variance is the highest of the n images, can be expressed as. ■. Pi = pas / n _ [Math 2], where i is an integer that is equal to n-1. In other words, i is between 0 and n-1. For the distance di = di pas, n is equal to 1 and i is equal to 0. Advantageously, an exponential variation between each of the focusing distances of each of the images can be used. In this case, and by way of non-limiting example, the focusing distance can be noted as being equal to di = dimin*exp(k*i / (nl)) with n greater than or equal to 2 and dimin is the minimum focusing distance and k is a real number such that k = log(dimax / dimin) where dimax is the maximum focusing distance.

[0281] According to any one of the aspects or alternatives, the depth information and / or the sharpness information contained in one or both of the source image pair 21, 22 or, respectively, the depth information contained in the depth map is refined or enhanced by processing one or both of the source image pair 21, 22 or, respectively, the depth map.

[0282] According to the description of the embodiment, the processing phase is applied to one of the two images of the pair of source images 21, 22 or to each of the two images of the pair of source images 21, 22 separately.

[0283] At the end of the processing phase, one or more images of the pair of source images 21, 22, whose depth and / or sharpness information or, respectively, a depth map whose depth information is refined or improved, is obtained. The processing phase comprises an iterative modification of one of the images of the pair of source images 21, 22 considered where the depth map is modified so as to minimize a function E comprising:

[0284] - a term D, called difference term, determined by comparison of the image of the scene considered 21, 22 during iterative processing, or respectively of the depth map during iterative processing, convolved with a point spread function (PSF), dependent on the depth map, with the image of the scene considered 21, 22, or respectively with the depth map, before convolution, the PSF describing the response of an imaging system from which the image of the scene considered 21, 22, or respectively the depth map, is obtained, and

[0285] - a term A, called anomaly term, representative of defects or anomalies within the image of the scene considered 21, 22 during iterative processing, or respectively of the depth map during iterative processing.

[0286] The depth-dependent PSF describes the response of an imaging system from which the image(s) of the source image pair 21, 22 are obtained.

[0287] In other words, the processing phase constitutes an iterative loop in which, the number of iterations is noted k, the image(s) of the scene of the pair of source images 21, 22 constitute(s) the image processed during the first iteration at k=1. The processed image is then modified at each iteration of the loop to minimize the error function E.

[0288] As a non-limiting example, the term D may comprise the sum of several terms Di. Obtaining the distances Di are obtained by comparing the image considered during processing at iteration k convolved by the PSFs with the image displayed at iteration k-1 or convolved with the image of the scene 21, 22 considered at iteration k=0. The distances Di may be, among other things, obtained by an evaluation of the local colors, or by evaluation of another parameter contained in the pixels or groups of pixels, of the image(s) considered of the scene 21, 22 at iteration k opposite the discretization grid of the PSF, so as to know the calculated colors, or to know another parameter contained in the pixels or groups of pixels, at the positions of the photosites of the image considered at iteration k to make the distance comparisons at the appropriate locations.

[0289] As a non-limiting example, the term A, called penalty, representative of defects or anomalies within the image being processed at iteration k, determined from the image of the scene being processed at iteration k-1, can be calculated concomitantly with the term D.

[0290] According to the non-limiting embodiment, the term A comprises:

[0291] - at least one component Al whose effect is minimized for small differences in intensity between neighboring pixels of the image of the scene 21, 22 considered during processing at iteration k, and / or

[0292] - at least one component A2 whose effect is minimized for small differences in hue between neighboring pixels of the image of the scene 21, 22 considered during processing at iteration k, and / or

[0293] - at least one component A3 whose effect is minimized for low frequencies of changes of direction between neighboring pixels of the image of the scene 21, 22 considered during processing at iteration k drawing an outline of an object of the scene.

[0294] The second term A can thus include a sum of several terms Ai.

[0295] Preferably, prior to the processing phase, the method comprises a convolution of the image of the scene 21, 22 by an inverse function of the PSF, called iPSF. The convolution of the image of the scene 21, 22 considered can be carried out as a first step of the processing phase (initialization of the image during iterative processing).

[0296] The iterative modification of the image 21, 22 considered ends when the function E, or a combination of partial derivatives of the function E with respect to the image, with respect to the image currently being processed at iteration k or with respect to the image of the scene 21, 22 considered, is less than a minimization threshold, or when a certain number of iterations of the iterative modification of the image of the scene 21, 22 considered is reached; preferably, the image considered currently being processed thus modified is restored.

[0297] According to a second aspect of the invention, the step of generating the depth map of the imaged scene 21, 22 is implemented from one of the two images of the pair of source images 21, 22 or from each of the two images of the pair of source images 21, 22 separately, by means of a convolutional neural network. Those skilled in the art will know how to choose and adapt the most suitable neural network. By way of non-limiting example, the convolutional neural network may comprise several layers of neurons. Among these layers of neurons, the network comprises processing layers comprising neurons arranged to process the displayed image with a convolution. The neural network may also comprise sub-sampling (or pooling) layers, interposed between two processing layers, comprising neurons arranged to combine or merge the output data of the processing layers and reduce their size.By way of non-limiting example, the convolutional neural network may further comprise a correction layer, interposed between two processing layers, comprising neurons arranged to operate an activation function on the output data of the processing layers.

[0298] Still according to the advantageous embodiment mentioned above, in which all of the pixels or groups of pixels of the images 21, 22 of the pair of source images 21, 22 are moved and / or all of the pixels or groups of pixels of the images 21, 22 of the pair of source images 21, 22 are spread, during the modification step, as a function of the depth information that the pixel considered or the group of pixels considered contains and of the change of point of view. With reference to FIGURE 8, the tree 8 at mid-depth of the scene has been chosen as the reference depth. The change of point of view is therefore carried out relative to the depth PI associated with the tree 8 in the scene.As for the embodiment described in FIGURE 7, the person skilled in the art will be able to use simple mathematical tools allowing the calculation of the displacement (spreading and / or offset) of each pixel or group of pixels of the images 21, 22 of the pair of source images 21, 22 (such as trigonometry and simple geometry (Thales' theorem, etc.)). In the case illustrated in FIGURE 8, the pixels or groups of pixels will be displaced proportionally to the depth difference separating a pixel to be modified considered and the reference depth associated with the tree 8. In practice, the character 9, having a depth difference with the tree 8 (the reference depth) smaller than the difference between the background 10 and the reference depth, will be less offset than the background 10.

[0299] With reference to FIGURE 8, and for any of the aspects or alternatives, the modification of one of the images of the pair of source images 21, 22 is detailed as a function of the change in point of view. According to the non-limiting embodiment illustrated in FIGURE 8, the observer can maintain his gaze focused on the same area of ​​attention 8 during the change of point of view. This embodiment makes it possible to take into account the scenario in which the modification of the pair of source images 21, 22 is carried out as a function of the area of ​​attention 8 on which the observer's gaze remains focused during the change of point of view.

[0300] The person skilled in the art will be able to adapt and transpose the description of this non-limiting embodiment to the case of a change of point of view during which the observer modifies his attention zone 8 during the change of point of view and / or during which his gaze is focused at infinity before and / or after the change of point of view. In the case where the observer modifies his attention zone 8 during the change of point of view and / or his gaze is focused at infinity before and / or after the change of point of view, as described previously, the entire image (pixels or group(s) of pixels) will be modified.

[0301] In this embodiment, the modification of the pair of source images 21, 22 is also a function of the attention zone 8. According to this embodiment, the attention zone 8 could remain unchanged, during the modification of at least one of the displayed images, by anchoring the attention zone relative to the observer or to a reference frame other than or external to the observer.

[0302] To facilitate the description, FIGURE 8 shows the perspective view from above of the imaged scene 21, 22, i.e. the scene with the notion of depth. Illustrated is the optical axis 6 of the observer before changing the point of view and the optical axis 7 of the observer after changing the point of view. In the case presented, the given change in the observer's point of view corresponds to a displacement parallel to the plane of the displayed image 21, 22, which would correspond to a horizontal displacement of the observer if he were actually observing the scene as imaged. The attention zone 8 of the observer, on which the observer's gaze remains focused during the change of point of view, is the tree 8 which is located behind the individual 9 but in front of the landscape 10 in the background of the image.During the modification step, all of the objects in the scene 9, 10, other than the one on which the observer's gaze is focused, will be translated in the image plane in response to the change in point of view obtained. In addition, some objects or parts of objects in the scene, corresponding to some pixels of the outline of the attention zone 8, that is to say pixels located in the immediate vicinity of the attention zone 8, may be hidden by the object on which the observer's gaze is focused and some objects or parts of objects in the scene, corresponding to some pixels of the outline of the attention zone 8, which were hidden before the change in point of view may appear on the modified image 31, 32 according to the method.

[0303] The above description concerning the modification of the pair of source images 21, 22 as a function of the attention zone 8 is directly transposable to FIGURE 7. In this case, the attention zone _8 could be considered as being an area, an object or a group of objects of the background 10 of the scene.

[0304] FIGURES 3A and 3B illustrate the modification of the image corresponding to a given change in the observer's point of view corresponding to a displacement dv parallel to the plane of the displayed image 21, 22, which would correspond to a horizontal displacement of the observer if he were actually observing the scene as imaged. The schematic representation on the left illustrates the displayed image 21, 22 before the change in point of view and the schematic representation on the right illustrates the modified image 31, 32 according to the method restoring the change in point of view.

[0305] The change of point of view comprises a displacement (du, dv) in a reference frame or relative to a reference frame of one of the displayed images 21, 22. The reference frame of one of the displayed images may be the plane or a surface, for example curved, on which the image 21, 22 is displayed. The reference frame may be that or may be associated with the user or may be that or may be associated with the head-mounted display or may be that or may be associated with the 3D display device or may be a reference frame other than or external to the user and / or the head-mounted display.

[0306] The method comprises a calculation of an angle a corresponding to the change of point of view. The angle a corresponds to the angle formed between the optical axis 6 of the observer before change of point of view and the optical axis 7 after change of point of view. The modification of the pair of source images 21, 22 is carried out as a function of the calculated angle a.

[0307] The modification of the pair of source images 21, 22 comprises, for each pixel or group(s) of pixels of each image of the pair of source images 21, 22, a calculation of a difference between a depth associated with the attention zone 8 and a depth associated with each pixel or group(s) of pixels of each image of the pair of source images 21, 22 located outside the attention zone 8. This difference is calculated from the data of the depth map. The modification of the pair of source images 21, 22 is carried out according to the calculated difference.In the case where the observer modifies his attention zone 8 during the change of point of view and / or his gaze focused at infinity before and / or after the change of point of view, as described previously, the modification of the pair of source images 21, 22 comprises, for each pixel or group(s) of pixels of each image of the pair of source images 21, 22, a calculation of a difference between a depth associated with each pixel or group(s) of pixels of each image of the pair of source images 21, 22 before change of point of view and a depth associated with each pixel or group(s) of pixels of each image of the pair of source images 21, 22 after change of point of view. This difference is calculated from the data of the depth map. The modification of the pair of source images 21, 22 is carried out according to the calculated difference.

[0308] In practice, the translation of a given object in an image of the pair of source images 21, 22, excluding the object of the attention zone 8, will be proportional to the calculated deviation and to the angle a. The translation will be carried out along the axis connecting the optical axis 6 of the observer before changing the point of view and the optical axis 7 after changing the point of view. The objects being translated in the plane of the displayed image, a trigonometric relationship between the calculated depth deviation AP, the angle a and the distance, noted m, by which the object must be moved to restore the change of point of view can be established. For example, in the case of individual 9 in the foreground of scene imaged 21, 22, the depth gap AP is equal to Pl-PO, or the absolute value of Pl-PO, where PO corresponds to the depth of individual 9 in the scene and PI corresponds to the depth of attention area 8 in the scene.In the case of background 10 in the image, the depth difference AP is equal to P1-P2, or the absolute value of Pl-PO, where P2 corresponds to the depth of background 10 in the scene. Thus, as a non-limiting example, the relationship m = sin(^) x 2AP can be established.

[0309] According to the embodiment, the method comprises filtering the head-mounted display data to limit an amplitude of the obtained viewpoint change, as a function of time, to a value less than a limit value. For example, the limit value may be the absolute value or the norm of the displacement vector determined from the head-mounted display data. Preferably, the filtering is linear.

[0310] The first purpose of the filtering step is to cut off changes in viewpoint that have a high amplitude. In other words, the filtering step aims to prevent sudden variations or gradients in changes in viewpoint from occurring, for example when the observer wearing the head-mounted display suddenly turns around or suddenly pivots 90°. Indeed, since the modification made to the pair of source images 21, 22 is a function of the change in viewpoint, beyond a certain amplitude of change in viewpoint obtained, the objects of the modified images 31, 32 would be too strongly deformed and / or spread out, which could make the images thus modified 31, 32 blurred or unobservable. Indeed, the invention does not aim to generate data or information not present in the source image but only to modify the source image.In other words, it is a matter of truncating or not taking into account movements with too great a variation in amplitude in a short period of time so that the changes in point of view obtained do not deviate too much from the initial position and orientation of the head of the observer wearing the head-mounted display.

[0311] The filtering step may also comprise a step of filtering or smoothing out small changes in amplitude such as small physiological movements of the head. This makes it possible to avoid rendering to the observer small frequent modifications of the images of the pair of source images 21, 22 which could induce a seasickness effect in the observer wearing the head-mounted display. In other words, the filtering step may be defined as comprising the application of a band-pass filter, in particular a high-pass filter, or a subtraction of an estimate of the average position of the head-mounted display, to the changes in viewpoint actually detected. This therefore makes it possible to smooth out the change in viewpoint actually detected.

[0312] Viewpoint change variations, which can be described as high amplitude viewpoint changes, are amplitude changes occurring at low frequencies. High amplitude can be defined as an amplitude greater than 1 cm. Low frequency amplitude changes can be defined as having a frequency less than 1 Hz. High frequency viewpoint change variations can be defined as low amplitude head-mounted display movements, typically less than 0.1 cm. High frequency viewpoint changes can be defined as having a frequency greater than 10 Hz. The notion of “low” and “high” amplitude change, frequency and amplitude are relative. According to the embodiment, the method comprises non-linear clipping of the head-mounted display data to limit an amplitude of the obtained viewpoint change to a value less than a limit value.Clipping can be defined as a cut applied to the actual viewpoint changes, or to the data relating to the actual viewpoint changes, so as to limit or maintain the amplitude of the resulting viewpoint change used to modify the pair of source images 21, 22 below a maximum amplitude. Indeed, since the modification made to the pair of source images 21, 22 is a function of the viewpoint change, beyond a certain amplitude of the resulting viewpoint change, the objects of the modified images 31, 32 would be too strongly distorted and / or spread, potentially rendering the thus modified images 31, 32 blurred or unobservable. This step aims to limit the high values ​​of actual viewpoint change to maintain them below predefined maximum values.

[0313] Finally, the method may comprise non-linear dynamic compression of the head-mounted display data to attenuate an amplitude variation of the obtained viewpoint change, as a function of time, such that said amplitude variation does not exceed an amplitude variation threshold value. This non-linear dynamic compression of the effective viewpoint change is complementary to the clipping of the effective viewpoint change. The non-linear dynamic compression may be defined as a modulation of a gain applied to the effective viewpoint changes, or to the data relating to the effective viewpoint changes, so as to limit the amplitude variation of the obtained viewpoint change used to modify the pair of source images 21, 22.Indeed, in addition to cutting off the sudden strong amplitude variations of the effective change of viewpoint, it may be judicious to limit the variation of change of viewpoint to be restored to the limitation of the modification of the pair of source images 21, 22 below a predefined value beyond which the modified images 31, 32 are no longer usable.

[0314] Of course, the invention is not limited to the examples which have just been described and numerous adjustments can be made to these examples without departing from the scope of the invention.Thus, in variants of the previously described embodiments that can be combined with each other: the change of viewpoint comprises a displacement (du, dv and / or dz) in a frame of reference of one or both of the displayed images 21, 22 and / or in a frame of reference of the user of the head-mounted display or of the 3D imaging device, and / or the step of modifying the pair of source images 21, 22 is carried out by adding and / or deleting pixels or groups of pixels in at least one image of the pair of source images 21, 22, and / or the method may comprise the acquisition of the pair of source images 21, 22 of the scene originating from the same mobile or stationary imaging system and / or of the at least two images of the scene obtained, preferably by means of the two separate imaging systems or the mobile imaging system, in two positions or two separate viewpoints, and / or.

[0315] - in the present description, the mobile or stationary imaging system and / or the two separate imaging systems may be one, the imaging system(s) of the VR / VA head-mounted display or of the 3D display device, and / or there is provided according to the invention a computer program comprising instructions which, when the program is executed by a computer, lead the latter to implement the method according to any one of the embodiments described, and / or there is provided according to the invention a readable medium, in particular, by a computer or by any apparatus comprising a processing unit comprising instructions which, when executed by said computer or said apparatus lead the latter to implement the method according to any one of the embodiments described.

[0316] In addition, the various features, forms, variations and embodiments of the invention may be combined with each other in various combinations to the extent that they are not incompatible or mutually exclusive.

Claims

CLAIMS 1. Method for modifying images of a scene to restore a change in point of view of the scene, said method comprises the steps of: - obtain a change of point of view in relation to the imaged scene, - modify a pair of stereoscopic 2D images of the scene, called a pair of source images, by moving or spreading pixels or group(s) of pixels of at least one of the images of the pair of source images and / or by adding and / or deleting pixels or groups of pixels in at least one of the images of the pair of source images, as a function of: • a depth map of the imaged scene, • the change of point of view obtained.

2. Method according to the preceding claim, in which the modification step, or the obtaining step and the modification step, are applied to a stream of source image pairs.

3. Method according to claim 1 or 2, comprising: prior to the modification step, the step of displaying, in a video headset or on a 3D display device, the pair of source images, and / or the step of displaying, in a video headset or on a 3D display device, the modified pair of images.

4. Method according to any one of the preceding claims, in which the modification of the pair of source images is also a function of an area of the scene, called the attention area.

5. Method according to any one of the preceding claims, comprising a step of generating the depth map of the imaged scene from depth information of the imaged scene and / or from sharpness information of at least one of the images of the pair of source images.

6. Method according to the preceding claim, in which the depth map is generated, from at least two images of the scene originating from the same stationary imaging system and each being acquired with a different focusing distance and / or from at least two images of the scene originating from the same mobile imaging system having acquired, at least in part, the scene in two distinct positions and / or from at least two images of the scene originating from at least two distinct imaging systems, arranged in two distinct positions so as to each acquire, at least in part, the scene: - by selection and / or identification of pixels or groups of pixels from the at least two images of the scene or by identification and / or selection of an image or image(s) from among the at least two images, - by combining, assembling or associating pixels or groups of pixels of at least two images of the scene.

7. Method according to the preceding claim, in which the depth map of the displayed scene is generated by: - according to an alternative, called alternative A: • for each point of the depth map to be generated, identification of the image of the scene, among the at least two images of the scene, for which the sharpness is maximum, and • for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene for which the sharpness is maximum, and / or - according to an alternative, called alternative B: • for each point of the depth map to be generated, interpolation or extrapolation, from the respective focusing distances of the at least two images, of an optimal focusing distance for which each point of the depth map to be generated considered would present the maximum local sharpness, and • for each point of the depth map to be generated, association of said optimal focusing distance.

8. The method of claim 6, wherein the depth map of the displayed scene is generated by: - according to an alternative, called alternative C: • for each point of the depth map to be generated, identification of the image of the scene, among the at least two images of the scene, for which a local variance is the highest, • for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene for which local variance is the highest, or - according to an alternative, called alternative D: • for each point of the depth map to be generated, interpolation or extrapolation, from the respective focusing distances of the least two images, of an optimal focusing distance for which the depth map would exhibit the highest local variance, and • for each point of the depth map to be generated, or association of said optimal focusing distance.

9. Method according to any one of claims 1 to 4, comprising a step of generating the depth map of the imaged scene from at least one of the images of the pair of source images, by means of a neural network.

10. Method according to one of claims 5 to 9, in which the depth information and / or the sharpness information contained in an image of the scene considered or, respectively, the depth information contained in the depth map is refined by processing the image of the scene considered or, respectively, the depth map; at the end of the processing phase an image of the scene of which the depth information and / or the sharpness information or, respectively, a depth map of which the refined depth information is obtained;the processing phase comprises an iterative modification of the image of the scene considered or, respectively, of the depth map so as to minimize a function E comprising: a term D, called a difference term, determined by comparing the image of the scene considered during iterative processing, or respectively of the depth map during iterative processing, convolved by a point spread function (PSF), dependent on the depth map, with the image of the scene considered before convolution, or respectively with the depth map before convolution, the PSF describing the response of an imaging system from which the image of the scene considered, or respectively the depth map, is obtained, and a term A, called anomaly term, representative of defects or anomalies within the image of the scene considered during iterative processing, or respectively of the depth map during iterative processing.; 11. Method according to any one of the preceding claims, in which the modification of the pair of source images is proportional to the change of point of view.

12. Method according to any one of the preceding claims, in which the modification of the pair of source images is inversely proportional to the depth associated with the pixel or group of pixels considered.

13. Method according to any one of the preceding claims, comprising, for each pixel or group(s) of pixels of one or each image of the pair of source images, a calculation of a difference between a depth associated with the attention zone and a depth associated with each pixel or group(s) of pixels located outside the attention zone; the modification of the pair of source images is carried out as a function of the calculated difference.

14. Method according to any one of the preceding claims, comprising a calculation of an angle corresponding to the change of point of view; the modification of the pair of source images is carried out as a function of the calculated angle.

15. A method according to claim 3, or according to any one of claims 4 to 14 taken in combination with claim 3, wherein: - the change of point of view obtained is a function of a change in position and / or orientation, in space, of the head-mounted display or a change in position and / or orientation, in space, of an observer relative to the 3D display device, and / or - the attention zone corresponds to an area of the displayed scene on which the observer's gaze is focused.

16. A method according to claim 3, or according to any one of claims 4 to 15 taken in combination with claim 3, wherein: - the change of viewpoint is obtained from position data of the head-mounted display and / or movement data of the head-mounted display, called head-mounted display data, and / or - the change of viewpoint is obtained from images of the observer and / or the observer's environment and / or the head-mounted display or the 3D display device, and / or - the attention zone is obtained from images of the observer's eyes.

17. Method according to claim 15 or 16, comprising filtering the data from the head-mounted display to limit an amplitude of the change in viewpoint obtained, as a function of time, to a value less than a limit value.

18. Method according to any one of claims 15 to 17, comprising non-linear clipping of the data from the head-mounted display to limit an amplitude of the change in viewpoint obtained to a value less than a limit value.

19. Method according to any one of claims 15 to 18, comprising a non-linear dynamic compression of the data of the head-mounted display to attenuate a variation in amplitude of the change of point of view obtained, as a function of time, such that said amplitude variation does not exceed a threshold amplitude variation value.

20. Method according to claim 3, or according to any one of claims 4 to 19 taken in combination with claim 3, comprising a step of acquiring images of the scene, and in particular of the pair of source images, and / or images of the observer and / or of at least part of an environment of the head-mounted display or of at least part of an environment of the 3D display device.

21. Data processing device comprising means arranged and / or programmed and / or configured to implement the method according to any one of claims 1 to 20.

22. Computer program comprising instructions which, when the program is executed by a computer, cause the latter to implement the method according to any one of claims 1 to 20.

23. A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 20.

24. Virtual reality / augmented reality head-mounted display, known as VR / VA head-mounted display, comprising means for displaying images in monocular or binocular vision and: - means arranged and / or programmed and / or configured to implement the method according to any one of claims 1 to 20, and / or - communication means arranged and / or programmed and / or configured for and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the method according to any one of claims 1 to 20.

25. VR / VA head-mounted display according to the preceding claim, comprising at least one imaging system arranged to image the eyes of an observer wearing said VR / VA head-mounted display.

26. VR / VA head-mounted display according to claim 24 or 25, comprising: - at least one imaging system arranged to image at least part of an environment of said VR / VA head-mounted display and / or an environment of the observer, and / or - means arranged to detect a movement and / or a relative position of said VR / VA head-mounted display.

27. 3D display device, arranged to display 3D images, comprising: - means arranged and / or programmed and / or configured to implement the method according to any one of claims 1 to 20, and / or - communication means arranged and / or programmed and / or configured for and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the method according to any one of claims 1 to 20.

28. 3D display device according to the preceding claim, comprising at least one imaging system arranged to image the eyes of an observer.

29. 3D display device according to claim 27 or 28, comprising: - at least one imaging system arranged to image at least part of an environment of said 3D display device and / or of the observer and / or of an environment of the observer, and / or - means arranged to detect a movement and / or a position of an observer relative to said 3D display device.

Citation Information

Patent Citations

  • Field of view (FOV) throttling of virtual reality (VR) content in a head mounted display

    US20180096518A1

  • Method for 3D scene dense reconstruction based on monocular visual slam

    US20200273190A1

  • Head-mountable display system

    US20200296354A1