Method for carrying out 3D rendering for head-mounted displays and 3D display devices based on an image of a scene
The method addresses the inefficiencies of existing 3D image generation by using a depth map and pupillary distance to create stereoscopic 2D images for real-time 3D rendering in head-mounted and 3D display devices, enhancing depth perception without high computational costs.
Patent Information
- Application Number
- PCT/FR2023/052129
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-07-03
AI Technical Summary
Existing methods for generating 3D images are costly in terms of time, energy, and computing resources, and are incompatible with real-time dynamic visualization of 2D images, particularly in portable devices.
A method for 3D rendering from a single 2D image using a depth map and pupillary distance to create modified images that simulate stereoscopic vision, allowing 3D rendering in head-mounted or 3D display devices.
Enables efficient and immersive 3D rendering from 2D images in real-time, reducing computational and resource requirements while providing a realistic depth perception.
Smart Images

Figure FR2023052129_03072025_PF_FP_ABST
Abstract
Description
[0001] DESCRIPTION
[0002] 3D rendering method for head-mounted displays and 3D display devices from an image of a scene
[0003] Technical field
[0004] The present invention relates to stereoscopic 2D images intended to be displayed in a head-mounted display or in a 3D display device in order to render a 3D rendering when they are displayed.
[0005] The invention relates to the dynamic observation of stereoscopic 2D images displayed in a head-mounted display or in a 3D display device. The invention also relates to the immersive observation of stereoscopic 2D images displayed in a head-mounted display or in a 3D display device.
[0006] The invention does not relate to 3D images.
[0007] State of the prior art
[0008] 3D images are known in the prior art which make it possible to introduce into displayed 2D images a perspective effect such as it would be perceived if the imaged scene were observed directly by the observer. According to the invention, the 3D images are not stereoscopic images but three-dimensional images offering a relief representation of the objects. The invention relates to stereoscopic 2D images making it possible to provide a 3D rendering, when they are displayed in a suitable display medium, by exploiting the binocular vision of the observer which gives him a perception of depth. The synthesis of 3D images can be carried out directly from a model of a real scene. The production of 3D images is costly in terms of time and the computer resources required for processing and storing these large images.3D images can be understood to mean relief images obtained, for example, by 3D surface modeling of space or from two images with different disparities, typically two images obtained from two distinct viewpoints. The load is increased tenfold when it comes to a stream of 3D images, in particular a stream of 3D images resulting from 3D surface modeling of space. In addition, the time required for processing is incompatible with dynamic visualization of 2D image(s) acquired in real time. Stereoscopic vision is also known in the state of the art. Binocular human vision, or more broadly animal vision, from two images that are compared by the brain, constitutes the most widespread example.The biological processing of the images obtained is extremely efficient, since it provides in real time a notion of depth of the observed scene, allowing for example to move knowing at what relative distance the observed objects are.
[0009] The observation of stereoscopic images is known in the state of the art, which is inspired by or mimics the principle of stereoscopic vision. The principle of stereoscopic observation consists of increasing the immersive nature of image observation by adding a depth effect by projecting a different image onto each eye of an observer. In practice, a stereoscopic image comprises two distinct images of the same scene acquired in two distinct positions in space, in particular by two cameras spaced apart by the pupillary distance. Human stereoscopic vision will make it possible, from a stereoscopic image, to reconstruct a scene including depth information. The time required to generate 3D images is incompatible with dynamic visualization of 2D image(s) acquired in real time or very energy-intensive on a portable device.Producing 3D images is costly in terms of time, energy, and computing resources needed to process and store these images.
[0010] One aim of the invention is to propose an image processing method enabling:
[0011] - to address problems with state-of-the-art processes, and / or
[0012] - to provide from a 2D image, in particular from a single 2D image, a 3D rendering from two 2D images in a head-mounted display device or a 3D display device, and / or
[0013] - to restore, to an observer, a perspective effect to a 2D image, and / or
[0014] - to restore a perspective effect to an observer solely by displaying 2D images. Presentation of the invention
[0015] For this purpose, a 3D rendering method is proposed from an image of a scene. The method comprises the step of modifying the image of the scene, preferably the single image of the scene, to obtain a modified image, preferably a single modified image, or respectively two modified images, by moving or spreading a part of the pixels or group(s) of pixels of the image of the scene and / or by adding and / or deleting pixels or groups of pixels in the image of the scene, depending on:
[0016] - a depth map of the scene,
[0017] - a change in the point of view of the scene which is a function of pupillary distance.
[0018] Preferably, at the end of the step of modifying the image of the scene, a pair of images is obtained. The pair of images comprises:
[0019] - two modified images of the scene, or
[0020] - an image of the modified scene and the image of the scene (unmodified).
[0021] Both images in the image pair are intended to provide a 3D rendering when displayed on a head-mounted display or 3D display device.
[0022] Preferably, the two images of the pair of images correspond to two distinct viewpoints of the scene, and in particular to two images of the scene observed from two viewpoints spaced apart by the pupillary distance.
[0023] Preferably, the method aims to restore two stereoscopic 2D images corresponding to two distinct viewpoints of the scene, in particular to two viewpoints of the scene spaced apart by the pupillary distance.
[0024] Preferably, the difference in point of view between the two images of the pair of images, restored or obtained by implementing the method, is equal to the pupillary distance.
[0025] It can be understood by point of view, a position or coordinates of space, preferably relative to the image of the scene.
[0026] Preferably, the step of modifying the image of the scene is carried out by means of a processing unit.
[0027] Preferably, each data processing step of the method according to the invention, comprising for example any operation carried out from or on data, for example a calculation, a determination or a comparison, is implemented by a processing unit or any device or system capable of processing data. Preferably, the modification of the image of the scene is carried out or implemented or carried out on at least a part of the image of the scene, preferably on the entire image of the scene. In other words, preferably, the displacement of a part of the pixels or group(s) of pixels of the image of the scene and / or the addition and / or deletion of pixels or groups of pixels in the image of the scene is carried out on the entire image of the scene.
[0028] Preferably, the displacement of a portion of the pixels or group(s) of pixels of the scene image makes it possible to reproduce a parallax effect induced by the change of viewpoint. Preferably, displacement of a portion of the pixels or group(s) of pixels of the scene image is understood to mean: a modification of the position and / or a translation and / or a spreading of a portion of the pixels or group(s) of pixels within the scene image, preferably in the plane of the scene image.
[0029] Preferably, each pixel or each group of pixels is moved and / or translated and / or spread according to a reference depth, more preferably relative to the reference depth. Preferably, the reference depth is associated with one or more objects in the scene or a group of objects.
[0030] The depth map may be defined as a map or image containing, preferably only, depth information. That is, the depth map may be an image or map comprising a single depth channel of the imaged scene. Preferably, the depth information constitutes a depth field of the imaged scene.
[0031] Preferably, a pixel corresponds to or is associated with a point in the imaged scene and a group of pixels corresponds to or is associated with an object in the imaged scene, or a part of an object in the scene located at a certain depth, or between two depth limits.
[0032] Depth information can be understood as data representing or corresponding to or providing information on the depth associated with a pixel, corresponding to a point in the imaged scene, or to a group of pixels, corresponding to an object in the imaged scene, in the image of the scene.
[0033] Preferably, the image of the scene was acquired by an imaging system.
[0034] Depth can be understood as a distance between a point or object in the imaged scene and the optical sensor of the imaging system from which the image of the scene was acquired. The depth field can comprise depth information of all or part of the points or objects in the scene, preferably a depth field of the imaged scene. The depth field can be contained in a channel of all or part of the pixels in the image of the scene.
[0035] Preferably, the change of point of view corresponds to the passage from a point of view, called the previous point of view, to a different point of view, called the new point of view.
[0036] Preferably, the modification of the scene image to obtain one (single) or two modified images aims to restore one or more images of the scene corresponding to a different point of view of the scene. In other words, preferably, the scene image and the modified image, or the two modified images, correspond to two distinct images of the scene as the observer would visualize them before and after changing point of view.
[0037] Preferably, in the present description, unless otherwise indicated, the term “modified image(s)” is understood to mean the modified image(s) of the scene.
[0038] Preferably, it is understood by "pupillary distance" to mean the average pupillary distance of a user of the head-mounted display.
[0039] Those skilled in the art will know that head-mounted displays are adapted to the morphology of a panel of users so that the display, in particular in binocular display, is suitable for most users. Therefore, by "pupillary distance" is meant the average pupillary distance provided by a head-mounted display.
[0040] Preferably, change of point of view is understood to mean: the change of point of view of the modified image and / or of each of the two modified images in relation to the scene and / or in relation to the imaged scene and / or in relation to the image of the scene.
[0041] Preferably, the change in viewpoint is different from or equal to the pupillary distance.
[0042] Preferably, the change in viewpoint between the imaged scene and the (single) modified image is equal to the pupillary distance. Preferably, the difference in viewpoint is understood to mean: the difference in viewpoint between: the scene and / or the imaged scene and / or the image of the scene and the (single) modified image, one of the two modified images and the other of the modified images.
[0043] Preferably, the difference in viewpoint between the scene image and the (single) modified image or between one of the two modified images and the other of the modified images is equal to the pupillary distance.
[0044] Preferably, to obtain the (single) modified image, the change in viewpoint between the scene image and the (single) modified image is equal to half the pupillary distance and the difference in viewpoint between the scene image and the (single) modified image is equal to the pupillary distance.
[0045] Preferably, to obtain the two modified images, the change in viewpoint of the scene, preferably between or from the image of the scene and each of the two modified images, is arbitrary and the difference in viewpoint, between the two images of the scene, is equal to the pupillary distance.
[0046] Preferably, to obtain the two modified images, the difference in viewpoint change is equal to the pupillary distance. Preferably, the viewpoint change, according to which the scene image is modified, between the scene image and each of the two modified images is arbitrary but a function of the pupillary distance so that the viewpoint difference between the two modified images is equal to the pupillary distance.
[0047] Preferably, the modification of the scene image is carried out to obtain:
[0048] - a modified image associated with or reproducing a change in point of view, in relation to the imaged scene, corresponding to or equal to half the pupillary distance and / or a difference in point of view between the image of the scene and the modified image corresponding to or equal to half the pupillary distance, and / or
[0049] - two modified images associated with or restoring a change of point of view, relative to the imaged scene, and / or so that the difference in change of point of view, between the two modified images, corresponds to or is equal to the pupillary distance. The person skilled in the art knows 3D display, monocular, binocular and stereoscopic vision. Also, it is known that the binocular display is implemented so that each of the two displayed images is seen only by one of the eyes of the observer / user of the head-mounted display.
[0050] Preferably, the modifying step, or the obtaining step and the modifying step, are applied to a stream of images of the scene.
[0051] The method can be implemented from a stream of images of the scene. The step of modifying the image of the scene can be iterated or implemented, preferably independently, for each image of the stream of images of the scene.
[0052] Preferably, at the end of the implementation of the modification step, or of the obtaining step and the modification step, preferably at the end of the implementation of the method, a stream of image pairs is obtained or restored.
[0053] The modification of the scene image may be a function, in addition, of an area of the scene image, called the attention area, when modifying the scene image.
[0054] Preferably, in the case where the modification of the image of the scene is a function of the area of attention, the modification of the image of the scene is carried out or implemented or carried out on at least a part of the image of the scene.
[0055] Preferably, in the case where the modification of the scene image is a function of the attention zone, the modification of the scene image is carried out or implemented or performed on the entire scene image. In other words, preferably, the displacement of a part of the pixels or group(s) of pixels of the scene image and / or the addition and / or deletion of pixels or groups of pixels in the scene image is carried out on the entire scene image.
[0056] It can be understood by attention zone: one or more pixels or one or more groups of pixels or an area of the image corresponding to an object or a group of objects in the scene.
[0057] Preferably, the image is modified, during the modification step, in relation to or relative to, in addition, the area of attention.
[0058] Preferably, the step of displaying, in a virtual reality / augmented reality (VR / AV) headset or in a 3D display device: - the image of the scene and the modified image, or
[0059] - the two images of the scene modified.
[0060] Preferably, in the case of a binocular head-mounted display, the scene image and the modified image or each of the two modified images are displayed so as to be seen by only one of the eyes to allow stereoscopic viewing of the scene.
[0061] Preferably, a VR / VA headset is understood to mean a monocular or binocular head-mounted display.
[0062] In the remainder of this description, unless otherwise indicated, the term “headset” used alone corresponds to the VR / VA headset or video headset according to the invention.
[0063] The display step, in the head-mounted display or in the 3D display device, may include the display of the stream of the pair of images.
[0064] Preferably, the method comprises a step of generating, producing, determining or calculating the depth map of the imaged scene from depth information, preferably from a depth field, of the scene of the image of the scene and / or from sharpness information of the image of the scene.
[0065] Preferably, the step of generating the depth map can be defined as enriching the image of the scene with depth information, preferably with a depth field, associated with the imaged scene such that the displayed image comprises a depth field channel, preferably a single depth field channel.
[0066] By way of non-limiting example, the information contained in a pixel or block of pixels may include an entropy value, energy, a color channel, typically R, G, B (Red, Green, Blue), intensity detected by the photosites, a gain and / or depth information or a depth field.
[0067] Preferably, the information contained in a pixel or block of pixels is a value or information, called sharpness, relating to or representative of sharpness.
[0068] According to the present invention, the depth information or the depth field can be obtained by comparing images, for example by photogrammetry, or by a distance measuring device, for example a device for detecting and estimating distance by light, called "LIDAR" for "light detection and ranging".
[0069] Preferably, according to a first aspect of the invention, the depth information and / or the depth field is determined or calculated and / or the depth map is generated, produced, calculated or determined from at least two images of the scene originating from:
[0070] - from the same stationary imaging system and each being acquired with a different focusing distance, and / or
[0071] - of the same mobile imaging system having acquired, at least in part, the scene in two distinct positions, and / or
[0072] - at least two separate imaging systems, arranged in two separate positions so as to each acquire, at least in part, the scene.
[0073] Preferably, the at least two images of the scene comprise the image of the scene. Preferably, the at least two images of the scene, from which the depth map is generated, comprise the image of the scene.
[0074] In the present description, unless otherwise indicated, the term "at least two images of the scene" used alone corresponds to the at least two images of the scene from which the depth map of the scene is generated and / or from which the depth information of the scene is extracted.
[0075] The step of generating the depth map, according to any of the alternatives, can be defined as consisting of determining and / or associating a distance between each of the points or each group(s) of point(s) of the observed scene and an optical sensor, preferably an optical sensor of an imaging system, for example a camera, having imaged the scene.
[0076] Preferably, according to the first aspect of the invention, the depth information and / or the depth field is determined or calculated and / or the depth map is generated, produced, calculated or determined:
[0077] - by selection and / or identification of pixels or groups of pixels from the at least two images of the scene or by identification and / or selection of an image or image(s) from among the at least two images, - by combination, assembly or association of pixels or groups of pixels from the at least two images of the scene.
[0078] The determination of the depth information or the depth field and / or the depth map may comprise, prior to the selection and / or identification step, a step of comparing image(s) of the scene, among the at least two images of the scene, or pixels or groups of pixels of the at least two images of the scene.
[0079] In this application, generating the depth map may be understood to mean producing, calculating or determining the depth map.
[0080] In the present application, it may be understood by determining the depth information or the depth field: calculating or deducing the depth information.
[0081] Preferably, according to a first alternative of the first aspect of the invention, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by:
[0082] - identification of a local maximum or maxima of sharpness or locals within each of the at least two images, and
[0083] - for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene presenting the maximum local sharpness.
[0084] Preferably, according to a second alternative of the first aspect of the invention, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by:
[0085] - for each point of the depth map to be generated, identification of the image of the scene, among the at least two images of the scene, for which the sharpness is maximum, and
[0086] - for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene for which the sharpness is maximum.
[0087] Preferably, according to a third alternative of the first aspect of the invention, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by:
[0088] - for each point of the depth map to be generated, interpolation or extrapolation, from the respective focusing distances of the at least two images, of an optimal focusing distance for which each point of the depth map to be generated considered would present the maximum local sharpness, and
[0089] - for each point of the depth map to be generated, association of said optimal focusing distance.
[0090] Preferably, according to a fourth alternative of the first aspect of the invention, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by:
[0091] - for each point of the depth map to be generated, identification of the image of the scene, among the at least two images of the scene, for which a local variance is the highest, and
[0092] - for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene for which local variance is the highest.
[0093] Preferably, according to a fifth alternative of the first aspect of the invention, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by:
[0094] - identification of a maximum or maxima of local variance or local variances within the at least two images of the scene, and
[0095] - for each point of the depth map to be generated, association of the local variance corresponding to the image of the scene presenting the highest local variance.
[0096] Preferably, according to a sixth alternative of the first aspect of the invention, the depth information is determined and / or the depth map is generated, from the at least two images of the scene, by:
[0097] - for each point of the depth map to be generated, interpolation or extrapolation, from the respective focusing distances of the at least two images, of an optimal focusing distance for which the depth map would present the highest local variance, and
[0098] - for each point of the depth map to be generated, or association of said optimal focusing distance.
[0099] It can be understood as "a point in the depth map": a pixel or group of pixels in the depth map.
[0100] Preferably, each pixel considered or group of pixels considered correspond(s) to an area of the scene or to an object of the scene. Preferably, the pixel considered or the group of pixels considered corresponding(s) or common(s) to the at least two images of the scene correspond(s) to the same area of the scene or to the same object of the scene.
[0101] The local maximum or maxima may correspond to an area of the scene or to an object in the scene.
[0102] Preferably, according to a second aspect of the invention, the depth information is determined and / or the depth map is generated, from an image of the scene, preferably from the image of the scene, by means of a neural network, preferably an acyclic neural network, more preferably a convolutional neural network.
[0103] Preferably, the third aspect of the invention is an alternative of the first and second aspect of the invention.
[0104] Preferably, the determination of the depth information is carried out and / or the generation of the depth map according to the third aspect of the invention is implemented from a single image of the scene, preferably from the image of the scene. Thus, preferably, a depth map is generated for each image of the scene.
[0105] Preferably, the determination of the depth information is carried out and / or the generation of the depth map according to the third aspect of the invention comprises a step of minimizing an objective function.
[0106] Preferably, the determination of the depth information is carried out and / or the generation of the depth map according to the third aspect of the invention is implemented from the R, G, B values of the image of the scene.
[0107] The characteristics described in this description apply, unless otherwise indicated, to each of the aspects of the invention and to each of the alternatives according to the invention.
[0108] Preferably, the depth information and / or the sharpness information contained in the image of the scene or, respectively, the depth information contained in the depth map is refined or improved by processing the image of the scene, or an image associated with the image of the scene, or, respectively, the depth map. At the end of the processing phase, an image of the scene whose depth and / or sharpness information or, respectively, a depth map whose depth information is refined or improved, is obtained.
[0109] The processing phase comprises an iterative modification of the image of the scene or, respectively, of the depth map so as to minimize a function E comprising: a term D, called a difference term, determined by comparing the image of the scene being iteratively processed, or respectively, the depth map being iteratively processed, convolved by a point spread function (PSF), dependent on the depth map, with the image of the scene before convolution, or respectively with the depth map, before convolution, the PSF describing the response of an imaging system from which the image of the scene, or respectively the depth map, is obtained, and a term A, called anomaly term, representative of defects or anomalies within the image of the scene being iteratively processed, or respectively of the depth map being iteratively processed.
[0110] Preferably, at least a portion of the pixels of the scene image comprises depth information and / or sharpness information.
[0111] The depth map, by definition, includes depth information.
[0112] The image associated with the scene image may be a reconstructed image. At least a portion of the pixels of the image associated with the scene image may include depth and / or sharpness information.
[0113] Preferably, the reconstructed image is obtained from the scene image.
[0114] Preferably, the reconstructed image can be enriched by the depth information, contained in at least one part of the pixels of said reconstructed image, obtained by a dual pixel sensor. Preferably, the reconstructed image can be enriched by the depth information, contained in at least one part of the pixels of said reconstructed image, obtained by comparing at least two images of the scene, preferably acquired simultaneously or successively, more preferably comprising the image of the scene, obtained in two positions or two distinct viewpoints. The at least two images of the scene can be obtained by two distinct imaging systems. The reconstructed image can be obtained by photogrammetry. Preferably, the change of viewpoint comprises a displacement (du, dv) in a frame of reference of the image of the scene or a displacement relative to the image of the scene.
[0115] Preferably, the modification of the scene image is proportional or inversely proportional to the depth associated with the pixel or group of pixels considered.
[0116] Preferably, the method comprises, for each pixel or group(s) of pixels of the image of the scene, a calculation of a difference between a depth associated with a pixel considered or with one or more group(s) of pixels considered of the image of the scene and a depth associated with each other pixel or each other group(s) of pixels of the image of the scene; the modification of the image of the scene is carried out as a function, in addition, of the calculated difference.
[0117] Preferably, the modification of the displayed image is carried out relatively and / or proportionally to the calculated deviation.
[0118] Preferably, the method comprises, for each pixel or group(s) of pixels of the image of the scene, a calculation of a difference between a depth associated with an area of the image of the scene considered, preferably corresponding to an object or a group of objects of the scene, and a depth associated with each pixel or group(s) of pixels of the displayed image located outside the area of the image of the scene considered; the modification of the image of the scene is carried out as a function, in addition, of the calculated difference.
[0119] Preferably, the modification of the scene image is carried out relatively and / or proportionally to the calculated angle.
[0120] Preferably, the method comprises calculating an angle corresponding to the change of viewpoint; the modification of the image of the scene is carried out as a function of the calculated angle.
[0121] Preferably, the modification of the displayed image is carried out according to a linear function, for example proportionally, preferably proportionally, to the calculated angle, and / or according to a non-linear function. The angle corresponding to the change of viewpoint can be defined as the angle formed between the optical axis of the observer before change of viewpoint and the optical axis of the observer of the observer after change of viewpoint.
[0122] Preferably, the change of viewpoint corresponds to the change of viewpoint, relative to the imaged scene, between, or from, one axis of observation of the imaged scene and, or towards, another axis of observation of the imaged scene.
[0123] Preferably, the change of point of view corresponds to or aims to restore a change of point of view, relative to the imaged scene, between, or from, an optical axis of one of the eyes of an observer and, or towards, the optical axis of the other eye of the observer.
[0124] The angle corresponding to the change of point of view can be defined as the angle formed between the observation axis of the image of the scene or of the imaged scene, before change of point of view, and, or towards, another observation axis of the image of the scene or of the imaged scene, after change of point of view.
[0125] The angle corresponding to the change of point of view can be defined as the angle formed between the optical axis of the observer, or the axis of observation before, or respectively after, the change of point of view, and the optical axis of the observer, or the axis of observation after, or respectively before, the change of point of view.
[0126] The viewpoint or observation axis or optical axis of the observer may correspond to the axis or straight line connecting the point located equidistant from each of the observer's eyes, for example equidistant from the fovea or the cornea or the iris of each of the eyes, to an image, for example the image of the scene or the modified scene image(s). The area of attention may be predetermined, selected, predefined or defined.
[0127] Preferably:
[0128] - the change of point of view obtained is a function of, or linked to or corresponds to, a change of position and / or orientation, in space, of the head-mounted display or a change of position and / or orientation, in space, of an observer relative to the 3D display device, and, preferably, preferably to a change of position and / or orientation, in space, of the observer's head, and / or - the attention zone corresponds to an area of the displayed scene on which the observer's gaze is focused.
[0129] The change of viewpoint can be obtained in the absence of a change in position and / or orientation, in space, of the head-mounted display and / or the observer. The change of viewpoint obtained can be predetermined. The change of viewpoint can be chosen by the observer.
[0130] The point of view can correspond to the center of the straight line segment connecting the same part of each of the observer's eyes, for example to the center of the straight line connecting the foveas or the cornea or the iris of each of the observer's eyes.
[0131] The change of the observer's point of view can be defined as the observer's passage from the previous point of view to the new point of view.
[0132] The observer's point of view can be defined as the position and / or orientation of the observer, preferably as the position and / or orientation of the observer's head and / or eyes.
[0133] The attention area may be predetermined, selected, predefined, or defined. Preferably, the attention area may be predetermined, selected, predefined, or defined based on information about the area of the source image or pair of displayed images on which the observer's gaze is focused or intended to be focused.
[0134] Preferably, the area of the image of the scene or of the pair of displayed images on which the observer's gaze is focused or intended to be focused corresponds to the point, the pixel or the group(s) of pixels of the source image or of the pair of displayed images corresponding to the same point, object or group(s) of object(s) of the imaged scene.
[0135] Preferably, the attention zone corresponds to the point, pixel or group of pixels or to the area of the image of the scene or pair of images displayed on which the observer's gaze is focused or intended to be focused.
[0136] Preferably, the attention zone corresponds to a point, for example to a pixel or a group of pixels, of the source image or of each of the images of the pair of displayed images located or intended to be located within the cone of the respective fields of vision of each of the observer's eyes. Preferably, the attention zone corresponds to the point of the image of the scene or to the point of the images of the pair of displayed images of the scene having the shallowest depth within the cone of the observer's field of vision. Preferably, the cone of the observer's field of vision, preferably extending around an axis of revolution of the optical axis of an eye of the observer, formed by a solid angle of between 3 and 5° relative to the optical axis of an eye of the observer.
[0137] Preferably:
[0138] - the change of viewpoint is obtained from position data of the head-mounted display and / or movement data of the head-mounted display, called head-mounted display data, and / or
[0139] - the change of viewpoint is obtained from images of the observer and / or the observer's environment and / or the head-mounted display or the 3D display device, and / or
[0140] - the attention area is obtained, measured, detected, determined or calculated from images of the observer's eyes.
[0141] Preferably, the at least one image of the observer's eyes comes from an imaging system capable of imaging the observer and / or the observer's eyes.
[0142] Determining the observer's point of view and / or area of attention can be done by eye tracking, also known as eye tracking.
[0143] The data from the head-mounted display or 3D display device may be measured or detected by the head-mounted display or 3D display device. The change in viewpoint may be determined or calculated from the data from the head-mounted display or 3D display device.
[0144] The images of the viewer and / or the head-mounted display or 3D display device may be acquired by an external imaging system, preferably not belonging to or being part of the head-mounted display or 3D display device. The images of the viewer's environment may be acquired by the external imaging system and / or by an imaging system of the head-mounted display or 3D display device.
[0145] The external imaging system may be arranged to image the viewer and / or the head-mounted display or 3D display device.
[0146] Preferably, the method comprises linear or non-linear filtering of the head-mounted display data to limit an amplitude of the obtained viewpoint change, as a function of time, to a value less than a limit value. Preferably, the method comprises non-linear clipping of the head-mounted display data to limit an amplitude of the obtained viewpoint change to a value less than a limit value.
[0147] Preferably, the method comprises non-linear dynamic compression of the head-mounted display data to attenuate an amplitude variation of the obtained viewpoint change, as a function of time, such that said amplitude variation does not exceed a variation amplitude threshold value.
[0148] The filtering and / or clipping and / or dynamic compression steps have the effect of limiting and / or reducing the amplitude and / or the variation in amplitude, as a function of time, of the change in point of view.
[0149] The change in position and / or orientation, in space, of the head-mounted display and / or the observer can be defined or qualified as a real or effective change of point of view.
[0150] The effective viewpoint change may be defined as the viewpoint change measured or detected by the head-mounted display or 3D display device, or determined or calculated from data from the head-mounted display or 3D display device.
[0151] The resulting viewpoint change may be equal to, correspond to, or be proportional to the actual viewpoint change. This is particularly the case, for example, when the amplitude and / or variation in amplitude of the viewpoint change is less than the limit value(s) or threshold value.
[0152] The resulting viewpoint change may differ from the actual viewpoint change. This is particularly the case, for example, when the amplitude and / or variation in amplitude of the viewpoint change is greater than the limit value(s) or threshold value.
[0153] Preferably, the method comprises a step of acquiring the image of the scene and / or the at least two images of the scene.
[0154] The method may comprise a step of acquiring the image of the scene and / or at least two images of the scene originating from the same stationary or mobile imaging system and / or the mobile imaging system or from the two separate imaging systems, preferably arranged in two positions or two separate viewpoints.
[0155] According to another aspect of the invention, a device is provided comprising configured means arranged and / or programmed and / or configured to implement all the steps of the method according to the invention.
[0156] According to the invention, a data processing device is also proposed comprising means arranged and / or programmed and / or configured to implement the method according to the invention.
[0157] According to the invention, there is also provided a computer program comprising instructions which, when the program is executed by a computer, cause the latter to implement the method according to the invention.
[0158] According to the invention, there is also provided a medium, for example a recording medium, readable by a computer comprising instructions which, when executed by a computer, cause the latter to implement the method according to the invention.
[0159] According to the invention, a computer-readable data carrier is also provided, on which the computer program according to the invention is recorded.
[0160] According to the invention, a virtual reality / augmented reality (VR / AV) head-mounted display device is proposed comprising means for displaying images in monocular or binocular vision.
[0161] Preferably, the VR / VA headset includes:
[0162] - means arranged and / or programmed and / or configured to implement the method according to the invention, and / or
[0163] - communication means arranged and / or programmed and / or configured for and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the method according to the invention. Preferably, the VR / VA head-mounted display comprises at least one imaging system arranged to image the eyes of an observer wearing said VR / VA head-mounted display.
[0164] Preferably, the VR / VA head-mounted display includes:
[0165] - at least one imaging system arranged to image at least part of an environment of said VR / VA head-mounted display and / or an environment of the observer, and / or
[0166] - means arranged to detect a movement and / or a relative position of said VR / VA head-mounted display.
[0167] According to the invention, a 3D display device is proposed, arranged to display 3D images, comprising:
[0168] - means arranged and / or programmed and / or configured to implement the method according to the invention, and / or
[0169] - communication means arranged and / or programmed and / or configured for and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the method according to the invention.
[0170] By way of non-limiting examples, the 3D display device may be a screen, for example LCD or plasma, or a video projector.
[0171] Preferably, the 3D display device comprises at least one imaging system arranged to image the eyes of an observer.
[0172] Preferably, the 3D display device comprises:
[0173] - at least one imaging system arranged to image at least part of an environment of said 3D display device and / or of the observer and / or of an environment of the observer, and / or
[0174] - means arranged to detect a movement and / or a position of an observer relative to said 3D display device. Preferably, the head-mounted display and the 3D display device according to the invention are arranged to provide or enable or provide a 3D rendering from the pair of stereoscopic 2D images of the displayed scene. Where appropriate, the head-mounted display may comprise, or the 3D display device may be combined with or used with:
[0175] - synchronous shutter means, active (or synchronous) glasses in the case of the 3D display device,
[0176] - polarizing means or polarizing glasses.
[0177] Preferably, according to the invention, stereoscopic rendering is understood to mean any display method or process making it possible to provide a 3D rendering from a pair of displayed 2D images.
[0178] A person skilled in the art will be able to envisage the different display modes capable of providing a 3D rendering from a pair of 2D images to be displayed.
[0179] Preferably, both images of the pair of displayed images may be displayed simultaneously in the binocular head-mounted display, with one displayed image being viewed by the right eye and one displayed image being viewed by the left eye.
[0180] Preferably, the two images of the pair of displayed images may be displayed simultaneously, by multiplexing or spatial separation, in the monocular head-mounted display or the 3D display device, for example in the form of a single composite or raster image comprising a portion of each of the two images of the pair of images. In this case, only a portion of each of the two displayed images is displayed in a single image.
[0181] Preferably, the two images of the pair of displayed images can be displayed simultaneously in the monocular head-mounted display or the 3D display device by superimposing the two images of the pair of images, each image of the pair of displayed images having a distinct polarization.
[0182] Preferably, the two displayed scene images may be displayed consecutively or alternately or iteratively, for example by time multiplexing, in a synchronous monocular head-mounted display or on a 3D display device, preferably matched by the user wearing synchronous glasses.
[0183] Preferably, the head-mounted display is a binocular head-mounted display in which each image of the scene is displayed so as to be seen by only one of the eyes so as to allow stereoscopic viewing of the scene.
[0184] Preferably, the head-mounted display is worn by the user when implementing the method. Preferably, the method according to the invention is suitable, more preferably is particularly suitable, more preferably is designed and particularly advantageously is specially designed, to be implemented by the VR / VA headset and by the 3D display device according to the invention. Also, any characteristic of the method according to the invention is directly transposable to the VR / VA headset and to the 3D display device according to the invention and vice versa.
[0185] Description of figures
[0186] Other advantages and particularities of the invention will appear on reading the detailed description of implementations and embodiments which are in no way limiting, and the following appended drawings:
[0187] - FIGURE 1 represents a flowchart of an embodiment of the method according to the invention,
[0188] - FIGURE 2A is a schematic representation illustrating the image of a scene and FIGURE 2B is a schematic representation of the image of the scene obtained by modifying the image of the scene illustrated in FIGURE 1A according to the method according to the invention,
[0189] - FIGURE 3A, identical to FIGURE 2A, is a schematic representation illustrating an image of the scene and FIGURES 3B and 3C are schematic representations of two images obtained by modifying the image of the scene of FIGURE 2A according to the method according to the invention,
[0190] - in FIGURE 4 is illustrated a schematic representation of a top view of the images displayed in a binocular head-mounted display,
[0191] - in FIGURE 5 is illustrated a schematic representation of a top view of the images displayed in a 3D display device or monocular head-mounted display,
[0192] - FIGURE 6 is a schematic representation of a depth map of the scene illustrated in FIGURES 2A and 3A,
[0193] - FIGURE 7 is a schematic representation of an object projected through a lens onto an optical sensor,
[0194] - in FIGURES 8 and 9 is illustrated a schematic representation of a perspective top view of the scene displayed in FIGURES 2A and 3A.
[0195] Description of the embodiments
[0196] The embodiments described below being in no way limiting, it will be possible in particular to consider variants of the invention comprising only a selection of the described characteristics, isolated from the other described characteristics (even if this selection is isolated within a sentence comprising these other characteristics), if this selection of characteristics is sufficient to confer a technical advantage or to differentiate the invention compared to the state of the prior art. This selection comprises at least one characteristic, preferably functional without structural details, or with only a part of the structural details if this part only is sufficient to confer a technical advantage or to differentiate the invention compared to the state of the prior art.
[0197] With reference to FIGURES 1 to 9, an embodiment of the method for modifying an image of a scene 20 is presented, referred to as the method in the remainder of the description. The method aims to modify an image of a scene 20 to restore a pair of stereoscopic 2D images 20, 21 or 22, 23 providing or intended to provide a 3D rendering. The method restores a single image 21 or two images 22, 23 of the scene. For each pair of stereoscopic 2D images 20, 21 or 22, 23 restored considered, the two images of a pair of images considered correspond to the observation of the scene from a different point of view separated by the pupillary distance. In other words, by pair of 2D stereoscopic images 20, 21 or 22, 23 restored or obtained, we mean two images of the scene corresponding to the observation of the scene from a different point of view separated by the pupillary distance.
[0198] In the case of the pair of images 20, 21, the change in viewpoint between the two images of the pair of images 20, 21 corresponds to or reproduces the scene as it would be observed from two distinct viewpoints spaced apart by the pupillary distance. In the case of the pair of images 22, 23, the change in viewpoint between each of the two images of the pair of images 22, 23 and the image of the scene 20 corresponds to or reproduces the scene as it would be observed from two distinct viewpoints spaced apart by a different distance from the pupillary distance (the difference in viewpoint between the two images of the pair of images 22, 23 corresponds to or reproduces the scene as it would be observed from two distinct viewpoints spaced apart by the pupillary distance). A simplified schematic representation of an imaged scene 20 is illustrated in FIGURES 1A and 2A.
[0199] The scene image may be a stored or available image, such as a stack of images or a video. As detailed elsewhere, the scene image may originate from or be transmitted by external means.
[0200] The method comprises the step of modifying the image of the scene 20. According to the non-limiting embodiment, the modification of the image is carried out by moving a portion of the pixels or group(s) of pixels of the image of the scene 20. In particular, according to a non-limiting embodiment, the method of moving at least a portion of the pixels or group(s) of pixels of the image of the scene 20 may consist of a spreading of certain pixels or groups of pixels, preferably a spreading of the pixels not occluded by objects of the scene after a change of viewpoint. More preferably, the displacement of a portion of the pixels or group(s) of pixels of the image of the scene 20 comprises a spreading of certain pixels or groups of pixels onto pixels occluded by objects of the scene after a change of viewpoint.
[0201] The modification step is implemented from a single scene 20.
[0202] FIGURE 1 shows, in dotted lines, data that can be transmitted or provided, preferably in real time, by:
[0203] - external means, for example storage means, data processing means or image acquisition means, or by means of the video headset or the 3D display device according to the invention.
[0204] As non-limiting examples, the images of the scene, in particular the pair of source images or the at least two images of the scene from which the depth map of the scene is generated and / or from which the depth information of the scene is extracted, the depth information and / or the depth map, may be transmitted or provided.
[0205] In FIGURES 2A and 3A, the solid black arrows illustrate the direction of movement, relative to the image of the scene 20, of the change of viewpoint and the dotted arrows illustrate the direction of movement and / or spreading of at least a portion of the pixels or group(s) of pixels of the image of the scene 20. Advantageously, the method is implemented for a stream of images of the scene. Also, the method can be implemented from a stream of images of the scene. The step of modifying the image of the scene can be iterated or implemented, preferably independently, for each image of the stream of images of the scene.
[0206] The modification of the image of the scene 20 is a function of: a depth map of the image of the scene 20, the change of point of view of the scene, or in relation to the scene, or to the imaged scene 20, which is a function of the pupillary distance.
[0207] With reference to FIGURES 2A, 2B, 3A, 3B and 3C, the change of point of view and the modification of the image of the scene 20 is illustrated. The change of point of view of the image of the scene 20 as a function of a pupillary distance makes it possible to obtain a single modified image 21 or two modified ones 22, 23 restoring a change of point of view with respect to the imaged scene. The change of point of view is a function of the pupillary distance.
[0208] The method therefore makes it possible, from a single image of a scene 20, to have a pair of stereoscopic 2D images: the image of the scene 20 and the image of the modified scene 21 or the two modified images of the scene 22, 23, the pair of images 20, 21 constituting a particular case, restoring a different change of viewpoint of the imaged scene. However, a difference in viewpoint between the two images of each pair of images is identical and is equal to the pupillary distance.
[0209] The person skilled in the art knows the pupillary distance. It is possible to use an average pupillary distance corresponding, for example, to the average pupillary distance of the users for whom the VA / VR 12 head-mounted displays are sized. As a non-limiting example, it can be considered that the pupillary distance is between 45 and 75 mm.
[0210] According to a preferred embodiment: the difference in point of view is equal to half the pupillary distance to obtain a single modified image 21 of the imaged scene, the change in point of view is arbitrary, to obtain two modified images 22, 23 of the imaged scene.
[0211] After implementing the method, one or more pairs of images 20, 21 and 22, 23 restoring a difference in point of view equal to the pupillary distance is available.
[0212] Also, with reference to FIGURES 2A, 2B, 3A, 3B and 3C, a schematic representation of pairs of stereoscopic 2D images obtained according to the invention intended to be displayed in a head-mounted display or a 3D display device is illustrated. Those skilled in the art will observe that the two images 20, 21 and 22, 23 correspond to two different points of view of the scene so that the observer is able to observe a perspective or relief view of the scene. It will be considered in the present description that the observer of the scene imaged in the head-mounted display wears the head-mounted display or observes the scene in a 3D display device, where appropriate with suitable observation means (synchronous glasses, polarizing glasses or glasses equipped with red / blue filters).
[0213] The method also comprises the step of obtaining a change of point of view when observing the pairs of images 20, 21 and 22, 23 of the displayed scene. The change of point will preferably be a function of a change in position and / or orientation, in space, of the head-mounted display, and therefore of a change in position and / or orientation of the observer, and in particular of the observer's head, in space.
[0214] In a particular case (not shown), the method can only return a 3D rendering to the observer without returning a change in viewpoint. This is the case in which the source image is modified so that each of the two modified images provides a change in viewpoint equal to half the pupillary distance (and ultimately a difference in viewpoint between them equal to the pupillary distance).
[0215] Furthermore, it is not excluded according to the invention that the change of point of view or the area of attention is predetermined. The change of point of view, for example, corresponds to a model or a sequence of given change(s) which will be implemented in the step of obtaining change of point of view to propose or offer or submit to the observer a change of point of view without the head-mounted display or the observer changing position and / or orientation. The change of point of view or the area of attention can be informed or provided, by a third party or preferably by the observer himself.By way of non-limiting examples, the change of point of view or the area of attention can be indicated by means of a touch zone (a smartphone or tablet screen for example), a movement of the observer's hand or limb, preferably detected by the head-mounted display or by the 3D display device, or any other means that a person skilled in the art will be able to envisage.
[0216] Alternatively, and by way of example, the observer may look at the image from a first position constituting the observer's point of view and may provide a second position different from his own, with the aim of obtaining the image as he would see it from the point of view associated with this second position.
[0217] The change of point of view can be defined as an angle formed between the observation axis 6 of the scene or the imaged scene 20 before change of point of view and the observation axis 7 of the scene or the imaged scene after change of point of view.
[0218] Therefore, the step of modifying the image of the scene 20 can be implemented according to this angle.
[0219] According to a non-limiting embodiment, the method may comprise a calculation of the angle corresponding to the change of point of view.
[0220] In FIGURES 2A and 3A illustrating the image of the scene 20, an imaged scene 20 is shown observed from a point of view given by a user. The scene is therefore represented as seen by the observer along this observation axis. According to the non-limiting embodiment, FIGURE 2B illustrates the modified image 21 according to the method as the observer would view it after a change of point of view equal to half the pupillary distance. According to the non-limiting embodiment illustrated in FIGURE 2B, the modified image 21 was obtained by a change of point of view relative to the reference frame of the image of the scene 20 illustrated in FIGURE 2A. The change of point of view consists of a displacement (of) to the left in the horizontal plane parallel to the imaged scene 20 illustrated in FIGURE 2A.In the case of obtaining a single modified image 21, the change in viewpoint, relative to the imaged scene, is equal to half the difference in viewpoint, between the image of the scene 20 and the modified image 21. In this case, the change in viewpoint and the difference in viewpoint are equal to the pupillary distance.
[0221] The change of point of view may comprise a displacement (du, dv) in a reference frame or relative to a reference frame of the image of the scene 20. The reference frame of the displayed image 20 may be a plane or a surface, for example curved. The reference frame may be that or may be associated with the user or may be that or may be associated with the VA / VR head-mounted display 12 or the 3D display device 13 or may be a reference frame other than or external to the user and / or the VA / VR head-mounted display 12 or the 3D display device 13.
[0222] FIGURE 3B illustrates one of the two modified images 21 and FIGURE 3C illustrates the other of the two modified images.
[0223] According to the non-limiting embodiment illustrated in FIGURE 3B, the modified image 22 corresponds to a change of point of view relative to the frame of reference of the image of the scene 20 illustrated in FIGURE 3A. The change of point of view consists of a displacement (of) to the right in the horizontal plane parallel to the imaged scene 20.
[0224] According to the non-limiting embodiment illustrated in FIGURE 3C, the modified image 23 corresponds to a change of point of view relative to the frame of reference of the image of the scene 20 illustrated in FIGURE 3A. The change of point of view consists of a displacement (of) to the left in the horizontal plane parallel to the imaged scene 20.
[0225] In the case of obtaining two modified images 22, 23, the change in point of view, relative to the imaged scene, is different from the difference in point of view, between the image of the scene 20 and each of the two modified images 22, 23. In this case, the difference in point of view, between the two modified images is equal to the pupillary distance. However, in this case, the change in point of view is different from the pupillary distance and the difference in point of view.
[0226] In the case of obtaining two modified images 22, 23, the change of point of view can be equal to any distance. Preferably, the change of point of view will be less than or equal to a maximum distance making it possible to avoid the objects, or the pixels or group(s) of pixels corresponding to the imaged objects, of the two modified images 22, 23 being too strongly deformed and / or spread out which could make the two modified images 22, 23 blurred or unobservable. A person skilled in the art will know how to adapt and fix, depending on the desired application and / or the specific case, a maximum value of change of point of view.
[0227] As a non-limiting example, the change in viewpoint may be greater than or equal to a minimum displacement value. Indeed, small changes in amplitude such as small physiological movements of the head may be filtered so as not to induce a seasickness effect in the user or observer. A person skilled in the art will be able to define this minimum value according to the specific case. The change in viewpoint may be less than or equal to a maximum distance. Equivalently or alternatively, a maximum angle and a minimum angle may be defined to limit the amplitude of the change in viewpoint used for implementing the method.
[0228] All pixels in the scene 20 image can be modified during the modification step.
[0229] According to a non-limiting embodiment, each pixel or each group of pixels will be moved inversely proportionally to the depth of the point, object or group of objects in the imaged scene 20.
[0230] According to the non-limiting embodiment, the modification of the image of the scene 20 is inversely proportional to the depth associated with the pixel or group of pixels considered. Thus, the more an object of the image of the scene 20 is located in the background, in other words the pixel or group of pixels corresponding to the object of the scene will have a high depth value (in the depth map), and the less the displacement or spreading of the pixel or group of pixels will be significant.
[0231] It might be preferable to modify only a single part of the pixels of the image of the scene 20 in order to save the resources necessary for processing the step of modifying the image of the scene 20. In practice, according to a non-limiting embodiment, the pixels or group(s) of pixels corresponding to the objects in the background 10 of the scene may not be moved or spread during the step of modifying the image of the scene 20.
[0232] According to this non-limiting embodiment, each pixel or each group of pixels will also be moved proportionally to the angle a corresponding to the change of viewpoint. In some cases, the displacement of pixel(s) or each group of pixels may be weighted according to sharpness or any other parameter known to those skilled in the art, such as, by way of non-limiting examples, the parameters R, G, B (for R (Red), VI (Green, or G (Green)), V2 (Green, or G (Green)), B (Blue)), a dilation parameter and / or an outline parameter.
[0233] Advantageously, the method comprises, for each pixel or group(s) of pixels of the image of the scene 20, a calculation of a difference between a depth associated with a pixel or group(s) of pixels considered in the image of the scene 20 and a depth associated with each other pixel or each other group(s) of pixels in the image of the scene 20. According to the non-limiting embodiment, the modification of the image of the scene 20 is carried out as a function, in addition, of the calculated difference.
[0234] The method may comprise a step of generating the depth map of the imaged scene 20 from depth information of the image of the scene 20 and / or from sharpness information of the image of the scene 20.
[0235] A person skilled in the art will be able to choose the most appropriate sharpness information depending on the situation. The sharpness information contained in a pixel or block of pixels may include the entropy, energy and / or variance of an image parameter between neighboring pixels.
[0236] According to a first aspect of the invention, the depth information and / or the depth map is determined from at least two images of the scene coming from:
[0237] - from the same stationary imaging system and each being acquired with a different focusing distance, or
[0238] - of the same mobile imaging system or of at least two separate imaging systems arranged, in two separate positions, so as to each acquire, at least in part, the scene.
[0239] Preferably, the depth information and / or the depth map is generated by selecting and / or identifying pixels or groups of pixels from the at least two images of the scene or by identifying and / or selecting one or more images from the at least two images. Alternatively, or in combination, the depth information and / or the depth map is generated by combining, assembling or associating pixels or groups of pixels from the at least two images of the scene from which the depth information and / or the depth map is determined.
[0240] Thus, according to one embodiment, the depth map according to the first aspect of the invention is obtained by aggregating pixels or groups of pixels from different images of the scene. Prior to their aggregation, the group or groups of pixels of the depth map to be generated are selected from the at least two images.
[0241] According to the first aspect of the invention, the depth map of the scene can be generated from:
[0242] - a maximum or maxima of local sharpness or locals identified within, or interpolated or extrapolated from, at least two images of the scene, or
[0243] - of a local variance or of a maximum or maxima of local variance or locals within, or interpolated or extrapolated from, the at least two images of the scene.
[0244] Preferably, according to the invention, any energy, entropy estimator, conventionally used in focus detection, can be used to determine the variance. Those skilled in the art know a set of techniques for calculating the local variance. As a non-limiting example, the local variance can be the variance of the intensity of the pixels, for example the averaged intensity of the R, G, B channels, calculated on a given set of pixel matrices, to be adjusted according to the resolution of the image and / or the processing power / capacity of the processing unit.
[0245] In this description, the term "imaging system" means a camera. The camera comprises an optical lens and an optical sensor.
[0246] The processing unit may be arranged and / or programmed and / or configured to determine and / or calculate the change of viewpoint and / or the attention zone 8 based on data from the at least one imaging system arranged to image the eyes of an observer. Advantageously, the at least two images of the scene may come from the stream of images of the scene.
[0247] In the case where the at least two images of the scene come from the same imaging system, the imaging system is stationary for each shot. The at least two images are each acquired with a different focusing distance. A shooting technique known in the state of the art is called "focus bracketing". It consists of acquiring a set of images of the same scene from the same stationary optical system, each image is acquired with a different focusing distance. The image acquired with the shortest focusing distance, noted Ido, will give the image with the largest viewing angle while the image acquired with the highest focusing distance, noted Idmax, will give the smallest viewing angle. Thus, some objects in the scene will not be on Idmax either because they are outside the field of the image or because they are obscured by objects in the scene.
[0248] In the case of a mobile imaging system or two separate imaging systems arranged in two separate positions, a depth map can be determined from at least two images of the scene. Preferably, the at least two separate imaging systems are stationary relative to each other. Furthermore, the objects in the scene imaged by the imaging systems are at different distances from one or more optical sensors (of the imaging systems) and therefore at a different depth.
[0249] In the case where the two separate imaging systems are arranged in two separate positions, a disparity map of the imaged scene 21, 22 can be determined. The depth map can be calculated and / or determined from the disparity map. Thus, any characteristic relating to the depth map can be transposed to the disparity map, and vice versa. The disparity map can be defined as a map or an image containing depth information. In other words, the disparity map can be an image or a map comprising at least one depth channel of the imaged scene. Preferably, the depth information constitutes a depth field of the imaged scene.
[0250] Also, a group or block of pixels corresponding to one of the objects of the imaged scene 20 will have depth information different from one or more other groups or blocks of pixels corresponding to the other or other objects of the imaged scene. Consequently, each of the at least two images acquired by the two separate imaging systems or by the mobile imaging system may be enriched by a depth field, that is to say that each pixel or group or block of pixels (of the matrix or table of pixels representing the image of the scene) corresponding to an object of the scene will have distinct depth information which corresponds to the distance between the object and the optical sensor; the depth of the different objects of the imaged scene 20 being, in most cases, different for at least some of the objects of the scene.
[0251] In the remainder of this description, the term “at least two images of the scene” used alone designates the at least two images from which the depth information and / or the depth map is determined.
[0252] Preferably, but not necessarily, the at least two scene images include scene image 20.
[0253] According to a second aspect of the invention, the step of generating the depth map is implemented from the image of the scene 20, by means of a convolutional neural network. Those skilled in the art will know how to choose and adapt the most suitable neural network. By way of non-limiting example, the convolutional neural network may comprise several layers of neurons. Among these layers of neurons, the network comprises processing layers comprising neurons arranged to process the displayed image with a convolution. The neural network may also comprise sub-sampling (or pooling) layers, interposed between two processing layers, comprising neurons arranged to combine or merge the output data of the processing layers and reduce their size.By way of non-limiting example, the convolutional neural network may further comprise a correction layer, interposed between two processing layers, comprising neurons arranged to operate an activation function on the output data of the processing layers.
[0254] According to an advantageous but non-limiting embodiment of the invention, the depth information and / or the sharpness information contained in the image of the scene 20 or in one of the at least two images of the scene (from which the depth map is generated) or, respectively, the depth information contained in the depth map is refined or improved by processing the image of the scene 20 or in one of the at least two images of the scene or, respectively, the depth map.
[0255] At the end of the processing phase, an image of the scene 20 whose depth and / or sharpness information or, respectively, a depth map whose depth information is refined or improved is obtained.
[0256] The processing phase comprises an iterative modification of the image of the scene 20 or, respectively, of the depth map so as to minimize a function E comprising:
[0257] - a term D, called difference term, determined by comparison of the image of the scene being iteratively processed, or respectively of the depth map being iteratively processed, convolved by a point spread function (PSF), dependent on the depth map, with the image of the scene 20 before convolution, or respectively with the depth map, before convolution, and
[0258] - a term A, called anomaly term, representative of defects or anomalies within the image of the scene 20 being iteratively processed, or respectively of the depth map being iteratively processed.
[0259] The PSF describes the response of an imaging system from which the scene image 20 or the at least two scene images (from which the depth map is generated) are obtained.
[0260] In other words, the processing phase constitutes an iterative loop in which, the number of iterations is noted k, the image of scene 20 constitutes the image processed during the first iteration at k=1. The processed image is then modified at each iteration of the loop.
[0261] As a non-limiting example, the term D may comprise the sum of several terms Di. The distances Di are obtained by comparing the scene 20 being processed at iteration k convolved by the PSFs with the image of the scene at iteration k-1 or convolved with the image of the scene 20 at iteration k=0. The distances Di may be, among other things, obtained by an evaluation of the local colors, or by evaluation of another parameter contained in the pixels or groups of pixels, of the image of the scene at iteration k opposite the discretization grid of the PSFs, so as to know the calculated colors, or to know another parameter contained in the pixels or groups of pixels, at the positions of the photosites of the image of the scene at iteration k to make the distance comparisons at the appropriate locations.As a non-limiting example, the term A, called penalty, representative of defects or anomalies within the image being processed at iteration k, determined from the image of the scene being processed at iteration k-1, can be calculated concomitantly with the term D.
[0262] According to the non-limiting embodiment, the term A comprises:
[0263] - at least one component Al whose effect is minimized for small intensity differences between neighboring pixels of the image of the scene being processed at iteration k, and / or
[0264] - at least one component A2 whose effect is minimized for small differences in hue between neighboring pixels of the image of the scene being processed at iteration k, and / or
[0265] - at least one component A3 whose effect is minimized for low frequencies of changes of direction between neighboring pixels of the image of the scene being processed at iteration k drawing an outline of an object of the scene.
[0266] The second term A can thus include a sum of several terms Ai.
[0267] Preferably, prior to the processing phase, the method comprises a convolution of the scene image by an inverse function of the PSF, called iPSF. The convolution of the scene image by the iPSF can be carried out as a first step of the processing phase.
[0268] The iterative modification of the image of the scene 20 ends when the function E, or a combination of partial derivatives of the function E with respect to the image, with respect to the image of the scene being processed at iteration k or with respect to the image of the scene, is less than a minimization threshold, or when a certain number of iterations of the iterative modification of the image of the scene is reached; preferably, the at least one image of the scene being processed thus modified is restored.
[0269] As a non-limiting example, and with reference to FIGURE 7, it is possible to establish a discrete relationship between each image among the at least two images of the scene and the focusing distance, denoted f, associated with each of the images. According to the geometric optical relationship + = j [Math 1], it is possible to relate the focusing distance of the lens of the optical system, that is, the physical distance between the object in the scene, corresponding to the pixel or group of pixels appearing as sharp in the selected image, and the optical sensor of the optical system of the selected image. Considering, for example, a set of n images each acquired with a distinct focusing distance di between dimin, for example 20 cm, and dimax, corresponding, for example, to an infinite distance or 50 meters. The number of images n is an integer between one and the number of at least two images.Considering also that there is a linear relationship between the focusing distances of each of the images, the physical distance or depth, denoted P, at which the point of the scene or the object of the scene is located, corresponding to the pixel or group of pixels whose local variance is the highest of the n images, can be expressed as: Pi = [Math 2], where i is an integer that is equal to n-1. In other words, i is. between 0 and n-1. For the distance di = di pas, n is equal to 1 and i is equal to 0. Advantageously, an exponential variation between each of the focusing distances of each of the images can be used. In this case, and by way of non-limiting example, the focusing distance can be noted as being equal to di = dimin*exp(k*i / (nl)) with n greater than or equal to 2 and dimin is the minimum focusing distance and k is a real number such that k = log(dimax / dimin) where dimax is the maximum focusing distance.
[0270] According to the embodiment, the depth map corresponds to an image comprising a single depth channel. By depth, it is understood the physical distance between a point or a physical object of the scene, corresponding to a pixel or a group of pixels of a considered image, and the sensor from which the considered image was acquired. By way of illustration, a simplified schematic depth map associated with the image of the scene 20 illustrated in FIGURES 2A and 3A is presented in FIGURE 6. On this depth map, the white part 4 corresponds to an object located, for example, five meters from the optical sensor of the camera and the gray zone 5 of the background of the image corresponds to the objects located beyond, for example, fifty meters from the optical sensor.The distance of five meters corresponds, for example, to the minimum focusing distance of the camera lens and the distance of fifty meters corresponds, for example, to the maximum focusing distance of the camera lens.
[0271] According to an advantageous embodiment, the method comprises the step of displaying, in a virtual reality / augmented reality (VR / AV) headset 12 or in a 3D display device 13: - the image of the scene 20 and the modified image 21, or
[0272] - the two modified scene images 22, 23.
[0273] It can be understood by observer: the user of the head-mounted display 12, preferably the user wearing the head-mounted display 12, or the user or observer of the 3D display device 13.
[0274] According to this embodiment, the method also comprises the display, in real time, of the pair of images 20, 21 or 22, 23. The pair of images 20, 21 or 22, 23, resulting from the iterative implementation of the method on the image of the scene 20 coming from the stream of images of the scene, can be displayed successively in the head-mounted display 12 or in the 3D display device and substituted for the previous pair of images previously displayed in the headset 12. Thus, the 3D rendering, by modification of the image of the scene 20 coming from the stream of images of the scene, is restored in real time to the observer. The method therefore provides the observer, from a stream of a single 2D image of a scene 20, with an immersive observation rendering with relief of a scene as if he were directly observing the scene.
[0275] In an advantageous embodiment making it possible to restore the most realistic rendering possible, all of the pixels or groups of pixels of the image of the scene 20 are moved and / or all of the pixels or groups of pixels of the image of the scene 20 are spread, during the modification step, as a function of the depth information that the pixel considered or the group of pixels considered contains and of the change of point of view.
[0276] According to this non-limiting embodiment, each pixel or each group of pixels will be moved inversely proportionally to the depth of the point, object or group of objects in the imaged scene 20. According to the invention, proportionally is understood to mean: proportionally or inversely proportionally. Furthermore, preferably, each pixel or each group of pixels will be moved according to a reference depth. The reference depth is advantageously associated with one or more objects in the scene or a group of objects. In practice, each pixel or each group of pixels will be moved relative to this reference depth.
[0277] According to the implementation illustrated in FIGURE 8 of this embodiment, the background 10 of the scene (or for example objects with which the highest depth of the depth map is associated) has been chosen as the reference depth. In this case, all of the objects located between the background 10 and the observer will be displaced inversely proportional to their depth (the objects in the foreground (character 9 of FIGURE 8 for example) will be displaced more than those at mid-depth (tree 8 of FIGURE 8 for example) which will be displaced more than the objects in the background 10. Also, according to the embodiment illustrated in FIGURE 8, the closer the pixel considered or a group of pixels considered is to the imaging system used to image the scene, the greater the displacement and / or the spread will be (relative to the pixels and group of pixels having a higher depth).
[0278] According to this non-limiting embodiment, each pixel or each group of pixels will also be moved proportionally to the angle a corresponding to the change of point of view.
[0279] The modification of the source image pair can be inversely proportional to the depth associated with the pixel or group of pixels considered.
[0280] Still according to the advantageous embodiment mentioned above, and with reference to FIGURE 9, the tree 8 at mid-depth of the scene has been chosen as the reference depth. The change of viewpoint is therefore carried out relative to the depth PI associated with the tree 8 in the scene. As for the embodiment described in FIGURE 9, the person skilled in the art will be able to use simple mathematical tools allowing the calculation of the displacement (spreading and / or offset) of each pixel or group of pixels of the image of the scene 20 (such as trigonometry and simple geometry (Thales' theorem, etc.)). In the case illustrated in FIGURE 9, the pixels or groups of pixels will be moved proportionally to the depth difference separating a pixel to be modified considered and the reference depth associated with the tree 8.In practice, character 9, having a depth difference with tree 8 (the reference depth) smaller than the difference between background 10 and the reference depth, will be less offset than background 10.
[0281] In certain particular cases, the displacement of a part of the pixels or group(s) of pixels of the image of the scene 20 and / or the addition and / or deletion of pixels or groups of pixels in the image of the scene 20 is carried out on the entire image of the scene, with the exception of the attention zone 8. With reference to FIGURE 8, according to a possible but in no way limiting embodiment, the modification of the image of the scene 20 is also a function of a zone 8 of the displayed image 20, called the attention zone 8, during the modification of the displayed image 20. The optical axis 6 of the observer before changing the point of view and the optical axis 7 of the observer after changing the point of view are presented.
[0282] The change in viewpoint of the scene illustrated by the difference between the observer's position before the change in viewpoint and the observer's position after the change in viewpoint is also shown in FIGURE 8.By way of non-limiting example, the change of point of view may, among other things, be defined from or correspond to: the variation between the position of the point located midway between the two pupils of the user before changing point of view and the position of the point located midway between the two pupils of the user after changing point of view (which may be related to the displacement (du, dv, dz) of the point located midway between the two pupils of the user before changing point of view relative to one of the reference frames as described in the present description), and / or the angle a formed between the observation axis 6 before changing point of view and the observation axis 7 after changing point of view (this angle may be zero, for example, but not only, for a displacement along the dz axis).
[0283] The attention zone 8 can be predetermined, selected, predefined or defined. This embodiment can be interesting for proposing or offering or submitting or forcing the observer to focus his gaze on the attention zone 8 when the method is implemented for a flow of images of the scene. Indeed, in this case, this embodiment will encourage the observer to focus his gaze on a predetermined or chosen area of the image of the scene. FIGURES 8 and 9 show the perspective in top view of the imaged scene 20, that is to say the scene with the notion of depth. The optical axis 6 of the observer before changing the point of view and the optical axis 7 of the observer after changing the point of view are illustrated.In the present case, the given change of viewpoint corresponds to a displacement parallel to the plane of the image of the scene 20, which would correspond to a horizontal displacement of the observer if he were actually observing the scene as imaged. The tree 8 which is located behind the individual 9 but in front of the landscape 10 in the background of the image. During the modification step, all of the objects of the scene 8, 9, 10, will be translated relative to the plane of the image. Furthermore, as mentioned previously, it may be advantageous, but not necessary, not to modify the landscape 10 in the background of the image corresponding to the background (at the highest depth(s) or greater than a certain value).
[0284] According to a particular embodiment, the translation of a given object of the image of the scene 20 will be proportional to the calculated deviation and to the angle corresponding to the change of point of view. The translation will be carried out along the axis connecting the observation axis before change of point of view and the observation axis after change of point of view. The objects being translated in the plane of the displayed image, a trigonometric relationship between the calculated depth deviation AP, the angle corresponding to the change of point of view and the distance, noted m, by which the object must be moved to restore the change of point of view can be established. For example, in the case of individual 9 in the foreground of the image of the scene 20, the depth deviation AP is equal to PI-PO, or to the absolute value of PI-PO, where PO corresponds to the depth of individual 9 in the scene and PI corresponds to the depth of tree 8 in the scene.In the case of background 10 in the image, the depth difference AP is equal to P0-P2, or the absolute value of P2-P0, where P2 corresponds to the depth of background 10 in the scene. Thus, as a non-limiting example, the relationship m = sin. x 2AP can be established.
[0285] According to the invention, a VR / VA head-mounted display 12 and a 3D display device 13 are also provided. The head-mounted display 12 comprises means for displaying images in monocular or binocular vision arranged to provide a 3D rendering from two stereoscopic 2D images. The 3D display device 13 comprises means for displaying images arranged to provide a 3D rendering from two stereoscopic 2D images.
[0286] FIGURE 4 illustrates the optional display step of the pair of images of the scene in a binocular head-mounted display. The image of the scene 20 and one 22 of the two modified images 22, 23 is displayed in monocular vision for the left eye 13 of the observer wearing the VR / VA headset and the modified image 21 or the other 23 of the two modified images 22, 23 is displayed in monocular vision for the right eye 14 of the observer wearing the VR / VA headset.
[0287] Referring to FIGURE 5, the optional display step of the scene image pair in a 3D display device is illustrated. The two images of the image pair 20, 21 or 22, 23 are displayed on a display means of the 3D display device 13 either simultaneously or alternately so that one 20, 22 of the two images of the image pair 20, 21 or 22, 23 or a part of one 20, 21 of the two images of the image pair 20, 21 or 22, 23 is seen only by one eye of the observer and the other 21, 23 of the two images of the image pair 20, 21 or 22, 23 or a part of the other 21, 23 of the two images of the image pair 20, 21 or 22, 23 is seen only by the other eye of the observer (if necessary using additional viewing means (e.g. glasses)). As detailed above, those skilled in the art will be able to understand, among other things, the concepts of display by spatial and temporal multiplexing as well as the display of polarized images.
[0288] A person skilled in the art will be able to understand the concepts of stereoscopy and monocular and binocular vision as well as their implementations. Also, with reference to FIGURE 4, a schematic representation of two images of the scene displayed 20, 21 or 22, 23 (as shown in FIGURES 2A and 2B, and in FIGURES 3B and 3C) in a head-mounted display 12 or in a 3D display device 13 is illustrated, a person skilled in the art will observe that the two images of the pair of images displayed 20, 21 and 22, 23 correspond to two different points of view of the scene so that the observer wearing the head-mounted display is able to observe a perspective or relief view of the scene.
[0289] Advantageously, the head-mounted display 12 and the 3D display device 13 comprise means arranged and / or programmed and / or configured to implement the method according to the invention. In practice, the means arranged and / or programmed and / or configured to implement the method according to the invention are a processing unit. In this case, the image of the scene 20 or the stream of images of the scene can be loaded into a storage means of the head-mounted display 12 and by the 3D display device 13. The method will therefore be implemented autonomously by the head-mounted display 12 and by the 3D display device 13.
[0290] However, the head-mounted display 12 and the 3D display device 13 may comprise communication means arranged and / or programmed and / or configured for and / or capable of communicating with external or remote means. In this case, the image of the scene 20 or the stream of images of the scene may be transferred, preferably in real time, from the external means to the head-mounted display 12 and from the 3D display device 13.
[0291] Alternatively, the external means may be arranged and / or programmed and / or configured to implement the method according to the invention. In this case, the head-mounted display 12 and the 3D display device 13 may not implement the method according to the invention, or may not implement all the steps of the method according to the invention. By way of non-limiting example, the method may, in this case, not implement the step of generating the depth or disparity map. In an extreme case, the head-mounted display 12 and the 3D display device 13 may not implement any of the steps of the method according to the invention and only display the image of the scene 20 and the single modified image 21 or the two modified images 22, 23.
[0292] According to an advantageous embodiment, the head-mounted display 12 and the 3D display device 13 comprise at least one imaging system arranged to image the eyes of an observer wearing said VR / VA head-mounted display.
[0293] The processing unit may be arranged and / or programmed and / or configured to determine and / or calculate the change in viewpoint based on data from the at least one imaging system arranged to image the eyes of an observer.
[0294] The VR / VA head-mounted display 12, and respectively the 3D display device 13, comprise means arranged to detect a movement and / or a relative position of the VR / VA head-mounted display 12, and respectively of the 3D display device 13. By way of non-limiting examples, the means arranged to detect a movement and / or a position may be a gyroscope or an accelerometer for the head-mounted display 12, and respectively one or more imaging systems for the 3D display device 13.
[0295] Also, advantageously, the change of viewpoint obtained is a function of or is calculated and / or determined from data, originating for example from the means arranged to detect a movement and / or a relative position of the VR / VA head-mounted display 12, called head-mounted display data, and from the imaging system(s) for the 3D display device 13.
[0296] The effective viewpoint change may be defined as the viewpoint change measured or detected by the head-mounted display 12 or the 3D display device 13, or determined or calculated from the data of the head-mounted display 12 or the 3D display device 13. In other words, the effective viewpoint change corresponds to the actual viewpoint change made by the head-mounted display 12 or by the user. The resulting viewpoint change may, in some cases, correspond to or be equal to the effective viewpoint change. However, in most cases, the resulting viewpoint change will be a function of, calculated from, or determined from the effective viewpoint change, but will not be equal to the effective viewpoint change.
[0297] The change of viewpoint is obtained from images of the observer and / or the observer's environment and / or the head-mounted display 12 and / or the 3D display device 13. The change of viewpoint may be obtained, alternatively or in combination, from data originating from images of the observer, for example from the at least one imaging system of the head-mounted display 12 or the 3D display device 13 arranged to image the eyes of an observer, and / or from the observer's environment, for example from the at least one imaging system of the head-mounted display 12 or the 3D display device 13 arranged to image at least part of the environment.
[0298] According to the embodiment, the method comprises filtering the data from the head-mounted display 12 to limit an amplitude of the change in viewpoint obtained, as a function of time, to a value less than a limit value. For example, the limit value may be the absolute value or the norm of the displacement vector determined from the data from the head-mounted display. Preferably, the filtering is linear.
[0299] The first purpose of the filtering step is to cut off changes in viewpoint that have a high amplitude. In other words, the filtering step aims to prevent sudden variations or gradients in changes in viewpoint from occurring, for example when the observer wearing the head-mounted display suddenly turns around or suddenly pivots 90°. Indeed, since the modification made to the image of the scene 20 is a function of the change in viewpoint, beyond a certain amplitude of change in viewpoint obtained, the objects of the modified images 21, 22, 23 would be too strongly deformed and / or spread out, which could make the images thus modified 21, 22, 23 blurred or unobservable. Indeed, the invention does not aim to generate data or information not present in the source image but only to modify the source image.In other words, it is a matter of truncating or not taking into account movements with too great a variation in amplitude in a short period of time so that the changes in point of view obtained do not deviate too much from the initial position and orientation of the head of the observer wearing the head-mounted display.
[0300] The filtering step may also comprise a step of filtering or smoothing out small changes in amplitude such as small physiological movements of the head. This makes it possible to avoid rendering to the observer frequent small changes in the image of the scene 20 which could induce a seasickness effect in the observer wearing the head-mounted display 12. In other words, the filtering step may be defined as comprising the application of a band-pass filter, in particular a high-pass filter, or a subtraction of an estimate of the average position of the head-mounted display, to the changes in viewpoint actually detected. This therefore makes it possible to smooth out the change in viewpoint actually detected.
[0301] Viewpoint shift variations, which can be referred to as high amplitude viewpoint shifts, are amplitude changes occurring at low frequencies. High amplitude can be defined as an amplitude greater than 1 cm. Low frequency amplitude changes can be defined as having a frequency less than 1 Hz. High frequency viewpoint shift variations can be defined as low amplitude head-mounted display movements, typically less than 0.1 cm. High frequency viewpoint shifts can be defined as having a frequency greater than 10 Hz. The notion of “low” and “high” amplitude change, frequency and amplitude are relative.
[0302] According to the embodiment, the method comprises non-linear clipping of the data from the head-mounted display 12 to limit an amplitude of the obtained viewpoint change to a value below a limit value. Clipping can be defined as a cut applied to the actual viewpoint changes, or to the data relating to the actual viewpoint changes, so as to limit or maintain the amplitude of the obtained viewpoint change used to modify the image of the scene 20 below a maximum amplitude. Indeed, since the modification made to the image of the scene is a function of the viewpoint change, beyond a certain amplitude of the obtained viewpoint change, the objects of the modified images 21, 22, 23 would be too strongly distorted and / or spread, which could make the thus modified images 21, 22, 23 blurred or unobservable.This step aims to limit the effective high viewpoint change values to keep them below predefined maximum values.
[0303] Finally, the method may comprise non-linear dynamic compression of the head-mounted display data to attenuate an amplitude variation of the obtained viewpoint change, as a function of time, such that said amplitude variation does not exceed an amplitude variation threshold value. This non-linear dynamic compression of the effective viewpoint change is complementary to the clipping of the effective viewpoint change. The non-linear dynamic compression may be defined as a modulation of a gain applied to the effective viewpoint changes, or to the data relating to the effective viewpoint changes, so as to limit the amplitude of the obtained viewpoint change used to modify the image of the scene 20.Indeed, in addition to cutting off the sudden large amplitudes of the effective change of viewpoint, it may be judicious to limit the variation of change of viewpoint to be restored to the limitation of the modification of the image of the scene 20 below a predefined value beyond which the modified images 21, 22, 23 are no longer usable.
[0304] Of course, the invention is not limited to the examples which have just been described and numerous adjustments can be made to these examples without departing from the scope of the invention.
[0305] Thus, in variants which can be combined with each other of the embodiments described above:
[0306] - the depth map of the scene is generated by:
[0307] • according to an alternative, called alternative A:
[0308] ■ identification of a local maximum or maxima of sharpness or locals within each of the at least two images of the scene, and ■ for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene presenting the maximum of local sharpness, or
[0309] • according to an alternative, called alternative B:
[0310] ■ for each point of the depth map to be generated, identification of the image of the scene, among the at least two images of the scene, for which the sharpness is maximum, and
[0311] ■ for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene for which the sharpness is maximum, or
[0312] • according to an alternative, called alternative C:
[0313] ■ for each point of the depth map to be generated, interpolation or extrapolation, from the respective focusing distances of the at least two images, of an optimal focusing distance for which each point of the depth map to be generated considered would present the maximum local sharpness,
[0314] ■ for each point of the depth map to be generated, association of said optimal focusing distance, or
[0315] • according to an alternative, called alternative D:
[0316] ■ for each point of the depth map to be generated, identification of the image of the scene, among the at least two images of the scene, for which a local variance is the highest,
[0317] ■ for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene for which local variance is the highest, or
[0318] • according to an alternative, called alternative E:
[0319] ■ identification of a maximum or maxima of local variance or locals within the at least two images of the scene,
[0320] ■ for each point of the depth map to be generated, association of the local variance corresponding to the image of the scene presenting the highest local variance, or
[0321] • according to an alternative, called alternative F: ■ for each point of the depth map to be generated, interpolation or extrapolation, from the respective focusing distances of the at least two images, of an optimal focusing distance for which the depth map would present the highest local variance, and
[0322] ■ for each point of the depth map to be generated, or association of said optimal focusing distance, and / or the step of modifying the displayed image 20 is carried out by moving a part of the pixels or group(s) of pixels of the displayed image 20 and / or by adding and / or by deleting pixels or groups of pixels in the displayed image 20, and / or when obtaining the two modified images 22, 23 by modifying the image of the scene 20, the change of point of view can be predetermined; it may, for example, correspond to a model or a sequence of given change(s) which will be implemented in the step of obtaining a change of point of view to propose or offer or submit to the observer a predetermined change of point of view, and / or when obtaining the two modified images 22, 23 by modifying the image of the scene 20, the change of point of view can be entered so that the two modified images 22, 23,offer a point of view from a position different from his own, with the aim of obtaining the image as he would see it from the point of view associated with this second position, and / or there is provided according to the invention a computer program comprising instructions which, when the program is executed by a computer, lead the latter to implement the method according to any one of the embodiments described, and / or there is provided according to the invention a medium readable, in particular, by computer or by any apparatus comprising a processing unit comprising instructions which, when executed by said computer or said apparatus lead the latter to implement the method according to any one of the embodiments described. In addition, the different characteristics, shapes,Variants and embodiments of the invention may be combined with each other in various combinations provided that they are not incompatible or mutually exclusive.
Claims
CLAIMS 1. Method for 3D rendering from an image of a scene, said method comprises the step of modifying the image of the scene, to obtain a modified image or respectively two modified images, by moving or spreading a part of the pixels or group(s) of pixels of the image of the scene and / or by adding and / or deleting pixels or groups of pixels in the image of the scene, depending on: - a depth map of the scene, - a change in the point of view of the scene which is a function of pupillary distance.
2. Method according to the preceding claim, in which the modification step, or the obtaining step and the modification step, are applied to a stream of images of the scene.
3. Method according to claim 1 or 2, comprising the step of displaying, in a virtual reality / augmented reality (VR / AV) headset or in a 3D display device: - the scene image and the modified image, or - the two modified scene images.
4. Method according to any one of the preceding claims, comprising a step of generating the depth map of the imaged scene from depth information of the image of the scene and / or from sharpness information of the image of the scene.
5. Method according to the preceding claim, in which the depth map is generated, from at least two images of the scene coming from the same stationary imaging system and being acquired, each, with a different focusing distance and / or coming from the same mobile imaging system having acquired, at least in part, the scene in two distinct positions and / or coming from at least two distinct imaging systems, arranged in at least two distinct positions so as to each acquire, at least in part, the scene: - by selection and / or identification of pixels or groups of pixels from the at least two images of the scene or by identification and / or selection of an image or image(s) from among the at least two images, - by combining, assembling or associating pixels or groups of pixels of at least two images of the scene.
6. Method according to the preceding claim, in which the depth map of the scene is generated by: - according to an alternative, called alternative A: • identification of a local maximum or maxima of sharpness or locals within each of the at least two images of the scene, and • for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene presenting the maximum local sharpness, or - according to an alternative, called alternative B: • for each point of the depth map to be generated, identification of the image of the scene, among the at least two images of the scene, for which the sharpness is maximum, and • for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene for which the sharpness is maximum, - according to an alternative, called alternative C: • for each point of the depth map to be generated, interpolation or extrapolation, from the respective focusing distances of the at least two images, of an optimal focusing distance for which each point of the depth map to be generated considered would present the maximum local sharpness, • for each point of the depth map to be generated, association of said optimal focusing distance.
7. Method according to claim 5, wherein the depth map of the scene is generated by: - according to an alternative, called alternative D: • for each point of the depth map to be generated, identification of the image of the scene, among the at least two images of the scene, for which a local variance is the highest, • for each point of the depth map to be generated, association of the focusing distance corresponding to the image of the scene for which local variance is the highest, or - according to an alternative, called alternative E: • identification of a maximum or maxima of local variance or locals within the at least two images of the scene, • for each point of the depth map to be generated, association of the local variance corresponding to the image of the scene presenting the highest local variance, or - according to an alternative, called alternative F: • for each point of the depth map to be generated, interpolation or extrapolation, from the respective focusing distances of the at least two images, of an optimal focusing distance for which the depth map would present the highest local variance, and • for each point of the depth map to be generated, or association of said optimal focusing distance.
8. Method according to claim 4, wherein the depth map of the scene is generated, from the image of the scene, by means of a neural network.
9. Method according to one of claims 4 to 8, in which the depth information and / or the sharpness information contained in the image of the scene or, respectively, the depth information contained in the depth map is refined by processing the image of the scene or, respectively, the depth map; at the end of the processing phase, an image of the scene is obtained whose depth information and / or the sharpness information is refined or, respectively, a depth map of the scene whose depth information is refined;the processing phase comprises an iterative modification of the image of the scene or, respectively, of the depth map so as to minimize a function E comprising: a term D, called a difference term, determined by comparing the image of the scene being iteratively processed, or respectively of the depth map being iteratively processed, convolved by a point spread function (PSF), dependent on the depth map, with the image of the scene before convolution, or respectively with the depth map before convolution, the PSF describing the response of an imaging system from which the image of the scene, or respectively from which the depth map, is obtained, and; a term A, called anomaly term, representative of defects or anomalies within the image of the scene being iteratively processed, or respectively of the depth map being iteratively processed.
10. Method according to any one of the preceding claims, in which the change of point of view comprises a displacement (du, dv) in a reference frame of the image of the scene or a displacement relative to the image of the scene.
11. Method according to any one of the preceding claims, in which the modification of the image of the scene is inversely proportional or proportional to the depth associated with the pixel or group of pixels considered.
12. Method according to any one of the preceding claims, comprising, for each pixel or group(s) of pixels of the image of the scene, a calculation of a difference between a depth associated with a pixel or group(s) of pixels considered in the image of the scene and a depth associated with each other pixel or each other group(s) of pixels of the image of the scene; the modification of the image of the scene is carried out as a function, in addition, of the calculated difference.
13. Method according to any one of the preceding claims, comprising, for each pixel or group(s) of pixels of the image of the scene, a calculation of a difference between a depth associated with an area of the image of the scene considered and a depth associated with each pixel or group(s) of pixels of the displayed image located outside the area of the image of the scene considered; the modification of the image of the scene is carried out as a function, in addition, of the calculated difference.
14. Method according to any one of the preceding claims, comprising a calculation of an angle corresponding to the change of point of view; the modification of the image of the scene is carried out as a function, in addition, of the calculated angle.
15. A method according to claim 3, or according to any one of claims 4 to 15 taken in combination with claim 3, wherein: - the change of point of view obtained is a function of a change in position and / or orientation, in space, of the video headset or of a change in position and / or orientation, in space, of an observer relative to the 3D display device, and / or - the attention zone corresponds to an area of the displayed scene on which the observer's gaze is focused.
16. Method according to claim 3, or according to any one of the claims 4 to 16 taken in combination with claim 3, wherein: - the change of viewpoint is obtained from position data of the head-mounted display and / or movement data of the head-mounted display, called head-mounted display data, and / or - the change of viewpoint is obtained from images of the observer and / or the observer's environment and / or the head-mounted display or the 3D display device, and / or - the attention area is obtained from images of the observer's eyes.
17. Method according to claim 15 or 16, comprising filtering the data from the head-mounted display to limit an amplitude of the change in viewpoint obtained, as a function of time, to a value less than a limit value.
18. A method according to any one of claims 15 to 17, comprising non-linear clipping of the head-mounted display data to limit an amplitude of the resulting viewpoint change to a value less than a limit value.
19. Method according to any one of claims 15 to 18, comprising a non-linear dynamic compression of the data of the head-mounted display to attenuate a variation in amplitude of the change of point of view obtained, as a function of time, so that said variation in amplitude does not exceed a threshold value of variation amplitude.
20. Method according to claim 3, or according to any one of claims 4 to 19 taken in combination with claim 3, comprising a step of acquiring the image of the scene, and in particular the pair of source images, and / or images of the observer and / or at least part of an environment of the head-mounted display or at least part of an environment of the 3D display device.
21. Data processing device comprising means arranged and / or programmed and / or configured to implement the method according to any one of claims 1 to 20.
22. Computer program comprising instructions which, when the program is executed by a computer, cause the latter to implement the method according to any one of claims 1 to 20.
23. A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to implement the method of any one of claims 1 to 20.
24. Virtual reality / augmented reality (VR / AV) head-mounted display comprising means for displaying images in monocular or binocular vision and: - means arranged and / or programmed and / or configured to implement the method according to any one of claims 1 to 20, and / or - communication means arranged and / or programmed and / or configured for and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the method according to any one of claims 1 to 20.
25. VR / VA head-mounted display according to the preceding claim, comprising at least one imaging system arranged to image the eyes of an observer wearing said VR / VA head-mounted display.
26. VR / VA head-mounted display according to claim 24 or 25, comprising: - at least one imaging system arranged to image at least part of an environment of said VR / VA head-mounted display and / or an environment of the observer, and / or - means arranged to detect a movement and / or a relative position of said VR / VA head-mounted display.
27. 3D display device, arranged to display 3D images, comprising: - means arranged and / or programmed and / or configured to implement the method according to any one of claims 1 to 20, and / or - communication means arranged and / or programmed and / or configured for and / or capable of communicating with external means arranged and / or programmed and / or configured to implement the method according to any one of claims 1 to 20.
28. 3D display device according to the preceding claim, comprising at least one imaging system arranged to image the eyes of an observer.
29. A 3D display device according to claim 27 or 28, comprising: - at least one imaging system arranged to image at least part of an environment of said 3D display device and / or of the observer and / or of an environment of the observer, and / or - means arranged to detect a movement and / or a position of an observer relative to said 3D display device.
Citation Information
Patent Citations
Field of view (FOV) throttling of virtual reality (VR) content in a head mounted display
US20180096518A1
Method for 3D scene dense reconstruction based on monocular visual slam
US20200273190A1
Head-mountable display system
US20200296354A1