Method for processing an image of a scene in order to render a change in the viewpoint of the scene
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- FOGALE OPTIQUE
- Filing Date
- 2023-07-08
- Publication Date
- 2026-05-13
AI Technical Summary
Current methods for generating 3D and stereoscopic images are costly in terms of time and computer resources, making them incompatible with dynamic visualization of 2D images in real time, and require specialized equipment for observation.
A method for modifying a 2D image to restore a change in point of view by moving or adding pixels based on a disparity map, allowing for a parallax effect without the need for additional equipment, using a processing unit to detect and implement the change in point of view.
Enables the restoration of a perspective or parallax effect in 2D images, enhancing the immersive experience without the computational burdens and equipment requirements of traditional 3D or stereoscopic methods, allowing for real-time dynamic visualization.
Smart Images

Figure FR2023051059_16012025_PF_FP_ABST
Abstract
Description
[0001] DESCRIPTION
[0002] TITLE: Method for processing an image of a scene to restore a change in point of view of the scene
[0003] Technical field
[0004] The present invention relates to the observation of displayed 2D images.
[0005] The invention relates to the dynamic observation of displayed 2D images. The invention also aims at the immersive observation of displayed 2D images.
[0006] The invention does not relate to 3D images or stereoscopic images.
[0007] State of the prior art
[0008] 3D images are known in the state of the art which make it possible to introduce into displayed images a perspective effect such as it would be perceived if the imaged scene were observed directly by the observer. The 3D images described here are not stereoscopic images but rather images allowing a perspective representation of objects. The synthesis of 3D images can be carried out directly from a model of a real scene. The production of 3D images is costly in terms of time and computing resources required for processing and storing these large images. The load is increased tenfold when it is a stream of 3D images. Furthermore, the time required for processing is incompatible with dynamic visualization of 2D image(s) acquired in real time.
[0009] Stereoscopic vision is also known in the state of the art. Binocular human vision, or more broadly animal vision, based on two images that are compared by the brain, is the most widespread example. The biological processing of the images obtained is extremely efficient, since it provides in real time a notion of depth of the observed scene, allowing, for example, to move knowing the relative distance of the observed objects.
[0010] The observation of stereoscopic images is known in the state of the art, which is inspired by or mimics the principle of stereoscopic vision. The principle of stereoscopic observation consists of increasing the immersive nature of image observation by adding a depth effect by projecting a different image onto each eye of an observer. In practice, a stereoscopic image comprises two distinct images of the same scene acquired in two distinct positions in space, in particular by two cameras spaced apart by the pupillary distance. Human stereoscopic vision will make it possible, from a stereoscopic image, to reconstruct a scene including depth information. As with 3D images, the time required to generate 3D images is incompatible with dynamic visualization of 2D image(s) acquired in real time.Producing stereoscopic images is costly in terms of time, computing resources, and processing and storage. In addition, viewing stereoscopic images requires special glasses or headsets.
[0011] One aim of the invention is to propose an image processing method enabling:
[0012] - to address problems with state-of-the-art processes, and / or
[0013] - to restore a perspective effect to an observer of a 2D image, and / or
[0014] - to restore a parallax effect to an observer of a 2D image, and / or
[0015] - to restore a change of viewpoint of the scene to an observer of a 2D image, and / or
[0016] - to restore a perspective and / or parallax and / or change of viewpoint effect to an observer solely from 2D images, and / or
[0017] - to restore a perspective and / or parallax and / or change of point of view effect to an observer solely by displaying 2D images.
[0018] Presentation of the invention
[0019] For this purpose, a method is proposed for modifying an image of a scene to restore a change in the point of view of the image of the scene. The method comprises the steps of:
[0020] - display the scene image,
[0021] - obtain or detect a change of point of view in relation to the displayed image, - modify the displayed image to restore a change of point of view of the scene image.
[0022] Preferably, the step of modifying the displayed image comprises:
[0023] - a displacement or spreading of a part of the pixels or group(s) of pixels of the displayed image, and / or
[0024] - an addition and / or deletion of pixels or groups of pixels in the displayed image based on: a disparity map of the displayed image, and / or the change of viewpoint obtained, and / or
[0025] - an area of the displayed image remaining unchanged, called the attention zone, when the displayed image is modified.
[0026] Preferably, the image of the scene is displayed by a display means, for example a screen or a projector.
[0027] Preferably, the detection of the change in viewpoint of the displayed image or the displayed scene is carried out by a detection means, for example an imaging system, for example a camera.
[0028] It can be understood as a point of view, a position or coordinates of space.
[0029] Preferably, the step of modifying the displayed image is carried out by means of a processing unit.
[0030] Preferably, the change of viewpoint corresponds to the passage from a viewpoint, called the previous viewpoint, to a different viewpoint, called the new viewpoint. Preferably, the image obtained by implementing the method or the image modified according to the method corresponds to the image observed from the new viewpoint.
[0031] Preferably, each data processing step of the method according to the invention, comprising for example any operation carried out from or on data, for example a calculation, a determination or a comparison, is implemented by a processing unit or any device or system capable of processing data.
[0032] Preferably, the modification of the displayed image is carried out or implemented or performed on at least a part of the displayed image, preferably on the entire displayed image, with the exception of the attention zone. In other words, preferably, the displacement of a part of the pixels or group(s) of pixels of the displayed image and / or the addition and / or deletion of pixels or groups of pixels in the displayed image is carried out outside the attention zone.
[0033] Preferably, the displacement of a portion of the pixels or group(s) of pixels of the displayed image makes it possible to reproduce a parallax effect induced by the change of viewpoint. Preferably, displacement of a portion of the pixels or group(s) of pixels of the displayed image is understood to mean a modification of the position and / or a translation and / or a spreading of a portion of the pixels or group(s) of pixels within the displayed image, preferably in the plane of the displayed image.
[0034] It can be understood by attention zone: one or more pixels of the displayed image.
[0035] Preferably, at least a portion of the pixels of the displayed image, more preferably the pixel(s) of the attention zone, more preferably only the pixel(s) of the attention zone, remain unchanged and / or immobile during the modification step, in particular when the change in the observer's point of view comprises or is a variation in the distance between an observer and the displayed image.
[0036] The disparity map may be defined as a map or image containing, preferably only, depth information. That is, the disparity map may be an image or map comprising a single depth channel of the imaged scene. Preferably, the depth information constitutes a depth field of the imaged scene.
[0037] Preferably, a pixel corresponds or is associated with a point in the imaged scene and a group of pixels corresponds or is associated with an object in the imaged scene.
[0038] Depth information can be understood as data representing or corresponding to or providing information on the depth associated with a pixel, corresponding to a point in the imaged scene, or to a group of pixels, corresponding to an object in the imaged scene, in the displayed image.
[0039] Preferably, the image of the displayed scene was acquired by an imaging system.
[0040] Depth can be understood as a distance between a point or object in the imaged scene and the optical sensor of the imaging system from which the image of the scene was acquired.
[0041] The depth field may comprise depth information of some or all of the points or objects in the scene, preferably a depth field of the imaged scene. The depth field may be contained in a channel of some or all of the pixels in the displayed image.
[0042] Preferably, the image is modified, during the modification step, with respect to or relative to the area of attention and / or the change in point of view obtained or detected.
[0043] Preferably, the method may comprise the step of displaying the modified image to restore the change in viewpoint, in other words the image of the scene as the observer would view it after changing viewpoint.
[0044] Preferably, the method comprises a step of generating, producing, determining or calculating the disparity map of the imaged scene from depth information, preferably from a depth field, of the scene of the displayed image and / or from sharpness information of the displayed image.
[0045] Preferably, the step of generating the disparity map can be defined as enriching the displayed image with depth information, preferably with a depth field, associated with the imaged scene such that the displayed image comprises a depth field channel, preferably a single depth field channel.
[0046] By way of non-limiting example, the information contained in a pixel or block of pixels may include an entropy value, energy, a color channel, typically R, G, B (Red, Green, Blue), intensity detected by the photosites, a gain and / or depth information or a depth field.
[0047] Preferably, the information contained in a pixel or block of pixels is a value or information, called sharpness, relating to or representative of sharpness.
[0048] According to the present invention, the depth information or the depth field can be obtained by comparing images, for example by photogrammetry, or by a distance measuring device, for example a device for detecting and estimating distance by light, called "LIDAR" for "light detection and ranging".
[0049] Preferably, according to a first aspect of the invention, the depth information and / or the depth field is determined or calculated and / or the disparity map is generated, produced, calculated or determined from at least two images of the scene coming from the same stationary imaging system and being acquired, each, with a different focal length.
[0050] Preferably, the at least two scene images comprise the displayed scene image.
[0051] Preferably, the depth information and / or the depth field is determined or calculated and / or the disparity map is generated, produced, calculated or determined:
[0052] - by selection and / or identification of pixels or groups of pixels from the at least two images of the scene or by identification and / or selection of an image or image(s) from among the at least two images,
[0053] - by combining, assembling or associating pixels or groups of pixels of at least two images of the scene.
[0054] The determination of the depth information or the depth field and / or the disparity map may comprise, prior to the selection and / or identification step, a step of comparing image(s) of the scene, among the at least two images of the scene, or pixels or groups of pixels of the at least two images of the scene.
[0055] In this application, the term “generated the disparity map” may be understood to mean: producing, calculating or determining the disparity map.
[0056] In the present application, it may be understood by determining the depth information or the depth field: calculating or deducing the depth information.
[0057] Preferably, according to a first alternative of the first aspect of the invention, the depth information is determined and / or the disparity map is generated, from the at least two images of the scene, by:
[0058] - identification of a local maximum or maxima of sharpness or locals within each of the at least two images, and
[0059] - for each point of the disparity map to be generated, association of the focal distance corresponding to the image of the scene presenting the maximum local sharpness.
[0060] Preferably, according to a second alternative of the first aspect of the invention, the depth information is determined and / or the disparity map is generated, from the at least two images of the scene, by: - for each point of the disparity map to be generated, identification of the image of the scene, among the at least two images of the scene, for which the sharpness is maximum, and
[0061] - for each point of the disparity map to be generated, association of the focal distance corresponding to the image of the scene for which the sharpness is maximum.
[0062] It can be understood as "a point of the disparity map": a pixel or a group of pixels of the disparity map.
[0063] Preferably, each pixel considered or group of pixels considered correspond(s) to an area of the scene or to an object of the scene.
[0064] Preferably, the pixel considered or the group of pixels considered corresponding to or common to the at least two images of the scene correspond(s) to the same area of the scene or to the same object of the scene.
[0065] The local maximum or maxima may correspond to an area of the scene or to an object in the scene.
[0066] Preferably the method comprises, according to a third alternative of the first aspect of the invention, the depth information is determined and / or the disparity map is generated, from the at least two images of the scene, by:
[0067] - for each image of the scene among the at least two images, calculation of a local variance, preferably a local variance of a data item or a parameter or information contained in at least a part of the pixels, more preferably in each pixel or block of pixels, of the at least two images of the scene,
[0068] - for each point of the disparity map to be generated, selection of the image of the scene with the highest variance,
[0069] - for each point of the disparity map to be generated, association of the focal distance corresponding to the selected image.
[0070] Preferably, according to a second aspect of the invention, the depth information is determined and / or the disparity map is generated, from an image of the scene, preferably from the displayed image of the scene, by means of a neural network, preferably an acyclic neural network, more preferably a convolutional neural network. Preferably, the determination of the depth information is carried out and / or the generation of the disparity map according to the second aspect of the invention is implemented from a single image of the scene, preferably from the displayed image of the scene. Thus, preferably, a disparity map is generated for each image of the scene.
[0071] Preferably, the determination of the depth information is carried out and / or the generation of the disparity map according to the second aspect of the invention comprises a step of minimizing an objective function.
[0072] Preferably, the determination of the depth information is carried out and / or the generation of the disparity map according to the second aspect of the invention is implemented from the R, G, B values of the image of the scene.
[0073] The characteristics described in this description apply, unless otherwise indicated, to each of the aspects of the invention and to each of the alternatives according to the invention.
[0074] Preferably, the depth information and / or the sharpness information contained in the image of the displayed scene or, respectively, the depth information contained in the disparity map is refined or improved by processing the image of the displayed scene, or an image associated with the image of the displayed scene, or, respectively, the disparity map. At the end of the processing phase, an image of the displayed scene whose depth and / or sharpness information or, respectively, a disparity map whose depth information is refined or improved is obtained.
[0075] The processing phase comprises an iterative modification of the image of the displayed scene or, respectively, of the disparity map so as to minimize a function E comprising: a term D, called a difference term, determined by comparing the image of the scene or, respectively, of the disparity map convolved by a point spread function (PSF) with the image of the displayed scene before convolution or, respectively, with the disparity map before convolution, the PSF describing the response of an imaging system from which the image of the displayed scene is obtained, and a term A, called anomaly term, representative of defects or anomalies within the image of the scene before convolution or, respectively, of the disparity map before convolution, determined from the image of the scene before convolution or, respectively, of the disparity map before convolution.
[0076] Preferably, at least a portion of the pixels of the displayed image comprises depth information and / or sharpness information.
[0077] The disparity map includes, by definition, depth information.
[0078] The image associated with the displayed scene image may be a reconstructed image. At least a portion of the pixels of the image associated with the displayed scene image may include depth and / or sharpness information.
[0079] Preferably, the reconstructed image is obtained from the displayed image.
[0080] Preferably, the reconstructed image can be enriched by the depth information, contained in at least a portion of the pixels of said reconstructed image, obtained by a dual pixel sensor, or by comparison of at least two images of the scene, preferably acquired simultaneously, more preferably including the displayed image, obtained in two positions or two distinct viewpoints. The at least two images of the scene can be obtained by two distinct imaging systems. The reconstructed image can be obtained by photogrammetry.
[0081] The step of generating the disparity map, according to any of the alternatives, can be defined as consisting of determining and / or associating a distance between each of the points or each group(s) of point(s) of the observed scene and an optical sensor, preferably an optical sensor of an imaging system, for example a camera, having imaged the scene.
[0082] Preferably:
[0083] - the given change of viewpoint includes a displacement (du, dv) parallel to a plane or a flat or curved surface of the displayed image, and / or
[0084] - the given change of viewpoint includes a displacement (dz) perpendicular to the plane or to the flat or curved surface of the image considered.
[0085] Preferably, the method comprises, for each pixel or group(s) of pixels of the displayed image, a calculation, preferably from the disparity map, of a difference between a depth, or an average depth, associated with the attention zone and a depth associated with each pixel or group(s) of pixels of the displayed image located outside the attention zone; the modification of the displayed image is carried out as a function of the calculated difference.
[0086] Preferably, the modification of the displayed image is carried out relatively and / or proportionally to the calculated deviation.
[0087] Preferably, the method comprises calculating an angle corresponding to the change in viewpoint; the modification of the displayed image is carried out as a function of the calculated angle.
[0088] Preferably, the modification of the displayed image is carried out relatively and / or proportionally to the calculated angle.
[0089] Preferably:
[0090] - the change of point of view obtained corresponds to a change of point of view of an observer of the displayed image, preferably to a change of position, in space, of the eyes of the observer, and
[0091] - the attention zone corresponds to the area of the displayed image on which the observer's gaze is focused.
[0092] The angle corresponding to the change of viewpoint can be defined as the angle formed between the optical axis of the observer before the change of viewpoint and the optical axis of the observer of the observer after the change of viewpoint.
[0093] The point of view can correspond to the center of the line connecting the same part of each of the observer's eyes, for example to the center of the line connecting the foveas or the cornea or the iris of each of the observer's eyes.
[0094] The change of the observer's point of view can be defined as the observer's passage from the previous point of view to the new point of view.
[0095] Preferably, the observer's optical axis extends between the observer and the area of attention.
[0096] The viewer's point of view can be defined as the position of the observer relative to the displayed image.
[0097] The attention area may be predetermined, selected, predefined, or defined. Preferably, the attention area may be predetermined, selected, predefined, or defined based on information about the area of the displayed image on which the observer's gaze is focused. Preferably, the attention area corresponds to the point, pixel, or group of pixels or the area of the displayed image on which the observer's gaze is focused.
[0098] Preferably, the attention zone corresponds to a point, for example a pixel or a group of pixels, of the displayed image within the cone of the observer's field of vision. Preferably, the attention zone corresponds to the point of the displayed image having the shallowest depth within the cone of the observer's field of vision or to the point located at the center of the cone of the observer's field of vision. Preferably, the cone of the observer's field of vision, preferably extending around an axis of revolution of the observer's optical axis, formed by a solid angle of between 3 and 5° relative to the observer's optical axis.
[0099] Preferably, the observer's viewpoint and / or area of attention is determined from at least one image of the observer's eyes.
[0100] Preferably, the at least one image of the observer's eyes comes from an imaging system capable of imaging the observer and / or the observer's eyes.
[0101] Determining the observer's point of view and / or area of attention can be done by eye tracking, also known as eye tracking.
[0102] Preferably, the method comprises a step of acquiring the displayed image.
[0103] The method may comprise a step of acquiring at least two images of the scene originating from the same stationary imaging system and / or at least two images of the scene obtained in two positions or two distinct points of view.
[0104] According to another aspect of the invention, a device is proposed comprising means configured to implement all the steps of the method according to the invention.
[0105] According to the invention, there is also provided a data processing device comprising means arranged and / or programmed and / or configured to implement the method according to the invention. According to the invention, there is also provided a computer program comprising instructions which, when the program is executed by a computer, cause the latter to implement the method according to the invention.
[0106] According to the invention, there is also provided a medium, for example a recording medium, readable by a computer comprising instructions which, when executed by a computer, cause the latter to implement the method according to the invention.
[0107] According to the invention, a computer-readable data carrier is also provided, on which the computer program according to the invention is recorded.
[0108] The device according to the invention can be, or be integrated into, any type of device such as a smartphone, a tablet, a computer, a calculator, a processor, a computer chip, programmed to implement the method according to the invention, for example by executing the computer program according to the invention.
[0109] The device may not include an image acquisition means. In this case, the device is used to display one or more images acquired by another device.
[0110] Alternatively, the apparatus may comprise an image acquisition means, such as a camera or a camera module. In this case, the apparatus may be used to display one or more images acquired by said apparatus or by another apparatus.
[0111] In particular, the device may be a user device such as a smartphone, tablet, etc. comprising a display screen. In this case, the detection means may be or may comprise the touch surface, in particular integrated into, or associated with, the display screen of said device.
[0112] In particular, the device may be a computer-type user device, comprising a display screen. In this case, the detection means may be or may comprise a touch-sensitive surface, in particular integrated into, or associated with, the display screen of said computer, or a pointer moved for example by a mouse, or a directional pad of said computer.
[0113] In particular, the device may be a television. In this case, the detection means may be a camera integrated into said television, detecting the gaze and the position of the head of the observer, or a pointer moved for example by a remote control of said television.
[0114] In particular, the device may be a virtual reality or augmented reality headset comprising a display screen or a projector associated with a projection surface onto which each image is projected. In this case, the detection means may be or may comprise a sensor, in particular an optical sensor, equipping said headset.
[0115] Of course, the apparatus according to the invention is not limited to the examples which have just been given.
[0116] In particular, the device may be a medical imaging device.
[0117] In particular, the device may be an endoscope, an ultrasound device, etc.
[0118] According to another aspect of the present invention, there is provided a vehicle comprising:
[0119] - a means of displaying an image, and
[0120] - at least detection of a target position or a change in the observer's point of view,
[0121] - at least one calculation means; configured to implement all the steps of the method according to the invention.
[0122] The vehicle may not include a means of image acquisition. In this case, the image(s) of the scene are provided by another device or another vehicle.
[0123] Alternatively, the vehicle may comprise an image acquisition means, such as a camera or a camera module. In this case, the image(s) of the scene are captured by said image acquisition means, or provided by another device or another vehicle.
[0124] According to embodiments, the vehicle may be a land vehicle, such as a car, autonomous or not.
[0125] According to embodiments, the vehicle may be a flying vehicle, such as a drone, an airplane, a helicopter, autonomous or not. According to embodiments, the vehicle may be a maritime vehicle, such as a boat or a submarine, autonomous or not.
[0126] According to embodiments, at least one image of the scene is a 2D image. In particular, the image of the scene is a 2D image. In particular, the modified image is a 2D image.
[0127] Where appropriate, the image stack comprises at least one 2D image. In particular, all the images in the image stack are 2D images. According to embodiments, at least one image of the scene is a 3D image. In particular, the image of the scene is a 3D image. In particular, the modified image is a 3D image.
[0128] If applicable, the image stack includes at least one 3D image. In particular, all images in the image stack are 3D images.
[0129] Description of figures
[0130] Other advantages and features of the invention will become apparent upon reading the detailed description of implementations and embodiments which are in no way limiting, and the following appended drawings: FIGURE 1A is a schematic representation illustrating an image of a displayed scene and FIGURE 1B is a schematic representation of a disparity map of the scene, FIGURE 2 is a schematic representation of an object projected through a lens onto an optical sensor, FIGURE 3 illustrates a schematic representation of a side perspective view of the scene displayed in FIGURE 1A, FIGURE 4 illustrates, on the left, a schematic representation of the displayed image and, on the right, a schematic representation of the modified image, restoring a change of viewpoint of the observer horizontally with respect to the displayed image, obtained by implementing the method,FIGURES 5a-5c are schematic representations of non-limiting exemplary embodiments of an apparatus according to the invention, FIGURE 6 is a schematic representation of a non-limiting exemplary embodiment of a vehicle according to the invention.,
[0131] Description of the embodiments
[0132] The embodiments described below being in no way limiting, it will be possible in particular to consider variants of the invention comprising only a selection of the described characteristics, isolated from the other described characteristics (even if this selection is isolated within a sentence comprising these other characteristics), if this selection of characteristics is sufficient to confer a technical advantage or to differentiate the invention compared to the state of the prior art. This selection comprises at least one characteristic, preferably functional without structural details, or with only a part of the structural details if this part only is sufficient to confer a technical advantage or to differentiate the invention compared to the state of the prior art.
[0133] With reference to FIGURES 1 to 4, an embodiment of the method for modifying a displayed image 20 of a scene is presented, called method in the remainder of the description. The method aims to modify the displayed image 20 to restore to the observer an image 30 corresponding to the observation of the scene from a different point of view. A simplified schematic representation of a scene is illustrated in FIGURE 1A. The method comprises the step of displaying the image of the scene, for example on a screen. The method also comprises the step of detecting a change in point of view with respect to the displayed image 20. Finally, the method comprises the step of modifying the displayed image 20. According to the non-limiting embodiment, the modification of the image is carried out by moving a portion of the pixels or group(s) of pixels of the displayed image 20.In particular, the displacement of a portion of the pixels or group(s) of pixels of the displayed image 20 consists of a spreading of certain pixels or groups of pixels, preferably a spreading of the pixels not occluded by objects in the scene after a change of viewpoint. More preferably, the displacement of a portion of the pixels or group(s) of pixels of the displayed image 20 comprises a spreading of certain pixels or groups of pixels over pixels occluded by objects in the scene after a change of viewpoint. The modification of the displayed image 20 is a function of: a disparity map of the displayed image 20, the change of viewpoint detected, a zone 8 of the displayed image 20 remaining unchanged, called the attention zone, during the modification of the displayed image 20. According to the non-limiting embodiment, the attention zone corresponds to the zone 8 of the displayed image 20 on which the observer's gaze is focused.Although the attention zone constitutes the data used according to the method, this attention zone can be deduced from the data relating to the observer's gaze. The data relating to the user's gaze can be obtained by oculometry or "eye tracking" without the acquisition step or even the step of determining the attention zone from this data necessarily being part of the method according to the invention. A person skilled in the art knows the techniques for recording eye movement and also knows the data that can be derived therefrom, such as, for example, the foveal path. The method can comprise the step of acquiring the data relating to the observer's gaze, for example, by means of one or more optical systems that may comprise one or more cameras.
[0134] The method also includes displaying, in real time, the modified image 30 obtained by the method following a change in the observer's point of view.
[0135] The change of viewpoint may be effective, that is to say it corresponds to a change in the position of the observer's eyes relative to the displayed image 20. However, the change of viewpoint may also be fictitious, that is to say the position of the observer's eyes relative to the displayed image 20 remains identical. When the change of viewpoint is fictitious, the new viewpoint therefore does not correspond to the position from which the observer is looking at the image. For example, the observer may look at the image from a first position constituting the observer's viewpoint but may provide a second position different from his own, with the aim of obtaining the image as he would see it from the viewpoint associated with this second position.
[0136] According to one embodiment, the detected change of viewpoint corresponds to an actual change in position of the eyes of an observer. The detection of the change of viewpoint can be carried out from one or more optical systems, for example one or more cameras.
[0137] The method may comprise the step of generating the disparity map of the displayed image 20 from depth information of the scene of the displayed image 20 and / or from sharpness information of the displayed image 20. Those skilled in the art will know how to choose the most suitable sharpness information depending on the case. The sharpness information contained in a pixel or block of pixels may comprise the entropy, the energy and / or the variance of a parameter of the image between neighboring pixels.
[0138] According to the embodiment, the disparity map corresponds to an image comprising a single depth channel. Depth is understood to mean the physical distance between a point or a physical object of the scene, corresponding to a pixel or a group of pixels of a considered image, and the sensor from which the considered image was acquired. By way of illustration, a simplified schematic disparity map associated with the image of the scene illustrated in FIGURE 1A is presented in FIGURE 1B. On this disparity map, the white part 4 corresponds to an object located, for example, five meters from the optical sensor of the camera and the gray area 5 of the background of the image corresponds to objects located beyond, for example, fifty meters from the optical sensor.The distance of five meters corresponds, for example, to the minimum focal length of the camera lens and the distance of fifty meters corresponds, for example, to the maximum focal length of the camera lens.
[0139] According to a first aspect of the invention, the depth information and / or the disparity map is determined from at least two images of the scene originating from the same imaging system. The imaging system is stationary and each of the images of the scene is acquired with a different focal length. The at least two images of the scene comprise the displayed scene image 20.
[0140] Preferably, the depth information and / or disparity map is generated determined:
[0141] - by selection and / or identification of pixels or groups of pixels from the at least two images of the scene or by identification and / or selection of an image or image(s) from among the at least two images,
[0142] - by combining, assembling or associating pixels or groups of pixels of at least two images of the scene.
[0143] Thus, the disparity map according to the first aspect of the invention is obtained by aggregating pixels or groups of pixels from different images of the scene. Prior to their aggregation, the group or groups of pixels of the disparity map to be generated are selected from the at least two images. According to a first alternative of the first aspect of the invention, the depth information is determined and / or the disparity map is generated, from the at least two images of the scene, by:
[0144] - identification of a local maximum or maxima of sharpness or locals within each of the at least two images, and
[0145] - for each point of the disparity map to be generated, association of the focal distance corresponding to the image of the scene presenting the maximum local sharpness.
[0146] According to a second alternative of the first aspect of the invention, the depth information is determined and / or the disparity map is generated, from the at least two images of the scene, by:
[0147] - for each point of the disparity map to be generated, identification of the image of the scene, among the at least two images of the scene, for which the sharpness is maximum, and
[0148] - for each point of the disparity map to be generated, association of the focal distance corresponding to the image of the scene for which the sharpness is maximum.
[0149] In other words, the generation of the disparity map can be described as consisting of combining groups of pixels, from several images of the scene, each presenting a maximum of sharpness and enriching each group of pixels with the depth information associated with the object of the scene to which they relate.
[0150] In other words, the first alternative of the first aspect of the invention can be defined as consisting of, from the at least two images of the scene:
[0151] - for each of the at least two images of the scene, identify a local maximum or maxima of sharpness within each of the at least two images,
[0152] - determining the depth information and / or generating the disparity map from the local sharpness maximum or maxima of each of the at least two images, for example by comparing the local sharpness maximum or maxima of each of the at least two images, and the focal length corresponding to the local sharpness maximum or maxima. In other words, the second alternative of the first aspect of the invention can be defined as consisting of, from the at least two images of the scene:
[0153] - for each pixel considered or group of pixels considered, preferably corresponding or common to the at least two images of the scene, of the at least two images of the scene, determine the image of the scene, among the at least two images of the scene, for which the sharpness of the pixel considered or of the group of pixels considered is maximum,
[0154] - for each pixel or group of pixels of the at least two images of the scene, determining the depth information and / or generating the disparity map from the image(s) of the scene for which the sharpness of a pixel considered or of a group of pixels considered is maximum and from the focal distance corresponding to the image of the scene for which the pixel considered or the group of pixels considered has the maximum sharpness.
[0155] According to a third alternative of the first aspect of the invention, the disparity map of the displayed image 20 is obtained from at least two images of the scene originating from the same imaging system, preferably stationary for each shot. The at least two images are each acquired with a different focal length. A shooting technique known in the prior art is called “focus bracketing”. It consists of acquiring a set of images of the same scene from the same stationary optical system, each image is acquired with a different focal length. The image acquired with the shortest focal length, noted Ido, will give the image with the largest viewing angle while the image acquired with the highest focal length, noted Idmax, will give the smallest viewing angle.Thus, some objects in the scene will not be found on Idmax either because they are outside the image field or because they are hidden by objects in the scene.
[0156] From these images, the disparity map is generated by calculating, for each pixel or group of pixels or objects of each image of the scene among the at least two images, a local variance. Preferably, any energy, entropy estimator, conventionally used in focus detection, can be used to determine the variance. Those skilled in the art know a set of techniques for calculating the local variance. As a non-limiting example, the local variance can be the variance of the intensity of the pixels, for example the averaged intensity of the R, G, B channels, calculated on a set of given pixel matrix, to be adjusted according to the resolution of the image and / or the processing power / capacity of the processing unit. Then, for each pixel, group of pixels or object of the image of the disparity map to be generated, the image of the scene whose variance is the highest is selected.In practice, this step can be defined as selecting the sharpest image for a given pixel, group of pixels or object. The focal length corresponding to the selected image is associated with the pixel, group of pixels or object in the image of the disparity map to be generated. When selecting the scene image with the highest variance, a comparison between the scene images and an identification of common objects and / or a correlation between the pixels or objects (groups of pixels) of the scene images can be performed.
[0157] As a non-limiting example, and with reference to FIGURE 2, it is possible to establish a discrete relationship between each image among the at least two images of the scene and the focusing distance, noted f, associated with each of the images. According to the geometric optics relationship + [Math 1], it is possible to connect the focal length of the lens 1, that is, the physical distance between the object 2 of the scene, corresponding to the pixel or group of pixels appearing as sharp in the selected image, and the optical sensor 3 of the camera at the selected image. Considering, for example, a set of n images each acquired with a distinct focal length f between dO, for example 20 cm, and dmax, corresponding, for example, to an infinite distance. The number of images n is an integer between 1 and the number of the at least two images. Considering also that there is a linear relationship between the focal lengths of each of the images, the physical distance or depth, noted P, at which the point of the scene or the object of the scene is located, corresponding to the pixel or group of pixels whose local variance is the highest of the n images, can be expressed as: > [Math 2], where i is an integer that is equal to n-1. In other words, i is between 0 and n-1. For the distance dO, n is equal to 1 and i is equal to 0.
[0158] According to any one of the aspects or alternatives, the depth information and / or the sharpness information contained in the image of the displayed scene 20 or, respectively, the depth information contained in the disparity map is refined or improved by processing the image of the displayed scene 20 or, respectively, the disparity map. At the end of the processing phase, an image of the displayed scene whose depth and / or sharpness information or, respectively, a disparity map whose depth information is refined or improved is obtained.
[0159] The processing phase comprises an iterative modification of the displayed scene image 20 or, respectively, of the disparity map so as to minimize a function E comprising:
[0160] - a term D, called difference term, determined by comparing the image of the displayed scene or, respectively, the disparity map convolved by a point spread function (PSF) with the image of the displayed scene before convolution or, respectively, with the disparity map before convolution, and
[0161] - a term A, called anomaly term, representative of defects or anomalies within the image displayed before convolution or, respectively, the disparity map before convolution, determined from the image displayed before convolution or, respectively, the disparity map before convolution.
[0162] The PSF describes the response of an imaging system from which the image of the displayed scene 20 is obtained.
[0163] In other words, the processing phase constitutes an iterative loop in which, the number of iterations is noted k, the image of the scene displayed 20 constitutes the image processed during the first iteration at k=1. The processed image is then modified at each iteration of the loop.
[0164] As a non-limiting example, the term D may comprise the sum of several terms Di. Obtaining the distances Di are obtained by comparing the image displayed during processing at iteration k convolved by the PSFs with the image displayed at iteration k-1 or convolved with the image of the scene displayed 20 at iteration k=0. The distances Di may be, among other things, obtained by an evaluation of the local colors, or by evaluation of another parameter contained in the pixels or groups of pixels, of the image displayed at iteration k opposite the discretization grid of the PSF, so as to know the calculated colors, or to know another parameter contained in the pixels or groups of pixels, at the positions of the photosites of the image displayed at iteration k to make the distance comparisons at the appropriate locations.
[0165] As a non-limiting example, the term A, called penalty, representative of defects or anomalies within the image of the scene displayed during processing at iteration k, determined from the image of the scene displayed during processing at iteration k, can be calculated concomitantly with the term D.
[0166] According to the non-limiting embodiment, the term A comprises:
[0167] - at least one component Al whose effect is minimized for small intensity differences between neighboring pixels of the image of the scene displayed during processing at iteration k, and / or
[0168] - at least one component A2 whose effect is minimized for small differences in hue between neighboring pixels of the image of the scene displayed during processing at iteration k, and / or
[0169] - at least one component A3 whose effect is minimized for low frequencies of changes of direction between neighboring pixels of the image of the scene displayed during processing at iteration k drawing an outline of an object of the scene.
[0170] The second term A can thus include a sum of several terms Ai.
[0171] Preferably, at each iterative modification of the image displayed during processing, the processing phase comprises a convolution of the image displayed during processing at iteration k by an inverse function of the PSF, called iPSF. Preferably, the convolution of the image displayed during processing at iteration k by the iPSF is carried out as the last step of an iteration k of the processing phase or as the first step of an iteration k> 1 of the processing phase.
[0172] The iterative modification of the displayed image 20 ends when the function E, or a combination of partial derivatives of the function E with respect to the image, with respect to the displayed image being processed at iteration k or with respect to the displayed image 20, is less than a minimization threshold, or when a certain number of iterations of the iterative modification of the displayed image 20 is reached; preferably, the at least one displayed image being processed thus modified is restored.
[0173] According to a second aspect of the invention, the step of generating the disparity map of the displayed image is implemented from the image of the displayed scene, by means of a convolutional neural network. Those skilled in the art will know how to adapt and use the most suitable neural network. By way of non-limiting example, the convolutional neural network may comprise several layers of neurons. Among these layers of neurons, the network comprises processing layers comprising neurons arranged to process the displayed image with a convolution. The neural network may also comprise sub-sampling (or pooling) layers, interposed between two processing layers, comprising neurons arranged to combine or merge the output data of the processing layers and reduce their size.By way of non-limiting example, the convolutional neural network may further comprise a correction layer, interposed between two processing layers, comprising neurons arranged to operate an activation function on the output data of the processing layers.
[0174] With reference to FIGURES 3 and 4, and for any of the aspects or alternatives, the modification of the displayed image 20 as a function of the change of point of view is detailed.
[0175] To facilitate the description, FIGURE 3 shows the perspective of the imaged scene, i.e. the scene with the notion of depth. Illustrated is the optical axis 6 of the observer before changing the point of view and the optical axis 7 of the observer after changing the point of view. In the case presented to the observer, the given change in the observer's point of view corresponds to a displacement from the parallel to the plane of the displayed image 20, which would correspond to a vertical displacement of the observer if he were actually observing the scene as imaged. The attention zone 8 of the observer, on which the observer's gaze remains focused during the change of point of view, is the tree 8 which is located behind the individual 9 but in front of the landscape 10 in the background of the image.When the observer's point of view changes, all the objects in the scene 9, 10, other than the one on which the observer's gaze is focused, will be translated into the image plane. In addition, some objects or parts of objects in the scene, corresponding to a few pixels of the outline of the attention zone 8, that is to say pixels located in the immediate vicinity of the attention zone 8, may be hidden by the object on which the observer's gaze is focused and some objects or parts of objects in the scene, corresponding to a few pixels of the outline of the attention zone 8, which were hidden before the change of point of view may appear on the modified image 30 according to the method.
[0176] FIGURE 4 illustrates the modification of the image corresponding to a given change in the observer's point of view corresponding to a displacement dv parallel to the plane of the displayed image 20, which would correspond to a horizontal displacement of the observer if he were actually observing the scene as imaged. The schematic representation on the left illustrates the displayed image 20 before the change in point of view and the schematic representation on the right illustrates the modified image 30 according to the method restoring the change in point of view.
[0177] The method comprises a calculation of an angle a corresponding to the change of point of view. The angle a corresponds to the angle formed between the optical axis 6 of the observer before changing the point of view and the optical axis 7 after changing the point of view. The modification of the displayed image 20 is carried out according to the calculated angle a.
[0178] The modification of the displayed image 20 comprises, for each pixel or group(s) of pixels of the displayed image 20, a calculation of a difference between a depth associated with the attention zone 8 and a depth associated with each pixel or group(s) of pixels of the displayed image 20 located outside the attention zone 8. This difference is calculated from the data of the disparity map. The modification of the displayed image 20 is carried out according to the calculated difference.
[0179] In practice, the translation of a given object in the displayed image 20, excluding the object of the attention zone 8, will be proportional to the calculated deviation and to the angle a. The translation will be carried out along the axis connecting the optical axis 6 of the observer before changing the point of view and the optical axis 7 after changing the point of view. The objects being translated in the plane of the displayed image, a trigonometric relationship between the calculated depth deviation AP, the angle a and the distance, noted m, by which the object must be moved to restore the change of point of view can be established. For example, in the case of the individual 9 in the foreground of the displayed image 20, the depth deviation AP is equal to PI-PO, or to the absolute value of PI-PO, where PO corresponds to the depth of the individual 9 in the scene and PI corresponds to the depth of the attention zone 8 in the scene.In the case of background 10 in the image, the depth difference AP is equal to P0-P2, or the absolute value of P2-P0, where P2 corresponds to the depth of background 10 in the scene. Thus, as a non-limiting example, the relationship m = sin. x 2AP can be established.
[0180] According to a particular embodiment of the invention, the method can be implemented on a smartphone. In this case, the displayed image 20 can be acquired by a camera of the smartphone. The camera or another camera(s) can image the observer to acquire data relating to the position of the observer relative to the displayed image, in particular the position of the eyes, and data relating to the gaze of the user to calculate the change of point of view and the area of attention from the data acquired by the camera(s) of the smartphone.
[0181] Of course, the invention is not limited to the examples which have just been described and numerous adjustments can be made to these examples without departing from the scope of the invention.
[0182] Thus, in variants of the previously described embodiments that can be combined with each other: the given change of viewpoint comprises a displacement dz perpendicular to the plane of the image considered, in other words a displacement along the optical axis 6 of the observer, and / or the step of modifying the displayed image 20 is carried out by moving a portion of the pixels or group(s) of pixels of the displayed image 20 and / or by adding and / or deleting pixels or groups of pixels in the displayed image 20, and / or there is proposed according to the invention a device, a system or any apparatus comprising a processing unit arranged to implement any one of the embodiments of the method according to the invention just described to restore a modified image 30 reproducing the change of viewpoint of an observer of the displayed image 20 from images acquired by one or more imaging systems of said device,of said system or said apparatus or to restore a modified image 30 reproducing the change of point of view of an observer of the displayed image 20 from images stored in said device, said system or said apparatus; said device, said system or said apparatus being able to be, by way of non-limiting examples: a smart mobile phone (or smartphone), a computer, a photographic device, a vehicle, a drone, a medical device or a satellite, and / or it is proposed according to the invention a device, a system, a machine (preferably motorized) or any apparatus comprising a processing unit implementing any of the embodiments of the method according to the invention just described for correcting / processing / restore images acquired by one or more imaging systems 2 of said device, said system, said machine or said apparatus or for correcting / processing / restore images stored in said device, said system,said machine or said apparatus, said device, said system, said machine or said apparatus being able to be, by way of non-limiting examples: a smart mobile phone (or smartphone), a tablet, a camera, a virtual reality headset, a vehicle, a drone, a medical device or a satellite, and / or a use of any of the embodiments of the device according to the invention just described and / or of any of the embodiments of the method according to the invention just described within a device, a system, a machine or an apparatus just described or within a processing unit from images, preferably a stream of images, coming from a device, a system, a machine or an apparatus just described, and / or it is proposed according to the invention a computer program comprising instructions which, when the program is executed by a computer,lead the latter to implement the method according to any one of the embodiments described, and / or there is provided according to the invention a support readable, in particular, by computer or by any device comprising a processing unit comprising instructions which, when executed by said computer or said device, lead the latter to implement the method according to any one of the embodiments described.,
[0183] FIGURE 5a is a schematic representation of a non-limiting exemplary embodiment of an apparatus according to the present invention.
[0184] The apparatus 910 of FIGURE 5a comprises means configured to implement the invention.
[0185] The apparatus 910 of FIGURE 5a may comprise a device according to the invention.
[0186] In the example shown in FIGURE 5a, the device 910 is a smartphone, or a tablet, comprising the device according to the invention. In particular, the device 910 comprises a display screen 810 equipped with a detection surface 812, for example capacitive, and at least one camera 814. FIGURE 5b is a schematic representation of another non-limiting exemplary embodiment of a device according to the present invention.
[0187] The apparatus 920 of FIGURE 5b comprises means configured to implement the invention.
[0188] The apparatus 920 of FIGURE 5b may comprise a device according to the invention, without the camera 814.
[0189] In the example shown in FIGURE 5b, the apparatus 920 is a virtual reality, VR, headset, or an augmented reality headset, comprising the device according to the invention. In particular, the headset 920 comprises a display screen 810, a sensor for detecting the change of point of view, i.e. the position aimed by an eye, or eyes, of the user on said display screen 810.
[0190] In the example shown in FIGURE 5b, the headset 920 does not include imaging means for capturing images of the scene. In this case, the image(s) of the scene to be displayed by the headset 920 are provided by another device to said headset 920.
[0191] Alternatively, the headset 920 may comprise at least one camera for capturing images of the scene in which it is located to display them on the screen 810, optionally after enriching said images, for example in the context of an augmented reality application.
[0192] FIGURE 5c is a schematic representation of a non-limiting exemplary embodiment of an apparatus according to the present invention.
[0193] The apparatus 930 of FIGURE 4c comprises means configured to implement the invention.
[0194] The apparatus 930 of FIGURE 4c may comprise a device according to the invention.
[0195] In the example shown in FIGURE 4c, the apparatus is a medical imaging apparatus, such as an endoscope, an ultrasound apparatus, etc. comprising the device according to the invention. In particular, the medical imaging apparatus 930 comprises a display screen 810 equipped with a detection surface 812, for example capacitive. The medical imaging apparatus 930 further comprises an imaging means formed by a distal objective connected to an imaging module (not shown).
[0196] FIGURE 6 is a schematic representation of a non-limiting exemplary embodiment of a vehicle according to the present invention. The vehicle 100 of FIGURE 6 comprises means configured to implement the invention.
[0197] The vehicle 1000 of FIGURE 6 may comprise a device according to the invention. In the example shown in FIGURE 6, the vehicle 1000 is a land vehicle, in particular a car, comprising the device according to the invention. In particular, the vehicle 1000 comprises a display screen 810 equipped with a detection surface 812, for example capacitive, arranged in the passenger compartment of the vehicle 1000. The vehicle 1000 further comprises at least one camera, for example arranged on the windshield of the vehicle 1000.
[0198] In addition, the various features, forms, variations and embodiments of the invention may be combined with each other in various combinations to the extent that they are not incompatible or mutually exclusive.
Claims
CLAIMS 1. Method for modifying an image of a scene to restore a change of point of view of the image of the scene, said method comprises the steps of: - display the scene image, - obtain a change of point of view in relation to the displayed image, - modify the displayed image, by moving or spreading part of the pixels or group(s) of pixels in the displayed image and / or by adding and / or deleting pixels or groups of pixels in the displayed image, depending on: - a disparity map of the displayed image, - the change of point of view obtained, - an area of the displayed image remaining unchanged, called the attention zone, when the displayed image is modified.
2. Method according to claim 1, comprising a step of generating the disparity map of the imaged scene from depth information of the scene of the displayed image and / or from sharpness information of the displayed image.
3. Method according to claim 2, in which the disparity map is generated, from at least two images of the scene coming from the same stationary imaging system and each being acquired with a different focal length: - by selection and / or identification of pixels or groups of pixels from the at least two images of the scene or by identification and / or selection of an image or image(s) from among the at least two images, - by combining, assembling or associating pixels or groups of pixels of at least two images of the scene.
4. The method of claim 3, wherein the disparity map of the displayed scene is generated by: - according to an alternative, called alternative A: - identification of a local maximum or maxima of sharpness or locals within each of the at least two images, and - for each point of the disparity map to be generated, association of the focal length corresponding to the image of the scene presenting the maximum local sharpness, or - according to an alternative, called alternative B: - for each point of the disparity map to be generated, identification of the image of the scene, among the at least two images of the scene, for which the sharpness is maximum, and - for each point of the disparity map to be generated, association of the focal distance corresponding to the image of the scene for which the sharpness is maximum.
5. The method of claim 3, wherein the disparity map of the displayed image is generated by: - for each image of the scene among the at least two images, calculation of a local variance, - for each point of the disparity map to be generated, selection of the image of the scene whose local variance is the highest, - for each point of the disparity map to be generated, association of the focal distance corresponding to the selected image.
6. Method according to claim 1 or 2, wherein the disparity map of the displayed image is generated, from the image of the displayed scene, by means of a neural network.
7. Method according to one of claims 2 to 6, in which the depth information and / or the sharpness information contained in the image of the displayed scene or, respectively, the depth information contained in the disparity map is refined by processing the image of the displayed scene or, respectively, the disparity map, at the end of the processing phase an image of the displayed scene of which the depth information and / or the sharpness information or, respectively, a disparity map of which the refined depth information is obtained;the processing phase comprises an iterative modification of the image of the displayed scene or, respectively, of the disparity map so as to minimize a function E comprising: a term D, called a difference term, determined by comparing the image of the scene or, respectively, of the disparity map convolved by a point spread function (PSF) with the image of the scene before convolution or, respectively, with the disparity map before convolution, the PSF describing the response of an imaging system from which the image of the displayed scene is obtained, and a term A, called anomaly term, representative of defects or anomalies within the image of the scene before convolution or, respectively, of the map; pre-convolution disparity, determined from the pre-convolution scene image or, respectively, from the pre-convolution disparity map.
8. Method according to any one of the preceding claims, in which: - the given change of viewpoint includes a displacement (du, dv) parallel to a plane of the displayed image, and / or - the given change of viewpoint includes a displacement (dz) perpendicular to the plane of the image considered.
9. Method according to any one of the preceding claims, comprising, for each pixel or group(s) of pixels of the displayed image, a calculation of a difference between a depth associated with the attention zone and a depth associated with each pixel or group(s) of pixels of the displayed image located outside the attention zone; the modification of the displayed image is carried out as a function of the calculated difference.
10. Method according to any one of the preceding claims, comprising a calculation of an angle corresponding to the change of point of view; the modification of the displayed image is carried out as a function of the calculated angle.
11. Method according to any one of the preceding claims, in which: - the change of point of view obtained corresponds to a change of position, in space, of the eyes of an observer, and - the attention zone corresponds to the area of the displayed image on which the observer's gaze is focused.
12. Method according to any one of the preceding claims, in which the point of view of the observer and / or the area of attention is determined from at least one image of the eyes of the observer.
13. Method according to any one of the preceding claims, comprising a step of acquiring the displayed image.
14. Data processing device comprising means arranged and / or programmed and / or configured to implement the method according to any one of claims 1 to 13.
15. Computer program comprising instructions which, when the program is executed by a computer, cause the latter to implement the method according to any one of claims 1 to 13.
16. A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 13.