Method and device for determining a depth map of a scene from dual-pixel data of the scene
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- FOGALE OPTIQUE
- Filing Date
- 2025-01-22
- Publication Date
- 2026-07-30
Smart Images

Figure FR2025050044_30072026_PF_FP_ABST
Abstract
Description
DESCRIPTION Title: Method and device for determining a depth map of a scene from dual-pixel data of said scene
[0001] The present invention relates to a method for determining a depth map of a scene from dual-pixel data of said scene. It also relates to a computer program and a device implementing such a method.
[0002] The domain of the invention is the domain of obtaining a depth map of a scene from image(s) of said scene. State of the art
[0003] The depth map of a scene, of which only two-dimensional (2D) image(s) are available, is an important piece of data. It can be used to perform various functions, for example to image the scene using a focus bracketing technique, or to provide a representation of the scene with a 3D effect, a parallax effect, etc., particularly in a virtual reality (VR) or augmented reality (AV) headset, but not exclusively.
[0004] Solutions based on artificial intelligence (AI) exist for estimating the depth map of a scene from a 2D image of that scene. In short, a 2D image of the scene is fed into an AI model, such as a convolutional neural network (CNN), which has been previously trained on a training dataset. This AI model then outputs an estimated depth map of the scene. Examples of known models available to professionals include the Marigold Depth Estimation Neural Network and the Depth Anything Neural Network, both of which can generate a depth map of a scene from a 2D image.
[0005] These solutions represent a significant advancement, but they still have drawbacks. In particular, the accuracy of these solutions depends on the training dataset used, as the images processed when using the AI model generally differ from those in the training dataset, leading to a decrease in the accuracy and performance of the AI model when used.
[0006] Added to this is another, more significant difficulty related to the use of a 2D image of the scene. A 2D image provides limited information about the scene, entirely dependent on the viewpoint used to image it, and more generally on the image acquisition conditions. These conditions can cause optical effects on the 2D image, leading to errors in depth map estimation, not to mention optical effects deliberately introduced into the imaged scene that are then reflected in the 2D image provided as input to the AI model.
[0007] One objective of the present invention is to remedy at least one of the aforementioned drawbacks.
[0008] Another aim of the invention is to provide a more accurate and efficient depth map estimation solution for a scene.
[0009] Another aim of the invention is to provide a simpler and less expensive solution for estimating the depth map of a scene. Description of the invention
[0010] The invention proposes to achieve at least one of the aforementioned goals by a method for determining a depth map of a scene, - receiving an image, called a dual-pixel image, comprising, and in particular consisting of, values provided by dual pixels of an image sensor, and - estimation of a depth map of said scene with a predetermined model taking said dual-pixel image as input.
[0011] The invention proposes to use the signals captured by the dual-pixels of an image sensor to determine, by estimation or by calculation, the depth map of a scene.
[0012] Currently, the dual pixels of an image sensor are used during image acquisition to perform autofocus on an imaging device, or to assist a user when manually adjusting the focus distance of their imaging device. This adjustment, whether manual or automated, of the focus distance is implemented before image acquisition and aims solely to find the correct focus distance for a given region of the scene.
[0013] The invention proposes using the dual-pixel functionality of an image sensor to determine the depth map of a scene as captured by the image sensor, specifically the depth map of the entire scene and not just for a region where focus needs to be adjusted. This depth map determination can be performed before capturing an image of the scene. In particular, this determination can be performed during or after capturing an image of the scene.
[0014] Thus, the invention proposes a solution for determining a depth map of a scene which is more precise since it is based on signals provided by dual-pixels which are more representative of the distance of objects in the scene, and above all simpler and less expensive to implement since dual-pixel technology is present on the majority, if not all, image sensors.
[0015] By image, we mean a digital image, and in particular a raster image, and more specifically an RGB raster image for example.
[0016] A dual-pixel image is defined as the data matrix formed by the values provided by the dual pixels. Preferably, but without loss of generality, the data matrix representing the dual-pixel image is the same size as the image sensor. For each location (u,v) of the image sensor containing a dual pixel, the matrix includes a set of values, in particular a set of two values, provided by that dual pixel. For each location (u,v) of the image sensor not containing a dual pixel, the matrix may include a set of predetermined values, for example, zeros. Alternatively, for each location (u,v) of the image sensor not containing a dual pixel, the matrix may include a set of values estimated based on the values of the other dual pixels, and in particular the surrounding dual pixels, or even of the other photosites, in particular photosites filtered to detect color information.
[0017] A dual-pixel is generally formed by two photodiodes, each with a lens attached. One half of the lens covers one of the photodiodes, and the other half covers the other. In this configuration: - a point in the scene, located at the focal distance, illuminates identically the two photodiodes situated at the same point on the sensor, and - a point in the scene, which is not at the focal distance, illuminates in a spatially offset manner on the sensor, at least two photodiodes of different dual-pixel type: the spatial offset of the at least two photodiodes is monotonic with respect to the distance to the distance of said point from the focal plane.
[0018] A dual-pixel, also called a multi-pixel, can be formed by several photodiodes, each covered by one or more lenses. For each photodiode, and specifically for the lens or lens portion(s) covering it, the photodiode is illuminated by an angular sector from the image focusing lens. This sector has a different, and generally complementary, spatial distribution compared to the other photodiodes in the multi-pixel. The photodiodes associated with the lens or lens portion may have identical or different light-collecting surfaces. The photodiodes themselves may also be identical or different. In this configuration: - a point in the scene, which is at the focal distance, illuminates identically, or with an intensity proportional to the collection of each respective angular sector, the several multi-pixel photodiodes located at the same point on the sensor, at the point of convergence of the optical beam and - A point in the scene, which is not at the focal distance, illuminates the sensor with a spatial offset, respectively for each different type of multi-pixel: the relationship of the spatial offset of at least two photodiodes is monotonic with respect to the distance of said scene point from the focal plane. - The set of photosites of a multi-pixel type illuminated by a scene point, which is not at the focal distance, corresponds to a luminous spot whose center moves fairly proportionally to the inverse of the focal distance from the inverse of the current depth.
[0019] Depending on the embodiment, the following relationship can be used to link the offset Ad observed for a point to the depth Zpoint of the image point that formed that area of the image. - where Zfocus is the focusing distance, F is the focal length of the lens, - Zpoint is the desired depth for the object or point area in the scene. - <D est le diamètre de collection de la lumière au niveau de l'objectif, - k est un nombre généralement compris entre 0 et 0.5, ou 0 et 1, lié à la géométrie des dual-pixels et de leur lentille(s) associée(s). Il peut devoir être modulé le long du champ de l'image. Il est généralement assez indépendant de la distance Zpoint, cette valeur de k peut être déterminée par calcul ou par étalonnage, - Ad is the displacement of the light spot, relative to the ideal focal point or center of the light spot for standard pixels of the image sensor, or the distance between two such spots in different sub-images.
[0020] By "depth of field" we mean the extent of the area of sharpness that appears on an image, that is to say the area between the first sharp plane and the last sharp plane of the image.
[0021] By "extent of sharpness" we mean the distance over which the sharp part of the image extends, that is to say the distance from the first sharp plane to the last sharp plane of the image.
[0022] The focal length refers to the distance from a point in the scene to the lens at which the optical lens is focused. The focal length for capturing an image is generally adjusted by changing the distance between the image sensor and the optical lens. Thus, a first image of a scene acquired at a first focal length will clearly depict one part of the scene, and a second image of a scene acquired at a second focal length will clearly depict another part of the scene.
[0023] Following a non-limiting example, given to illustrate the definitions indicated above, the focusing distance can be adjusted to 15 meters. The depth of field can be 1.50 meters and the depth of field can be 14.50 meters to 16 meters. In this case, the image will sharply depict all objects, or parts of the scene, located between 14.5m and 16m from the optical lens.
[0024] According to embodiments, the method according to the invention may further include a step of obtaining the dual-pixel (IC) image from a raw image provided by the image sensor, by selecting in said raw image and preserving the values provided by the dual-pixels of said image sensor.
[0025] In this case, the raw image as provided by the image sensor, such as a CCD or CMOS sensor, is processed, for example by calculation, to eliminate the pixels representing the color data, known as the standard image, of the scene. Thus, in the raw image provided by the image sensor, only the data from the dual-pixels is retained.
[0026] In this case, the dual-pixel image contains only the data provided by the dual pixels. Thus, the dual-pixel image is smaller and requires fewer resources to be stored.
[0027] Of course, depending on the embodiment, the raw image containing the dual-pixel data can be used as input for the depth map estimation or calculation model. In this case, it is the raw image as provided by the image sensor that is given as input to the depth map estimation or calculation model.
[0028] According to embodiments, the process according to the invention may include a step of acquiring the raw image with the image sensor.
[0029] In this case, a device including a camera module is used to acquire the raw image. Such a device can be any type of device, such as a smartphone, tablet, camera, etc.
[0030] Of course, according to specific embodiments, the method according to the invention may not include a raw image acquisition step. In this case, the raw image may be captured beforehand by the device implementing the method according to the invention, or by another device.
[0031] Depending on the embodiment, the dual-pixel image can be formed by several sub-images, each sub-image being formed by the data provided by a photodiode of each dual-pixel.
[0032] For example, when each dual-pixel is formed by two photodiodes, called the left photodiode and the right photodiode, without loss of generality, then the dual-pixel image is formed by two sub-images: a left sub-image formed by the data provided by the left photodiodes of the dual-pixels of the image sensor, and a right sub-image formed by the data provided by the right photodiodes of the dual-pixels of the image sensor.
[0033] In general terms, a dual-pixel image can be formed by as many sub-images as there are photodiodes in each dual-pixel of the image sensor. For example, if each dual-pixel contains N photodiodes, then the dual-pixel image can potentially comprise N sub-images, each sub-image being formed by the data provided by one photodiode of the dual-pixels.
[0034] Each sub-image should preferably be formed with dual-pixel photodiodes of the same type, i.e., for example, with all left-handed dual-pixels. By photodiode type or dual-pixel, we mean a set of photodiodes or pixels associated with dual-pixels, or multi-pixels, offset in a similar direction with respect to a reference direction of the sensor, any axis on the sensor being able to serve as a reference.
[0035] Depending on the embodiment, the standard [RGB] image can itself form a sub-image also used by the process.
[0036] According to some embodiments, the predetermined model can also receive as input the focusing distance used for the dual-pixel image acquisition.
[0037] In these embodiments, the predetermined model receives as input both the dual-pixel image and the focal distance used to acquire said dual-pixel image.
[0038] The focal length information is useful for determining the depth map of the scene from the dual-pixel image. Indeed, at a given point in the scene, the focal length used to image the scene and the information provided by the dual-pixels for that point allow us to deduce the depth of that point within the scene.
[0039] According to embodiments, the predetermined model can further receive the lens-sensor distance of an imaging module comprising said image sensor associated with an optical lens and used for dual-pixel image acquisition.
[0040] This distance information can be useful for determining the depth of each point in the scene by taking into account the lens-sensor distance used for dual-pixel image acquisition.
[0041] In some embodiments, the predetermined model receives as input the dual-pixel image, the focal length, and the lens-to-sensor distance. From this data, the predetermined model determines the depth map of the imaged scene for at least certain points in the scene, specifically each point in the scene corresponding to each dual-pixel.
[0042] The dual-pixel image can be provided to the model as a single matrix comprising the data provided by the dual-pixels, or as multiple matrices, each comprising the data provided by the same type of photodiode of the dual-pixels.
[0043] Depending on the embodiment, the model used may be a previously trained artificial intelligence (AI) model.
[0044] In this case, the artificial intelligence model can provide by estimation a depth map of the scene, including depth data for at least some, in particular all, points of the scene corresponding to dual-pixels.
[0045] Using an artificial intelligence model allows for faster estimation of a scene's depth map from a dual-pixel image. Of course, in this case, it is necessary to train the artificial intelligence model beforehand.
[0046] The artificial intelligence model can be any type of artificial intelligence model.
[0047] Depending on the implementation, the AI model can be a neural network, and in particular a pre-trained convolutional neural network. Such neural networks are well known to those skilled in the art and are widely used in image processing, for example in object recognition in images, or in object tracking.
[0048] Non-limiting examples of neural networks that can be used in the present invention may include, for example, a deep learning convolutional neural network. Such a model may include: - at least as many inputs as the number of photosites in the dual-pixel image, i.e., the left image + right image in the case of a dual-pixel image with two photosites; and optionally, one input for each additional focus data point, such as, for example, one input for the focus distance, and / or one input for the lens-to-sensor distance, etc. Such a model may include one or more convolution layers. Such a model may include outputs for depth data. The number of outputs preferably corresponds to the number of depth data points provided.
[0049] According to embodiments, the method according to the invention may further include a step of training the AI model with a training database comprising training sets.
[0050] Each training game may include: - a dual-pixel image of a scene, possibly formed by several sub-images; - a depth map of said scene; - optionally, but preferably, the focusing distance used for dual-pixel image acquisition and / or optionally, but preferably, the lens-sensor distance used for dual-pixel image acquisition.
[0051] The training dataset used may include a multitude of images, one part of which is used for training the AI model, and a second part used for validating the learning of the AI model.
[0052] Following a non-imitative implementation example, the AI model can be trained using any suitable training algorithm. In one example, the training algorithm could be the backpropagation of error algorithm.
[0053] A training set to train the AI model can be obtained in different ways.
[0054] For example, a training game can be obtained for a known training scene.
[0055] Given the known training scene, the depth of each point in the scene can be determined, for example, by measurement with a sensor such as a LiDAR or time-of-flight camera. Alternatively, the depth of each point in the training scene can be determined by construction; that is, the training scene is built based on a depth map. This depth map can then be used as the depth map for a training dataset. A dual-pixel image of the training scene can be acquired with a camera module. This dual-pixel image can also be used as the dual-pixel image for the training dataset. Optionally, the focal length and / or the sensor-to-lens distance used to acquire the dual-pixel image is / are known.Thus, the training set can be formed by the depth map thus obtained, by the dual-pixel image thus obtained, and optionally by at least one of the distances thus obtained.
[0056] By repeating this operation for different training scenes, it is possible to build a training database to train the AI model.
[0057] Depending on the embodiment, the predetermined model used can be a computational model or a mathematical model, that is to say a model that is not an artificial intelligence model.
[0058] In this case, the depth determination for at least one point in the scene can be performed sequentially, or simultaneously with at least one other point in the scene. Specifically, the depth determination for all points in the scene can be performed simultaneously or in parallel.
[0059] The non-AI model can be any type of model.
[0060] According to some embodiments, in the case where a non-AI model is used, the depth map estimation step may include a step of dividing the dual-pixel image into several zones, and in particular each sub-image composing the dual-pixel image into several zones.
[0061] At least one area can correspond to an object.
[0062] At least one area can correspond to a part of the scene, or to a part of an object.
[0063] The clipping step can be carried out using any known technique, such as for example by edge detection, by texture detection, based on visual properties of objects / part of the scene, etc.
[0064] Thus, the cropping step provides at least two areas for each sub-image composing the dual-pixel image. For example, when the dual-pixel image is formed from a left sub-image and a right sub-image, then each sub-image is cropped into several areas.
[0065] The estimation step may further include, for at least one, and in particular each, area, a step of searching for a difference in depth between the sub-images forming the dual-pixel image.
[0066] Depending on the embodiment, the depth difference search step can be carried out by correlation between the left image and the right image.
[0067] For example, correlation scores can be calculated between sub-images for several position difference values, Adi-Adm, with m>l. The position difference value Adi providing the highest correlation score, or greater than a given threshold, can be retained as corresponding to the position difference between the left and right sub-images, for the same area.
[0068] For a given area, knowing the difference in position between the sub-images, for example between the left sub-image and the right sub-image, it is possible to determine the depth of said area in the scene, optionally as a function of the focal length used to obtain the dual-pixel image.
[0069] Following an example implementation, for an image area Z and for a given position difference value Ad, the correlation score between two sub-images IMi and IM2 can be calculated as follows: - the value of the pixel at position (i,j) of the image area Z(IMi) is multiplied by the value of the pixel at position (i,j)+Ad in the image area Z(IMz); - this operation is performed for all pixels of the image area Z(IMi) and Z(IMz) respectively; - the values obtained are summed and the correlation score corresponds to said sum obtained.
[0070] Of course, this example is by no means limiting and other correlation techniques can be used.
[0071] It may be preferable, but by no means limiting, to consider the Ad in a direction defined by the arrangement of the dual-pixels. For example, for left and right dual-pixels, generating a shifted image projection in the direction of index i, without loss of generality, Ad is a displacement to be considered as a positive or negative shift depending on index i. In the case where there are several dual-pixels producing a shift in directions i and / or j, Ad is to be considered as a positive or negative shift depending on index i and / or j.
[0072] Depending on embodiments, the predetermined model used may be a combination of at least one AI model and at least one non-AI model, such as those described above.
[0073] According to another aspect of the same invention, a computer program is proposed comprising executable instructions which, when executed by a computer device, implement all the steps of the process according to the invention.
[0074] The computer program can be in any computer language, such as for example machine language, C, C++, JAVA, Python, etc.
[0075] Such a computer program can take the form of a standalone application. Alternatively, such a computer program can be integrated into a photo or video application, or even into an image or video playback application.
[0076] The computer program can be stored in a non-transient, or non-volatile, manner in a storage medium.
[0077] According to another aspect of the same invention, a device is proposed comprising means configured to implement all the steps of the process according to the invention.
[0078] The device according to the invention can be a computer, a processor, a computer chip, etc. programmed to implement the method according to the invention, for example by executing the computer program according to the invention.
[0079] The device according to the invention can be integrated into any type of device such as a smartphone, a tablet, a computer, a calculator, a processor, a computer chip, a medical imaging device, etc.
[0080] The device according to the invention may include, in terms of technical means, at least one, or any combination of at least two, of the characteristics described above with reference to the method according to the invention, and which are not repeated here exhaustively for the sake of brevity.
[0081] According to embodiments, the device according to the invention may further include an image sensor to obtain the dual-pixel image.
[0082] The image sensor can be an image sensor from a camera module equipping the device according to the invention. In this case, the image sensor is associated with an optical lens within a camera module.
[0083] Alternatively, the device according to the invention may not include an image sensor, particularly when the device according to the invention is integrated into a device that already includes an image sensor or a camera module. For example, when the device according to the invention is integrated into a smartphone or a camera, the latter is already equipped with an image sensor. In this case, it is not necessary for the device according to the invention to include an image sensor.
[0084] According to some embodiments, the image sensor may comprise a matrix of dual-pixels distributed within said image sensor.
[0085] For example, an image sensor can include patterns of photosites used to acquire the image of the scene and dual pixels inserted between these patterns, for example, at a given pitch. Such sensors are widely used in camera modules found in many devices, such as smartphones. In particular, the CMOS sensors used already include dual pixels inserted between the patterns of photosites used to image the scene.
[0086] Alternatively, the image sensor may include a pattern array of photosite patterns, at least some, or even all, photosite patterns incorporating dual-pixel functionality.
[0087] In this case, each photosite pattern with dual-pixel functionality provides a dataset used to generate the scene image and dual-pixel data representing the scene's depth at the point in the scene corresponding to that photosite pattern. It is then possible to obtain a higher-resolution scene depth map, since each photosite pattern provides data for determining depth.
[0088] According to another aspect of the same invention, an apparatus is proposed comprising a device according to the invention, or comprising means configured to implement the method according to the invention.
[0089] Such a device can be of any type.
[0090] In particular, the device according to the invention can be a user device such as a smartphone, tablet, etc.
[0091] The user device may also include a display screen.
[0092] The user device according to the invention may further include at least one camera module for imaging a scene.
[0093] In particular, the device according to the invention can be a computer-type user device.
[0094] The computer-type user device may also include a display screen.
[0095] The computer-type user device according to the invention may further include at least one camera module for imaging a scene.
[0096] In particular, the device according to the invention can be a television.
[0097] Television may also include a display screen.
[0098] The television according to the invention may further include at least one camera module for imaging a scene.
[0099] In particular, the device according to the invention may be a virtual reality headset or glasses or an augmented reality headset or glasses.
[0100] The helmet, or glasses respectively, according to the invention may further comprise at least one display screen.
[0101] The helmet, or glasses respectively, according to the invention may further include at least one camera module for imaging a scene.
[0102] In particular, the device according to the invention can be a medical imaging device.
[0103] In particular, the medical imaging device can be an endoscope, an ultrasound machine, etc.
[0104] The medical imaging device may also include at least one display screen.
[0105] The medical imaging device according to the invention may further include at least one camera module for imaging a scene.
[0106] Of course, the device according to the invention is not limited to the examples of devices that have just been given.
[0107] According to another aspect of the present invention, a vehicle is proposed comprising a device according to the invention, or comprising means configured to implement the method according to the invention.
[0108] The vehicle according to the invention may further include a display screen, for example arranged in a passenger compartment of the vehicle, or a projector for projecting at least one image onto a display surface generally known as a "head-up display".
[0109] The vehicle according to the invention may further include at least one camera module for imaging a scene.
[0110] According to embodiments, the vehicle according to the invention can be a land vehicle, such as a car, autonomous, semi-autonomous or non-autonomous. [YES] According to embodiments, the vehicle according to the invention can be a flying vehicle, such as a drone, an airplane, a helicopter, autonomous, semi-autonomous or non-autonomous.
[0112] According to embodiments, the vehicle according to the invention can be a maritime vehicle, such as a boat or a submarine, autonomous, semi-autonomous or non-autonomous. Description of the figures and methods of implementation
[0113] Other advantages and features will become apparent upon examination of the detailed description of non-limiting embodiments and the accompanying drawings, in which: - FIGURE 1 is a schematic representation of a non-limiting example of an embodiment of a process according to the invention; - FIGURE 2 is a schematic representation of a non-limiting example of an embodiment of a device according to the invention; - FIGURES 3 to 5 are schematic representations of non-limiting example embodiments of a device according to the invention; and - FIGURE 6 is a schematic representation of a non-limiting example embodiment of a vehicle according to the invention.
[0114] It is understood that the embodiments described below are by no means exhaustive. In particular, variants of the invention may be conceived comprising only a selection of the features described below, isolated from the other features described, if this selection of features is sufficient to confer a technical advantage or to differentiate the invention from the prior art. This selection includes at least one preferably functional feature without structural details, or with only a portion of the structural details if this portion alone is sufficient to confer a technical advantage or to differentiate the invention from the prior art.
[0115] In particular, all the variants and embodiments described can be combined with each other if there are no technical obstacles to this combination.
[0116] In the figures and in the rest of the description, elements common to several figures retain the same reference.
[0117] FIGURE 1 is a schematic representation of a non-limiting example of an embodiment of a method according to the present invention.
[0118] Method 100 of FIGURE 1 can be used to determine the depth map of a scene from a dual-pixel image of the scene, denoted IMDP hereafter, without loss of generality, i.e., an image comprising at least, or only, the data provided by the dual pixels of an image sensor used to image the scene.
[0119] The dual-pixel IMDP image can comprise a single image including the data provided by all photodiodes of all dual pixels of the image sensor.
[0120] Alternatively, the dual-pixel IMDP image can comprise multiple images, called sub-images, each sub-image containing the data provided by one photodiode of each dual-pixel of the image sensor. Following a non-limiting embodiment, where each dual-pixel comprises a left and a right photodiode, then the dual-pixel IMDP image can comprise a left sub-image, denoted IMDPG hereafter without loss of generality, and a right sub-image, denoted IMDPD hereafter without loss of generality. Following a more general formulation, where each dual-pixel comprises N photodiodes, then the dual-pixel IMDP image can comprise N sub-images, denoted IMDPI-IMDPN.
[0121] In the following, without loss of generality, to simplify the description of a non-limiting example of realization, we consider that the dual-pixel IMDP image includes a left image IMDPG and a right image IMDPD.
[0122] The process 100 may include an optional step 102 for acquiring an image of the scene with an image sensor comprising dual pixels.
[0123] The image sensor may contain dual pixels distributed between the photosite patterns used to image the scene. Such sensors are widely used in the camera modules found in most devices. The function of these dual pixels is to provide data to control an autofocus mechanism.
[0124] Alternatively, each photosite pattern of the image sensor can incorporate dual-pixel functionality.
[0125] In the following, without loss of generality, to simplify the description of a non-limiting example of implementation, we consider that the image sensor comprises a matrix of dual-pixels distributed in the image sensor between the patterns of photosites used to image the scene, each dual-pixel comprising a left photodiode and a right photodiode.
[0126] Alternatively, the dual-pixel image can be the raw image provided by the image sensor as captured in step 102.
[0127] Alternatively, the IMDP dual-pixel image can be formed from the data provided by the dual-pixels, i.e., the photodiodes forming the dual-pixels. In this case, the raw image provided by the image sensor in step 102 is processed in an optional step 104 to remove, from said image, the data provided by the photosite patterns used for imaging that do not have a dual-pixel function, and to keep only the data provided by the dual-pixels.
[0128] The process 100 provides, after step 106, the dual-pixel IMDP image by optional steps 102 and 104.
[0129] Alternatively, process 100 may not include optional steps 102 and 104. In this case, the dual-pixel IMDP image may be obtained beforehand, independently of process 100, and supplied to process 100.
[0130] Process 100 includes a step 106 taking as input the dual-pixel IMDP image.
[0131] Optionally, step 106 may also include consideration of the focal length, denoted DOS, used to capture the raw image in step 102, and more generally the focal length with which the dual-pixel IMDP image was obtained.
[0132] Optionally, step 106 may also include consideration of the lens-to-sensor distance (DOC) used to capture the raw image in step 102, and more generally, the lens-to-sensor distance (DOC) with which the dual-pixel IMDP image was obtained. This DOC is also referred to as the focal length in the relevant technical field.
[0133] During step 106, the dual-pixel IMDP image, and optionally the DOS distance and / or the DOC distance, is / are provided to a predetermined model to determine the depth map of the scene, denoted CP, without loss of generality.
[0134] This step 106 therefore provides a depth map of the scene, for several, and in particular each point, of the scene corresponding to a duel-pixel.
[0135] Optionally, process 100 may include a step 108 in which, for at least one point in the scene, a depth data is calculated by interpolating the depth data of one or more adjacent, or surrounding, points in the CP depth map.
[0136] For example, for at least one point, such an interpolation may include, or may be, the calculation of an average of the depth data of adjacent, or surrounding, points in the CP depth map.
[0137] It is possible to use different models during step 106 to determine the CP depth map.
[0138] Depending on the embodiment, the model used in step 106 may be a mathematical model, that is, a model that is not an artificial intelligence model. Such a model may, for example, take the form of a computer program and / or electronic components programmed to perform the estimation of the CP depth map.
[0139] In this case, step 106 may include a step to slice the dual-pixel image into several areas. Specifically, each IMDPG and IMDPD sub-image may be sliced into multiple areas.
[0140] At least one area can correspond to an object.
[0141] At least one area can correspond to a part of the scene, or to a part of an object.
[0142] The clipping step can be carried out using any known technique, such as for example by edge detection, by texture detection, based on visual properties of objects / part of the scene, etc.
[0143] Thus, the clipping step provides, for each IMDPG and IMDPD sub-image composing the dual-pixel image, at least two zones.
[0144] Step 106 of the depth map estimation may then include, for at least one, and in particular each, zone, a step to search for a depth difference between the IMDPG and IMDPD sub-images. This depth difference search may be carried out using any suitable technique.
[0145] Depending on the embodiment, the depth difference search step can be performed by correlation. For example, correlation scores can be calculated between sub-images for several depth difference values, Adi-Adm, with m>l. The depth difference value Adi providing the highest correlation score, or exceeding a predetermined threshold, can be retained as corresponding to the depth difference between sub-images for the same area.
[0146] For a given area, knowing the difference in depth between the IMDPG and IMDPD sub-images, it is possible to determine the depth of said area in the scene, possibly as a function of the DOS focusing distance used to obtain the dual-pixel IMDP image.
[0147] Following an example implementation, for an image area Z and for a given depth difference value Ad, the correlation score between the two sub-images IMDPG and IMDPD can be calculated as follows: - the value of the pixel at position (i,j) in the image area Z(IMDPG) is multiplied by the value of the pixel at position {(i,j)+Ad} in the image area Z(IMDPD); - this operation is performed for all pixels in the image area Z(IMDPG); - the values thus obtained for all pixels in the image area Z(IMDPG) are summed and the correlation score corresponds to said sum obtained. Of course, this example is by no means limiting and other correlation techniques can be used.
[0148] In the case described above where the depth map estimation is carried out using a mathematical model, the determination of the depth can be carried out in turn for each point of the scene or of an image area, or simultaneously for at least two, in particular all, points of the scene or of an image area.
[0149] In alternative embodiments, the model used in step 106 can be an artificial intelligence (AI) model, previously trained to determine the CP depth map as a function of the dual-pixel image, or sub-images forming the dual-pixel image, and optionally of the DOS distance and / or the DOC distance.
[0150] Such an artificial intelligence model can be any type of model, such as, for example, a previously trained neural network.
[0151] In this case, process 100 may optionally include a training phase 110 of said AI model with a training dataset comprising training sets. Each training set may include: - a dual-pixel training image of a training scene, - a depth map of said training scene, - optionally, the DOS focusing distance used to obtain said dual-pixel training image and / or the DOC lens-sensor distance used to obtain said dual-pixel training image.
[0152] For example, a training set can be obtained for a known training scene. Since the training scene is known, it is possible to determine the depth of each point in the scene, for example, by measurement with a sensor such as a LiDAR or a time-of-flight camera. Alternatively, the depth of each point in the training scene can be determined by construction; that is, the training scene is built based on a predetermined depth map. The resulting depth map of the training scene can then form the training depth map for a training set. A dual-pixel image of the training scene can be obtained with a camera module. This dual-pixel image can then form the dual-pixel training image for the training set.Optionally, the focusing distance and / or the sensor-lens distance used for acquiring the dual-pixel training image is / are known. Thus, the training set can be formed by: - the training depth map thus obtained, - the dual-pixel training image thus obtained, and - optionally by at least one of the DOS and DOC distances thus obtained. By repeating this operation for a multitude of training scenes, it is possible to build a training database comprising a multitude of training games to train the AI model.
[0153] The AI model can be trained using any known technique, for example, backpropagation of the error gradient. One part of the training dataset can be used to train the AI model, and a second part can be used to validate the AI model's training.
[0154] Such an AI model can be a single AI model. Alternatively, such an AI model can be a combination of several AI models.
[0155] According to yet another alternative, the predetermined model used in step 106 can be a combination of at least one mathematical model, i.e. non-AI, and at least one AI model.
[0156] Following a non-limiting example of implementation, the dual-pixel image can be cut into zones and each zone can then be given as input to an AI model to estimate the depth map of said zone.
[0157] FIGURE 2 is a schematic representation of a non-limiting example embodiment of a device according to the invention.
[0158] The device 200 of FIGURE 2 can be used to implement a method according to the invention, and in particular the method 100 of FIGURE 1.
[0159] Device 200 may include an optional module 202 for training the AI model. This module 202 can, for example, be configured / programmed to perform training phase 110 of FIGURE 1.
[0160] Device 200 may include an optional module 204 for obtaining a dual-pixel image from a raw image provided by an image sensor. This optional module 204 can, for example, be configured / programmed to perform step 104 of process 100 in Figure 1.
[0161] Device 200 includes a module 206 for running a model, denoted MOD, to provide a depth map from the dual-pixel IMDP image, given as input to said model, and optionally from a focal length DOS used to obtain the dual-pixel IMDP image and / or a focal length DOC used to obtain the dual-pixel IMDP image. This module 206 can, for example, be configured / programmed to perform step 106 of process 100 in FIGURE 1.
[0162] Device 200 may include an optional module 208 for calculating, by interpolation, the depth of a point in the scene from the depth of one or more adjacent or surrounding points in the depth map. This optional module 208 can, for example, be configured / programmed to perform step 108 of process 100 in Figure 1.
[0163] At least one of the modules 202-208 can be a module independent of the others.
[0164] At least two of the 202-208 modules can be integrated within the same module. In particular, the 202-208 modules can be integrated within a computing unit 210.
[0165] At least one of the 202-208 modules can be a hardware module, such as a processor, an electronic chip, etc.
[0166] At least one of the modules 202-208 may be a software module, such as a computer program.
[0167] At least one of the modules 202-208 may be a combination of at least one software module and at least one hardware module.
[0168] In particular, at least one of the 202-208 modules can be integrated into an electronic chip, or into an application installed in a user device.
[0169] In particular, the 210 computing unit can be, or can be integrated, into an electronic chip, or into an application installed in a user device.
[0170] Optionally, device 200 may also include at least one storage means 212. The storage means 212 is optional because it may already exist in a device incorporating device 200.
[0171] Optionally, device 200 may also include at least one camera module 214 for acquiring an image of the scene. Camera module 214 includes an optical lens associated with an image sensor comprising dual pixels, or dual-pixel functionality in at least some of the photosite patterns. The at least one camera module 214 is optional because it may already exist in a device incorporating device 200.
[0172] FIGURE 3 is a schematic representation of a non-limiting example embodiment of a device according to the present invention.
[0173] The apparatus 300 of FIGURE 3 includes means configured to implement the invention, and in particular the method 100 of FIGURE 1.
[0174] The device 300 of FIGURE 3 may include a device according to the invention, and in particular the device 200 of FIGURE 2.
[0175] In the example shown in FIGURE 3, device 300 is a smartphone, or a tablet, comprising device 200 from FIGURE 2.
[0176] Optionally, the device 300 can also include a display screen 302, optionally equipped with a sensing surface 304, for example capacitive.
[0177] Of course, the 300 device may include other components than those indicated above.
[0178] FIGURE 4 is a schematic representation of another non-limiting embodiment of a device according to the present invention.
[0179] The apparatus 400 of FIGURE 4 includes means configured to implement the invention, and in particular the method 100 of FIGURE 1.
[0180] The device 400 of FIGURE 4 may include a device according to the invention, and in particular the device 200 of FIGURE 2.
[0181] In the example shown in FIGURE 4, device 400 is: - a virtual reality (VR) headset or glasses, or - an augmented reality (AR) headset or glasses; including device 200 of FIGURE 2.
[0182] Optionally, the 400 device may also include a 402 display screen, for example in / on a visor of said 400 helmet.
[0183] Of course, the 400 helmet may include other components than those listed above.
[0184] FIGURE 5 is a schematic representation of a non-limiting example embodiment of a device according to the present invention.
[0185] The apparatus 500 of FIGURE 5 includes means configured to implement the invention, and in particular the method 100 of FIGURE 1.
[0186] The device 500 of FIGURE 5 may include a device according to the invention, and in particular the device 200 of FIGURE 2.
[0187] In the example shown in FIGURE 5, device 500 is a medical imaging device, such as an endoscope, an ultrasound machine, etc.
[0188] Optionally, the device 500 may also include a display screen 502, optionally equipped with a sensing surface 504, for example capacitive.
[0189] Of course, the 500 medical imaging device may include other organs than those indicated above.
[0190] FIGURE 6 is a schematic representation of a non-limiting example embodiment of a vehicle according to the present invention.
[0191] The vehicle 600 of FIGURE 6 includes means configured to implement the invention, and in particular the method 100 of FIGURE 1.
[0192] The vehicle 600 of FIGURE 6 may include a device according to the invention, and in particular the device 200 of FIGURE 2.
[0193] In the example shown in FIGURE 6, vehicle 600 is a land vehicle, in particular a car, comprising device 200 of FIGURE 2.
[0194] Optionally, the vehicle 600 may also include a display screen 602, optionally equipped with a sensing surface 604, for example capacitive, arranged in the passenger compartment of the vehicle 600.
[0195] Of course, the 600 vehicle may include other components than those indicated above.
[0196] Of course, the invention is not limited to the examples that have just been described.
Claims
DEMANDS 1. A method (100) for determining a depth map of a scene, said method (100) comprising the following steps: - reception of an image (IMDP), known as a dual-pixel image, comprising, and in particular consisting of, values provided by dual pixels of an image sensor, and -estimation (106) of a depth map of said scene with a predetermined model (MOD) taking said dual-pixel image (IMDP) as input.
2. A method (100) according to the preceding claim, characterized in that it comprises a step (104) of obtaining the dual-pixel image (IMDP) from a raw image provided by the image sensor, by selecting in said raw image and preserving the values provided by the dual-pixels of said image sensor.
3. Method (100) according to the preceding claim, characterized in that it comprises a step (102) of acquiring the raw image with the image sensor.
4. Method (100) according to any one of the preceding claims, characterized in that the predetermined model (MOD) further receives as input the focusing distance (DOS) used for dual-pixel image acquisition (IMDP).
5. Method (100) according to any one of the preceding claims, characterized in that the predetermined model (MOD) further receives the lens-to-sensor distance (DOC) from an imaging module comprising said image sensor associated with an optical lens and used for dual-pixel image acquisition (IMDP).
6. Method (100) according to any one of the preceding claims, characterized in that the model (MOD) used is a previously trained artificial intelligence (AI) model.
7. Method (100) according to the preceding claim, characterized in that the AI model (MOD) is a neural network, and in particular a previously trained convolutional neural network.
8. A method (100) according to any one of claims 6 or 7, characterized in that it further comprises a step (110) of training the AI model (MOD) with a training set comprising training sets, each training set comprising: - a dual-pixel image of a scene, and - a depth map of said scene.
9. Method (100) according to any one of claims 1 to 5, characterized in that the predetermined model (MOD) used is a mathematical calculation model.
10. Computer program comprising executable instructions which, when executed by a computing device, implement all the steps of the process (100) according to any one of the preceding claims.
11. Device (200) comprising means configured to carry out all the steps of the process (100) according to any one of claims 1 to 9.
12. Device (200) according to the preceding claim, characterized in that it comprises an image sensor for obtaining the dual-pixel image (IMDP).
13. Device (200) according to the preceding claim, characterized in that the image sensor comprises: - a matrix of dual-pixels distributed in said image sensor; or - a matrix of photosite patterns, at least some, and in particular all, photosite patterns incorporating a dual-pixel functionality.
14. Apparatus (300;400;500) comprising a device (200) according to any one of claims 11 to 13.
15. Vehicle (600) comprising a device (200) according to any one of claims 11 to 13.