Method and device for determining a depth map of a scene

By using multiple AI models and consolidating depth maps from diverse source images, the method improves the accuracy and efficiency of depth map estimation, addressing the limitations of single-model reliance and imaging condition sensitivity.

WO2025253055A1PCT designated stage Publication Date: 2025-12-11FOGALE OPTIQUE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/FR2024/050734
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-06
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing methods for estimating depth maps from 2D images using artificial intelligence models are inaccurate due to dependence on training datasets and sensitivity to imaging conditions, leading to errors in depth estimation.

Method used

The method involves generating multiple depth maps using different AI models from diverse source images and consolidating these maps using a second AI model or mathematical averaging, considering image parameters to reduce sensitivity to single-model inaccuracies.

Benefits of technology

This approach enhances the accuracy and efficiency of depth map estimation by reducing reliance on a single AI model and compensating for errors caused by varying imaging conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FR2024050734_11122025_PF_FP_ABST
    Figure FR2024050734_11122025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method (100) for determining a depth map of a scene, said method (100) comprising the following steps: - obtaining (104) a plurality of estimated depth maps (PCP1- PCPn) of said scene, referred to as first maps, each one obtained:  with a respective previously trained artificial intelligence, AI, model (PM1-PMn), referred to as the first model, and  taking as input a respective image of said scene, referred to as the source image (IM1- IMn); - and obtaining (108) a depth map, referred to as a consolidated depth map (CPC), from at least two of said first maps (PCP1-PCPn). The invention also relates to a computer program, a device, an apparatus and a vehicle implementing such a method.
Need to check novelty before this filing date? Find Prior Art

Description

DESCRIPTION Title: Method and device for determining a depth map of a scene

[0001] The present invention relates to a method for determining a depth map of a scene from at least one source image of said scene. It also relates to a computer program and a device implementing such a method.

[0002] The domain of the invention is the domain of obtaining a depth map of a scene from image(s) of said scene. State of the art

[0003] The depth map of a scene, when only two-dimensional (2D) image(s) are available, is important data. It can be used to perform various functions, for example, to image the scene using a focus bracketing technique, or to provide a 3D representation of the scene, etc., for example in a virtual reality (VR) or augmented reality (AV) headset.

[0004] We know of solutions for estimating the depth map of a scene from a 2D image of that scene, based on the use of artificial intelligence. In summary, a 2D image of the scene is fed into an artificial intelligence (AI) model, for example a convolutional neural network, previously trained on a training dataset: this AI model provides as output an estimated depth map of the scene.

[0005] These solutions represent a significant advancement, but they still have weaknesses. Their accuracy depends on the training dataset used, as the images processed when using the AI ​​model generally differ from those in the training dataset, leading to a decrease in the accuracy and performance of the AI ​​model.

[0006] Added to this is another difficulty related to the use of a 2D image of the scene. A 2D image provides limited information about the The scene is entirely dependent on the viewpoint used to image it, and more generally on the imaging conditions. These conditions can cause optical effects on the 2D image, leading to errors in depth map estimation, not to mention optical effects deliberately introduced into the imaged scene that are reflected in the 2D image.

[0007] One objective of the present invention is to remedy at least one of the aforementioned drawbacks.

[0008] Another objective of the invention is to provide a more accurate and efficient solution for estimating the depth map of a scene from at least one source image of said scene. Description of the invention

[0009] The invention proposes to achieve at least one of the aforementioned goals by a method for determining a depth map of a scene, said method comprising the following steps: - generation of several estimated depth maps of said scene, called first maps, each obtained: ■ with a respective artificial intelligence (AI) model, referred to as the first model, previously trained, and ■ taking as input an image, called the source image, respective to said scene; - obtaining a consolidated depth map from at least two of the aforementioned first maps.

[0010] Thus, the invention proposes to determine the depth map of a scene by using artificial intelligence models that take a source image as input to provide a depth map.

[0011] However, unlike current techniques, the invention proposes to use several first respective artificial intelligence models, each providing a respective depth map of the scene from a respective source image of said scene, and then to determine a consolidated depth map from said first depth maps. provided by the aforementioned first AI models. Thus, the invention enables a more efficient estimation of a scene's depth map from source images. The use of multiple AI models makes the determination of the scene's depth map less dependent on the sensitivity of a single AI model and compensates for the shortcomings of a single AI model, thereby increasing the performance and accuracy of the depth map estimation.

[0012] By image, we mean a digital image, and in particular a raster image, and more specifically an RGB raster image for example.

[0013] Each source image can be a 2D image.

[0014] By "depth of field" we mean the extent of the area of ​​sharpness that appears on an image, that is to say the area between the first sharp plane and the last sharp plane of the image.

[0015] By "extent of sharpness" we mean the distance over which the sharp part of the image extends, that is to say the distance from the first sharp plane to the last sharp plane of the image.

[0016] Focal length refers to the distance at which an optical lens is focused, relative to the lens's position. The focal length for capturing an image is typically adjusted by changing the distance between the image sensor and the optical lens. Thus, a first image of a scene acquired at a first focal length will clearly depict one part of the scene, and a second image of a scene acquired at a second focal length will clearly depict another part of the scene.

[0017] Following a non-limiting example, given to illustrate the definitions indicated above, the focusing distance can be adjusted to 15 meters. The depth of field can be 1.50 meters and the depth of field can be 14.50 meters to 16 meters. In this case, the image will sharply depict all objects, or parts of the scene, located between 14.5m and 16m from the optical lens.

[0018] By "angle of view", we mean the angular opening of the module allowing the image of the scene to be captured.

[0019] By "viewing angle", we mean two angles which allow us to characterize the direction in which the image is taken in relation to the scene.

[0020] The consolidated depth map can be obtained using any type of tool that takes at least two of the first maps as input.

[0021] Depending on the embodiment, the consolidated depth map can be obtained with a predefined mathematical relationship taking as input at least two of the first maps.

[0022] For example, the consolidated depth map can be a mathematical average of the first depth maps, for example at each point.

[0023] Optionally, it is possible to exclude one or more extreme depth maps to eliminate results that are furthest from other depth maps.

[0024] Depending on the embodiment, the consolidated depth map can be obtained with an artificial intelligence (AI) model, referred to as a second model, which has been previously trained: - taking at least two initial depth maps as input, and - providing the consolidated depth map as output.

[0025] Such a second AI model could, for example, be a neural network, in particular a deep learning neural network, in particular a convolutional neural network or a regression neural network.

[0026] In this case, depending on the embodiment, the second AI model may include: - as many neurons as there are points in each first depth map, or in the largest first depth map one wishes to obtain; and - each neuron has as many inputs as there are depth maps. Of course, the number of neurons input to the second AI model may differ from what is indicated above, for example depending on the definition of the input depth maps, the spatial extent with which each respective source image represents the scene, etc.

[0027] Alternatively, the second AI model can be a machine vector support, a decision tree, etc., or any other AI model.

[0028] The second AI model can take as input only the first depth maps and generate the consolidated depth map based on said first depth maps.

[0029] According to embodiments, the second AI model can take as input, for at least one, in particular each, first depth map, at least one image parameter, or a set of several image parameters, relative to the source image from which said first depth map was generated, in particular when at least two source images represent the scene differently.

[0030] In this case, the second AI model takes into account not only the first depth maps, but also the parameters of the source images used to generate said depth maps.

[0031] For example, an image parameter relative to an image may be relative to how said source image represents the scene.

[0032] Such an image parameter can be of any type.

[0033] Based on implementation examples, such an image parameter can be any one, or a combination of any one, of the following parameters: - focal distance with which the source image represents the scene; - field angle with which the source image represents the scene; - the shooting angle from which the source image represents the scene; - position from which the source image represents the scene; - brightness with which the source image represents the scene; - etc. In this case, for example, the second AI model takes the following data as input: {(PCPi ; PI^-PIi" 1 ), ..., (PCPn ; P -PIn 171 )} with : - PCPi, ..., PCPn: the first depth maps, with n>l; and - Plj^PIj 171 : the image parameters relating to the source image IMj with which the first depth map PCPj was obtained.

[0034] According to embodiments, the process according to the invention may further include a step of training the second AI model.

[0035] Such training can be carried out with a training database, called a second training database, comprising a multitude of training games, each training game including: - several initial depth maps, known as training maps, and - a consolidated depth map, known as a training map.

[0036] Each training game can be obtained in different ways.

[0037] Following an example implementation, a training set can be obtained from a scene, called a training scene. In this case, each initial training depth map can be obtained by inputting a source image of the training scene into a first AI model: thus, it is possible to obtain several initial training depth maps by inputting several source images of the scene into several first AI models. Furthermore, the consolidated training depth map can be determined: - either by designing / constructing said training scene, where - either by measurement with at least one measurement sensor, such as a LIDAR or a time-of-flight camera: By using a multitude of different training scenes, it is possible to obtain a multitude of training games comprising the second training database used to train the second AI model.

[0038] The second AI model can be trained with any known and suitable training algorithm. For example, the second AI model can be trained using a backpropagation algorithm. Of course, this example is only illustrative and is by no means exhaustive.

[0039] Of course, when image parameters are taken into account by the second AI model for the generation of the consolidated depth map, each training set may further include said image parameter(s), for each first training depth map.

[0040] Each initial AI model can be any type of AI model.

[0041] Following an example implementation, at least one, in particular each, first model can be a convolutional neural network.

[0042] Such neural networks are well known to those in the field, such as, for example: - the Marigold Depth Estimation model, see https: / / github.com / prs-eth / Marigold for more information; - the Depth Anything model, see https: / / depth-anything.github.io / for more information, - etc.

[0043] Depending on the embodiments, at least two, in particular all the first models can be identical.

[0044] For example, such identical models can be obtained by duplicating the same model.

[0045] Identical AI models can take as input identical or different respective source images

[0046] Depending on the embodiment, at least two, in particular all the first models, can be different

[0047] Different AI models can take identical or different respective source images as input.

[0048] Two different initial AI models can have different architectures. In this case, said different AI models can be trained with an identical training set, in particular with identical images, or with different training sets, in particular with different images.

[0049] Two different initial AI models can have the same architecture. In this case, the AI ​​models differ from each other in that they were trained with different training datasets.

[0050] According to embodiments, at least two, in particular of all the respective source images given as input to the respective first models may be identical.

[0051] Depending on the embodiment, at least two, in particular of all the respective source images given as input to the respective first models may be different.

[0052] In other words, each first depth map is obtained with a respective source image different from those used to obtain the other first depth maps.

[0053] At least two respective source images, used to generate two initial depth maps, may differ from each other in various ways.

[0054] Following examples of implementation, at least two source images may differ from each other in that they are acquired at two different focal distances.

[0055] Thus, these source images represent the scene with different areas of sharpness. Using source images with different focal lengths increases and varies the scene information taken into account when creating the depth map of said scene, thereby reducing errors due to optical effects that may be present in a given source image and that would be related to the use of a particular localization distance.

[0056] Alternatively, or in addition, at least two source images may differ from each other in that they are acquired with different field angles.

[0057] Thus, one of the source images is better defined than the other image, but the other source image provides a broader view of the scene. Using source images with different field angles increases and varies the scene information used when creating the scene depth map, thereby reducing errors due to optical effects that may be present in any given source image and that are related to the use of a particular field angle.

[0058] Following examples of implementation, at least two source images may differ from each other in that they are acquired from two different shooting angles.

[0059] Thus, the source images represent the scene from different shooting angles. Using source images acquired from different shooting angles increases and varies the information about the scene taken into account when creating the depth map of said scene, which reduces errors due to optical effects that may be present in a particular source image and that would be related to the use of a particular shooting angle.

[0060] Following examples of implementation, at least two respective source images may differ from each other in that they are acquired from two different positions.

[0061] Thus, the source images represent the scene from different positions. Using source images acquired from different positions increases and varies the information about the scene taken into account when creating the depth map of said scene, which reduces errors due to optical effects that may be present in a given source image and that would be related to the use of a particular acquisition position.

[0062] Following examples of implementation, at least two source images may differ from each other in that they are acquired with different brightness levels.

[0063] Thus, the source images represent the scene with different brightness levels for part or all of the scene. Using Using source images with different brightness levels increases and varies the information about the scene taken into account when establishing the depth map of said scene, particularly with regard to areas that could be overexposed or underexposed on one of the source images, which helps to reduce errors due to a brightness defect on one of the source images or to optical effects related to such brightness defects.

[0064] Depending on the embodiment, the consolidated depth map can have the same definition as the source images.

[0065] In this case, for each pixel of the source image(s) of the scene, the consolidated depth map indicates a depth value for the point in the scene to which each pixel corresponds.

[0066] If the source images do not all have the same resolution, it is possible to: - to keep the resolution of the image with the highest resolution, - to keep the resolution of the image with the lowest resolution, or - choose another resolution.

[0067] In this case, for each pixel of the source image(s) of the scene, the consolidated depth map indicates a depth value for the point in the scene represented by said pixel.

[0068] Depending on the embodiment, the consolidated depth map may have a lower resolution than the source image(s). This is because a depth map is generally used to determine the distance between the image sensor and each object in the scene. Therefore, in some cases, it may be sufficient for the consolidated depth map to indicate a depth for the different areas of the scene, each corresponding to an object in the scene: in this case, the consolidated depth map will have a lower resolution than the source images.

[0069] In all cases, the depth map can be indexed to points in the scene, to indicate

[0070] Depending on the embodiments, at least two, in particular all the first cards can be obtained at least partially simultaneously, or in parallel.

[0071] This speeds up the process of obtaining the consolidated depth map.

[0072] Depending on the embodiment, at least two, in particular all the first cards can be obtained at least partially in turn.

[0073] Depending on the implementation methods, the source images may have the same definition.

[0074] Alternatively, at least two source images of the scene may have different resolutions: in this case, the consolidated depth map may have the same resolution as the source image with the highest resolution, or the lowest resolution.

[0075] At least one source image can be an image of a real scene and acquired with a device equipped with at least one camera module.

[0076] At least one source image can be a computationally generated image.

[0077] According to embodiments, the process according to the invention may further include a spatial registration step of at least one source image.

[0078] This registration process eliminates any potential discrepancies between source images of the same scene. For example, the registration step eliminates differences in viewpoint, viewpoint, slight movements, etc., that may exist between the different source images.

[0079] The registration of a source image can be performed relative to another source image. In particular, one of the source images can be chosen as the reference image, and the other source image, and specifically each of the other source images, can be spatially registered relative to said reference image. Following a non-limiting example, The source image chosen as the reference image can be the one representing the scene with the widest field of view.

[0080] Alternatively, the registration of a source image can be performed against a reference format that is independent of the source images. In this case, the source image, and in particular each of the source images, can be registered against said reference format.

[0081] The registration of a source image can be performed using any known technique, for example, by using points of interest or objects located within the source image. Alternatively, or in addition, the registration of a source image can be performed using velocity or acceleration data associated with the pixels of the source image.

[0082] The registration step can be carried out before the generation of the first depth maps.

[0083] All source images can have the same size.

[0084] Alternatively, at least two source images can have different sizes, and in particular represent the scene from different angles. In this case, the smaller source image can be supplemented with predetermined values, especially constant ones, such as zero. Thus, there is no data loss.

[0085] Of course, other implementations are possible. For example, the larger source image can be reduced, for instance, by removing the data located at the periphery of the source image (this data is generally the least defined), to adjust its size to that of the smaller source image. Other implementation examples are possible.

[0086] All source images can have the same resolution.

[0087] Alternatively, at least two source images may have different definitions.

[0088] In this case, the least defined source image can be supplemented with predetermined values, specifically identical values ​​such as null or undefined, to adjust the aspect ratio of said less defined source image to the aspect ratio of the most defined source image. In particular, all images Sources can be adjusted to the format of the most defined source image.

[0089] Alternatively, the higher-resolution source image can be adjusted to the aspect ratio of the lower-resolution source image, for example, by removing some pixels or merging pixels within the higher-resolution source image. In particular, all source images can be adjusted to the aspect ratio of the lower-resolution source image.

[0090] Alternatively, all source images can be adjusted to another resolution, for example to a resolution equal to the average resolution of the source images.

[0091] According to another aspect of the same invention, a computer program is proposed comprising executable instructions which, when executed by a computer device, implement all the steps of the process according to the invention.

[0092] The computer program can be in any computer language, such as for example machine language, C, C++, JAVA, Python, etc.

[0093] Such a computer program can take the form of a standalone application. Alternatively, such a computer program can be integrated into a photo or video application, or even into an image or video playback application.

[0094] The computer program can be stored in a non-transient, or non-volatile, manner in a storage medium.

[0095] According to another aspect of the invention, a device is proposed comprising means configured to implement all the steps of the process according to the invention.

[0096] The device according to the invention can be a computer, a processor, a computer chip, etc. programmed to implement the method according to the invention, for example by executing the computer program according to the invention.

[0097] The device according to the invention can be integrated into any type of device such as a smartphone, a tablet, a computer, a calculator, a processor, a computer chip, etc.

[0098] The device according to the invention may include, in terms of technical means, at least one, or any combination of at least two, of the characteristics described above with reference to the method according to the invention, and which will not be repeated here exhaustively for the sake of brevity.

[0099] In particular, the system may include: - several initial modules, each running a respective initial AI model, to generate several initial depth maps of a scene, each with a respective source image, and - a second module to obtain a consolidated depth map of the scene from said first depth maps, for example with a mathematical relationship or with a second AI model executed by said second module.

[0100] According to another aspect of the invention, an apparatus is proposed, and in particular a user apparatus, comprising a device according to the invention.

[0101] In particular, the device can be a user device such as a smartphone, tablet, etc.

[0102] The user device may also include a display screen.

[0103] Alternatively, or in addition, the user device may further include at least one camera module to image a scene and obtain one or more source images of said scene.

[0104] In particular, the device can be a user device such as a computer.

[0105] The computer-type user device may also include a display screen.

[0106] Alternatively, or in addition, the computer-type user device may further include at least one camera module to image a scene and obtain one or more source images of said scene.

[0107] In particular, the device could be a television.

[0108] Television may also include a display screen.

[0109] Alternatively, or in addition, the television may further include at least one camera module to image a scene and obtain source images of said scene.

[0110] In particular, the device can be a virtual reality headset or glasses, or an augmented reality headset or glasses. [YES] The helmet, or glasses respectively, according to the invention may further comprise at least one display screen.

[0112] Alternatively, or in addition, the helmet, or glasses respectively, according to the invention may further comprise at least one camera module for imaging a scene and obtaining one or more source images of said scene.

[0113] In particular, the device may be a medical imaging device.

[0114] In particular, the medical imaging device can be an endoscope, an ultrasound machine, etc.

[0115] The medical imaging device may also include at least one display screen.

[0116] Alternatively, or in addition, the medical imaging device may further include at least one camera module to image a scene and obtain one or more source images of said scene.

[0117] Of course, the device according to the invention is not limited to the examples of devices that have just been given.

[0118] According to another aspect of the present invention, a vehicle comprising a device according to the invention is proposed.

[0119] The vehicle may further include a display screen, for example arranged in a passenger compartment of the vehicle, or a projector to project at least one image onto a display surface generally known as a "head-up display".

[0120] Alternatively, or in addition, the vehicle may further include at least one camera module to image a scene and obtain one or more source images of said scene.

[0121] Depending on the embodiment, the vehicle can be a land vehicle, such as a car, autonomous, semi-autonomous or non-autonomous.

[0122] Depending on the embodiment, the vehicle can be a flying vehicle, such as a drone, an airplane, a helicopter, autonomous, semi-autonomous or non-autonomous.

[0123] Depending on the embodiment, the vehicle can be a maritime vehicle, such as a boat or a submarine, autonomous, semi-autonomous or non-autonomous. Description of the figures and methods of realization

[0124] Other advantages and features will become apparent upon examination of the detailed description of non-limiting embodiments and the accompanying drawings, in which: - FIGURES aa and lb are schematic representations of a non-limiting example of an embodiment of a process according to the invention; - FIGURES 2a and 2b are schematic representations of another non-limiting example of an embodiment of a method according to the invention; - FIGURE 3 is a schematic representation of a non-limiting example of an embodiment of a device according to the invention; - Figures 4a-4c are schematic representations of non-limiting examples of embodiments of a device according to the invention; and - FIGURE 5 is a schematic representation of a non-limiting example embodiment of a vehicle according to the invention.

[0125] It is understood that the embodiments described below are by no means exhaustive. In particular, variants of the invention may be conceived comprising only a selection of the features described below, isolated from the other features described, if this selection of features is sufficient to confer a technical advantage or to differentiate the invention from the prior art. This selection includes at least one preferably functional feature without structural details, or with only a portion of the structural details if this portion alone is sufficient to confer a technical advantage or to differentiate the invention from the prior art.

[0126] In particular, all the variants and embodiments described can be combined with each other if there are no technical obstacles to this combination.

[0127] In the figures and in the rest of the description, elements common to several figures retain the same reference.

[0128] FIGURES aa and lb are schematic representations of a non-limiting example embodiment of a method according to the present invention.

[0129] The method 100 of FIGURES 1a and 1b can be used to estimate the depth map of a scene using N AI models, denoted PMi-PM n , with N>1, called first models in the following, each model taking as input an image, called the source image, of the respective scene, denoted IMi-IM n .

[0130] At least one, in particular each, source image can be a 2D image of the scene.

[0131] Each first model (PMi) can be any type of AI model, and in particular a neural network. Specifically, at least one, and in particular each, first model (PMi) can be a neural network. For example, at least one, and in particular each, first model (PMi) can be the Marigold Depth Estimation model, the Depth Anything model, etc.

[0132] Depending on embodiments, at least two of the, in particular all, first PMi-PM models n may be identical. In this case, the first identical models can be obtained by duplicating the same AI model, called the source model.

[0133] Alternatively, at least two, especially all, of the first models can be different. Two different first models can have: - different architectures: in this case, the aforementioned first different models can be trained with an identical training dataset, in particular with identical images, or with different training datasets, in particular with different images; or - the same architecture: in this case, the said models differ from each other in that they have been trained with different training bases so that they have different weights or coefficients.

[0134] As mentioned above, each first PMi model takes as input a respective source image IMi of the scene.

[0135] At least two of them, specifically all of them, IMi-IM images n They can be identical, that is, represent the scene in the same way. In this case, it is preferable to use different first models.

[0136] Alternatively, at least one of the IMi-IM images nmay represent the scene differently from another IMi-IM image n In particular, each of the IMi-IM images n can represent the scene differently from other IMi-IMn images. For example, at least two source images differ from each other in that they are acquired: - at different focusing distances, - with different field angles - with different viewing angles, - from two different positions, - with different lighting for the scene, - etc. Thus, each of these source images provides different information about the scene. In this case, it is possible to use identical or different initial models.

[0137] In what follows, and without loss of generality, we consider that: - the first models are identical, but their respective source images are different; or - the first models are different and the respective source images may be identical or different. Thus, a first depth map, PCPi, generated by a first model PMi taking as input a respective image IMi will be different from another first depth map, PCPj, generated by another first model PMj taking as input a respective image IMj because either the first models PMi and PMj are different, or the images IMi and IMj are different.

[0138] The process 100 includes a phase 102 of estimating the depth map of the scene from said source images of the scene.

[0139] The estimation phase 102 includes a step 104 which determines a first depth map, PCPi, using a first model PMi that takes as input a first source image IMi of the scene. This first depth map PCPi is stored in a step 106.

[0140] Steps 104 and 106 are repeated, as many times as there are first models, i.e. N times, to generate N first depth maps, PCPi-PCPn, each with a respective first model and taking as input a respective source image.

[0141] At least two of the, and in particular all, of the first PCPi-PCPn depth charts can be generated simultaneously, or at least in parallel. In this case, for each depth chart, steps 104 and 106 are performed in parallel with the other depth charts. Alternatively, at least two of the, and in particular all, of the first PCPi-PCPn depth charts can be generated sequentially. In this case, for each first depth chart, steps 104 and 106 are performed sequentially.

[0142] Phase 102 of the depth map estimation further includes a step 108 generating a consolidated depth map, denoted CPC, from at least two of the aforementioned, in particular from all the first PCPi-PCPn depth maps previously generated by the first models

[0143] The consolidated CPC depth map can be generated in different ways.

[0144] In some embodiments, the consolidated CPC depth map can be generated using a predetermined mathematical relationship. For example, the consolidated CPC depth map can be obtained by averaging the initial PCPi-PCPn depth maps, specifically point by point. In another example, the consolidated CPC depth map can be obtained by averaging, for each point, the extreme depth values, {min;max}, obtained for all the initial PCPi-PCPn depth maps. In yet another embodiment, the consolidated CPC depth map can be obtained by averaging the initial PCPi-PCPn depth maps, specifically point by point, after excluding the extreme values, {min;max}, for each point. These embodiments are by no means exhaustive.

[0145] Depending on embodiments, the consolidated depth map CPC can be generated using a second artificial intelligence model, AI, denoted DM, previously trained.

[0146] This second AI model can be any type of suitable AI model.

[0147] Following one embodiment example, the second AI model can be a convolutional neural network pre-trained to provide the consolidated CPC depth map as a function of at least two of, and in particular all of, the first PCPi-PCPn depth maps.

[0148] Optionally, the second AI model can also take as input, for at least one and in particular each, first depth map, at least one image parameter relating to the source image used to generate said first depth map. This image parameter can be related to how the scene is represented in said source image, for example, the focal length of the source image, the angle of view of the source image, the angle of view of the source image, the position in which the source image represents the scene, the direction in which the source image represents the scene, the brightness of the image, etc.

[0149] As mentioned above, at least two of, and in particular all of, the source images used to generate the initial depth maps may be identical or different.

[0150] For example, at least two source images may differ from each other in that they are acquired: - at two different focusing distances, - with different field angles, - with different viewing angles, - from two different positions, and / or - with different lighting levels for the scene. Thus, each of these source images provides different information about the scene, which allows for the generation of initial depth maps that are potentially different as well.

[0151] When the consolidated CPC depth map is generated by a second AI model, in particular by a convolutional neural network, the process 100 may optionally include a phase 110 of training said second model.

[0152] The training phase 110 of the second model can be carried out just before the scene depth map estimation phase 102, so that the estimation phase 102 is carried out immediately after the training phase 110.

[0153] Alternatively, training phase 110 can be carried out well before scene depth map estimation phase 102, so that estimation phase 102 is not carried out immediately after training phase 110 and a non-negligible time lapse occurs between training phase 110 and estimation phase 102.

[0154] The training phase 110 can be common to several iterations of the estimation phase 102, so that the second model trained during the training phase 110 is used during several iterations of the estimation phase 102, for the same scene or for different scenes. Alternatively, the training phase 110 can be specific to an estimation phase and be repeated before each estimation phase 102, for example to update the second AI model, depending on such and such a type of scene for example.

[0155] When the second model is a convolutional neural network, the training phase 110 can be carried out in all possible ways.

[0156] In particular, training can be carried out using a training database, called a second training database, comprising a multitude of training games. Each training game may include: - several initial depth maps, known as training maps, labeled PCPEi-PCPEn; - a consolidated depth chart, known as a training chart, denoted CPCE; and - optionally, for each first PCPi depth map, one or more image parameters relating to the source image with which said first depth map was generated.

[0157] Each training game can be obtained in different ways.

[0158] Following an example implementation, a training set can be obtained from a known training scene. In this case, each first training depth map (PCPEj) can be obtained by inputting a respective source image of the training scene into a respective first model. The consolidated training depth map (CPCE) can be known: - either by design / construction of said training scene, which, we recall, is known; - either by measurement with at least one measurement sensor, such as a LIDAR or a time-of-flight camera. By using a multitude of different training scenes, it is possible to obtain a multitude of training sets for the second model, each training set including: - several first training depth maps, PCPEi-PCPEn, each obtained by entering a source image of the training scene into a respective first model; - a consolidated training depth map, CPCE, obtained either by design or by measurement; and - optionally, for each first PCPEi training depth map, one or more image parameters relating to the source image with which said first depth map was generated.

[0159] The second AI model can be trained with any known and suitable training algorithm. For example, the second AI model can be trained using a backpropagation algorithm. Of course, this example is only illustrative and is by no means exhaustive.

[0160] The process according to the invention may optionally include a training phase 120 of at least and in particular of each, first AI model, carried out according to any known method.

[0161] In particular, training one, or each, initial AI model can be done with a training dataset, called the first training dataset, comprising a multitude of training sets, each training set including: - a source image for training, of a scene, and - a training depth map, of said scene.

[0162] Each training game can be obtained in different ways.

[0163] Following an example implementation, a training set can be obtained from a known scene, called the training scene. In this case, a source image of the training scene is acquired: this source image constitutes the training source image that makes up a training set. Furthermore, the depth map of the training scene can be determined. - either by design / construction of said training scene, which, we recall, is known; - either by measurement with at least one measurement sensor, such as a LIDAR or a time-of-flight camera. This depth map can serve as the training depth map for the training set of the, or each, first model.

[0164] By using a multitude of different training scenes, it is possible to obtain a multitude of training games for the first AI model, each training game including: - a source image of a training scene; and - a depth map of said training scene.

[0165] The first model, or each of them, can be trained with any known and suitable training algorithm. For example, the first model (or each of them) can be trained using a backpropagation algorithm. Of course, this example is only illustrative and is by no means exhaustive.

[0166] Optionally, process 100 may include an optional step 132 of spatial registration of at least one source image, in particular when said source image represents the scene differently from at least one other source image.

[0167] In particular, the registration step 132 allows for the spatial registration of all source images.

[0168] This registration process eliminates discrepancies between source images representing the same scene. For example, the registration step eliminates differences in viewpoint, shooting angle, slight movements, etc., that may exist between different source images of the same scene.

[0169] In the example of realization of FIGURE 1, in no way limitingly, at least one, in particular each, source image is spatially recalibrated with respect to the source image representing the scene with the largest field of view.

[0170] Registration can be achieved by detecting the same objects of interest in the source images and using their position to register the source images to each other.

[0171] The process 100 may further include an optional step 134 of adjusting at least one source image, in particular when said source image represents the scene differently from at least one other source image.

[0172] In particular, adjustment step 104 can adjust the size of at least one source image, especially when the source images are of different sizes. In the example shown in Figure 1, the size of the source images is adjusted, without limitation, to match the size of the largest source image, for example, the source image representing the scene with the widest field of view. In this case, each smaller source image (IMi) is padded with predetermined values, specifically constant values ​​such as zero or undefined values, to make it the same size as the largest source image.

[0173] Alternatively, or in addition, step 104 can adjust the resolution of the source images, particularly when the source images have different resolutions. In the example shown in Figure 1, the resolution of the source images is adjusted, without limitation, to match the resolution of the highest-resolution source image, for example, the image representing the scene at the greatest zoom level. In this case, each lower-resolution source image (IMj) is supplemented with predetermined, or constant, values, such as zero or undefined values, so that it has the same resolution as the highest-resolution source image.

[0174] FIGURES 2a and 2b are schematic representations of another non-limiting embodiment of a method according to the present invention.

[0175] The method 200 of FIGURES 1a and 1b can be used to estimate the depth map of a scene using N AI models, denoted PMi-PM n , with N>1, called first models in the following.

[0176] Process 200 includes all the steps of process 100 in FIGURES 1a and 1b, except for the differences below.

[0177] Process 200 differs from process 100 in that each of the first PMi-PM models n takes as input the same source image of the scene, denoted IM. In other words, the source image of the scene given as input to a first PMi model is identical to that given as input to each of the other first PMi-PM models. n Thus, the 200 method uses a single source image of the scene to estimate the scene's depth map.

[0178] For these reasons, process 200 does not include source image registration steps 132 and source image adjustment steps 134 of process 100.

[0179] In process 200 of FIGURES 2a and 2b, it is preferable to use early PMi-PM models n which are different from each other. Of course, alternatively it is possible to use first models which are identical in process 200, particularly for the sake of redundancy.

[0180] Based on examples of implementation not shown, it is possible to generate a depth map using: - the same source image for at least two of the first PMi-PMn models. In this case, the said first models taking said image as input may preferably be different; - different source images for at least two of the early PMi-PM models nIn this case, the first models taking the images as input may be identical or different.

[0181] FIGURE 3 is a schematic representation of a non-limiting example embodiment of a device according to the present invention.

[0182] The 300 device in FIGURE 3 can be used to obtain, by estimation, a depth map of a scene from: - of a single image of the scene; or - of several images of the scene, representing the scene in different ways from each other.

[0183] The device 300 of FIGURE 3 can be used to implement a method according to the invention, and in particular any one of the methods 100 or 200.

[0184] The 300 device may include an optional 302 module to drive at least one, in particular each, first PMi-PM model nThis module 302 is, for example, configured / programmed to perform training phase 120.

[0185] Device 300 can include an optional module 304 to train the second AI model, DM. This module 304 is, for example, configured / programmed to perform training phase 110.

[0186] Device 300 includes a 306 module, running the early PMi-PM models n Simultaneously or sequentially, to provide, with each first PMi model, a first PCPi depth map with a respective IMi source image given as input to said first PMi model. This module 306 is, for example, configured / programmed to perform step 104 of processes 100 and 200.

[0187] Device 300 includes module 308 to provide a consolidated depth map from several initial depth maps. For example, module 308 can execute the second AI DM model. This module 308 is, for instance, configured / programmed to perform step 108 of processes 100 and 200.

[0188] Optionally, device 300 can include a module 310 for spatially registering at least one source image. This module 308 is, for example, configured / programmed to perform step 132 of process 100.

[0189] Optionally, device 300 can include a module 312 to adjust at least one source image. This module 312 is, for example, configured / programmed to perform step 132 of process 100.

[0190] At least one of the 302-312 modules can be a module independent of the others.

[0191] At least two of the 302-312 modules can be integrated within the same module. In particular, the 302-312 modules can be integrated within a 320 computing unit.

[0192] At least one of the 302-312 modules can be a hardware module, such as a processor, an electronic chip, etc.

[0193] At least one of the 302-312 modules can be a software module, such as a computer program.

[0194] At least one of the 302-312 modules can be a combination of at least one software module and at least one hardware module.

[0195] In particular, at least one of the 302-312 modules can be integrated into an electronic chip, or into an application installed in a user device.

[0196] In particular, the 320 computing unit can be, or can be integrated, into an electronic chip, or into an application installed in a user device.

[0197] Device 300 may also optionally include storage means 316.

[0198] FIGURE 4a is a schematic representation of a non-limiting example embodiment of a device according to the present invention.

[0199] The apparatus 400 of FIGURE 4a includes means configured to implement the invention, and in particular any one of the methods 100 or 200.

[0200] The device 400 may include a device according to the invention, and in particular the device 300 of FIGURE 3.

[0201] In the example shown in FIGURE 4a, device 400 is a smartphone, or a tablet, comprising device 300 from FIGURE 3.

[0202] Optionally, the 400 unit may also include: - a 402 camera module, in particular for taking one or more source images of the scene which can be used to estimate the depth map of said scene, according to the invention; and - a display screen 404, optionally equipped with a detection surface 406, for example capacitive.

[0203] Of course, the 400 device may include other components than those indicated above.

[0204] FIGURE 4b is a schematic representation of another non-limiting embodiment of a device according to the present invention.

[0205] The apparatus 410 of FIGURE 4b includes means configured to implement the invention, and in particular any one of the methods 100 or 200.

[0206] The device 410 may include a device according to the invention, and in particular the device 300 of FIGURE 3.

[0207] In the example shown in FIGURE 4b, device 410 is: - a virtual reality (VR) headset or glasses, or - an augmented reality headset, or glasses; including device 300 of FIGURE 3.

[0208] Optionally, the 410 unit may also include: - a camera module 412, in particular for taking one or more source images of the scene which can be used to estimate the depth map of said scene, according to the invention; and - a display screen 414, for example in / on a visor of said helmet 410.

[0209] Of course, the 410 helmet may include other components than those listed above.

[0210] FIGURE 4c is a schematic representation of a non-limiting example embodiment of a device according to the present invention.

[0211] The apparatus 420 of FIGURE 4c includes means configured to implement the invention, and in particular any one of the methods 100 or 200.

[0212] The apparatus 420 of FIGURE 4c may include a device according to the invention, and in particular the device 300 of FIGURE 3.

[0213] In the example shown in FIGURE 4c, device 420 is a medical imaging device, such as an endoscope, an ultrasound machine, etc.

[0214] Optionally, the 420 unit may also include: - a 422 camera module, in particular to take one or more source images of the scene which can be used to estimate the depth map of said scene, according to the invention; - a display screen 424, optionally equipped with a detection surface 426, for example capacitive.

[0215] Of course, the 420 medical imaging device may include other organs than those indicated above.

[0216] FIGURE 5 is a schematic representation of a non-limiting example embodiment of a vehicle according to the present invention.

[0217] The vehicle 500 of FIGURE 5 includes means configured to implement the invention, and in particular any one of the methods 100 or 200.

[0218] The vehicle 500 of FIGURE 5 may include a device according to the invention, and in particular the device 300 of FIGURE 3.

[0219] In the example shown in FIGURE 5, vehicle 500 is a land vehicle, in particular a car, comprising device 300 of FIGURE 3.

[0220] Optionally, the 500 vehicle can also include: - a 502 camera module, in particular to take one or more source images of the scene which can be used to estimate the depth map of said scene, according to the invention; - a display screen 504, optionally equipped with a sensing surface 506, for example capacitive, arranged in the passenger compartment of the vehicle 500.

[0221] Of course, the 500 vehicle may include other components than those indicated above.

[0222] Of course, the invention is not limited to the examples that have just been described.

Claims

DEMANDS 1. A method (100;200) for determining a depth map of a scene, said method (100;200) comprising the following steps: - generation (104) of several estimated depth maps (PCPi- PCPn) of said scene, called first maps, each obtained: ■ with a respective artificial intelligence, AI, model (PMi-PM n ), said first model, previously trained, and ■ taking as input an image, called the source image, respective (IM i- IM n ;IM) of said scene; - obtaining (108) a consolidated depth map (CPC) from at least two of the said first maps (PCPi-PCPn).

2. Method (100;200) according to the preceding claim, characterized in that the consolidated depth map (CPC) is obtained with an artificial intelligence (DM) model, AI, referred to as the second model, in particular a neural network, previously trained: - taking input at least two initial maps (PCPi-PCPn), and - providing output the consolidated depth map (CPC).

3. Method (100;200) according to the preceding claim, characterized in that it further comprises a step (110) of training the second AI model (DM).

4. A method (100;200) according to any one of the preceding claims, characterized in that at least two first models (PMi-PM n ) are identical.

5. A method (100;200) according to any one of the preceding claims, characterized in that at least two first models (PMi-PM n ) are different.

6. A method (100;200) according to any one of the preceding claims, characterized in that at least two respective source images (IM) are given as input to respective first models (PMi-PM n ) are identical.

7. A method (100;200) according to any one of the preceding claims, characterized in that at least two respective source images (PMi-PM n ) input data of the respective first models are different.

8. A method (100;200) according to the preceding claim, characterized in that at least one source image differs from at least one, or each, of the other source images (IMi-IM n ) Regarding : - the focusing distance, - the field angle, - the angle of view, - the imaging position, and / or - the lighting for the scene.

9. Method (100;200) according to any one of the preceding claims, characterized in that the consolidated depth map (CPC) has the same definition as the source images.

10. Method (100;200) according to any one of the preceding claims, characterized in that at least two first cards (PCPi- PCPn) are obtained at least partially simultaneously, or in parallel.

11. Method (100;200) according to any one of the preceding claims, characterized in that at least two first cards (PCPi- PCPn) are obtained in turn.

12. Computer program comprising executable instructions which, when executed by a computing device, implement all the steps of the process (100;200) according to any one of the preceding claims.

13. Device (300) comprising means configured to carry out all the steps of the process (100;200) according to any one of claims 1 to 11.

14. Apparatus (400;410;420) comprising a device (300) according to the preceding claim.

15. Vehicle (500) comprising a device (300) according to claim 13.

Citation Information

Patent Citations

  • Method and device for depth image completion

    CN112001914B

  • Diner monitoring method based on mutual attention neural network

    CN112418160A

  • Generation of full-scale 3D models from 2d images produced by a single-eye imaging device

    WO2021245290A1