Method and device for determining a depth map of a scene from images of the scene

By using multiple source images with associated parameters, the method enhances AI-based depth map estimation accuracy by addressing limitations in existing 2D image-based solutions, improving depth map estimation through diverse scene information.

WO2026052904A1PCT designated stage Publication Date: 2026-03-12FOGALE OPTIQUE
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-03
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing AI-based methods for estimating depth maps from 2D images are inaccurate due to reliance on training datasets and limited scene information, leading to errors from optical effects and variations in image acquisition conditions.

Method used

Utilize multiple source images of a scene, each with associated image parameters such as focusing distance, field of view, and optical transfer function, to enhance the AI model's understanding of the scene, thereby improving depth map estimation accuracy.

Benefits of technology

The method provides a more accurate and efficient estimation of depth maps by leveraging diverse scene information from multiple images, reducing errors associated with optical effects and variations in image acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FR2024051156_12032026_PF_FP_ABST
    Figure FR2024051156_12032026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method (100) for determining a depth map (CP) of a scene, the method (100) comprising a step (102) of estimating the depth map (CP) of the scene by means of a previously trained artificial intelligence model (MIA), AI model, taking as input: - a plurality of images (IM1-IMn) which are referred to as source images and are different from the scene, each source image (IMi) representing the scene differently to the other source images; and - for each source image, at least one image parameter relating to the source image (IM1-IMn). The invention also relates to a computer program, to a device, to an apparatus and to a vehicle implementing such a method.
Need to check novelty before this filing date? Find Prior Art

Description

DESCRIPTION Title: Method and apparatus for determining a depth map of a scene from images of said scene

[0001] The present invention relates to a method for determining a depth map of a scene from several images, called source images, of said scene. It also relates to a computer program and a device implementing such a method.

[0002] The domain of the invention is the domain of obtaining a depth map of a scene from several images of said scene. State of the art

[0003] The depth map of a scene, when only two-dimensional (2D) image(s) are available, is important data. It can be used to perform various functions, for example, to image the scene using a focus bracketing technique, or to provide a representation of the scene with a 3D effect, etc., particularly in a virtual reality (VR) or augmented reality (AV) headset.

[0004] Solutions based on artificial intelligence (AI) exist for estimating the depth map of a scene from a 2D image of that scene. In short, a 2D image of the scene is fed into an AI model, such as a convolutional neural network (CNN), which has been previously trained on a training dataset. This AI model then outputs an estimated depth map of the scene. Examples of known models available to professionals include the Marigold Depth Estimation Neural Network and the Depth Anything Neural Network, both of which can generate a depth map of a scene from a 2D image.

[0005] These solutions represent a significant advance, but still have drawbacks. In particular, the accuracy of these solutions depends on the training dataset used because the images processed during use AI model data typically differ from those of the training dataset, leading to a decrease in the accuracy and performance of the AI ​​model when used during the inference phase.

[0006] Added to this is another, more significant difficulty related to the use of a 2D image of the scene. A 2D image provides limited information about the scene, entirely dependent on the viewpoint used to image it, and more generally on the image acquisition conditions. These conditions can cause optical effects on the 2D image, leading to errors in depth map estimation, not to mention optical effects deliberately introduced into the imaged scene that are then reflected in the 2D image provided as input to the AI ​​model.

[0007] One objective of the present invention is to remedy at least one of the aforementioned drawbacks.

[0008] Another objective of the invention is to provide a more accurate and efficient solution for estimating the depth map of a scene from images of said scene. Description of the invention

[0009] The invention proposes to achieve at least one of the aforementioned goals by a method for determining a depth map of a scene, said method comprising a step of estimating said depth map of said scene by an artificial intelligence model, AI model, previously trained, taking as input: - several images (IMi-IM) n ), said source images, different from said scene, each source image (IMi) representing the scene differently from the other source images; and - for each source image, at least one image parameter relative to said source image.

[0010] Thus, like current solutions, the invention proposes to determine the depth map of a scene using an AI model.

[0011] Unlike current solutions, the invention proposes using several different source images of the scene, rather than a single image. of the scene. Furthermore, the invention proposes to provide the AI ​​model with at least one image parameter for each source image, and in particular an image parameter providing information on how the source image represents the scene or how it was acquired. Thus, the invention enables a more efficient and accurate estimation of the depth map of a scene since the AI ​​model has more information than just one or more images of the scene.

[0012] Indeed, using multiple images of the scene provides more information about it, leading to a more accurate estimation of the depth map. For example, using several different images of the scene makes the depth map estimation less susceptible to optical effects, whether intentionally introduced into the scene or not, and present in one of the source images.

[0013] Furthermore, providing at least one image parameter relating to how each source image is acquired, or represents the scene, allows the AI ​​model to have even more information, going beyond just the information available in the source images, which improves the performance and accuracy of the AI ​​model for estimating the scene's depth map.

[0014] By image, or source image, we mean a digital image, and in particular a raster image, and more specifically an RGB raster image for example.

[0015] At least one source image can be a 2D image.

[0016] By "depth of field" we mean the extent of the area of ​​sharpness that appears on an image, that is to say the area between the first sharp plane and the last sharp plane of the image.

[0017] By "extent of sharpness" we mean the distance over which the sharp part of the image extends, that is to say the distance from the first sharp plane to the last sharp plane of the image.

[0018] Focusing distance refers to the distance at which an optical lens is focused, relative to the lens's position. The focusing distance for capturing an image is generally adjusted by changing the distance between the image sensor and the optical lens. Thus, a first image of a scene acquired at a first focal distance will clearly represent a first part of the scene, and a second image of a scene acquired at a second focal distance will clearly represent a second part of the scene.

[0019] Following a non-limiting example, given to illustrate the definitions indicated above, the focusing distance can be adjusted to 15 meters. The depth of field can be 1.50 meters and the depth of field can be 14.50 meters to 16 meters. In this case, the image will sharply depict all objects, or parts of the scene, located between 14.5m and 16m from the optical lens.

[0020] By "angle of view", we mean the angular opening of the imaging module allowing the image of the scene to be captured.

[0021] By "angle of view", or "angle of shot", we mean two angles which allow us to characterize the direction in which the image is taken in relation to the scene.

[0022] By "angle of rotation", we mean the angle of rotation of the image capture module around its optical axis.

[0023] At least one image parameter relating to a source image can be any type of parameter that provides more information about how said source image was acquired or represents the scene.

[0024] It may be relative to at least one condition in which said source image of the scene was acquired / obtained.

[0025] It may be relative to at least one condition in which said source image of the scene represents the scene.

[0026] It can relate to at least one imaging device, or to a camera module, or at least one optical lens, used to acquire the image of the scene.

[0027] According to a particularly advantageous characteristic, for at least one source image, at least one image parameter can be, or can understand, an optical transfer function, OTF, of the optical lens, or camera module, used for the acquisition of said source image.

[0028] This FTO indicates the optical behavior of the optical lens, or camera module, used to image the scene and obtain the source image. This FTO is therefore particularly useful for determining the depth map of the scene.

[0029] The FTO can be the PSF, for "Point Spread Function" in English, or "Function d'Etalement du Point" in French.

[0030] FTO can be MTF for "Modulation Transfer Function" in English, or "Fixture de Modulation de Transfert" in French.

[0031] It should be noted that the FTO can be a constant function for each point in the scene. In this case, a single function can be provided as an image parameter. The FTO can then be represented, for each point of the optical lens, by a matrix of values, for example, 10x10 values, centered on that point of the lens. Thus, for the entire optical lens, the FTO is formed by a collection of matrices, each matrix corresponding to a point of that optical lens.

[0032] The FTO can depend, at each point in the scene, on the distance between the optical lens and that point in the scene, that is, on the scene depth: thus, using the FTO of the optical lens, or the camera module, provides important information in determining the scene depth map. In this case, for each point of the optical lens, the FTO is formed by a set of several functions: - each function corresponding to a distance between the optical lens and the scene; - each function being represented by a matrix of values, for example a matrix of 10x10 values. For example, if there are potentially 100 distances, for each point of the optical objective, the FTO can be represented by a set of 100 functions, each function being, for example, a matrix of values.

[0033] The inventors also noted that the FTO can depend on the distance between the optical lens and the image sensor: thus, using the FTO of the optical lens, or the camera module, provides information important in determining the depth map of the scene. In this case, for each point of the optical lens, the FTO is formed by a set of several functions: - each function corresponding to a distance between the optical lens and the sensor; - each function being represented by a matrix of values, for example a matrix of 10x10 values.

[0034] For example, if there are potentially 10 distances between the optical lens and the sensor, for each point of the optical lens, the FTO can be represented by a set of 10 functions, each function being, for example, a matrix of values.

[0035] Thus, if the FTO depends on both the distance between the optical lens and the scene and the distance between the optical lens and the sensor, then for each point on the optical lens, the FTO is represented by a set of functions, each function being a matrix of values, for example, a 10x10 matrix of values. Therefore, if there are: - 100 possible distances between the optical lens and the scene, and - 10 possible distances between the optical lens and the sensor; the FTO is represented, for each point of the optical lens, by a set of 10x100 functions, each function being able for example to be a matrix of values, for example a matrix of 10x10 values.

[0036] In some embodiments, for at least one source image, at least one image parameter may be, or may include, any of the following parameters: - a focusing distance associated with said source image, - a field angle of view of said source image, - a shooting angle of said source image, - an acquisition position for said source image, and / or - a brightness associated with said image.

[0037] Thus, additional information is provided to the AI ​​model about how the source image represents the scene, for example the focal length at which the source image represents the scene, the shooting angle at which the scene is represented in the source image, or This also includes the field of view angle with which the scene is represented in the source image, etc. This information is particularly useful for the AI ​​model to estimate the depth map of the scene.

[0038] As mentioned above, the respective source images entered into the AI ​​model are all different.

[0039] Following examples of implementation, at least two source images of the scene may differ from each other in that they are acquired at two different focal distances.

[0040] Thus, the source images represent the scene with different areas of sharpness. Using source images with different focal lengths increases and varies the scene information taken into account when creating the depth map of said scene, thereby reducing errors due to optical effects that may be present in a given source image and that would be related to the use of a particular focal length.

[0041] Alternatively, or in addition, at least two source images of the scene may differ from each other in that they are acquired with different field angles.

[0042] Thus, one of the source images is sharper than the other, but the other provides a wider view of the scene. Using source images with different field of view increases and diversifies the scene information considered when creating the depth map of that scene. This reduces errors due to optical effects that may be present in a particular source image and related to the use of a specific field of view.

[0043] Alternatively, or in addition, at least two source images of the scene may differ from each other in that they are acquired with different viewing angles, or shooting angles.

[0044] Thus, the two images represent the scene from different viewing angles, that is, different acquisition directions. Using source images with different viewing angles increases and varies the information about the scene taken into account when creating the depth map of that scene, thereby reducing errors due to optical effects that may be present in a given source image and that would be related to the use of a particular viewing angle.

[0045] Alternatively, or in addition, at least two source images of the scene may differ from each other in that they are acquired with different angles of rotation.

[0046] Thus, the two images represent the scene at different rotation angles, meaning different rotations of the camera module. Using source images with different rotation angles increases and varies the scene information considered when creating the scene's depth map. This reduces errors due to optical effects that may be present in a given source image and related to the use of a particular rotation angle.

[0047] Following examples of implementation, at least two source images of the scene may differ from each other in that they are acquired from two different positions.

[0048] Thus, the source images represent the scene from different positions. Using source images acquired from different positions increases and varies the information about the scene taken into account when creating the depth map of said scene, which reduces errors due to optical effects that may be present in a given source image and that would be related to the use of a particular acquisition position.

[0049] Following realization examples, at least two source images of the scene may differ from each other in that they are acquired with different brightness levels for the scene.

[0050] Thus, the source images represent the scene with varying brightness levels for part or all of the scene. Using source images with different brightness levels increases and diversifies the scene information considered when creating the scene's depth map, particularly regarding areas that might be overexposed or underexposed in one of the source images. This helps reduce errors due to brightness deficiencies in one of the source images or optical effects related to such brightness deficiencies.

[0051] Depending on the implementation methods, the source images may have the same definition.

[0052] Alternatively, at least two source images of the scene may have different definitions.

[0053] Depending on the embodiment, the consolidated depth map may have the same definition as the, or at least one of the, source images, IMl-IMn.

[0054] In this case, for each pixel of the source image(s) of the scene, the depth map indicates a depth value for the point in the scene to which each pixel corresponds.

[0055] If the source images do not all have the same resolution, it is possible to: - to keep the resolution of the image with the highest resolution, - to keep the resolution of the image with the lowest resolution, or - choose another resolution.

[0056] In this case, for each pixel of the source image(s) of the scene, the depth map indicates a depth value for the point in the scene represented by said pixel.

[0057] Depending on the embodiment, the depth map may have a lower resolution than the source image(s). Indeed, a depth map is generally used to determine the distance between the image sensor and each object in the scene. Therefore, in some cases, it may be sufficient for the depth map to indicate a depth for the different areas of the scene, each corresponding to an object in the scene: in this case, the depth map will have a lower resolution than the source images.

[0058] Depending on the embodiment, the depth map may have a higher resolution than the source image(s).

[0059] The depth map can be provided for a format corresponding to a union of at least two source images, for example spatially, in this case, the depth map will cover a larger spatial extent than each of the source images and may have a higher definition than each source image taken individually.

[0060] Depending on the embodiment, the depth map can be indexed to points in the scene, to indicate a depth at each of these points, on at least one of the source images for example.

[0061] Indeed, it may be possible to provide the depth map, with an indexing table relating to at least one of the source images, so that the depth values ​​can be read in relation to said source image.

[0062] In this case, during the training stage, the AI ​​model can be trained to produce an indexing map relative to one of the source images.

[0063] According to embodiments, the process according to the invention may further include a spatial registration step of at least one source image.

[0064] This registration process eliminates any potential discrepancies between source images of the same scene. For example, the registration step eliminates differences in viewpoint, viewpoint, slight movements, etc., that may exist between the different source images.

[0065] Such recalibration can be used particularly effectively in the case of micro, or small, offset(s).

[0066] The registration of a source image can be performed relative to another source image. In particular, one of the source images can be chosen as the reference image, and the other source image, and indeed each of the other source images, can be spatially registered relative to said reference image. Following a non-limiting example, the source image chosen as the reference image could be the one representing the scene with the widest field of view.

[0067] Alternatively, the registration of a source image can be performed against a reference format that is independent of the source images. In this case, the source image, and in particular each of the source images, can be registered against said reference format.

[0068] The registration of a source image can be performed using any known technique, for example, by using points of interest or objects located within the source image. Alternatively, or in addition, the registration of a source image can be performed using velocity or acceleration data associated with the pixels of the source image.

[0069] In some embodiments, the spatial registration step may not be performed by the AI ​​model. In this case, the result of said registration step can be provided as input to the AI ​​model, for example in the form of an indexing table, for use in estimating the scene's depth map. Thus, the AI ​​model can identify the same point in the scene across multiple source images and determine the scene's depth at that point.

[0070] In some embodiments, the spatial registration step can be performed by the AI ​​model, for example in input layers of said AI model. In this case, the AI ​​model can provide as output, in addition to the scene depth map, an indexing map for at least one, and in particular each, source image, allowing the identification of each point of said source image in said depth map.

[0071] All source images can have the same size.

[0072] Alternatively, at least two source images may have different sizes, and in particular represent the scene from a different field of view. In this case, the method according to the invention may further Prior to the estimation step, the process includes a step to adjust the size of at least one source image. For example, the smaller source image can be augmented with predetermined values, particularly constant values ​​such as zero. This way, there is no data loss.

[0073] Of course, other implementations are possible. For example, the larger source image can be reduced, for instance, by removing the data at its periphery (this data is generally the least defined), to adjust its size to that of the smaller source image. Other implementation examples are possible.

[0074] All source images can have the same resolution.

[0075] Alternatively, at least two source images may have different definitions.

[0076] In this case, the method according to the invention may further include, prior to the estimation step, a step for adjusting the resolution of at least one source image. For example, the lowest-resolution source image may be augmented with predetermined values, in particular identical values, for example, zero or undefined, to adjust the aspect ratio of said lowest-resolution source image to the aspect ratio of the highest-resolution source image. In particular, all source images may be adjusted to the aspect ratio of the highest-resolution source image.

[0077] Alternatively, the higher-resolution source image can be adjusted to the aspect ratio of the lower-resolution source image, for example, by removing some pixels or merging pixels within the higher-resolution source image. In particular, all source images can be adjusted to the aspect ratio of the lower-resolution source image.

[0078] Alternatively, all source images can be adjusted to another resolution, for example to a resolution equal to the average resolution of the source images.

[0079] According to embodiments, the method according to the invention may include, for at least one source image, providing an indexing map of said source image relative to a reference image.

[0080] The reference image can be another source image.

[0081] The reference image can be a combination of at least two, and in particular of all the source images.

[0082] For at least one source image, the indexing map can be provided by the spatial registration step described above, said registration step being performed: - either upstream of the AI ​​model: in this case, the said indexing map is given as input to the AI ​​model for the estimation of the depth map; - either directly through the AI ​​model. In all cases, the indexing map allows, for a source image, the identification of the points of said source image in the estimated depth map.

[0083] The AI ​​model can be a machine vector support, a decision tree, etc., or any other AI model.

[0084] Depending on the embodiment, the AI ​​model can be a pre-trained neural network, in particular a deep learning convolutional neural network, DLCNN, for "Deep Learning Convolutional Neural network" in English.

[0085] The neural network can take the following as input: - several source images, IMi-IM n , - an image parameter, or a set of image parameters, for each source image; - and optionally, for at least one source image, an indexing map allowing the points of said image to be indexed relative to a reference image, in the case where said indexing map is determined upstream of said AI model.

[0086] For example, the neural network can take the following data as input: {(IMi ; Pl^-Pir ;CIPi), ..., (IMn ; Pl^-PIn" 1 ;CIPm)} with: - IMi, ..., IMn: the source images, with n>l; and - Pl^-PIi 171 the image parameter, or set of image parameters, relative / associated with the source image IMi; - and optionally, CIPi the indexing map of the source image IMi, in the case where said indexing map is determined upstream of said AI model.

[0087] The neural network can output the depth map. Optionally, the neural network can output, in addition to the depth map, for at least one source image, an indexing map of said depth map relative to said source image, in the case where said indexing map is not determined upstream of the AI ​​model and it is the AI ​​model that determines it.

[0088] For example, the neural network can provide the following output data: - CP: the depth map; and - optionally, CIPi: the depth indexing map, for the source image IMi.

[0089] Depending on the embodiment, the neural network can include, as input, a layer comprising: - at least as many neurons as there are points in an input source image, or at least as many neurons as the number of points in the depth map that one wishes to generate; and - each neuron can have at least as many inputs as there are different source images and image parameters. The neural network can include several layers of neurons, for example 3 to 10.

[0090] According to embodiments, the method according to the invention may further include a step of training the AI ​​model, and in particular the neural network forming the AI ​​model.

[0091] Such training can be carried out using a training platform comprising a multitude of training games. Each training game may include: - several source images, called training images, IMEi-IME n , And - for each training source image, IMEi, a set of at least one training image parameter, PIEi, associated / relative to said source image; - a depth chart, also known as a training chart, CPE; - and possibly a CIPE training depth indexing map, for each source image.

[0092] Each training game can be obtained in different ways.

[0093] Following an example implementation, a training set can be obtained from a known training scene. In this case, several different source images of the training scene are acquired: these source images constitute the training source images of a training set.

[0094] For each source image in the training scene, the image parameter set(s) are determined and stored in association with the source image. This image parameter set(s) is known / determinable because the conditions under which the scene is imaged are known: - we know the camera module used and therefore the FTO of said camera module or optical lens, - we know the focal distance used to image the scene, - we know the viewing angle to visualize the scene, - we know the field of view from which the scene is captured, - we know the rotation angle of the module with which the scene is imaged, - We know the position from which the scene is imaged, - We know the brightness with which the scene is imaged. - etc.

[0095] Furthermore, the depth map of the training scene can be determined: - either by design / construction of said training scene, which, we recall, is known; - either by measurement with at least one measurement sensor, such as a LIDAR or a time-of-flight camera. This depth map can serve as the training depth map for the model's training set.

[0096] Finally, knowing the training depth map it is possible to construct a training indexing map with respect to one, or each, of the training source images or another format.

[0097] Thus, for a given training scene, we have a training game comprising: - several different source images for training, IMEi-IME n ; - for training source image, IMEi, a set of training image parameter(s), PIEi; - a CPE training depth map, and - optionally one or more CIPE training index card(s). By using a multitude of different training scenes, it is possible to obtain a multitude of training games for the AI ​​model.

[0098] The AI ​​model can be trained with any known and suitable training algorithm. For example, the AI ​​model can be trained using a backpropagation algorithm. Of course, this example is only illustrative and is by no means exhaustive.

[0099] According to another aspect of the same invention, a computer program is proposed comprising executable instructions which, when executed by a computer device, implement all the steps of the process according to the invention.

[0100] The computer program can be in any computer language, such as for example machine language, C, C++, JAVA, Python, etc.

[0101] Such a computer program can be presented as a standalone application. Alternatively, such a computer program can be integrated into a photo or video application, or even into an image or video playback application.

[0102] The computer program can be stored in a non-transient, or non-volatile, manner in a storage medium.

[0103] According to another aspect of the invention, a device is proposed comprising means configured to implement all the steps of the process according to the invention.

[0104] The device according to the invention can be a computer, a processor, a computer chip, etc. programmed to implement the method according to the invention, for example by executing the computer program according to the invention.

[0105] The device according to the invention can be integrated into any type of device such as a smartphone, a tablet, a computer, a calculator, a processor, a computer chip, etc.

[0106] The device according to the invention may include, in terms of technical means, at least one, or any combination of at least two, of the characteristics described above with reference to the method according to the invention, and which will not be repeated here exhaustively for the sake of brevity.

[0107] According to another aspect of the invention, an apparatus is proposed, and in particular a user apparatus, comprising a device according to the invention.

[0108] In particular, the device can be a user device such as a smartphone, tablet, etc.

[0109] The user device may also include a display screen.

[0110] Alternatively, or in addition, the user device may further include at least one camera module to image a scene and obtain source images of said scene. [YES] In particular, the device can be a computer-type user device.

[0112] The computer-type user device may also include a display screen.

[0113] Alternatively, or in addition, the computer-type user device may further include at least one camera module to image a scene and obtain source images of said scene.

[0114] In particular, the device could be a television.

[0115] Television may also include a display screen.

[0116] Alternatively, or in addition, the television may further include at least one camera module to image a scene and obtain source images of said scene.

[0117] In particular, the device can be a virtual reality headset or glasses, or an augmented reality headset or glasses.

[0118] The helmet, or glasses respectively, according to the invention may further comprise at least one display screen.

[0119] Alternatively, or in addition, the helmet, or glasses respectively, according to the invention may further comprise at least one camera module for imaging a scene and obtaining source images of said scene.

[0120] In particular, the device could be a medical imaging device.

[0121] In particular, the medical imaging device can be an endoscope, an ultrasound machine, etc.

[0122] The medical imaging device may also include at least one display screen.

[0123] Alternatively, or in addition, the medical imaging device may further include at least one camera module to image a scene and obtain source images of said scene.

[0124] Of course, the device according to the invention is not limited to the examples of devices that have just been given.

[0125] According to another aspect of the present invention, a vehicle comprising a device according to the invention is proposed.

[0126] The vehicle may further include a display screen, for example arranged in a passenger compartment of the vehicle, or a projector to project at least one image onto a display surface generally known as a "head-up display".

[0127] Alternatively, or in addition, the vehicle may further include at least one camera module to image a scene and obtain source images of said scene.

[0128] Depending on the embodiment, the vehicle can be a land vehicle, such as a car, autonomous, semi-autonomous or non-autonomous.

[0129] Depending on the embodiment, the vehicle can be a flying vehicle, such as a drone, an airplane, a helicopter, autonomous, semi-autonomous or non-autonomous.

[0130] Depending on the embodiment, the vehicle can be a maritime vehicle, such as a boat or a submarine, autonomous, semi-autonomous or non-autonomous. Description of the figures and methods of implementation

[0131] Other advantages and features will become apparent upon examination of the detailed description of non-limiting embodiments and the accompanying drawings, in which: - FIGURE 1 is a schematic representation of a non-limiting example of an embodiment of a process according to the invention; - FIGURE 2 is a schematic representation of a non-limiting example of an embodiment of a device according to the invention; - Figures 3-5 are schematic representations of non-limiting examples of embodiments of a device according to the invention; and - FIGURE 6 is a schematic representation of a non-limiting example embodiment of a vehicle according to the invention.

[0132] It is understood that the embodiments described below are by no means limiting. In particular, variants of the invention may be conceived comprising only a selection of the features described below, isolated from the other features described, if this selection of features is sufficient to confer a technical advantage or to differentiate the invention from the prior art. This selection includes at least one preferably functional feature without structural details, or with only a part of the structural details if it is this part which alone is sufficient to confer a technical advantage or to differentiate the invention from the prior art.

[0133] In particular, all the variants and embodiments described can be combined with each other if there are no technical obstacles to this combination.

[0134] In the figures and in the rest of the description, elements common to several figures retain the same reference.

[0135] FIGURE 1 is a schematic representation of a non-limiting example of an embodiment of a method according to the present invention.

[0136] Method 100 of FIGURE 1 can be used to determine, by estimation, the depth map of a scene from different source images of the scene.

[0137] The process 100 includes a phase 102 of estimating the depth map of the scene from: - at least two different source images of the scene, and in particular a PIL stack of IMi-IM source images n different from the scene, with n>l; and - for each source image IMi, a set of at least one image parameter, JPIi=(PI i 1 -PI i m ).

[0138] As mentioned above, at least two of, and in particular all of, the source images used to generate the depth map are different from each other.

[0139] For example, at least two source images differ from each other in that they are acquired: - at different focusing distances, - with different field angles, - with different viewing angles, - from different positions, and / or - with different lighting for the scene. Thus, each of these source images provides different information about the scene, which allows for the generation of initial depth maps that are potentially different as well.

[0140] As mentioned above, the image parameter set(s) associated with a source image can include one or more parameters relating to how the source image represents the scene, and / or how the source image was acquired. Thus, the image parameter set(s) can include at least one, or any combination of at least two, of the following parameters: - the FTO, and in particular the PSF, of the optical lens used for the acquisition of said source image; - the focal distance associated with said source image, - the field of view angle with which the source image represents the scene, - the shooting angle with which the source image represents the scene, - the position from which said source image represents the scene, and / or - the brightness associated with said image.

[0141] The CP depth map is generated using a pre-trained artificial intelligence (AI) model.

[0142] This AI model can be any type of suitable AI model.

[0143] Following an example implementation, the AI ​​model can be a deep learning convolutional neural network pre-trained to provide the CP depth map and taking the following as input: - the different source images IMi-IM n ; And - for each source image IMi, the image parameter set JPIi = (PIi 1 -PIi m ).

[0144] At the output of phase 102, the AI ​​model provides a depth map, CP, and optionally a depth indexing map, CIP.

[0145] Optionally, the AI ​​model can also provide, for each source IMi image, a CIPi depth indexing map, in the case where the A neural network was trained to determine these depth indexing maps.

[0146] Alternatively, as will be described later, ClPi-CIPn depth indexing maps can be obtained during a spatial registration step carried out prior to phase 102 of depth map estimation.

[0147] When the CP depth map is generated by an AI model, in particular by a convolutional neural network, the process 100 may optionally include a phase 110 of training said neural network.

[0148] Training phase 110 can be carried out just before scene depth map estimation phase 102, so that estimation phase 102 is carried out immediately after training phase 110.

[0149] Alternatively, training phase 110 can be carried out well before scene depth map estimation phase 102, so that estimation phase 102 is not carried out immediately after training phase 110 and a non-negligible time lapse occurs between training phase 110 and estimation phase 102.

[0150] The training phase 110 is common to several iterations of the estimation phase 102, so the AI ​​model trained during training phase 110 is used in several iterations of estimation phase 102, for the same scene or for different scenes. Alternatively, training phase 110 can be specific to an estimation phase and be repeated before each estimation phase 102, for example, to update the AI ​​model based on a particular type of scene.

[0151] When the AI ​​model is a deep learning convolutional neural network, the 110 training phase can be carried out in all possible ways.

[0152] In particular, neural network training can be performed using a training dataset comprising a multitude of training sets. Each training set can include: - several different source images, called training images IMEi- IMEm; - for each training source image IMEi, a set of at least one image parameter associated / relative to said source image, JPIEi = (PIi 1 -PIi k ) ; - a depth chart, also known as a training chart, CPE; and - optionally, for each training source image IMEi-IMEm, a depth indexing map, called training map, CIPEl-CIPEm.

[0153] Each training game can be obtained in different ways.

[0154] Following an example implementation, a training set can be obtained from a known training scene, denoted SE. In this case, several different source images of the training scene SE are acquired: these source images constitute the training source images IMEi-IMEm of a training set.

[0155] For each training source image IMEi, the training image parameter set(s) J PIEi = ( PL 1 - PL k ) is determined and stored in association with the training source image IMEi. This set of training image parameters JPIEi is known / determinable because the conditions under which the scene is imaged in the training source image IMEi are known: - the camera module used and therefore the FTO of said camera module or optical lens, - the focusing distance used to image the scene, - the viewing angle to visualize the scene, - the field of view from which the scene is captured, - the rotation angle of the module with which the scene is imaged, - the position from which the scene is depicted, - the brightness with which the scene is depicted, - etc.

[0156] Furthermore, the depth map of the SE training scene can be determined: - either by design / construction of said training scene, which, we recall, is known; - either by measurement with at least one measurement sensor, such as a LIDAR or a time-of-flight camera. This depth map can serve as the CPE training depth map for the model training set.

[0157] Finally, for each IMEi training source image, the depth indexing map can be determined by determining the spatial offset between said source image and a reference image, which allows us to obtain the training depth indexing map, CIPEi.

[0158] Thus, for a given SE training scene, we have a training set comprising: - several different training source images IMEi-IME m ; - for each source training image IMEi, a set of training image parameter(s), J PIEi = ( PL 1 - PL k ) ; - a training depth map, CPE; and - optionally, for each training source image IMEi, a training depth indexing map, CIPEi. By using a multitude of different SEI-SEI training scenes, it is possible to obtain a multitude of training sets for the AI ​​model. Preferably, l > 100, or even 1000 or 10000.

[0159] The AI ​​model can then be trained with any known and suitable training algorithm. For example, the AI ​​model can be trained using a backpropagation algorithm. Of course, this example is only illustrative and is by no means exhaustive.

[0160] Optionally, process 100 may include an optional step 122 of spatial registration of at least one source image.

[0161] In particular, registration step 122 allows for the spatial registration of source images relative to a reference image. The reference image can be one of the source images. Alternatively, The reference image can be an image corresponding to the spatial union of all the source images.

[0162] This registration process eliminates discrepancies between source images representing the same scene. For example, the registration step eliminates differences in viewpoint, shooting angle, slight movements, etc., that may exist between different source images of the same scene.

[0163] In the example of realization of FIGURE 1, in no way limitingly, at least one, in particular each, source image is spatially recalibrated with respect to the source image representing the scene with the largest field of view.

[0164] Registration can be achieved by detecting the same objects of interest in the source images and using their position to register the source images to each other.

[0165] This optional spatial registration step 122 can be performed by the AI ​​model. Alternatively, this optional spatial registration step 122 may not be performed by the AI ​​model.

[0166] This optional spatial registration step 122 can provide, for at least one source image, a depth indexing table.

[0167] If step 122 is not performed by the AI ​​model, then the depth indexing table can be given as input to the AI ​​model for estimating the scene depth map during phase 102.

[0168] If step 122 is performed by the AI ​​model, then the depth indexing table is determined by the AI ​​model and given as output from the AI ​​model.

[0169] The process 100 may further include an optional step 124 of adjusting at least one source image.

[0170] In particular, adjustment step 124 can adjust the size of at least one source image, especially when the source images are of different sizes. In the implementation example in Figure 1, the size of the source images is adjusted, but not limited to, to match the size of the largest source image, for example, the source image representing the scene with the largest dimensions. wide field of view. In this case, each smaller IMi source image is supplemented with predetermined values, in particular constants, for example zero or undefined, so that it has the same size as the largest source image.

[0171] Alternatively, or in addition, step 124 can adjust the resolution of the source images, particularly when the source images have different resolutions. In the example shown in Figure 1, the resolution of the source images is adjusted, without limitation, to match the resolution of the highest-resolution source image, for example, the image representing the scene at the greatest zoom level. In this case, each lower-resolution source image (IMi) is supplemented with predetermined, or constant, values, such as zero or undefined values, so that it has the same resolution as the highest-resolution source image.

[0172] FIGURE 2 is a schematic representation of a non-limiting example embodiment of a device according to the present invention.

[0173] The device 200 in FIGURE 2 can be used to obtain, by estimation, a depth map of a scene from several source images of said scene, different from each other.

[0174] The device 200 of FIGURE 2 can be used to implement a method according to the invention, and in particular the method 100 of FIGURE 1.

[0175] Device 200 may include an optional module 202 for training the AI ​​model with a training dataset. This module 202 is, for example, configured / programmed to perform training phase 110 in FIGURE 1.

[0176] Device 200 includes a module 204, running the AI ​​model, MIA, to provide an estimated depth map from several different source images input to said AI model, MIA. This module 204 is, for example, configured / programmed to perform phase 102 of process 100 in FIGURE 1.

[0177] Device 200 may include an optional module 206 to perform spatial registration of at least one source image. This module 206 is, for example, configured / programmed to perform step 122 of process 100 in FIGURE 1. This module 206 is optional because this registration may not be performed, or may be performed by the AI ​​model.

[0178] Device 200 may include an optional module 208 for adjusting the size and / or resolution of at least one source image. This module 208 is, for example, configured / programmed to perform step 124 of process 100 in Figure 1.

[0179] Device 200 may also optionally include means of memorization (not shown).

[0180] At least one of the modules 202-208 can be a module independent of the others.

[0181] At least two of the 202-208 modules can be integrated within the same module. In particular, the 202-208 modules can be integrated within a computing unit 210.

[0182] At least one of the 202-208 modules can be a hardware module, such as a processor, an electronic chip, etc.

[0183] At least one of the modules 202-208 may be a software module, such as a computer program.

[0184] At least one of the modules 202-208 may be a combination of at least one software module and at least one hardware module.

[0185] In particular, at least one of the 202-208 modules can be integrated into an electronic chip, or into an application installed in a user device.

[0186] In particular, the 210 computing unit can be, or can be integrated, into an electronic chip, or into an application installed in a user device.

[0187] FIGURE 3 is a schematic representation of a non-limiting example embodiment of a device according to the present invention.

[0188] The apparatus 300 of FIGURE 3 includes means configured to implement the invention, and in particular the method 100 of FIGURE 1.

[0189] The device 300 of FIGURE 3 may include a device according to the invention, and in particular the device 200 of FIGURE 2.

[0190] In the example shown in FIGURE 3, device 300 is a smartphone, or a tablet, comprising device 200 from FIGURE 2.

[0191] Optionally, the 300 unit may also include: - a 302 camera module, in particular for taking source images of the scene which can be used to estimate the depth map of said scene, according to the invention; and - a display screen 304, optionally equipped with a detection surface 306, for example capacitive.

[0192] Of course, the 300 device may include other components than those indicated above.

[0193] FIGURE 4 is a schematic representation of another non-limiting embodiment of a device according to the present invention.

[0194] The apparatus 400 of FIGURE 4 includes means configured to implement the invention, and in particular the method 100 of FIGURE 1.

[0195] The device 400 of FIGURE 4 may include a device according to the invention, and in particular the device 200 of FIGURE 2.

[0196] In the example shown in FIGURE 4, device 400 is: - a virtual reality (VR) headset or glasses, or - an augmented reality headset, or glasses; including device 200 of FIGURE 2.

[0197] Optionally, the 400 unit may also include: - a 402 camera module, in particular for taking source images of the scene which can be used to estimate the depth map of said scene, according to the invention; and - a 404 display screen, for example in / on a visor of said 400 helmet.

[0198] Of course, the 400 helmet may include other components than those listed above.

[0199] FIGURE 5 is a schematic representation of a non-limiting example embodiment of a device according to the present invention.

[0200] The apparatus 500 of FIGURE 5 includes means configured to implement the invention, and in particular the method 100 of FIGURE 1.

[0201] The device 500 of FIGURE 5 may include a device according to the invention, and in particular the device 200 of FIGURE 2.

[0202] In the example shown in FIGURE 5, device 500 is a medical imaging device, such as an endoscope, an ultrasound machine, etc.

[0203] Optionally, the 500 unit may also include: - a 502 camera module, in particular for taking source images of the scene which can be used to estimate the depth map of said scene, according to the invention; - a display screen 504, optionally equipped with a sensing surface 506, for example capacitive.

[0204] Of course, the 500 medical imaging device may include other organs than those indicated above.

[0205] FIGURE 6 is a schematic representation of a non-limiting example embodiment of a vehicle according to the present invention.

[0206] The vehicle 600 of FIGURE 6 includes means configured to implement the invention, and in particular the method 100 of FIGURE 1.

[0207] The vehicle 600 of FIGURE 6 may include a device according to the invention, and in particular the device 200 of FIGURE 2.

[0208] In the example shown in FIGURE 6, vehicle 600 is a land vehicle, in particular a car, comprising device 200 of FIGURE 3.

[0209] Optionally, the 600 vehicle can also include: - a 602 camera module, in particular for taking source images of the scene which can be used to estimate the depth map of said scene, according to the invention; - a display screen 604, optionally equipped with a sensing surface 606, for example capacitive, arranged in the passenger compartment of the vehicle 600.

[0210] Of course, the 600 vehicle may include other components than those indicated above.

[0211] Of course, the invention is not limited to the examples that have just been described.

Claims

1. DEMANDS 1. A method (100) for determining a depth map (DM) of a scene, said method (100) comprising an estimation step (102) of said depth map (DM) of said scene by an artificial intelligence model (AIM), AI model, previously trained, taking as input: - several images (IMi-IM) n ), said source images, different from said scene, each source image (IMi) representing the scene differently from the other source images; and - for each source image, at least one image parameter related to said source image (IMi-IM n ).

2. A method (100) according to the preceding claim, characterized in that, for at least one source image (IMi-IM n), at least one image parameter is, or includes, an optical transfer function, OTF, of the optical lens, or camera module, used for the acquisition of said source image (IMi-IMn).

3. A method (100) according to any one of the preceding claims, characterized in that, for at least one source image (IMi-IM n ), at least one image parameter is, or includes, any of the following parameters: - a focusing distance associated with said source image (IMi-IM n ), - a field angle of said source image (IMi-IM n ), - a shooting angle of said source image (IMi-IM n ), - a rotation angle of the camera module used for capturing said source image (IMi-IM) n ), - an acquisition position for said source image (IMi-IM n ), and / or - a brightness associated with said source image (IMi-IM n ).

4. A method (100) according to any one of the preceding claims, characterized in that at least two source images (IMi-IM n ) of the scene differ from one another in that they are acquired: - at different focusing distances, - with different field angles, - with different viewing angles, - with different rotation angles, - from different positions, and / or - with different brightness levels for the scene.

5. A method (100) according to any one of the preceding claims, characterized in that the depth map (CP) has the same definition as the source images (IMi-IM n ).

6. Method (100) according to any one of the preceding claims, characterized in that it further comprises, prior to the estimation step (102), a spatial registration step (122) of at least one source image.

7. Method (100) according to any one of the preceding claims, characterized in that it further comprises, prior to the estimation step (102), a step (124) of adjusting the size of at least one source image.

8. Method (100) according to any one of the preceding claims, characterized in that it further comprises, prior to the estimation step (102), a step (124) of adjusting the definition of at least one source image.

9. Method (100) according to any one of the preceding claims, characterized in that the AI ​​model (AIM) is a pre-trained neural network, in particular a deep learning convolutional neural network.

10. Method (100) according to any one of the preceding claims, characterized in that it further comprises a step (110) of training the AI ​​model (AIM).

11. Computer program comprising executable instructions which, when executed by a computing device, implement all the steps of the process (100) according to any one of the preceding claims.

12. Device (200) comprising means (202-208) configured to carry out all the steps of the process (100) according to any one of claims 1 to 10.

13. Apparatus (300;400;500) comprising a device (200) according to the preceding claim.

14. Vehicle (600) comprising a device according to claim 12.

Citation Information

Patent Citations

  • Visual-inertial fusion localization and mapping method and device for pedestrian motion constraints

    CN115235454B

  • Image deep analysis method based on end-to-end deep learning

    CN118429748A

  • Determining a three-dimensional representation of a scene

    US20220068024A1

  • Generation and rendering of extended-view geometries in video see-through (VST) augmented reality (AR) systems

    US20240257475A1

  • Employing three-dimensional data predicted from two-dimensional images using neural networks for 3D modeling applications

    WO2020069049A1