Method and device for determining depth information on a scene with images of said scene taken by at least two cameras
Patent Information
- Application Number
- EP2023736409
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-21
- Publication Date
- 2026-01-28
AI Technical Summary
Existing methods for determining depth information from images taken by multiple cameras are computationally intensive, time-consuming, and often produce inaccurate results, especially when dealing with complex scene compositions.
The method involves identifying a target point in one image by comparing speed data from that image to speed data in another image, reducing computational resources required and enhancing precision by utilizing speed information to accurately locate the point across images.
This approach reduces computational burden and improves the accuracy of depth determination, providing more precise and reliable results by leveraging speed data to identify points across images from different capture positions.
Smart Images

Figure FR2023050396_26092024_PF_FP
Abstract
Description
DESCRIPTION Title: Method and device for determining depth information of a scene with images of said scene taken by at least two cameras
[0001] The present invention relates to a method for determining depth information of a scene from at least two images of said scene taken by at least two cameras. It also relates to a device and a system implementing such a method. It further relates to a use of the depth information in different applications, such as for example the piloting of a vehicle, or the identification of objects in the scene or the tracking of objects in said scene.
[0002] The field of the invention is the field of determining depth information associated with one or more points of a scene from several images of said scene. State of the art
[0003] The depth of a point in a scene in an image of said scene corresponds to the distance between said point in the scene and the observation point from which the image of the scene was acquired. This distance, or depth, data is important and useful information in different applications. For example, it can be used for driver assistance, or autonomous piloting, of a vehicle based on the position of objects in its environment. It can also be used for object detection, or for tracking objects over time.
[0004] There is a solution that allows this depth information to be determined by calculation from at least two images of a scene taken by cameras positioned at different capture positions. By identifying the position of the same point of the scene on said images, and knowing the capture positions of the cameras, it is then possible to calculate, by triangulation, the depth of said point on each image of the scene.
[0005] However, these solutions are very computationally intensive, time-consuming and produce results whose accuracy is not always satisfactory, particularly when the composition of the imaged scene is complex.
[0006] An aim of the present invention is to remedy at least one of the drawbacks of the state of the art.
[0007] Another aim of the invention is to propose a solution for determining depth information from images of a scene, which consumes less computing resources and / or is less time-consuming and / or produces more precise results. Statement of the invention
[0008] The invention proposes to achieve at least one of the aforementioned aims by a method for determining depth information of a scene in a first image of said scene captured by a first camera positioned in a first capture position, said method comprising, for a target point of said scene located in said first image, a processing phase comprising the following steps: - identification of said target point on a second image captured by a second camera positioned at a second capture position different from said first position, and - calculating a depth associated with said target point as a function of a position of said target point on each of said first and second images; characterized in that the identification step comprises a comparison of a speed data item of said target point in said first image with a speed data item of at least one point in said second image, said speed data being previously calculated.
[0009] Thus, in a manner similar to the state of the art, the invention proposes to determine the depth associated with a target point of a scene by using the positions of this target point on at least two images of said scene captured by cameras positioned at different capture positions. [OO1O] However, and in a manner that is entirely innovative and different from the state of the art, the invention proposes to use speed data of said target point in the first image to identify it in the second image. Thus, the identification of the target point in the second image requires fewer computing resources and is less time-consuming. In addition, the speed information associated with the target point makes it possible to reduce, or even avoid, errors when identifying the target point in the second image, which makes it possible to obtain more precise and more reliable results.
[0011] In other words, for a point of a scene located on an image, and in particular for a pixel of an image corresponding to said point of the scene, the speed data enriches the data associated with said point / pixel. Thus, we have more data to identify this point of the scene on different images of said scene captured from different capture positions.
[0012] In the present application, the term "image" means a digital image and in particular a matrix digital image. Each point of the matrix corresponds to a pixel of the image and comprises one or more digital values representing different components of the image for said pixel. For example, in the context of an RGB image, each pixel may comprise three values, each representing one of the RGB components of the digital image. Of course, each pixel may comprise other values representing other quantities, such as brightness, hue, saturation, etc.
[0013] In the present application, a point of a scene corresponds to a pixel of an image of said scene. When the scene is imaged by two cameras positioned in different capture positions, spaced by a distance preferably as large as possible, the same point of the scene can correspond to two pixels located in different positions on each of the two images. Preferably, the planes of the cameras can be parallel and the central axes of the cameras common in order to limit the calculations, although this is not essential.
[0014] In the following, we denote by UV the reference frame linked / associated with the image sensor. This reference frame is a two-dimensional reference frame defining a plane, and in particular the plane of the image sensor used for image acquisition.
[0015] In the following, we denote by XY the frame linked / associated with the scene. In addition, we denote by Z the direction perpendicular to the image plane, namely to the XY plane, in the frame linked to the scene.
[0016] For a point of an image, that is to say for the target point and / or for at least one point of the second image, the speed data used during the step of identifying the target point comprises at least the speed of said point in the plane of the image.
[0017] The speed of a point in the image plane can be expressed in a frame of reference, denoted UV, linked to the image sensor. In this case, the speed of said point in the image plane, denoted V uv , is a two-dimensional velocity vector comprising a first value, denoted Vu, giving the velocity in the U direction in the image plane, and a second value, denoted V v , giving the velocity in the V direction in the image plane.
[0018] According to the invention, the first image can be captured at a first instant, called the first current instant, and the second image is captured at a second instant, called the second current instant.
[0019] The first current instant and the second current instant are identical, or different instants but very close in time so that the displacement of the points of the scene between said instants is negligible.
[0020] Without loss of generality, and to avoid editorial heaviness, we consider in the following that the first current instant and the second current instant are identical and are called current instant.
[0021] The processing phase, or the method according to the invention, may comprise a step of calculating the speed of the target point in the first image.
[0022] Alternatively, the speed of the target point in the first image may be calculated during a prior step not forming part of the method according to the invention. For example, the first image may be a so-called enriched image, already comprising the speed associated with the target point. The speed data of the target point may be part of the data of the pixel corresponding to said target point in the first image.
[0023] For at least one point of the second image, the processing phase, or the method according to the invention, may comprise a step of calculating the speed of said point in said second image.
[0024] Alternatively, for at least one point of the second image, the speed of said point in the second image may be calculated during a prior step not forming part of the method according to the invention. For example, the second image may be an image, called enriched, already comprising the speed associated with said point. For example, the speed data of this point may be part of the data of the pixel corresponding to said point in the second image.
[0025] According to the invention, the speed data associated with a point of an image can be calculated in any known manner and the invention is not limited to a specific way of calculating said speed data.
[0026] According to embodiments, the speed data of a point of an image can be calculated from said image and at least one image, called past, captured by the same camera at a past instant preceding the current instant at which said image was captured.
[0027] For example, the velocity of the target point in the first image may be calculated using an image captured by the first camera at a past time preceding the current time of capture of the first image. The past image may or may not be the image immediately preceding the first image, for example in a stream of images captured by the first camera.
[0028] Alternatively, or in addition, the velocity of a point in the second image may be calculated using an image captured by the second camera at a past time preceding the current time of capture of the second image. The past image may or may not be the image immediately preceding the second image, for example in a stream of images captured by the first camera.
[0029] According to a non-limiting exemplary embodiment, the step of calculating the speed data of a point in an image may comprise the following steps: - identification of said point in the past image; - calculating a distance separating the positions of said point in said image and in the past image; and - calculation of speed data based on said distance and the duration separating the current and past moments.
[0030] For example, the speed of the target point in the first image can be calculated by dividing the distance between the positions of the target point in the first image and in the past image, captured by the first camera, by the time between the instants of capture of the first image and the past image. This provides a two-dimensional speed vector giving the speed of the target point in the plane of the first image.
[0031] Alternatively, or in addition, for at least one point of the second image, the speed of said point can be calculated by dividing the distance separating the positions of said point in the second image and in the past image, captured by the second camera, by the duration separating the instants of capture of the second image and the past image. This makes it possible to have a two-dimensional speed vector giving the speed of said point in the plane of the second image.
[0032] The identification, in a past image, of a point of a current image, can take into account velocity data associated with points in said past image, when these velocity data are known. Indeed, by knowing the velocities of at least some of the points in the past image, it is possible to determine which of these points can potentially be found in the current image at the location of said point. Thus, it is possible to determine a probable position of the point in the past image. The search for the point in the past image can start at the said probable position.
[0033] This solution can be used when determining the speed of the target point in the first frame.
[0034] This solution can be used when determining the speed of a point in the second image.
[0035] Generally speaking, this solution can be used to identify a point of a current image captured by a camera in a past image captured by said camera at a past time.
[0036] Furthermore, during the step of calculating the speed of a point in a current image, the identification of said point of the current image in a past image can be carried out by comparing the pixel corresponding to this point in said current image, to at least one pixel in said past image. The comparison can relate to at least one of the following respective data of said pixels: - color data, - texture data, - brightness data, - color data, - saturation data, - RGB data, - quantity calculated from at least one of these data, - etc.
[0037] Furthermore, optionally, the pixel's surroundings may also be taken into account during the comparison, to increase confidence in the comparison, such as, for example, the pixel's belonging to a similar geometric pattern (such as a line, of comparable direction) between the images, and / or a similar color configuration around the pixel.
[0038] According to embodiments, the step of identifying the target point in the second image may further comprise a comparison of acceleration data associated with the target point in said first image to an acceleration data of at least one point in said second image, said acceleration data being previously calculated.
[0039] The processing phase, or the method according to the invention, may comprise a step of calculating the acceleration of the target point in the first image.
[0040] Alternatively, the acceleration of the target point in the first image may be calculated during a prior step not forming part of the method according to the invention. For example, the first image may be a so-called enriched image, already comprising the acceleration associated with the target point. The data on the acceleration of the target point may, for example, be part of the data of the pixel corresponding to said target point in the first image.
[0041] The processing phase, or the method according to the invention, may comprise, for at least one point of the second image, a step of calculating the acceleration of said point in said second image.
[0042] Alternatively, for at least one point of the second image, the acceleration of said point in the second image may be calculated during a prior step not forming part of the method according to the invention. For example, the second image may be an image, called enriched, already comprising the acceleration associated with said point. For example, the data on the acceleration of this point may be part of the data of the pixel corresponding to said point in the second image.
[0043] For a point in a current image, the acceleration data can be determined by any known means.
[0044] According to embodiments, for a point of a current image, the acceleration of the point in said current image can be determined as a function of: - a speed of said point in said current image; - of a speed of said point in a past image captured by the same camera, at a past time relative to the capture time, called the current time, of said current image; and - of a duration elapsed between the current instant and the past instant
[0045] The velocity of the point in the current image can be calculated as described above. The velocity data of the point in the past image can be calculated in a similar way, using a so-called prior image captured by the same camera at a so-called prior time before the past time at which the past image was captured.
[0046] In other words, the acceleration of the target point in the first frame can be calculated as a function of: - the speed of said target point in a past image captured by the first camera at a past time preceding the current time of capture of the first image; - the speed of said target point in the first image; and - the duration separating the said past and current moments.
[0047] Alternatively, or in addition, for at least one point in the second image, the acceleration of said point in the second image may be calculated as a function of: - the speed of said point in a past image captured by the second camera at a past instant preceding the instant, called current, of capture of the second image; - the speed of said point in the second image; and - the duration separating the said past and current moments.
[0048] As indicated above, according to the invention, the target point is identified in the second image by comparing the speed data, and possibly the acceleration data, of said target point in the first image with speed data, and possibly acceleration data, of one or more points of the scene located in the second image.
[0049] According to embodiments, the step of identifying the target point in the second image may further comprise a comparison of the pixel corresponding to said target point in the first image, with at least one pixel in the second image, said comparison relating to at least one of the following data: - color data, - texture data, - brightness data, - color data, - saturation data, - RGB data, - data deduced from these data, - etc
[0050] Furthermore, optionally, the environment of the target point can also be taken into account during the comparison, to increase confidence in the comparison, such as for example the membership of the target point in a similar geometric pattern (such as a line, of comparable direction) between the images, and / or a similar color configuration around the target point.
[0051] So, when the target point is compared to a point in the second image, this comparison is made on the respective data: - speed, - and possibly acceleration and / or any combination of the data listed above. Thus, the identification of the target point on the second image can be performed more accurately.
[0052] According to embodiments, the search for the target point in the second image may begin at an arbitrary position in the second image.
[0053] Alternatively, the search for the target point in the second image can start at the position that this target point has in the first image.
[0054] According to other embodiments, the processing phase may comprise a selection of a target area in which said target point may be located in the second image, said selection being carried out according to: - a speed in the image plane, associated with said target point at an instant preceding the current instant, for example in a past image taken by the first camera or in a past image taken by the second camera; and - of a position of said target point in the first image, and / or in an image captured by the first camera an instant preceding the current instant and / or in an image captured by the second camera an instant preceding the current instant In this case, the search for the target point in the second image can start with the said area of interest.
[0055] Indeed, knowing the position of the target point in the first image, it is possible to find it in a past image taken by the first camera, at a past instant. It is possible that the target point was identified on a past image taken by the second camera at said past instant, for example during a previous iteration of the method according to the invention. It is furthermore possible that the speed of the target point was determined in the past image taken by the first camera, or in the past image taken by the second camera, for example during a previous iteration of the method according to the invention. Each of these data can be used, alone or in combination with another of said data, to estimate a probable position of the target point in the second image. During the identification step, the search for the target point in the second image can begin at said probable position, or in a target area defined around said probable position.
[0056] The invention has just been described with reference to a target point of a scene appearing in the first image.
[0057] Of course, the method according to the invention may comprise an iteration of the processing phase, in turn or simultaneously, for several target points of the first image. Thus, depth data may be determined for several target points of a scene appearing in the first image.
[0058] In particular, the processing phase can be implemented for each point of the first image. In this case, each point of the first image is a target point.
[0059] Alternatively, the processing phase can be implemented for previously selected target points within the first image. In this case, the method according to the invention may comprise a phase of selecting the target points in the first image.
[0060] The selection of target points in the first image can be done according to any technique / solution, or relationship.
[0061] According to embodiments, the phase of selecting the target points in the first image may comprise the following steps: - extraction at each point of said first image of a spatial gradient, in particular by derivation of said image, and even more particularly by Sobel filtering; - calculation of a norm of the gradient of each point, in particular the Euclidean norm; and - selection, as target points, of points whose gradient norm is greater than a predefined threshold.
[0062] Thus, the points of the first image for which the processing phase is carried out are the points of interest according to the norm of the spatial gradient. These points generally constitute the limits of the different objects of the scene and provide good information on the evolution of the scene. Thus, the method according to the invention avoids the processing of all the points of the scene located on the first image, but only the points of interest, which allows optimized processing of the first image in terms of time and computer resources.
[0063] The spatial gradient can be calculated on any of the quantities stored for each pixel. Alternatively, the spatial gradient can be calculated on a quantity itself deduced from at least one of the quantities stored for each pixel. For example, the spatial gradient can be calculated on the brightness, the latter being able to correspond to a weighted sum of the R, G and B components for each pixel, in the context of an RGB image.
[0064] According to embodiments, the method according to the invention may comprise, for at least one target point, a storage of the depth data of said target point, in association with said target point, in the first image and / or in the second image.
[0065] The depth data can be stored with the data of the pixel corresponding to said target point, in the first image and / or in the second image.
[0066] Furthermore, according to an advantageous optional characteristic, the method according to the invention can comprise a storage of: - speed data, and / or - the acceleration data, possibly calculated for the target point during the processing phase implemented for said target point, in the first image and / or in the second image, for example with the data of the pixel corresponding to said target point.
[0067] Of course, the method according to the invention can be implemented to determine depth information of a scene, at several instants in time, from at least two image streams of said scene taken by at least two cameras positioned in different positions, each image stream comprising at least one image of said scene taken for each of said several instants.
[0068] At at least one current time, the method according to the invention can be implemented to determine a depth of a point for which no depth data has been determined at past times. This makes it possible to take into consideration new points of the scene.
[0069] Alternatively or in addition, at at least one current instant, the method according to the invention can be implemented to determine a depth of a target point of the scene for which depth data has been determined at at least one past instant. Thus, at said current instant, the depth of this target point can be known at said current instant but also at at least one past instant, which makes it possible to follow the evolution of the depth of this target point over time.
[0070] Furthermore, advantageously, by knowing the depth of the same target point at different times, it is possible to calculate: - the speed of said speed target point in the direction perpendicular to the image plane, for example in the frame linked to the scene; - and possibly the acceleration of said target point in said direction. The processing phase may further comprise a step of calculating such a speed, and / or a step of calculating such an acceleration, for at least one target point. The method according to the invention, or the processing phase, may further comprise a storage of this speed, and / or this acceleration, in the first image, and / or in the second image, in particular in the data of the pixel corresponding to said point in said image.
[0071] According to another aspect of the invention, an enriched image of a scene is provided, stored on a storage means, obtained by the method according to the invention.
[0072] The enriched image according to the invention is a digital image, preferably a matrix image.
[0073] The storage medium can be of any type, portable or not, such as a memory card, a USB key, a computer, a server, a telephone, a camera, etc.
[0074] According to another aspect of the invention, a use of enriched image(s) according to the invention is proposed for at least one of the following applications: - detection of an object in a scene at a given moment, - tracking an object from a scene in time, - assistance with driving, or for autonomous or semi-autonomous driving, of a vehicle.
[0075] Indeed, at least one depth data associated with a target point can be used to detect an object in this image because it is likely that the points having the same depth, and possibly a same speed and / or the same acceleration, belong to the same object. It is therefore possible to discriminate and / or detect objects on an image enriched according to the invention based on depth data, and possibly speed data and / or acceleration data.
[0076] Furthermore, depth data associated with a target point at multiple times can be used for time tracking of that target point, and thus for time tracking of the object to which it belongs.
[0077] Furthermore, as indicated above, the depth data of a target point at two different times can be used to calculate at least the speed of this point in the depth direction, that is to say in the direction perpendicular to the image plane. This speed data can be used to estimate, or predict, the probable position of a target point, and therefore of the object to which it belongs, in the near future, and in particular at a following time, in said depth direction. This makes it possible to improve the tracking of an object in time and space. In addition, this makes it possible to provide assistance in piloting a vehicle or a robot.
[0078] According to another aspect of the invention, there is provided a computer program comprising executable instructions which, when executed by computer, implement all the steps of the method according to the invention.
[0079] The computer program can be in any computer language, such as machine language, C, C++, JAVA, Python, etc.
[0080] According to another aspect of the invention, a device is provided comprising means configured to implement all the steps of the method according to the invention.
[0081] The device according to the invention can be any type of device such as a server, a computer, a tablet, a calculator, a processor, a computer chip, a camera, programmed to implement the method according to the invention.
[0082] For example, the device according to the invention may be provided with the computer program according to the invention.
[0083] In particular, the processing device may be integrated into, or may be, one of the first and second cameras, and more generally, one of the cameras used to image the scene and providing the processed image(s).
[0084] According to another aspect of the present invention, there is provided a system comprising: - at least two camera modules designed to image a scene from two different capture positions, and - a device comprising means configured to implement all the steps of the method according to the invention, or a device according to the invention.
[0085] Each camera module may include at least one lens and at least one image sensor such as a CCD or CMOS sensor.
[0086] According to another aspect of the present invention, there is provided a vehicle equipped with an imaging system according to the invention.
[0087] According to embodiments, the vehicle may be a land vehicle, such as a car, autonomous or not.
[0088] According to embodiments, the vehicle may be a flying vehicle, such as a drone, an airplane, a helicopter, autonomous or not.
[0089] According to embodiments, the vehicle may be a maritime vehicle, such as a boat or a submarine, autonomous or not.
[0090] According to yet another aspect of the present invention, there is provided a method for assisting in driving a vehicle, comprising at least one iteration of the following steps: - determination, by the method according to the invention, of depth data or distance data, associated with at least one target point, and in particular to an object to which said at least one target point belongs, of a scene imaged from said vehicle, and - generation of at least one instruction relating to the driving of said vehicle as a function of said depth data and / or said distance data.
[0091] According to another aspect of the present invention, there is provided a user apparatus equipped with an imaging system according to the invention.
[0092] The user device can be a camera, a smartphone, a tablet, a virtual reality headset, an augmented reality headset, etc.
[0093] According to another aspect of the present invention, there is provided a medical imaging apparatus equipped with an imaging system according to the invention.
[0094] The medical imaging device can be an endoscope, especially a disposable one. Description of figures and embodiments
[0095] Other advantages and characteristics will appear on examining the detailed description of non-limiting embodiments, and the attached drawings in which: - FIGURE 1 is a schematic representation of a non-limiting exemplary embodiment of a phase of selecting target points in an image, which can be implemented in the present invention; - FIGURE 2 is a schematic representation of a non-limiting exemplary embodiment of a speed calculation step that can be implemented in the present invention; - FIGURE 3 is a schematic representation of a non-limiting exemplary embodiment of a depth calculation that can be implemented in the present invention; - FIGURE 4 is a schematic representation of a non-limiting exemplary embodiment of a method according to the present invention; - FIGURE 5 is a schematic representation of a non-limiting exemplary embodiment of a method according to the invention applied to image streams; - FIGURE 6 is a schematic representation of a non-limiting exemplary embodiment of a device according to the invention; - FIGURE 7 is a schematic representation of a non-limiting exemplary embodiment of an imaging system according to the invention; and - FIGURE 8 is a schematic representation of a non-limiting exemplary embodiment of a car according to the invention.
[0096] It is understood that the embodiments which will be described below are in no way limiting. In particular, it is possible to imagine variants of the invention comprising only a selection of characteristics described below isolated from the other characteristics described, if this selection of characteristics is sufficient to confer a technical advantage or to differentiate the invention compared to the state of the prior art. This selection includes at least one preferably functional characteristic without structural details, or with only part of the structural details if it is this part which is only sufficient to confer a technical advantage or to differentiate the invention compared to the state of the prior art.
[0097] In particular, all the variants and embodiments described can be combined with each other if there is no technical obstacle to this combination.
[0098] In the figures and in the rest of the description, the elements common to several figures retain the same reference.
[0099] FIGURE 1 is a schematic representation of a non-limiting exemplary embodiment of a phase of selecting target points in an image, which can be implemented in the present invention.
[0100] The selection phase 100 of FIGURE 1 can be used to select target points, in a first image of a scene taken by a first camera arranged in a first capture position, for which a depth is determined by the method according to the invention using a second image of said scene taken by a second camera, arranged in a second capture position.
[0101] The selection phase 100 of FIGURE 1 aims to select the points of interest in the first image. The points of interest are generally the boundaries of the different objects in the scene and provide good information on the evolution of the scene. Thus, the selection phase 100 aims to avoid processing all the points in the scene located in the first image, in the rest of the process, which allows for optimized processing of the first image in terms of time and computing resources.
[0102] The point(s) of the scene in the image can be selected according to any logic, or relationship, during the selection phase 100.
[0103] In the following, the first image of the scene is noted IM1.
[0104] According to embodiments, the selection phase 100 comprises a step 102 of extracting at each point of the first image IM1 a spatial gradient by Sobel filtering, for example with respect to the brightness.
[0105] Then, during a step 104, for each point of the first image IM1, a norm, in particular a Euclidean norm, is calculated for each gradient calculated in step 102.
[0106] Then, in a step 106, each point whose gradient norm is greater than a predefined selection threshold is selected as a target point.
[0107] The selection threshold may be determined by trial. Alternatively, the selection threshold may be calculated based on the Euclidean norms calculated for all points in the first image. For example, the selection threshold may be an average of the Euclidean norms calculated in step 104.
[0108] Thus, the points of the first image for which the processing phase will be carried out are those of interest depending on the spatial gradient norm. As indicated above, these points generally constitute the limits of the different objects in the scene and provide good information on the evolution of the scene.
[0109] FIGURE 2 is a schematic representation of a non-limiting exemplary embodiment of a speed calculation step that may be implemented in the present invention.
[0110] Step 200 of FIGURE 2 can be used to calculate, for a target point of a scene appearing in an image, called the current image, denoted IMC, a speed associated with said target point in the image plane, denoted UV plane in the frame associated with the image sensor. The current image IMC is captured by a camera at an instant, called the current instant. [YES] The speed calculation uses an image, called a past image, noted IMP, captured by the same camera at a time in the past relative to the current time. Preferably, the camera is stationary. In the case where the camera is mobile between the past and current times, then its movement is taken into account in the speed calculation.
[0112] The calculation step 200 comprises a step 202 of identifying the target point of the current image, in the past image. The identification is carried out by comparing the pixel corresponding to the target point in the current image to at least one pixel of the past image. The comparison relates to at least one of the following data: - color data, - texture data, - brightness data, - color data, - saturation data, - RGB data, - at least one quantity calculated from at least one of these data, - etc.
[0113] The target point can be compared to all points in the past image, to identify it in the past image.
[0114] Alternatively, the identification step 202 may comprise an optional step 204 of determining a target area in the past image and in which the target point may be located. The identification of the target area may be carried out from the speed data associated with points in said past image, when these speed data are known. Indeed, by knowing the speeds of at least some of the points in the past image, it is possible to determine which of these points may potentially be located in the current image at the location of the target point in said current image. Thus, it is possible to determine a probable position of the target point in the past image. A target area, for example of a predetermined size / surface, may be defined around said probable position. The search for the target point in the past image may begin at said probable position, or in said target area.
[0115] In a step 206, the target point is compared to one or more points in the past image: - starting from an arbitrary position, if optional step 204 is not performed, or - starting with the probable position, or with the target area, if optional step 204 is carried out.
[0116] When the comparison does not allow the target point to be identified, for example because the past image does not contain any point identical, or sufficiently similar, to the target point, it is terminated at step 200.
[0117] Otherwise, step 200 is continued.
[0118] During a step 208, the speed, noted VCuv, of the target point in the image plane, expressed in the UV reference frame associated with the image sensor, is calculated, for example using the following relationship: VCuv = (PPCuv-PPPuv) / (Ic-Ip) with - PPCuv the position of the target point in the current image, expressed in the UV reference associated with the image sensor; - PPPuv the position of the target point in the past image, expressed in the UV reference associated with the image sensor; - Here the current moment, and - Ip the moment passed.
[0119] When the speed, noted VP uv , of the target point in the UV image plane is known for the past image IMP, then step 200 can comprise an optional step 210 of calculating the acceleration, noted ACuv, of the target point, in the UV image plane, expressed in the frame linked to the image sensor, for the current image IMC, for example using the following relation: ACuv = (VCuv-VPuv) / (Ic-Ip)
[0120] During an optional step 212, the speed VCuv and / or the acceleration ACuv, associated with the target point in the current image, can be stored, for example directly in the current image, for example in the data of the pixel corresponding to said target point in the current image IMC.
[0121] FIGURE 3 is a schematic representation of a non-limiting exemplary embodiment of a step of calculating the depth of a point of a scene appearing on two images of this scene, which can be implemented in the present invention.
[0122] FIGURE 3 shows a configuration in which a scene is imaged with a first camera 302 providing a first image IM1 of the scene and a second camera 304 providing a second image IM2 of the scene. Each of the cameras 302 and 304 is represented very schematically.
[0123] The respective positions of the cameras 302 and 304 are known so that the distance, noted Dcam, between the positions of the cameras is known at the time of capturing the images IM1 and IM2.
[0124] A point 306 of a scene is imaged by the camera 302, is located at a position Pl = {ul,vl}, expressed in the UV reference associated with the image sensor, on the sensor of the camera 302: the point 306 therefore corresponds to the pixel located at the position PI in the image IM1. When the scene is imaged by the camera 304, said point 306 of the scene is located at a position P2 = {u2,v2} on the sensor of camera 304: point 306 therefore corresponds to the pixel located at position P2 in image IM2. We note d = Pl-P2.
[0125] Assuming that the focal length f of cameras 302 and 304 is the same, and making the approximation that: - the respective sensors of the cameras 302 and 304 are positioned in the same plane, and - the focal length f is very small compared to the distance D z between each of the cameras 302-304 and the point 306, in the Z direction perpendicular to the plane of the image sensor; said distance D z can be calculated according to the following relationship: D z = Dcam.f / d Since the focal length f is very small compared to D z , then D z corresponds to the distance between the plane of each sensor and point 306 in the scene, and therefore to the depth of said point on each of the images IM1 and IM2.
[0126] The technique described with reference to FIGURE 3 is in no way limiting. It can be used to calculate the depth associated with a target point in a scene using a first image and a second image of said scene.
[0127] FIGURE 4 is a schematic representation of a non-limiting exemplary embodiment of a method according to the invention.
[0128] The method 400 of FIGURE 4 can be used to determine depth information for at least one point, called target point, of a scene appearing in a first image, denoted IM1, of said scene captured by a first camera from a first capture position, using a second image, denoted IM2, of said scene taken by a second camera from a second capture position, different from said first capture position.
[0129] Images IM1 and IM2 can be captured at different times, but close together so that the movements of the objects in the scene are negligible between said different times. Preferably, images IM1 and IM2 are captured at the same time, called the current time in the following.
[0130] The method 400 may optionally comprise a step 402 of calculating a speed, and possibly an acceleration, in the UV image plane, associated with at least one point, and in particular with several points, of the scene appearing in the second image IM2, expressed in the reference frame associated with the image sensor. For at least one point of the second image IM2, the speed, and possibly the acceleration, in the UV image plane may be calculated by any technique. According to a non-limiting exemplary embodiment, step 402 may be step 200 of FIGURE 2. In this case, for at least one point of the second image IM2, the calculation of the speed, and possibly the acceleration, is carried out using a past image captured by the second camera at a past time relative to the current time of capture of the second image.
[0131] Alternatively, the speed, and possibly the acceleration, in the UV image plane associated with at least one point of the second image IM2 may be previously calculated, or measured, and provided in association with said second image, for example stored in said second image, in particular in the data of the pixel corresponding to said point.
[0132] Optionally, but particularly advantageously, the method 400 may comprise a phase 404 for selecting, in the first image IM1, one or more target points for which depth information is determined by the method according to the invention. For example, phase 404 may be phase 100 of FIGURE 1.
[0133] Alternatively, depth information can be determined for each of the points in the first image IM1.
[0134] The method 400 comprises a processing phase 410 for calculating, for a target point of the scene appearing in the first image IM1, depth information of said target point in the first image IM1, using the second image IM2.
[0135] Phase 410 is performed individually for each target point. When several target points are to be treated, a treatment phase 410 is carried out for each of said target points in turn, or simultaneously.
[0136] The processing phase 410 may comprise an optional step 412 for determining a speed VCuv of the target point in the UV image plane in the first image IM1. The speed VCuv may be determined using any known technique. According to a non-limiting exemplary embodiment, the speed VCuv may be determined by / according to step 200 of FIGURE 2. In this case, the calculation of the speed of said target point in the UV image plane is carried out using a past image captured by the first camera at a past time relative to the current time of capture of the first image.
[0137] Alternatively, the velocity VCuv of the target point in the UV image plane may be previously calculated, or measured. In this case, said velocity VCuv may be provided in association with said first image, for example stored in said first image, in particular in the data of the pixel corresponding to said target point in the first image.
[0138] The processing phase 410 may comprise an optional step 414 for determining an acceleration, denoted ACuv of the target point in the UV image plane in the first image IM1. The acceleration ACuv may be determined using any known technique. According to a non-limiting exemplary embodiment, the acceleration ACuv may be determined as described above with reference to step 210 of FIGURE 2.
[0139] Alternatively, the acceleration velocity ACuv of the target point in the UV image plane may be previously calculated, or measured. In this case, said acceleration ACuv may be provided in association with said first image, for example stored in said first image, in particular in the data of the pixel corresponding to said target point in the first image.
[0140] In some embodiments, optional steps 412 and 414 may be combined into a single step such that the same step calculates both the velocity and acceleration of the target point in the first image. Such a step may be identical to step 200 of FIGURE 2.
[0141] The processing phase 410 then comprises a step 416 of identifying the target point in the second image.
[0142] The identification step 416 may comprise an optional step 418 for identifying a target area in the second image IM2, in which the target point may be located. This identification may be carried out according to: - a speed associated with the target point at a time preceding the current time, for example in a past image taken by the first camera or in a past image taken by the second camera; and - a position of said target point in the first image, and / or in a past image captured by the first camera at a past time preceding the current time and / or in an image captured by the second camera at a time preceding the current time.
[0143] Indeed, knowing the position of the target point, it is possible to find it in a past image taken by the first camera, at a past time. If in addition, this target point has already been identified on a past image taken by the second camera at said past time, for example during a previous iteration of the method. It is furthermore possible that the speed of the target point was determined in the past image taken by the first camera, or in the past image taken by the second camera, for example during a previous iteration of the method according to the invention. At least one of these data can be used to estimate a probable position of the target point in the second image.
[0144] A target area of a predetermined size may be defined around said probable position thus determined.
[0145] The identification step 416 comprises a step 420, during which the target point is searched for in the second image IM2. This search is carried out by comparing the target point to at least one point in the second image.
[0146] According to the invention, the comparison comprises a comparison of the velocity VCuv of the target point in the first image to the velocity VCuv of the point in the second image IM2.
[0147] Optionally, the comparison may include a comparison of the acceleration ACuv of the target point in the first image to that ACuv of the point in the second image IM2.
[0148] Optionally, but preferably, the comparison may comprise a comparison of the pixel corresponding to the target point in the first image, to the pixel associated with the point in the second image compared to said target point, said comparison relating to at least one of the following data: - color data, - texture data, - brightness data, - color data, - saturation data, - RGB data, - data deduced from these data, - etc.
[0149] When the search carried out during step 420 does not identify, in the second image IM2, any point which is identical to, or sufficiently close to, the target point of the first image IM1, then the search has not been able to identify the target point in the second image IM2. The processing phase 410 for this target point is then stopped.
[0150] When the search carried out during step 420 identifies, in the second image, a point identical to, or sufficiently close to, the target point of the first image, then the search can be stopped and the target point is identified in the second image IM2.
[0151] When the optional step 418 is carried out, the search for the target point in the second image IM2, during step 420, can begin with the probable position, or in the target area, identified during said optional step 418.
[0152] When the optional step 418 is not performed, the search for the target point in the second image IM2, during step 420, may start with a random position, or the same position as the target point has in the first image IM1, etc.
[0153] At the end of the identification step 416, and when the target point has been identified in the second image IM2, the position of said target point is known: - in the first image IM1, expressed in a reference associated with the image sensor of the first image IM1, and - in the second image IM2, expressed in a reference associated with the image sensor of the second image IM2
[0154] The processing phase 410 then comprises a step 424 of calculating the depth associated with the target point in the first image IM1, as a function of the position of the target point in the first image IM1 and the position of the target point in the second image IM2.
[0155] This depth calculation can be carried out by a classic calculation technique, and in particular by triangulation, using: - the position of the target point in the first image IM1, - the position of the target point in the second image IM2, - the distance between the first camera and the second camera when taking the first image IM1 and the second image IM2, and - the focal length of said cameras.
[0156] In particular, the calculation of the depth can be carried out according to the technique described with reference to FIGURE 3.
[0157] When the depth associated with the target point is calculated for the target point, the processing phase 410 can end for this target point. A new iteration of the processing phase 410 can be carried out for another target point in the first image IM1. And so on.
[0158] The method 400 may comprise an optional step 430 of memorization, in the first image IM1, for at least one target point: - the depth data calculated in step 424 for said target point, - and possibly the speed data and / or the acceleration data, calculated in optional step 412 and / or optional step 414. This, or these, data may be stored in particular with the data of the pixel corresponding to the target point. Thus, step 430 may provide a first enriched image, denoted IM1*, with said data for one or more target points in said image.
[0159] The optional step 430 can also carry out a storage, in the second image IM2, for at least one target point, of at least one of the following data: - depth data calculated in step 424 for said target point, - and possibly the speed data and / or the acceleration data, calculated in optional step 412 and / or 414 for said target point. This, or these, data may be stored in particular with the data of the pixel corresponding to said target point in the second image.
[0160] The optional step 430 can also carry out a storage, in the second image IM2, for at least one point of the second image IM2, of at least one of the following data: - speed data, and / or - acceleration data; calculated in optional step 402 for said point of the second image IM2. This, or these, data may be stored in particular with the data of the pixel corresponding to said point in the second image. Thus, step 430 can provide a second enriched image, denoted IM2*, with said data.
[0161] FIGURE 5 is a schematic representation of a non-limiting exemplary embodiment of determining depth information with image streams.
[0162] In other words, FIGURE 5 shows an example of application of a method, and in particular of the method 400 of FIGURE 4, for determining depth information from two image streams 502 and 504, and in particular two videos, provided by two separate cameras for the same scene.
[0163] The image stream 502 comprises an image 502i captured at a time tl, an image 5022 captured at a time t2 following the time tl, an image 502s captured at a time t3 following the time t2, and so on up to an image 502m captured at a time tm.
[0164] The image stream 504 comprises an image 504i captured at time tl, an image 5042 captured at time t2 following time tl, an image 504s captured at time t3 following time t2, and so on up to an image 504m captured at time tm.
[0165] In the example shown in FIGURE 5, image 502i is not processed by the method according to the invention. Similarly, image 504i is not processed by the method according to the invention.
[0166] A first iteration of the 400 process, noted 400i, is carried out by considering: - time t2 as current time and time tl as past time; - image 5022 as the first image, and image 502i as the past image relative to the first image 5022; and - image 5042 as the second image, and image 504i as the image passed relative to the first image 5042. This first iteration 400i of the method 400 provides depth data, each associated with a target point of the scene appearing in the first image 5022, and in the second image 5042. Optionally, the depth data, and possibly speed data, associated with the target points are stored in a first enriched image noted 5022*, and / or a second enriched image noted 5042*.
[0167] A second iteration 4002 of the method 400 is carried out by considering: - time t3 as current time and time t2 as past time; - image 502s as the first image, and image 5022 or preferably the enriched image 5022*, as the past image relative to the first image 502s; and - image 504s as the second image, and image 5042 or preferably the enriched image 5042*, as the past image relative to the second image 504s. This second iteration 4002 of the method 400 provides depth data, each associated with a target point of the scene appearing in the first image 502s and in the second image 504s. Optionally, the depth data, and possibly speed data and / or acceleration data, associated with the target points can be stored in a first enriched image denoted 502s*, and / or in a second enriched image denoted 504s*.
[0168] The method 400 can thus be repeated until, potentially, the remaining images of the streams 502 and 504 have been processed.
[0169] Optionally, 502i, 5022*-502 images m * can be stored as the first enriched stream 502*. Images 504i, 5042*-504 m * can be stored as a second 504 enriched stream*.
[0170] Each enriched image obtained by the method according to the invention, namely image IM1*, or IM2*, or each of the images 5022*-502 m * or each of the images 5042*-504 m *, can be used to identify, in said image, objects using the depth data calculated for said image.
[0171] Additionally, images 5022*-502 m * from stream 502*, or images 5042*- 504m* from stream 504*, can be used to track objects in these images using the calculated depth data.
[0172] Furthermore, in the image stream 502*, respectively in the image stream 504*, the variation, over time, of the depth of a target point can be used to predict the depth of said target point, and therefore of an object to which this target point belongs, at a future time. This prediction can be used to generate a command or instruction, for example within a vehicle. Such a command or instruction can be a driving instruction, emergency braking or an object avoidance maneuver, or even a trajectory modification for example.
[0173] FIGURE 6 is a schematic representation of a non-limiting exemplary embodiment of a device according to the present invention.
[0174] The device 600 is configured to implement the method according to the invention, and in particular the method 400 of FIGURE 4, for determining depth information of at least one target point of a scene appearing in a first image of the scene captured by a first camera, using a second image of said scene captured by a second camera.
[0175] The device 600 may optionally comprise a module 602 for processing the second image IM2 and calculating the speed, and possibly an acceleration, in the XY image plane, of at least one point of the scene appearing in said second image. In particular, the module 602 may be configured to implement step 402 of the method 100 of FIGURE 1.
[0176] The device 600 may optionally comprise a module 604 for pre-processing the first image IM1, in order to select the target points which are of interest in said first image IM1. In particular, the module 604 may be configured to implement step 404 of the method 400 of FIGURE 4.
[0177] The device 600 comprises a module 610 for processing at least one target point in the first image to determine depth data associated with said target point. The processing module 610 can be configured to implement the processing phase of the method according to the invention and in particular phase 410 of the method 400 of FIGURE 4.
[0178] The module 610 may optionally comprise a module 612 for determining a speed associated with the target point in the image plane. In particular, the module 612 may be configured to implement step 412 of the method 400 of FIGURE 1.
[0179] The module 610 may optionally comprise a module 614 for determining an acceleration associated with the target point in the image plane. In particular, the module 614 may be configured to implement step 414 of the method 400 of FIGURE 1.
[0180] The module 610 may optionally comprise a module 618 for determining a target area in the second image IM2. In particular, the module 618 may be configured to implement step 418 of method 400 of FIGURE 4.
[0181] The module 610 comprises a module 620 for identifying the target point in the second image IM2. In particular, the module 620 can be configured to implement step 420 of the method 400 of FIGURE 4.
[0182] The processing module 610 comprises a module 622 for calculating the depth of the target point in the image IM1. In particular, the module 622 can be configured to implement step 424 of the method 400 of FIGURE 4.
[0183] The device 600 may optionally comprise an enrichment module 630 for adding, in the first image IM1, and / or in the second image IM2, depth data and possibly speed data and / or acceleration data, associated with at least one point of the scene appearing in said image. In particular, the module 630 may be configured to implement step 430 of the method 400 of FIGURE 4.
[0184] At least one of the modules 602-604, 610-614, 618-622 and 630 may be an individual module, independent of the other modules. Alternatively, at least two of these modules may be integrated into a single module. In particular, all the modules of the device 600 may be integrated into a single module.
[0185] At least one of the modules 602-604, 610-614, 618-622 and 630 may be a hardware module, such as a processor, an electronic chip, a graphics card, a calculator, a computer, a server, etc.
[0186] Alternatively, at least one of the modules 602-604, 610-612, 618-622 and 630 may be a software module executed by at least one computer means.
[0187] According to yet another alternative at least one of the modules 602-604, 610-614, 618-622 and 630 may be a combination of at least one hardware module with at least one software module.
[0188] The configuration of one, or each, of the modules 602-604, 610-614, 618-622 and 630 may be a software configuration, such as an application, or a computer program, loaded into said module, and / or a hardware configuration such as a specific hardware architecture.
[0189] FIGURE 7 is a schematic representation of a non-limiting exemplary embodiment of an imaging system according to the present invention.
[0190] The system 700 may be configured to implement a method according to the invention, and in particular the method 400 of FIGURE 4.
[0191] The system 700 comprises a first camera module 702 comprising at least one lens and at least one image sensor, for example a CCD or CMOS sensor. The first camera module 702 makes it possible to capture a first image, or a first stream of images, of a scene.
[0192] The system 700 comprises a second camera module 704 comprising at least one lens and at least one image sensor, for example a CCD or CMOS sensor. The second camera module 704 makes it possible to capture a second image, or a second stream of images, of the scene.
[0193] The imaging system 700 further comprises a processing unit 706, which may for example be the device 600 of FIGURE 6, for determining depth information for at least one target point of a scene appearing in the images captured by the modules 702-704.
[0194] At least one of the camera modules 702-704 and the processing unit 706 may be integrated within a single unit. Alternatively, at least one camera module 702-704 and the processing unit 706 may not be integrated into a single unit and may be independent of each other. In this case, said at least one camera module and the processing unit are in communication, direct or indirect, wired or wireless.
[0195] FIGURE 8 is a schematic representation of a non-limiting exemplary embodiment of a vehicle according to the present invention.
[0196] The vehicle 800 of FIGURE 8 may be any type of land vehicle, such as, for example, a car, a truck, etc.
[0197] The vehicle 800 may be a vehicle, partially or totally, autonomous or not.
[0198] In FIGURE 8, the vehicle 800 is shown viewed from the front.
[0199] The vehicle 800 is equipped with an imaging system according to the invention, and in particular by the system 700 of FIGURE 7. In general, the vehicle 800 can be equipped with: - at least two camera modules, for example camera modules 702 and 704; - at least one processing unit 706.
[0200] Of course, the camera modules 702-704 and the processing unit 706 can be arranged at any suitable location in the vehicle. In the example shown, the camera modules 702-704 are arranged in the front part of the vehicle and make it possible to image the scene in front of the vehicle 800.
[0201] The processing unit 706 may be an individual and independent unit. Alternatively, the processing unit 706 may be integrated into a computer or a management unit (not shown) of the vehicle 800.
[0202] The vehicle 800 may include a driving assistance unit 802, receiving the enriched images from the processing unit 706, or the depth values, or any other quantity determined from the depth values, and providing driving instructions either to the driver of the vehicle, or directly to components of the vehicle.
[0203] Of course, the invention is not limited to the examples which have just been described.
Claims
CLAIMS 1. Method (400) for determining depth information of a scene in a first image (IM1) of said scene captured by a first camera (302) positioned in a first capture position, said method (400) comprising, for a target point of said scene located in said first image (IM1), a processing phase (410) comprising the following steps: - identification (416) of said target point on a second image (IM2) captured by a second camera (304) positioned at a second capture position different from said first position, and - calculation (424) of a depth associated with said target point as a function of a position of said target point on each of said first and second images (IM1, IM2); characterized in that the identification step (416) comprises a comparison of a speed data item of said target point in said first image (IM1) with a speed data item of at least one point in said second image (IM2), said speed data being previously calculated.
2. Method (400) according to the preceding claim, characterized in that the speed data associated with a point of an image is calculated from said image and at least one image, called past, captured by the same camera at a past instant preceding the instant, called current, of capture of said image.
3. Method (400) according to the preceding claim, characterized in that the step (200) of calculating the speed data of a point in an image comprises the following steps: - identification (202) of said target point in the past image; - calculating a distance separating the positions of said point in said image and in the past image; and - calculation of speed data based on said distance and the duration separating the current and past moments.
4. Method (400) according to any one of the preceding claims, characterized in that the step of identifying (416) the target point in the second image comprises a comparison of an acceleration data item associated with said target point in said first image (IM1) with an acceleration data item of at least one point in said second image (IM2), said acceleration data being previously calculated.
5. Method (400) according to any one of the preceding claims, characterized in that the step of identifying (416) the target point on the second image (IM2) further comprises a comparison of a pixel corresponding to said target point in the first image (IM1), with at least one pixel in the second image (IM2), said comparison relating to at least one of the following data: - color data, - texture data, - brightness data, - color data, - saturation data, - RGB data, - data deduced from these data, - etc.
6. Method (400) according to any one of the preceding claims, characterized in that the processing phase comprises a step of identifying (416) the target point on the second image (IM2) comprises a selection of a target area in which said target point may be located in the second image (IM2), said selection being carried out as a function of: - a speed in the image plane, associated with said target point at an instant preceding the current instant; and - of a position of said target point in the first image (IM1), and / or in an image captured by the first camera (302) an instant preceding the current instant and / or in an image captured by the second camera (304) an instant preceding the current instant 7. Method (400) according to any one of the preceding claims, characterized in that it comprises an iteration of the processing phase (410) for several target points of the first image (IM1).
8. Method (400) according to any one of the preceding claims, characterized in that it comprises a phase (100; 404) of selection in the first image (IM1) of at least one target point, said selection step (100; 404) comprising the following steps: - extraction (102) at each point of said first image (IM1) of a spatial gradient, in particular by derivation of said image, and even more particularly by Sobel filtering; - calculation (104) of a norm of the gradient of each point, in particular the Euclidean norm; - conservation (106), as target points, the points whose gradient norm is greater than a predefined threshold.
9. Method (400) according to any one of the preceding claims, characterized in that it is implemented to determine a depth for at least one point of a scene, at several instants in time, from at least two image streams (502,504) of said scene taken by two cameras positioned in different positions, each image stream comprising at least one image of said scene taken for each of said several instants.
10. Method (400) according to any one of the preceding claims, characterized in that it further comprises, for at least one target point, a storage (430) of the depth data of said target point, in association with said target point, in the first image (IM1) and / or in the second image (IM2).
11. Image (IM1*,IM2*) enriched with a scene, stored on a storage means, obtained by the method according to the preceding claim.
12. Use of enriched image(s) (IMl*,IM2*;5022*-502 m * ;504z*- 504m*) of a scene according to claim 11 for at least one of the following applications: - detection of an object in a scene at a given moment, - tracking an object from a scene in time, - assistance with driving, or for autonomous or semi-autonomous driving, of a vehicle.
13. A computer program comprising executable instructions which, when executed by a computer, implement all the steps of the method (400) according to any one of claims 1 to 10.
14. Device (600) comprising means configured to implement all the steps of the method (400) according to any one of claims 1 to 10.
15. System (700) comprising: - at least two camera modules (702,704) designed to image a scene from two different capture positions, and - a device (706) comprising means configured to implement all the steps of the method (400) according to any one of claims 1 to 10.
16. User apparatus comprising an imaging system (700) according to the preceding claim.
17. Vehicle (800) comprising an imaging system (700) according to claim 15.