Method and device for rapidly determining the distance between an object and a vehicle by using a bounding box.

The method enhances ADAS systems by using a stereoscopic vision system with iterative resolution adjustment for fast and accurate depth prediction, addressing the challenge of high-speed object detection in vehicles.

FR3164307B1Active Publication Date: 2026-05-22STELLANTIS AUTO SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
STELLANTIS AUTO SAS
Filing Date
2024-07-05
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing Advanced Driver-Assistance Systems (ADAS) in vehicles face challenges in accurately and quickly determining distances and depths from objects due to high-speed movement, requiring improved image processing times and depth prediction models to ensure rapid and effective vehicle reactions.

Method used

A method using a stereoscopic vision system with two cameras to acquire images of a three-dimensional scene, reducing image resolution for nearby objects and increasing resolution only where necessary for accurate depth prediction, combined with a depth prediction model to iteratively enhance resolution for distant objects, ensuring fast and precise depth determination.

Benefits of technology

This approach allows for rapid and accurate depth prediction across various conditions, enhancing the responsiveness and reliability of ADAS systems by minimizing computation time while maintaining high accuracy, especially for images with distortion.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A method or device for determining the distance between an object and a vehicle equipped with a vision system comprising several cameras observing an object in a three-dimensional scene. The method includes receiving (21) and reducing (22) the resolution of the images of a pair of images, and predicting (23) the depths associated with the pixels of one of the images whose resolution is reduced. In areas where a predicted depth is greater than a threshold depth, the resolution is iteratively increased (25) to predict (26) depths more accurately, and the threshold depth is updated at each iteration based on the increased resolution. Once the depths have been predicted for all the pixels of an image, pixels corresponding to the observed object are detected (27) in the image, and the distance is determined (28) based on the depths associated with the pixels corresponding to the object. Figure 2 (for the abstract)
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for rapidly determining the distance between an object and a vehicle using a bounding box. Technical field

[0001] The present invention relates to methods and devices for determining the distance separating an object from a vision system mounted in a vehicle, for example in a motor vehicle.

[0002] The present invention also relates to a method and device for predicting the depth associated with a pixel of an image acquired by a vision system embedded in a vehicle. Technological background

[0003] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). These ADAS are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS use advanced technologies to assist the driver while driving and thus improve performance. ADAS use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.

[0004] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.

[0005] The ADAS systems installed in a vehicle are powered by data obtained from one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or any obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle's speed; - to act on the braking system in case of risk of impact with an object.

[0006] The quality of the data emitted by a vision system therefore determines the proper functioning of the driving assistance devices using this data.

[0007] A vision system installed in a vehicle makes it possible to determine the depth or distance separating the vehicle from an object in the vehicle's environment. However, when a vehicle is moving at high speed At speeds such as on the highway, the prediction time for depths and / or distances is crucial. Indeed, the reaction time of an ADAS (Advanced Driver Assistance System) depends on image processing time and the depth prediction time of the model associated with the vision system. This latter time must therefore be minimized to allow the vehicle, via the ADAS controls, to react quickly, while maintaining its effectiveness by ensuring sufficient accuracy in the predicted depths or distances. Summary of the present invention

[0008] One object of the present invention is to solve at least one of the problems of the technological background described above.

[0009] Another object of the present invention is to improve the quality of data from the processing of an image acquired by a vision system, in particular by a depth prediction model implemented by a neural network associated with a stereoscopic vision system, which is used in such a way as to be both fast and accurate for any type of image received, including for an image exhibiting strong distortion.

[0010] Another object of the present invention is to improve the processing speed of an image acquired by a vision system, in particular by limiting the resolution of a region of the processed image to predict a depth associated with pixels or a distance associated with an object. The vision system is then more efficient, making ADAS (Automatic Detection and Analysis System) systems using data from this vision system more responsive.

[0011] Another object of the present invention is to improve road safety, in particular by improving the reliability of AD AS systems powered by data obtained from a vision system.

[0012] According to a first aspect, the present invention relates to a method for determining the distance separating an object from a vehicle carrying a vision system comprising a first and a second camera arranged so as to each acquire an image of a three-dimensional scene including said object, the method being implemented by at least one processor, and being characterized in that the determination of the distance comprises the following steps: - reception of a pair of images comprising a first image acquired by the first camera and a second image acquired by the second camera, the first and second images being acquired at the same time instant of acquisition and having the same definition, called original definition; - decrease in the definitions of the first and second images of a first factor; - depth prediction for a first set of pixels of the first image from the pixels of the first set of pixels and pixels of the second image corresponding to the pixels of the first set of pixels by a depth prediction model; - as long as at least one depth associated with a pixel in the first image is greater than a first threshold depth and no area of ​​the image has a definition equal to the original definition: • determination of a mask comprising pixels from the first image for which a predicted depth is greater than the first threshold depth; • increasing the definition of a second set of pixels from the first image by a second factor, the pixels of the second set of pixels being called first pixels, including the pixels of the mask, and of pixels from a set of pixels of the second image, called second pixels, including pixels corresponding to the first pixels and updating the first threshold depth according to the increased definition; • prediction of depths associated with the first pixels from the first and second pixels by the depth prediction model; - determination of a third set of pixels from the first image including pixels corresponding to the object; - determination of the distance associated with the object from the depths associated with pixels of the third set of pixels.

[0013] Such a method makes it possible to quickly predict depths associated with pixels in the first received image. Indeed, the predicted depths for a large portion of the pixels in this image are predicted with reduced resolution and are therefore particularly fast to predict. For pixels where the predicted depth is close to or greater than the threshold depth, the lack of precision in depth prediction is compensated for by a local increase in the image resolution. The predicted depths are then more accurate for these pixels. Thus, the predicted depths for all pixels in the first image are accurate, with the higher resolution only in the areas of the first image where it is necessary.

[0014] Indeed, for nearby objects, it is not necessary to use the first image with its original resolution because the disparity is very large compared to the error of the depth prediction model. It is therefore possible to reduce the image resolution for the pixels corresponding to these nearby objects. The resulting depth map is called panoptic; however, it also includes areas containing pixels associated with great depths where the accuracy is not satisfactory. The great depth mask is used to enclose these areas in a rectangle. Increasing the image resolution in this rectangle improves the prediction accuracy for these great depths. Although there are several depth predictions, the total computation time is less than the time required to predict the depths associated with all the pixels of the first image with its original definition because the area of ​​the first image showing high definition is very constrained.

[0015] According to one variant of the method, the first factor is a power of the second factor. Thus, by iteratively increasing the definition with the second factor, the original definition is reached in a few iterations.

[0016] According to another variant of the method, the first threshold depth is determined as a function of a focal length associated with the first camera 11, a distance separating the first camera from the second camera and a minimum disparity defined in pixel units.

[0017] According to yet another variant of the method, the first threshold depth is determined from a simulation of the projection of a pixel from the first image into the second image as a function of its position in the first image, the definition of an area of ​​the first image in which the pixel is positioned and a test depth, the threshold depth being the shallowest depth from which the pixel of the first image is projected onto the same pixel of the second image for any test depth greater than or equal to the threshold depth.

[0018] According to a further variant of the method, the second set of pixels of the first image is a set of pixels contained in a box encompassing the pixels of the first image contained in the mask.

[0019] According to yet another variant, the method further includes determining a second threshold depth greater than the first threshold depth, the second threshold depth being determined according to the original definition, pixels whose associated depth is greater than the second threshold depth not being taken into account in the step of determining the mask.

[0020] According to yet another variant of the method, the third set of pixels is a box encompassing the pixels corresponding to the object.

[0021] According to another variant of the method, the distance is an average value of depths associated with pixels of the third set of pixels.

[0022] According to a second aspect, the present invention relates to a device for determining the distance separating an object from a vehicle by means of a vision system, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the method according to the first aspect of the present invention.

[0023] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0024] According to a fourth aspect, the present invention relates to a computer program that includes instructions adapted for executing the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0025] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0026] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.

[0027] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.

[0028] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.

[0029] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures

[0030] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 5, in which:

[0031] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a a particular and non-limiting example of a realization of the present invention;

[0032] [Fig.2] illustrates a flowchart of the different steps of a method for determining a distance separating an object from the vehicle in [Fig.1] by a depth prediction model, according to a particular and non-limiting embodiment of the present invention;

[0033] [Fig.3] schematically illustrates a device configured to determine a distance separating an object from the vehicle in [Fig.1], according to a particular and non-limiting embodiment of the present invention;

[0034] [Fig.4] schematically illustrates an image processed during the implementation of the the method of [Fig. 2], according to a particular and non-limiting embodiment of the present invention; and

[0035] [Fig.5] schematically illustrates areas of different definitions in the image of [Fig.4], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements

[0036] A method and device for determining the distance between an object and a vehicle carrying a vision system will now be described in the following with joint reference to Figures 1 to 5. The same elements are identified with the same reference signs throughout the following description.

[0037] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0038] Fig. 1 schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.

[0039] Such an environment 1 corresponds, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.

[0040] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.

[0041] The vehicle 10 advantageously comprises at least two on-board cameras, a first camera 11 and a second camera 12, configured to acquire images of a three-dimensional scene unfolding in the environment of the vehicle 10 from distinct observation positions. The first camera 11 and the second camera 12 form a stereoscopic vision system when used together, as illustrated in [Fig. 1]. The first camera 11 forms a monoscopic vision system when used alone, and similarly, the second camera 12 forms another monoscopic vision system when used alone. The present invention, however, extends to any vision system comprising at least two cameras, for example, 2, 3, or 5 cameras.

[0042] The intrinsic parameters of the first camera 11 characterize the transformation which associates, for an image point, hereafter called "point", its three-dimensional coordinates in the reference frame of the first camera 11 with the pixel coordinates in an image acquired by the first camera 11. These parameters do not change if the first camera 11 is moved. The intrinsic parameters of the first camera 11 include in particular a first focal length fl associated with the first camera 11.

[0043] The intrinsic parameters of the second camera 12 characterize, for their part, the transformation which associates, for an image point, its three-dimensional coordinates in the reference frame of the second camera 12 with the pixel coordinates in an image acquired by the second camera 12. These parameters do not change if the second camera 12 is moved. The intrinsic parameters of the second camera 12 include in particular a second focal length f2 associated with the second camera 12.

[0044] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in the curvature and parallelism of the lenses and the coaxiality of the optical axes.

[0045] According to a particular embodiment, the first camera 11 and / or the second camera 12 is of the "wide-angle" type, a wide-angle camera being, for example, equipped with a lens designed to acquire a representative image of a three-dimensional scene seen over a wider field of view than a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens makes it possible to capture a larger portion of the three-dimensional scene unfolding in front of or around the camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by that camera. The angle α of the field of view of the first camera 11 is, for example, equal to 120°, 145°, 180°, or 360°, whereas a standard camera offers, for example, a field of view open at an angle of 45° or less.Such a first camera 11 corresponds, for example, to a camera equipped with mirrors or a "fisheye" camera. Wide-angle lenses have a shorter focal length compared to standard lenses, making them suitable for capturing images of landscapes, architecture, road intersections, or any other subject requiring a wide perspective. Wide-angle cameras are, for example, . used to capture immersive and dynamic images with extended depth of field.

[0046] These two cameras 11, 12 are arranged so that each acquires an image of a scene from a different viewpoint. The first viewpoint is, for example, located on or in the left-hand rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. The second viewpoint is, for example, located on or in the right-hand rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. If both cameras are located at the top of the windshield of the vehicle, they are then positioned at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, and the second camera 12 is located in the right-hand rearview mirror of the vehicle 10.

[0047] A first marker is associated with the first camera 11: - the direction of the x-axis is defined as horizontal and normal to the optical axis Cl of the first camera 11. The distance B separating the optical center of the first camera 11 from the projection of the optical center of the second camera 12 onto the horizontal plane passing through the optical center of the first camera 11 is called the reference base (in English "baseline"); - the direction of the y-axis is defined as vertical and normal to the optical axis Cl of the first camera 11; - The direction of the z-axis is defined as orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal coordinate system.

[0048] The extrinsic parameters related to the position of cameras 11, 12 are the following parameters: - three translations in the x, y, and z directions: Tx, Ty, and Tz, constituting the translation vector T; and - three rotations in the x, y and z directions: 0x, 0y and 0z representing the relative orientation of the optical axis Cl of the first camera 11 and the optical axis C2 of the second camera 12.

[0049] An extrinsic matrix of the vision system then includes the extrinsic parameters previously defined.

[0050] The extrinsic parameters are determined, for example, during a calibration phase of the stereoscopic vision system comprising the first camera 11 and the second camera 12.

[0051] A key constraint of stereoscopic vision systems used in automobiles is, for example, the large distance between the two cameras. Indeed, to cover a measurement range of 200 meters, the reference base must be 60 cm for cameras commonly used in this field.

[0052] The two cameras 11, 12 acquire images of a scene located in front of the vehicle 10, the first camera 11 alone covering a first acquisition field 13, the second camera 12 alone covering a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic view of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic view of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic view of the scene by the stereoscopic vision system composed of the two cameras 11, 12.

[0053] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.

[0054] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system comprising the second camera 12.

[0055] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable using the monoscopic vision system comprising the first camera 11.

[0056] Finally, field 19 is not visible to any of the cameras. The part of the scene present in this field 19 is therefore not observable.

[0057] According to one particular embodiment, the field of view of the second camera 12 covers at least half of the field of view of the first camera 11.

[0058] It is evident that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently.

[0059] The images acquired by the cameras 11, 12 at a given acquisition time are in the form of data representing pixels characterized by: - ​​coordinates in each image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric coordinates (from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).

[0060] According to a particular embodiment, an image acquired by the first camera 11 and / or by the second camera 12 includes a distortion equal to 0.5%, 0.8% or greater than 1%. The measurement of such distortion corresponds to determining a ratio between: - the maximum spacing of a pixel in the image of a straight line in the first three-dimensional scene whose image is a line touching the longest edge of the first image, either at the center of the edge of the image, or at the corners of the edge of the image, and - the length of this edge.

[0061] In the world of photography, distortion is commonly considered to be: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • bothersome if it is greater than or equal to 1% or more.

[0062] Barrel distortion is characterized by a positive percentage, while crescent distortion is characterized by a negative percentage.

[0063] Each pixel of the acquired image represents an object in the three-dimensional scene present in the camera's field of view. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a point of light, or image point, resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, which are captured by a photosensitive sensor of the camera after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image that faithfully represents the physical object observed by the camera. An image point, as previously described, is therefore a point on the surface of an object in the three-dimensional scene.

[0064] The images acquired by cameras 11, 12 represent views of the same scene taken from different viewpoints, the camera positions being distinct. This scene includes, for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.

[0065] According to a particular embodiment, a field of view of the first camera 11 covers at least half of a field of view of the second camera 12, and a field of view of the second camera 12 covers at least half of a field of view of the first camera 11. In other words, more than half of the pixels of an image acquired by the first camera 11 correspond to an object in the three-dimensional scene seen by the second camera 12, pixels of an image acquired by the second camera 12 also corresponds to this object in the three-dimensional scene. Similarly, more than half of the pixels in an image acquired by the second camera 12 correspond to an object in the three-dimensional scene seen by the first camera 11, and pixels in an image acquired by the first camera 11 also correspond to this object in the three-dimensional scene.

[0066] The images acquired by the first camera 11 and by the second camera 12 are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.

[0067] A method for determining a distance separating an object from a vehicle carrying a vision system is advantageously implemented by the vehicle 10, i.e. by a processor, a computer or a combination of computers of the vehicle 10's on-board system, for example by the computer or computers in charge of the vehicle 10's vision system.

[0068] Figure 2 illustrates a flowchart of the different steps of a method 2 for determining the distance between an object and a vehicle equipped with a vision system, for example, vehicle 10 of Figure 1. The vision system onboard vehicle 10 comprises a first and a second camera arranged to each acquire an image of a three-dimensional scene, i.e., one taking place in the vicinity of vehicle 10. Such a three-dimensional scene includes, in particular, an object located in the field of vision of the first and second cameras, i.e., in the third acquisition field 15 shown opposite Figure 1. Method 2 is implemented, for example, by a device of the vision system onboard vehicle 10 or by device 3 of Figure 3 and comprises the following steps.

[0069] In a step 21, a pair of images comprising a first image acquired by the first camera 11 and a second image acquired by the second camera 12 are received. The reception of these images consists of receiving data representative of these images. These first and second images are acquired at the same time instant and represent the same three-dimensional scene observed from two distinct viewpoints corresponding to the respective positions of the first and second cameras.

[0070] According to a particular embodiment, the first and second images have the same resolution, that is, they have the same number of pixels, the same number of pixels in height, and the same number of pixels in width, and a large part of each image represents the same portion of the three-dimensional scene. The resolution of the first and second images is then called the original resolution.

[0071] According to another specific embodiment, the first and second images are not of the same resolution and / or only a small portion of each image represents the same part of the three-dimensional scene. An additional step then consists of resizing or cropping them to obtain a first and second image of the same resolution, the pixels of which correspond predominantly to the same part of the three-dimensional scene; for example, more than 80% of the pixels in each image represent the same part of the three-dimensional scene as seen by both cameras. The resolution of the first and second resized or cropped images is called the original resolution.

[0072] If the cameras are calibrated, for example by following a calibration method known to those skilled in the art such as that presented in the document "Single View Point Omnidirectional Camera Calibration from Planar Grids", written by Christopher Mei and Patrick Rives and published in April 2007, then the intrinsic parameters of the first and second cameras 11,12 are known. Edges of the common field of view can be identified in the first and second images by simulation using a back-projection and projection program, which are described below.

[0073] The rear projection program performs the following operations: • The pixel coordinates are corrected to remove distortion defined by distortion parameters: the parameter of the model presented in the cited document, called the Mei model, and distortion coefficients kh k2, k3, pi and p2, the definition of which can be found in the free library OpenCV® for example. • The coordinates are projected onto the unit sphere according to the Mei model. • Projection lines are calculated, always according to the Mei model. • Projection lines are multiplied by a depth.

[0074] The projection program then proceeds by performing the following operations: • The coordinates of the points in space are transformed into the coordinate system of the other image. • The coordinates of the points in space are modified via the distortion function defined by the same parameters as those mentioned previously. • Multiplication with the inverse of the intrinsic matrix similar to the pinhole camera lens.

[0075] By associating a depth corresponding to the midpoint of a measurement range, for example 100 m (one hundred meters) for a measurement range from 10 m (ten meters) to 200 m (two hundred meters), with the pixels of the first image and by performing backprojection and projection of the first image onto the second image with this depth, the simulation makes it possible to determine the average edge of the common field of view on the second image. The first image is to be cropped with the same definition as the second cropped image and on the opposite side along a diagonal to the cropping side of the second image.

[0076] If the first and second cameras are not calibrated, the projection and reprojection models are defined by a neural network such as the one described in the paper "Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion" by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich, and Adrien Gaidon, published in August 2020. This particular embodiment is well-suited to images acquired by the vision system cameras and exhibiting significant distortion. It is not possible to determine the exact cropping before training. The common field of view in the first and second images can be found using a trial-and-error method, starting with cropping by observing the images and training the projection and reprojection models to improve their effectiveness.If the pixels at the edge of the first and second images do not have a reasonable prediction, for example with a value that is too high or too low, or when a depth map does not match the texture of an original image, the size of the cropped image must be reduced. If some of the pixels, preferably half, have reasonable predictions while the others do not, the correct cropping is achieved.

[0077] In step 22, the definitions of the first and second images are reduced by a first factor. The first factor is, according to a particular embodiment, a power of a second factor; for example, the second factor is equal to 2 (two), while the first factor is equal to 8 (eight). Other values ​​for the first and second factors are also possible; for example, the second factor is equal to 3 (three) and the first factor is equal to 9 (nine). Similarly, the power can be equal to 3 (three), as in the examples presented previously, but can also be equal to 2 (two) or 4 (four), the first factor then being equal to 4 (four) or 16 (sixteen), respectively, when the second factor is equal to 2 (two).

[0078] The first factor is specifically defined to reduce the time required to predict depths associated with the pixels of the first image. Thus, the lower the resolution of the first image, the shorter the time required to predict the depths associated with the pixels of the first image. However, the accuracy of the predicted depths is thereby reduced, particularly for distant objects where the disparity associated with the corresponding pixels is small.

[0079] In a step 23, depths are predicted by a depth prediction model for a first set of pixels of the first image 4 from the pixels of the first set of pixels and pixels of the second image corresponding to the pixels of the first set of pixels.

[0080] According to a particular embodiment, the depth prediction model determines, for example, disparities and then determines the associated depth.

[0081] According to another particular embodiment, the prediction model uses a neural network to predict the depth associated with a pixel in the first image based on the position of that pixel in the first image and the position of a corresponding pixel in the second image. Such a prediction model using a neural network is particularly useful when the depth is not proportional to the disparity, for example, when the images exhibit strong distortion and / or the optical axes C1, C2 of the two cameras are not in parallel planes.

[0082] The depths predicted at this reduced definition are called panoptic depths. They are determined quickly by reducing the resolution of an image acquired by one of the cameras of the vision system on board the vehicle 10 and for a very large part of the first image, or even for all the pixels of the first image.

[0083] A threshold depth representing the maximum measurable depth at the highest level of definition in the first image is determined or received. This threshold depth is defined according to one of the methods described below, depending on the vision system.

[0084] According to a first particular embodiment, when the depth is proportional to a disparity, the depth prediction model determines the disparity in pixel units, which is the displacement of the pixel representing the same image point of the three-dimensional scene between the first and second images. If the error of the prediction model is approximately 0.7 pixels, which is the error of a model such as the model presented in the paper "Correlate-and-Excite: Real-Time Stereo Matching via Guided Cost Volume Excitation" by Antyanta Bangunharcana, Jae Won Cho, Seokju Lee, In So Kweon, Kyung-Soo Kim, and Soohyun Kim, published in August 2021, then to maintain a relative error of less than 20%, the predicted disparity must be greater than 3.5 pixels. The maximum measurable distance is then defined by the following function: [Math.l] / xB — -___ ^■max Dnti„

[0085] With: • Zmax is the maximum measurable depth, the index n indicating the level of image resolution reduction, 1 corresponding to the original resolution, • f is the focal length, • B the distance separating the first camera 11 from the second camera 12, and • ^min the minimum disparity determined by the level of error of the stereo model as well as a prerequisite on the relative error.

[0086] Assuming a focal length of 3500 pixels for an image resolution of 3000 x 2000, and a distance of 0.5 meters between the two cameras, the maximum depth measurable by this equation is 500 m (five hundred meters). Such a maximum depth is not necessary for most objects present in the three-dimensional scene, i.e., objects in the vicinity of vehicle 10. Therefore, the reduction in image resolution is reasonable.

[0087] Reducing the image resolution necessitates a change in the focal length in the equation, while the other parameters remain unchanged. Consequently, the maximum measurable depth is inversely proportional to the image resolution reduction factor, indicated by the subscript r. For example, if the image resolution reduction factor is 4, the focal length in the preceding equation is divided by 4, and the maximum measurable depth is then also divided by 4, reaching 125 m (one hundred and twenty-five meters) in an image whose resolution is divided by 4. This maximum depth is sufficient when the vehicle 10 is in an urban environment. The reduced image resolution of 750 x 500 thus allows for a much faster calculation.

[0088] Thus, according to this first particular embodiment, the first threshold depth is determined as a function of a focal length associated with the first camera 11, a distance separating the first camera 11 from the second camera 12 and a minimum disparity defined in pixel units.

[0089] According to a second particular embodiment, when the depth is not proportional to the disparity, for example when the first and second images exhibit strong distortion, the Mei model defines the pixel in the second image corresponding to the pixel in the first image as a function of a depth. By varying the depth value within the measurement range with an interval of 1 meter, each pixel in the first image has a series of corresponding pixel abscissas in the second image. From a certain depth, the difference between the abscissa in the second image of a pixel corresponding to the pixel in the first image for a first depth and the abscissa in the second image of a pixel corresponding to the pixel in the first image for a second depth remains unchanged, the second depth being equal to the first depth plus an incremental value, for example 1m (one meter).This is the maximum measurable depth for that pixel in the first image. This check must be performed for all pixels in the first image, provided that the pixel corresponding to the pixel in the first image is indeed present in the second image. The maximum measurable depth for the entire first image with a given resolution is then equal to the... minimum value of the maximum measurable depths determined for all pixels of the first image.

[0090] According to this second particular embodiment, the first threshold depth is then determined from a simulation of the projection of a pixel from the first image into the second image as a function of its position in the first image, the definition of an area of ​​the first image in which the pixel is positioned and a test depth, the threshold depth being the shallowest depth from which the pixel of the first image is projected onto the same pixel of the second image for any test depth greater than or equal to the threshold depth.

[0091] According to a particular embodiment, the method further includes determining a second threshold depth greater than the first threshold depth, the second threshold depth being determined according to the original definition and allowing not to take into consideration pixels corresponding in particular to the sky.

[0092] In a step 24 and as illustrated in [Fig. 4], a mask M comprising pixels from the first image 4 for which a predicted depth is greater than the first threshold depth is determined. This mask M includes, for example, one or more distinct areas depending on the observed three-dimensional scene.

[0093] In the case of taking into account the second threshold depth, pixels whose depth is greater than or equal to the second threshold depth are not included in the mask M. Indeed, they do not correspond to objects located at a certain distance but to the sky, the depth of the pixels linked to the sky not being useful in this process.

[0094] A second set of pixels is then determined and encompasses all the pixels within the mask M. This second set is, for example, rectangular in shape and corresponds to a box E encompassing the pixels of the first image within the mask M. Thus, the corners of the rectangle are defined by the following functions when the origin of the coordinate system associated with the pixel coordinates is the upper left corner of the first image:

[0095] [Math.2] x41 = min{P^.} i •> [Math.2] y4] = min{^.} [Math.2] x42 = max{py i And [Math.2] F 43= max{P^} i

[0096] With: • x4i the x-coordinate of the upper left corner 41 and the lower left corner 44, • 3^41 the y-coordinate of the upper left corner 41 and the upper right corner 42, • x42 the x-coordinate of the upper right corner 42 and the lower right corner 43, • the y-coordinate of the lower right corner 43 and the lower left corner 44, • p* the x-coordinate of a pixel1 included in the mask M, and • the ordinate of a pixel ' included in the mask M.

[0097] According to a particular embodiment, the second set of pixels covers the entire width of the image; for example, on a highway, without any vehicles very close to the sides of vehicle 10, it is possible to see the three-dimensional scene far into the distance. In this case, this bounding box is applied to both the first and second images for cropping.

[0098] If the second set of pixels covers only part of the image width, for example, if there are vehicles close together on the left and / or right, and the first camera 11 is to the left of the second camera 12, the width of the bounding box E must be increased so that the pixels at its left edge are visible in a second bounding box in the second image corresponding to the right image. For a stereoscopic vision system in which the optical axes of the cameras are parallel and the distance between the cameras is small, this enlargement can be defined by a maximum disparity determined by:

[0099] [Math.3] ^mask fx B

[0100] With: * iC"6 the width of the translation of the left edge of box E encompassing mask M, • f the focal length, • B the distance separating the first camera 11 from the second camera 12, and • the maximum measurable depth, the index n signifying the level of reduction of image definition.

[0101] Thus, the widening of the bounding box E is to be done only on the left side of the mask: [Math.4] T - min IJ* I -•*41 — WW1 rx J ï

[0102] If the depth prediction model requires an image height and width divisible by 32, then the width and height of the bounding box E must be adapted as follows:

[0103] [Math.5] w'=tv+ (32- w%32) And [Math.5] h' = h^32-h%32) •>

[0104] With: • w' the corrected width of the bounding box and divisible by 32, • w the initial width of the bounding box E, • h' the corrected height of the bounding box and divisible by 32, and • h the initial height of the bounding box E.

[0105] Note that the bounding boxes including the first and second pixels are necessarily of the same size.

[0106] Subsequently, when the bounding box E is resized, the term bounding box E applies to the resized bounding box.

[0107] In a step 25, the definition of the second set of pixels is increased by the second factor, as is the definition of the set of pixels of the second image corresponding to the pixels of the second set of pixels of the first image 4. In other words, the part of the first image contained in the bounding box E and the part of the second image contained in the second bounding box corresponding to the bounding box E have their definition increased by the second factor.

[0108] The pixels of the first image 4 included in the bounding box E whose definition has been increased are called first pixels, the pixels of the second bounding box included in a second bounding box corresponding to the bounding box E or to the corrected bounding box are called second pixels.

[0109] In a step 26, depths associated with the first pixels are predicted from the first and second pixels by the depth prediction model.

[0110] The first depth is then updated, that is to say, it is determined for the increased definition, therefore with the focal length corresponding to the definition increased. Indeed, with each increase in definition, the depth prediction is more precise and then allows reliable depth predictions for depths less than an updated threshold depth greater than the previously defined threshold depth.

[0111] As long as at least one depth associated with a pixel of the first image 4 is greater than the first threshold depth and no area of ​​the image 4 has a definition equal to the original definition, steps 24 to 26 are repeated, the threshold depth being updated at the end of step 26.

[0112] Thus, in the example illustrated in [Fig. 5], the first image, after a few iterations, comprises several areas in which the definitions are different. Thus, a first area E3 has a definition equal to the minimum definition, a second area E2 has a definition equal to the minimum definition divided by the second factor, a third area E3 has a definition equal to the minimum definition divided twice by the second factor, and the remainder of the first image 4 has a definition equal to the original definition.

[0113] Once the depth predictions are complete, that is, when no predicted depth exceeds the first threshold value, or when the enhanced resolution matches the original resolution, the predicted depth values ​​are aggregated. The original resolution is then applied to the first image for which the resolution is heterogeneous.

[0114] Note that before the agglomeration of the predicted depths, the enlargement of the rectangular bounding box must be removed where appropriate, both for visibility and for division by 32, namely returning to each bounding box initially determined via the function for determining the positions of the angles of the bounding box.

[0115] Agglomeration is achieved by increasing the resolution of the depth map with the lowest definition to the original image resolution, as well as the depth maps of the bounding boxes at all levels. In case of overlapping depth maps, the higher-resolution depth map takes precedence because it offers the best prediction accuracy. Thus, the final depth map exhibits satisfactory accuracy for the entire first image.

[0116] In step 27, as illustrated in [Fig. 4], a third set of pixels from the first image is determined. This third set of pixels comprises O pixels corresponding to the observed object. The O pixels are determined, for example, via an object detector used for the first image. Note that it is possible for several objects to be observable in the three-dimensional scene, in which case each relative distance to each object is determined independently.

[0117] According to a particular embodiment, the third set of pixels is a box B encompassing the pixels O corresponding to the object.

[0118] In a step 28, the distance associated with the object is determined from the depths associated with pixels of the third set of pixels.

[0119] According to a particular embodiment, the distance is an average value of depths associated with pixels of the third set of pixels, that is to say, the distance is determined by averaging the pixels of the depth map in its bounding box.

[0120] To avoid the impact of the environment, according to a particular embodiment, the width and height of the bounding box are reduced by a certain percentage, for example 50%, while keeping the position of the center of the bounding box.

[0121] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another road user, the ADAS is then able to accurately determine this distance and act accordingly. Indeed, the predicted depths for a large part of the image are panoptic depths, and the resolution is greater only for pixels associated with depths that cannot be measured with a resolution that is too low. Thus, if the ADAS's function is to activate the braking system of the vehicle 10 in the event of a risk of collision with another road user, and the distance between the vehicle 10 and that same road user decreases sharply, then the ADAS is able to detect this sudden approach and activate the braking system of the vehicle 10 to avoid a potential accident.

[0122] Figure 3 schematically illustrates a device 3 configured for determining the distance between an object and a vehicle (10) equipped with a vision system and / or for learning a local depth prediction model associated with a vision system mounted in a vehicle 10, according to a particular and non-limiting embodiment of the present invention. The device 3 corresponds, for example, to a device mounted in the first vehicle 10, for example, a computer associated with the stereoscopic vision system.

[0123] Device 3 is, for example, configured to carry out the steps described opposite [Fig. 2]. Examples of such a device 3 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 3, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 3 can be implemented in the form of electronic circuits or software (or computer) modules or a combination of electronic circuits and software modules.

[0124] The device 3 comprises one (or more) processor(s) 30 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 3. The processor 30 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 3 further comprises at least one memory 31, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.

[0125] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 31.

[0126] According to various particular and non-limiting embodiments, the device 3 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.

[0127] According to a particular and non-limiting embodiment, the device 3 comprises a block 32 of interface elements for communicating with external devices. The interface elements of the block 32 comprise one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HD MI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").

[0128] According to another particular and non-limiting embodiment, the device 3 includes a communication interface 33 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 330. The communication interface 33 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 330. The communication interface 33 corresponds, for example, to a wired network of the CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3) type.

[0129] According to a particular and non-limiting embodiment, the device 3 can provide output signals to one or more external devices, such as a display screen 340, touch or not, one or more speakers 350 and / or other peripherals 360 via the output interfaces 34, 35, 36 respectively. According to a variant, one or more of the external devices is integrated into the device 3.

[0130] Of course, the present invention is not limited to the embodiments described above but extends to a method for measuring the distance between an object and a vehicle equipped with a vision system and / or to a method for determining the depths associated with pixels in an image, which would include secondary steps without falling outside the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0131] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 3 of [Fig.3].

Claims

1. Demands Method for determining the distance separating an object from a vehicle (10) carrying a vision system comprising a first and a second camera arranged so as to each acquire an image of a three-dimensional scene including said object, said method being implemented by at least one processor, and being characterized in that the determination of said distance comprises the following steps: - reception (21) of a pair of images comprising a first image (4) acquired by the first camera (11) and a second image acquired by the second camera (12), the first and second images being acquired at the same time instant of acquisition and having the same definition, called original definition; - reduction (22) of the definitions of the first and second images of a first factor; - prediction (23) of depths for a first set of pixels of the first image (4) from the pixels of said first set of pixels and pixels of the second image corresponding to the pixels of said first set of pixels by a depth prediction model; - as long as at least one depth associated with a pixel of said first image (4) is greater than a first threshold depth and no area of ​​the image (4) has a definition equal to said original definition: • determination (24) of a mask (M) comprising pixels of the first image (4) for which a predicted depth is greater than said first threshold depth; • increasing (25) the definition of a second set of pixels of the first image by a second factor, the pixels of said second set of pixels being called first pixels, comprising the pixels of said mask and pixels of a set of pixels of the second image, called second pixels, comprising pixels corresponding to the first pixels and updating said first threshold depth according to the increased definition; • prediction (26) of depths associated with the first pixels from the first and second pixels by the depth prediction model; - determination (27) of a third set of pixels of the first image including pixels (0) corresponding to said object; - determination (28) of the distance associated with said object from the depths associated with pixels of the third set of pixels.

2. A method according to claim 1, wherein the first factor is a power of the second factor.

3. A method according to claim 1 or 2, wherein the first threshold depth is determined as a function of a focal length associated with the first camera 11, a distance separating the first camera (11) from the second camera (12) and a minimum disparity defined in pixel units.

4. A method according to claim 1 or 2, wherein the first threshold depth is determined from a simulation of the projection of a pixel of the first image (4) into the second image as a function of its position in the first image, the definition of an area of ​​the first image in which the pixel is positioned and a test depth, said threshold depth being the shallowest depth from which the pixel of the first image is projected onto the same pixel of the second image for any test depth greater than or equal to the threshold depth.

5. A method according to any one of claims 1 to 4, wherein the second set of pixels of the first image is a set of pixels contained in a box (E) encompassing the pixels of the first image contained in the mask (M).

6. A method according to any one of claims 1 to 5, further comprising determining a second threshold depth greater than the first threshold depth, the second threshold depth being determined according to the original definition, pixels whose associated depth is greater than said second threshold depth not being taken into account in the determination step (24) of said mask (M).

7. A method according to any one of claims 1 to 6, wherein said third set of pixels is a box (B) encompassing the pixels (0) corresponding to said object.

8. A method according to any one of claims 1 to 7, wherein said distance is an average value of depths associated with pixels of the third set of pixels.

9. Device (3) for determining a distance separating an object from a vehicle (10) carrying a vision system, said device (3) comprising a memory (31) associated with at least one processor (30) configured for carrying out the steps of the method according to any one of claims 1 to 8.

10. Vehicle (10) comprising the device (3) according to claim 9.