Method and device for determining depth using a stereoscopic vision system mounted in a vehicle

The method for determining depth using a stereoscopic vision system in vehicles addresses the processing time challenge by calculating depths without image rectification, improving ADAS system responsiveness and accuracy.

FR3158382B1Active Publication Date: 2025-11-28STELLANTIS AUTO SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2024000235
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-11-28
Estimated Expiration
2044-01-11

AI Technical Summary

Technical Problem

Existing stereoscopic vision systems in vehicles require significant processing time and resources for image rectification, which hampers the rapid and accurate determination of depth, crucial for ADAS systems, due to limited computational resources in embedded systems.

Method used

A method for determining depth using a stereoscopic vision system comprising a set of cameras, involving reprojection of pixels into a three-dimensional scene, determination of optical flux, and calculation of depths using intrinsic and extrinsic parameters, without the need for image rectification, thereby reducing processing time.

Benefits of technology

This method enables faster and more accurate depth prediction, enhancing the responsiveness and reliability of ADAS systems by providing rapidly updated depth data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000024_0000
    Figure 00000024_0000
  • Figure 00000024_0001
    Figure 00000024_0001
  • Figure 00000025_0000
    Figure 00000025_0000
Patent Text Reader

Abstract

A method or device implementing a depth determination process using a stereoscopic vision system. The method comprises receiving data representing a first image (31) and a second image (32) acquired respectively by a first and second camera at the same acquisition time. Coordinates of a second pixel (321) in the second image, corresponding to a pixel (311) of the first image, are determined by reprojecting the first pixel into the three-dimensional scene to obtain a point (312) projected into the second image. An optical flux (F) associated with the first pixel and representing a displacement vector between the first pixel and a third pixel of the second image is determined. Second depths associated with the first pixel are then determined from the coordinates and the optical flux, with the second and third pixels being considered as one. (See Figure 3 for abbreviations.)
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for determining depth using a stereoscopic vision system mounted in a vehicle technical field

[0001] The present invention relates to methods and devices for determining depth using a stereoscopic vision system mounted in a vehicle, for example, in a motor vehicle. The present invention also relates to a method and device for measuring the distance between an object and a vehicle equipped with a stereoscopic vision system. The present invention further relates to a method and device for controlling one or more ADAS systems mounted in a vehicle based on the determined depth. Technological background

[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS systems use advanced technologies to assist the driver while driving and thus improve performance. ADAS systems use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.

[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.

[0004] Vehicle-mounted ADAS systems are powered by data obtained from one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or any obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other road users; - to automatically regulate the vehicle's speed; - to act on the braking system in case of risk of impact with an object.

[0005] Determining depth or distance from images acquired by a stereoscopic vision system is performed using images acquired by that vision system. To enable the prediction of great depths or distances, for example, greater than 200 meters, a stereoscopic vision system is used. A stereoscopic vision system typically includes at least two cameras positioned several tens of centimeters apart. This distance creates a difference in the field of view of the cameras, and the images they produce must then be processed for usability. This processing most often involves rectifying the images.

[0006] Image processing, however, requires considerable time and resources, whereas an ADAS powered by depth or distance data requires rapid updating of this data while maintaining its accuracy. The resources allocated to the stereoscopic vision system are not unlimited, however, as an embedded system must be compact and / or lightweight to be integrated into a vehicle.

[0007] Thus, the quality of the data emitted and the speed of image processing by a vision system determine the proper functioning of driving aid devices using this data. Summary of the present invention

[0008] One object of the present invention is to solve at least one of the problems of the technological background described above.

[0009] Another object of the present invention is to reduce the processing time of images acquired by a stereoscopic vision system embedded in a vehicle.

[0010] Another object of the present invention is to improve road safety, in particular by improving the reliability of AD AS systems powered by data obtained from at least one camera.

[0011] According to a first aspect, the present invention relates to a method for determining depth by a stereoscopic vision system embedded in a vehicle, the method being implemented by a processor, the stereoscopic vision system comprising a set of cameras of at least two cameras arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the process being characterized in that it comprises the following steps: - reception of representative data from a first image acquired by the first camera at a given acquisition time and from a second image acquired by the second camera at said acquisition time - reprojection of a first pixel from the first image into the three-dimensional scene to obtain a point corresponding to the first pixel based on the initial coordinates of the first pixel in the first image, a first depth associated with the first pixel, and intrinsic parameters of the first camera, the second coordinates of the point being defined in a reference frame associated with the first camera; - determination of third coordinates of the point in a reference frame associated with the second camera as a function of the second coordinates and extrinsic parameters of the stereoscopic vision system; - determination of fourth coordinates of a second pixel in the second image corresponding to the projection of the point in the second image, the fourth coordinates being determined as a function of the third coordinates and intrinsic parameters of the second camera; - determination of an optical flux associated with said first pixel, the optical flux being representative of a displacement vector between the first pixel and a third pixel of the second image, the first and third pixels corresponding to the same object of the three-dimensional scene; - determination of second depths associated with the first pixel from the fourth coordinates and the optical flux, the second and third pixels being confused.

[0012] Such a method thus makes it possible to predict depths associated with pixels in an image, that is, to predict the distance separating the vehicle equipped with the stereoscopic vision system from an object present in the field of vision of the first and second cameras. This method is notably faster than a method requiring heavy image processing, such as rectification of the first and second images, when the first and second cameras are separated by several tens of centimeters. An ADAS controlled by data generated by this method, that is, by the depths associated with the pixels of the first image, is therefore more responsive.

[0013] According to one variant, the method includes a step of determining a third depth associated with the first pixel from components along a horizontal axis of the second image of the fourth coordinates and the optical flow, the second depth being determined from the third depth.

[0014] According to another variant, the method includes a step of determining a fourth depth associated with the first pixel from components along a vertical axis of the second image of the fourth coordinates and the optical flow, the second depth being determined from the fourth depth.

[0015] According to yet another variant of the process, the second depth is determined from the third and fourth depths.

[0016] According to a further variant of the method, the second depth is equal to an average of the third and fourth depths.

[0017] According to another variant of the method, the optical flux is predicted by a convolutional neural network.

[0018] According to a second aspect, the present invention relates to a device for determining depth by means of a stereoscopic vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.

[0019] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0020] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0021] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0022] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.

[0023] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.

[0024] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.

[0025] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures

[0026] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 4, in which:

[0027] [Fig-1] schematically illustrates a stereoscopic vision system embedded in a vehicle, according to a particular and non-limiting example of the present invention;

[0028] [Fig.2] illustrates a flowchart of the different stages of a process for determining depth by a stereoscopic vision system mounted in the vehicle of [Fig.1], according to a particular and non-limiting example of the present invention;

[0029] [Fig.3] schematically illustrates images acquired by the stereoscopic vision system on board the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention;

[0030] [Fig.4] schematically illustrates a device configured for determining depth by a stereoscopic vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements

[0031] A method and device for determining depth by means of a stereoscopic vision system mounted in a vehicle will now be described in what follows with joint reference to Figures 1 to 4. The same elements are identified with the same reference signs throughout the following description.

[0032] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0033] According to a particular and non-limiting example of an embodiment of the present invention, a method for determining depth by a stereoscopic vision system embedded in a vehicle is for example implemented by a processor or a computer of the vehicle's embedded system controlling this stereoscopic vision system.

[0034] The stereoscopic vision system comprises a set of cameras, at least two of which are arranged so that each acquires an image of a three-dimensional scene from a different viewpoint, the optical axes representing an orientation of the field of view of each camera being contained in two parallel planes, for example oriented non-parallel. The stereoscopic vision system is thus said to be non-parallel.

[0035] To this end, the method for determining depth using a stereoscopic vision system installed in a vehicle comprises receiving data representing a first image and a second image acquired respectively by a first and second camera at the same acquisition time. The ordinates of a second pixel in the second image corresponding to a pixel of the first image are determined via the reprojection of the first pixel in the three-dimensional scene to obtain a point projected in the second image.

[0036] An optical flux associated with the first pixel and representative of a displacement vector between the first pixel and a third pixel of the second image is determined.

[0037] Second depths associated with the first pixel are then determined from the coordinates and the optical flux, the second and third pixels being coincident.

[0038] Fig. 1 schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.

[0039] Such an environment 1 corresponds, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.

[0040] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.

[0041] The vehicle 10 advantageously comprises several onboard cameras 11, 12, each configured to acquire images of a scene in the environment of the vehicle 10. This set of cameras 11, 12 forms the stereoscopic vision system. Two cameras 11 and 12 are illustrated in [Fig. 1]. The present invention is not limited, however, to a stereoscopic vision system comprising two cameras but extends to any stereoscopic vision system comprising two or more cameras, for example, two, three, four, or five cameras.

[0042] The two cameras 11, 12 have known intrinsic parameters. These parameters include, in particular: - the focal length of the first camera 11; - the focal length of the second camera 12; - distortions which are due to imperfections in the optical system of each camera; - the direction Cl of the optical axis of the first camera 11; - the C2 direction of the optical axis of the second camera 12; and - the respective resolutions of cameras 11, 12.

[0043] The intrinsic parameters characterize the transformation that associates, for an image point, the camera coordinates to the pixel coordinates, in each camera. These parameters do not change if the camera is moved.

[0044] The intrinsic matrix of the first camera 11 is defined by:

[0045] [Math.l] L 0 0 1 J

[0046] With: • K, the intrinsic matrix of the first camera 11, • f, the focal length of the first camera 11, • 4 the abscissa of the optical center of an image acquired by the first camera 11, and • cJv the ordinate of the optical center of an image acquired by the first camera 11.

[0047] Similarly, the intrinsic matrix of the second camera 12 is defined by:

[0048] [Math.2]

[0049] With: • K the intrinsic matrix of the second camera 12, • f the focal length of the second camera 12, • the abscissa of the optical center of an image acquired by the second camera 12, and • Cy the ordinate of the optical center of an image acquired by the second camera 12.

[0050] The optical axes of the first and second cameras are arranged horizontally, that is to say in two parallel horizontal planes. These two planes are separated by a height H, H having a value of zero if the planes coincide.

[0051] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in lens curvature, parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are either not taken into account or their correction is addressed during image acquisition.

[0052] These two cameras 11, 12 are arranged so that each acquires an image of a scene from a different viewpoint; the first viewpoint is, for example, located on or in the left-hand rearview mirror of vehicle 10 or at the top of the windshield of vehicle 10; the second viewpoint is, for example, located on or in the right-hand rearview mirror of vehicle 10 or at the top of the windshield of vehicle 10. In the case where The two cameras are located at the top of the vehicle's windshield, and are therefore positioned at a certain distance. In this example, the first camera 11 is located at the top of the windshield of vehicle 10, and the second camera 12 is located in the right-hand side mirror of vehicle 10.

[0053] A first marker is associated with the first camera 11: - the direction of the x-axis is defined as horizontal and normal to the optical axis of the first camera 11. The distance B separating the optical center of the first camera 11 from the projection of the optical center of the second camera 12 onto the horizontal plane passing through the optical center of the first camera 11 is called the reference basis (in English "baseline"); - the direction of the y-axis is defined as vertical and normal to the optical axis of the first camera 11; - the direction of the z-axis is defined as orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal coordinate system.

[0054] The extrinsic parameters related to the position of cameras 11, 12 are the following parameters: - three translations in the x, y, and z directions: Tx, Ty, and Tz, constituting the translation vector T, with Tx = B, Ty = H, and Tz = 0; and - a rotation around the y-axis by an angle 0 called the yaw angle.

[0055] A matrix for going from the reference frame of the first camera 11 to the reference frame of the second camera 12 is thus defined by:

[0056] [Math.3] [Æ r] = ' COS0 0 . - sin# 0 1 0 sin(? 0 cos# B' H 0.

[0057] With: • [R, T] the matrix for going from the frame of reference of the first camera 11 to the frame of reference of the second camera 12, • 0 the yaw angle, • B the reference base, and • H the difference in height between the first 11 and second 12 cameras.

[0058] The extrinsic parameters are determined, for example, during a calibration phase of the stereoscopic vision system.

[0059] A key constraint of stereoscopic vision systems used in automobiles is, for example, the large distance between the two cameras. Indeed, to cover a measurement range of 200 meters, the reference base must be 60 cm for cameras commonly used in this field.

[0060] The two cameras 11, 12 acquire images of a scene located in front of the vehicle 10, the first camera 11 alone covering a first acquisition field 13, the second camera 12 alone covering a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic view of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic view of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic view of the scene by the stereoscopic vision system composed of the two cameras 11, 12.

[0061] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.

[0062] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system composed of the second camera 12.

[0063] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable with the monoscopic vision system composed of the first camera 11.

[0064] Finally, field 19 is not visible from any of the cameras. The part of the scene present in this field 19 is therefore not observable.

[0065] It is evident that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently.

[0066] The images acquired by cameras 11, 12 at a given acquisition time are in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colors and brightness of objects in the observed scene, for example in the form of RGB (Red Green Blue) or HSL (Hint, Saturation, Luminosity) colorimetric coordinates.

[0067] The images acquired by cameras 11 and 12 represent views of the same scene taken from different viewpoints, the camera positions being distinct. For example, this scene includes: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.

[0068] These images are sent to a computer in a device fitted to the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.

[0069] A process for determining depth by a stereoscopic vision system embedded in the vehicle 10 is advantageously implemented by the vehicle 10, i.e. by a processor, a computer or a combination of computers of the embedded system of the vehicle 10, for example by the computer or computers in charge of the stereoscopic vision system of the vehicle 10.

[0070] In a first operation, the computer receives initial data representative of: • a first image 31 acquired by the first camera 11 at a first temporal instant of acquisition, and • a second image 32 acquired by the second camera 12 at the first temporal instant of acquisition.

[0071] The first and second images 31, 32 form a pair of stereoscopic images, that is to say images acquired by two separate cameras at the same time instant and representative of the same three-dimensional scene.

[0072] As illustrated in [Fig. 3], the first image 31 comprises a set of pixels, including a first pixel 311, each pixel representing an object in the three-dimensional scene unfolding in the environment of the vehicle 10 and present in the field of view of the first camera 11. Indeed, a pixel in the first image 31 or the second image 32 is the smallest visible unit and corresponds to a point of light resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes an object, photons are emitted or reflected, captured by the first camera 11 or the second camera 12, each equipped with a photosensitive sensor. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details.The combination of millions of pixels creates a first or second image that faithfully represents the physical object recorded by the first or second camera.

[0073] In a second operation, a first pixel 311 of the first image 31 is reprojected into the three-dimensional scene to obtain a point 312 corresponding to the first pixel as a function of first coordinates of the first pixel 311 in the first image 31, a first depth associated with the first pixel and intrinsic parameters of the first camera 11. This second operation is equivalent to determining the spatial coordinates of the point corresponding to those of the object which emitted or reflected the light captured by the first camera 11 and corresponding to the first pixel 311.

[0074] The first pixel 311 of the first image 31 is, for example, reprojected into the scene

[0075]

[0076]

[0077]

[0078]

[0079]

[0080]

[0081] three-dimensional by the following function: [Math.4] With : • the coordinates of the point in the three-dimensional scene in the y! reference frame 1 IV J the first camera 11, • K]1 the inverse matrix of the intrinsic matrix of the first camera 11, • [-^ / 1 a homogeneous vector comprising the coordinates of the first pixel 311 of the 1. first image 31 having xi as its abscissa and ordinate, • a second depth associated with the pixel of the first image, • Clx is the abscissa of the optical center of an image acquired by the first camera 11, and • Cy is the ordinate of the optical center of an image acquired by the first camera 11. It should be noted that the abscissa xi is the component of the coordinates of pixel 311 in the first image 31 along a horizontal axis and the ordinate 3^ is the component of the coordinates of pixel 311 in the first image 31 along a vertical axis. Second coordinates of point 312 are thus defined in a reference frame associated with the first camera 11. It should also be noted that if the origin of a reference frame associated with the first image 31 is located in the top left corner of the first image 31, then the coordinates of the optical center of the first image 31 are defined by: [Math.5] And [Math.5] rl - h Ly — 2 With : • 4 the abscissa of the optical center of the first image 31, • Cy, the ordinate of the optical center of the first image 31,

[0082]

[0083] • wi the width of the first image 31, and • hi the height of the first image 31. Similarly, for the second image 32, if the origin of a reference frame associated with the second image 32 is located in the upper left corner of the second image 32, then the coordinates of the optical center of the second image 32 are defined by: [Math.6] cr = 42 And [Math.6]

[0084]

[0085] With : • At the abscissa of the optical center of the second image 32, • Cj is the ordinate of the optical center of the second image 32, • wr the width of the second image 32, and • hr the height of the second image 32. In a third operation, third coordinates of point 312 are determined in a reference frame associated with the second camera 12 as a function of the second coordinates and extrinsic parameters of the stereo vision system-

[0086]

[0087] scopic. The transition from second coordinates expressed in the frame of reference of the first camera 11 to third coordinates expressed in the frame of reference of the second camera 12 is obtained, for example, by the following function: [Math.7] Yw Z' yl 2 W yl 1 ' cos fl 0 - sinfl Z^sin^ + B 1 WTU - X^sin# + Z ! w cosd either [Math.7] + Z W / Sin# + B 2 IV 7r

[0088] With : fi

[0089]

[0090]

[0091]

[0092] the coordinates of point 312 of the three-dimensional scene in the reference frame from the second camera 12, • [R, T] the matrix for going from the frame of reference of the first camera 11 to the frame of reference from the second camera 12, yl I 7l a homogeneous vector including coordinates of point 312 of the scene three-dimensional in the frame of reference of the first camera 11, • 0 the yaw angle, • B the reference base, and • H the difference in height between the first 11 and second 12 cameras. In a fourth operation, fourth coordinates of a second pixel 321 are determined in the second image 32, the second pixel 321 corresponding to the projection of point 312 in the second image 32. The fourth coordinates are determined as a function of the third coordinates and intrinsic parameters of the second camera 12. The fourth coordinates are, for example, obtained by the following function: [Math. 8] Xw y" -Kr ¥rw With : 7W fr 0 Ci 0 Y and 0 0 1 Xw Xw )cos° + Z^sing+g) +cj( +Z[ycos0 j " \ HAS / a homogeneous vector comprising the homogeneous coordinates of the second pixel 321 of the second image 32 having X? for homogeneous abscissa, j'.1 for homogeneous ordinate and zr for homogeneous depth, • Kr the intrinsic matrix of the second camera 12, • *X'w coordinates of a point in the three-dimensional scene in the reference frame Y\ Z\ WJ from the second camera 12, • f the focal length of the second camera 12, • clx the abscissa of the optical center of the first image 31, • the ordinate of the optical center of the first image 31, • cî the abscissa of the optical center of the second image 32, • Cy, the ordinate of the optical center of the second image 32, • 0 the yaw angle, • B the reference base, • H the difference in height between the first 11 and second 12 cameras, and • Zlw a depth of point 312 in the frame of the first camera 11.

[0093] The following function converts homogeneous coordinates into coordinates in the second image 32:

[0094] [Math.9]

[0095] With : • the coordinates of the second pixel 321 in the second image 32 having xr v - r for x-coordinate and yr for y-coordinate, • x]1 a homogeneous abscissa, a homogeneous ordinate and zr a depth homogeneous, • f the focal length of the second camera 12, • dx the abscissa of the optical center of the first image 31, • Cy, the ordinate of the optical center of the first image 31, • c>: the abscissa of the optical center of the second image 32, • the ordinate of the optical center of the second image 32, • 0 the yaw angle, • B the reference base, • H the difference in height between the first 11 and second 12 cameras, and • Z!w a depth of point 312 in the frame of the first camera 11.

[0096] In a fifth operation, an optical flux F associated with the first pixel 311 is determined, the optical flux being representative of a displacement vector between the first pixel 311 and a third pixel of the second image 32, the first and third pixels corresponding to the same object in the three-dimensional scene.

[0097] The fifth operation consists in particular of determining the third pixel in the second image 32 corresponding to the first pixel 311 of the first image 31; this operation is called "stereo matching" (in English, "feature matching"). Stereo matching, or disparity estimation, is the process of finding pixels in stereoscopic views that correspond to the same object in the three-dimensional scene.

[0098] The stereo matching operation is performed by implementing a method known as optical flow computing. Such a method is described in particular in "PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume" by Deqing Sun, Xiaodong Yang, Ming-Yu Liu and Jan Kautz, September 2017.

[0099] The optical flow calculation method is performed, for example, by a convolutional neural network (CNN). This type of tool is commonly used in image processing.

[0100] The output data of this fifth operation is a vector representing a displacement between the first pixel 311 of the first image 31 and the third pixel corresponding to the first pixel 311 in the second image 32.

[0101] According to a particular embodiment, the optical flux F is approximated as a vector with two components: a first component Fx along a horizontal axis of the second image corresponding to the x-axis, and a second component Fy along a vertical axis of the second image corresponding to the y-axis. Thus, the coordinates of the third pixel in the second image are defined by:

[0102] [Math. 10] (x'r-C() = (Xi-cQ + FX And [Math. 10]

[0103] With: • x'r the coordinates of the third pixel in the second image 32 having x'r for v' -7 r abscissa and y' for ordinate, • the abscissa of the optical center of the first image 31, • the ordinate of the optical center of the first image 31, • the abscissa of the optical center of the second image 32, • Cy, the ordinate of the optical center of the second image 32,

[0104]

[0105]

[0106]

[0107]

[0108]

[0109]

[0110] [YES] • Fx, the first component of the optical flux, and • Fy the second component of the optical flux. Since the first and second images 31, 32 are not rectified, it is not possible to directly predict a depth associated with the first pixel 311. Thus, in a sixth operation, second depths associated with the first pixel 311 are determined from the fourth coordinates and the optical flux, the second and third pixels being coincident. Indeed, the second 321 and third pixels of the second image 32 correspond to the projection of the same object in the second image 32, itself coincident with point 312. Thus, in the absence of anomaly, the coordinates of the second pixel 321 and the third pixel are equal. The following equations are then verified: [Math. 11] x'r = xr And [Math. 11] y' = y - rsf With : x'r the coordinates of the third pixel in the second image 32 having x'r for y',- abscissa and y' for ordinate, and the coordinates of the second pixel 321 in the second image 32 having xr for abscissa and yr for ordinate. From the equations determined in [Math 9] and [Math 11], we then obtain the following equations allowing us to determine different depths of the point 312 in the frame of the first camera 11, a third depth along the x-axis of the second image 32 and a fourth depth along the y-axis of the second image 32. The third depth is obtained by the following functions: [Math. 12] 7l =____________f£____________ vr , ( vo isin(î \ ( )—; TT,—COS0- r And [Math. 12] = (xr4) + FX^S With : • xr an abscissa of the second pixel 321, • f the focal length of the second camera 12, • dx the abscissa of the optical center of the first image 31, • dy the ordinate of the optical center of the first image 31, • d* the abscissa of the optical center of the second image 32, • Cy, the ordinate of the optical center of the second image 32, • 0 the yaw angle, • B the reference base, • H is the difference in height between the first 11 and second 12 cameras, • Fx, the first component of the optical flux, and • Y the third depth of point 312 in the frame of the first camera 11.

[0112] The fourth depth is obtained by the following functions:

[0113] [Math. 13] 7i =__________fF__________ [ ] (cosf?-—j-— ) And [Math. 13]

[0114] With: • yr an ordinate of the second pixel 321, • / the focal length of the second camera 12, • d, the abscissa of the optical center of the first image 31, • c < the ordinate of the optical center of the first image 31, • d, the abscissa of the optical center of the second image 32, • Cy, the ordinate of the optical center of the second image 32, • 0 the yaw angle, • B the reference base, • H is the difference in height between the first 11 and second 12 cameras, • Fy, the second component of the optical flux, and * Zlw v 'a fourth depth of point 312 in the first camera 11 frame.

[0115] Thus it is possible to determine the second depth associated with the first pixel 311 from at least one of the two preceding equations.

[0116] According to a first particular embodiment, the second depth is equal to the third depth.

[0117] According to a second particular embodiment, the second depth is equal to the fourth depth.

[0118] According to a third particular embodiment, the second depth is determined from the third and fourth depths, for example by the following function:

[0119] [Math. 14] Z^y = G^Z\y^ + (1 - (X^'Zw v

[0120] With: • Zlw the second depth of point 312 in the frame of the first camera 11, • H is a weighting coefficient between 0 and 1, • Zlw x the third depth of point 312 in the coordinate system of the first camera 11, and * Z!w_y the fourth depth of point 312 in the frame of the first camera 11.

[0121] According to a particular embodiment of the third particular embodiment, the coefficient a is equal to 0.5, the second depth is then equal to the average of the third and fourth depths.

[0122] It should be noted that the first particular embodiment corresponds to a value of a equal to 1 and the second particular embodiment corresponds to a value of a equal to 0.

[0123] Thus, the depth determination process using a stereoscopic vision system embedded in the vehicle 10 allows for the rapid and accurate determination of the depth associated with a pixel in a stereoscopic image pair. Indeed, the images do not need to be processed, resulting in time savings, as image rectification requires significant machine power and lengthy processing time.

[0124] An AD AS receiving the depth data predicted by this process of Depth prediction, combined with the vehicle's onboard stereoscopic vision system, provides rapidly updated information and is therefore more responsive. Its operation is thus safer and more precise.

[0125] Figure [Fig. 2] illustrates a flowchart of the different steps of a method 2 for determining depth by a stereoscopic vision system mounted in a vehicle, for example in the vehicle 10 of [Fig. 1], according to a particular and non-limiting embodiment of the present invention.

[0126] The method 2 is for example implemented by one or more processors of one or more computers embedded in the vehicle 10, for example by a computer controlling the stereoscopic vision system.

[0127] In a first step 21, representative data of: • a first image 31 acquired by the first camera 11 at a given time acquisition, and • a second image 32 acquired by the second camera 12 at the same time instant of acquisition are received.

[0128] In a second step 22, a first pixel 311 of the first image 31 is reprojected into the three-dimensional scene to obtain a point 312 corresponding to the first pixel as a function of first coordinates of the first pixel 311 in the first image 31, a first depth associated with the first pixel and intrinsic parameters of the first camera 11. Second coordinates of the point 312 are defined in a reference frame associated with the first camera 11.

[0129] In a third step 23, third coordinates of point 312 are determined in a reference frame associated with the second camera 12 as a function of the second coordinates and extrinsic parameters of the stereoscopic vision system.

[0130] In a fourth step 24, fourth coordinates of a second pixel 321 are determined in the second image 32, the second pixel 321 corresponding to the projection of point 312 in the second image 32. The fourth coordinates are determined as a function of the third coordinates and intrinsic parameters of the second camera 12.

[0131] In a fifth step 25, an optical flux F associated with the first pixel 311 is determined, the optical flux being representative of a displacement vector between the first pixel 311 and a third pixel of the second image 32, the first and third pixels corresponding to the same object of the three-dimensional scene.

[0132] In a sixth step 26, second depths associated with the first pixel 311 are determined from the fourth coordinates and the optical flow, the second and third pixels being confused.

[0133] Figure 4 schematically illustrates a device 4 configured for determining depth by a stereoscopic vision system mounted in a vehicle, for example in the vehicle 10 of Figure 1, according to a particular and non-limiting embodiment of the present invention. The device 4 corresponds, for example, to a device mounted in the vehicle 10, for example a computer.

[0134] Device 4 is, for example, configured to carry out the operations described opposite Figures 1 and 3 and / or the steps described opposite [Fig. 2]. Examples of such a device 4 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 4, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 4 may be implemented in the form of electronic circuits or software (or computer) modules or a combination of electronic circuits and software modules.

[0135] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.

[0136] The computer code of the embedded software(s), including the instructions to be loaded and executed by the processor, is for example stored on memory 4L

[0137] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.

[0138] According to a particular and non-limiting embodiment, the device 4 includes a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 include one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HDMI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").

[0139] According to another particular and non-limiting embodiment, device 4 includes a communication interface 43 which allows communication with other devices (such as other computers in the embedded system) via a communication channel 430. The communication interface 43 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds, for example, to a wired network of the type CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3).

[0140] According to a particular and non-limiting embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch or not, one or more speakers 450 and / or other peripherals 460 via the output interfaces 44, 45, 46 respectively. According to a variant, one or more of the external devices is integrated into the device 4.

[0141] Of course, the present invention is not limited to the embodiments described above but extends to a method for measuring the distance between an object and a vehicle equipped with a stereoscopic vision system, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0142] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 4 of [Fig.4].

Claims

1. Demands Method for determining depth by means of a stereoscopic vision system mounted in a vehicle (10), said method being implemented by a processor, the stereoscopic vision system comprising an array of at least two cameras (11, 12) arranged so as to each acquire an image of a three-dimensional scene from a different viewpoint, said method being characterized in that it comprises the following steps: - reception (21) of data representing a first image (31) acquired by the first camera (11) at a given acquisition time and a second image (32) acquired by the second camera (12) at said acquisition time - reprojection (22) of a first pixel (311) of the first image (31) in the three-dimensional scene to obtain a point (312) corresponding to said first pixel as a function of first coordinates of said first pixel in the first image, of a first depth associated with said first pixel and of intrinsic parameters of the first camera (11), the second coordinates of the point (312) being defined in a reference frame associated with the first camera (11); - determination (23) of third coordinates of the point (312) in a reference frame associated with the second camera (12) as a function of the second coordinates and extrinsic parameters of the stereoscopic vision system; - determination (24) of fourth coordinates of a second pixel (321) in the second image (32) corresponding to the projection of the point (312) in the second image (32), the fourth coordinates being determined as a function of the third coordinates and intrinsic parameters of the second camera (12); - determination (25) of an optical flux (F) associated with said first pixel (311), said optical flux being representative of a displacement vector between the first pixel (311) and a third pixel of the second image, the first and third pixels corresponding to the same object of the three-dimensional scene; - determination (26) of second depths associated with the first pixel (311) from the fourth coordinates and the optical flux, the the second and third pixels being indistinguishable.

2. A method according to claim 1, further comprising a step of determining a third depth associated with the first pixel from components along a horizontal axis of the second image of said fourth coordinates and said optical flux, the second depth being determined from the third depth.

3. A method according to claim 1 or 2, further comprising a step of determining a fourth depth associated with the first pixel from components along a vertical axis of the second image of said fourth coordinates and said optical flux, the second depth being determined from the fourth depth.

4. A method according to claim 3 depending on claim 2, wherein the second depth is determined from the third and fourth depths.

5. A method according to claim 4, wherein the second depth is equal to an average of the third and fourth depths.

6. A method according to any one of claims 1 to 5, wherein said optical flux is predicted by a convolutional neural network.

7. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by a processor.

8. Computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to claims 1 to 6.

9. Device (4) for determining depth by means of a stereoscopic vision system mounted in a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for carrying out the steps of the method according to any one of claims 1 to 6.

10. Vehicle (10) comprising the device (4) according to claim 9.