Method and device for determining a depth by means of a stereoscopic vision system on board a vehicle
The method and device for determining depth using a stereoscopic vision system on vehicles address the challenge of rapid and accurate depth prediction, enhancing ADAS system responsiveness by utilizing reprojection and optical flow calculations, reducing processing time and improving image rectification requirements.
Patent Information
- Application Number
- PCT/FR2025/000003
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2025-01-09
- Publication Date
- 2025-07-17
AI Technical Summary
Existing stereoscopic vision systems in vehicles face challenges in processing images quickly and efficiently due to the need for heavy image rectification, which hinders the responsiveness of ADAS systems by delaying the update of depth data, while maintaining accuracy and scalability is limited by the compact and lightweight requirements of on-board systems.
A method and device for determining depth using a stereoscopic vision system on a vehicle that involves reprojection of pixels, determination of optical flow, and calculation of depths associated with pixels, utilizing intrinsic and extrinsic parameters of multiple cameras, and optionally employing a convolutional neural network for optical flow prediction, to rapidly predict depths without extensive image processing.
This approach significantly reduces processing time, enabling ADAS systems to have rapidly updated and accurate depth data, enhancing their responsiveness and safety by improving the operational efficiency of ADAS systems.
Smart Images

Figure FR2025000003_17072025_PF_FP_ABST
Abstract
Description
DESCRIPTION Title: Method and device for determining a depth by a vision system in a vehicle Technical field
[0001] The present invention claims priority from French application 2400235 filed on January 11, 2024, the content of which (text, drawings and claims) is incorporated herein by reference.
[0002] The present invention relates to methods and devices for determining a depth using a stereoscopic vision system on board a vehicle, for example in a motor vehicle. The present invention also relates to a method and a device for measuring a distance separating an object from a vehicle carrying a stereoscopic vision system. The present invention also relates to a method and a device for controlling one or more ADAS systems on board a vehicle based on the determined depth. Technological background
[0003] Many modern vehicles are equipped with so-called ADAS (Advanced Driver Assistance System). ADAS are passive and active safety systems designed to eliminate human error in the operation of all types of vehicles. ADAS uses advanced technologies to assist the driver while driving and thus improve their performance. ADAS uses a combination of sensor technologies to perceive the environment around a vehicle, then provides information to the driver or influences certain vehicle systems.
[0004] There are several levels of ADAS, such as rearview cameras and blind spot sensors, lane departure warning systems, adaptive cruise control, and automatic parking systems.
[0005] ADAS systems embedded in a vehicle are powered by data obtained from one or more on-board sensors such as, for example, cameras. These cameras are used to detect and locate other road users or possible obstacles around a vehicle in order to, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle speed; - to act on the braking system in the event of a risk of impact with an object.
[0006] Determining depth or distance from images acquired by a stereoscopic vision system is performed using images acquired by this vision system. In order to enable the prediction of great depths or distances, for example greater than 200 meters, a stereoscopic vision system on board a vehicle generally comprises at least two cameras separated by several tens of centimeters. This distance causes a difference in the field of vision of the cameras of the stereoscopic vision system; the images obtained by the latter must then be processed to enable their exploitation. Such processing most often consists of rectifying these images.
[0007] Image processing, however, requires a lot of time and resources, while an ADAS powered by depth or distance requires rapid updating of this data while maintaining its accuracy. The resources allocated to the stereoscopic vision system are not scalable, however, as an on-board system must be compact and / or lightweight in order to be integrated into a vehicle.
[0008] Thus, the proper functioning of the driving assistance devices using this data depends on the quality of the data transmitted and the speed of image processing by a vision system. Summary of the present invention
[0009] An object of the present invention is to solve at least one of the problems of the technological background described above.
[0010] Another object of the present invention is to reduce the processing time of images acquired by a stereoscopic vision system on board a vehicle.
[0011] Another object of the present invention is to improve road safety, in particular by improving the operational safety of ADAS systems supplied by data obtained from at least one camera.
[0012] According to a first aspect, the present invention relates to a method for determining a depth by a stereoscopic vision system on board a vehicle, the method being implemented by a processor, the stereoscopic vision system comprising a set of cameras of at least two cameras arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the method being characterized in that it comprises the following steps: - reception of data representative of a first image acquired by the first camera at an acquisition time instant and of a second image acquired by the second camera at said acquisition time instant - reprojection of a first pixel of the first image into the three-dimensional scene to obtain a point corresponding to the first pixel as a function of first coordinates of the first pixel in the first image, a first depth associated with the first pixel and intrinsic parameters of the first camera, second coordinates of the point being defined in a reference system associated with the first camera; - determination of third coordinates of the point in a reference system associated with the second camera as a function of the second coordinates and extrinsic parameters of the stereoscopic vision system; - determination of fourth coordinates of a second pixel in the second image corresponding to the projection of the point in the second image, the fourth coordinates being determined as a function of the third coordinates and intrinsic parameters of the second camera; - determination of an optical flow associated with said first pixel, the optical flow being representative of a displacement vector between the first pixel and a third pixel of the second image, the first and third pixels corresponding to the same object of the three-dimensional scene; - determination of second depths associated with the first pixel from the fourth coordinates and the optical flow, the second and third pixels being the same.
[0013] Such a method thus makes it possible to predict depths associated with pixels of an image, that is to say to predict a distance separating the vehicle carrying the stereoscopic vision system from an object present in the field of vision of the first and second cameras. This method is notably faster than a method requiring heavy image processing such as rectification of the first and second images while the first and second cameras are several tens of centimeters apart. An ADAS controlled by data generated by this method, that is to say by the depths associated with the pixels of the first image, is then more responsive.
[0014] According to a variant, the method comprises a step of determining a third depth associated with the first pixel from components along a horizontal axis of the second image of the fourth coordinates and of the optical flow, the second depth being determined from the third depth.
[0015] According to another variant, the method comprises a step of determining a fourth depth associated with the first pixel from components along a vertical axis of the second image of the fourth coordinates and of the optical flow, the second depth being determined from the fourth depth.
[0016] According to a further variant of the method, the second depth is determined from the third and fourth depths.
[0017] According to a further variant of the method, the second depth is equal to an average of the third and fourth depths.
[0018] In another variant of the method, optical flow is predicted by a convolutional neural network.
[0019] According to a second aspect, the present invention relates to a device for determining a depth by a stereoscopic vision system on board a vehicle, the device comprising a memory associated with at least one processor configured for implementing the steps of the method according to the first aspect of the present invention.
[0020] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0021] According to a fourth aspect, the present invention relates to a computer program which comprises instructions adapted for executing the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0022] Such a computer program may use any programming language and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0023] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to the first aspect of the present invention.
[0024] On the one hand, the recording medium may be any entity or device capable of storing the program. For example, the medium may include a storage medium, such as a ROM memory, a CD-ROM or a ROM memory of microelectronic circuit type, or even a magnetic recording medium or a hard disk.
[0025] Furthermore, this recording medium may also be a transmissible medium such as an electrical or optical signal, such a signal being able to be conveyed via an electrical or optical cable, by conventional or hertzian radio or by self-directed laser beam or by other means. The computer program according to the present invention may in particular be downloaded from a network such as the Internet.
[0026] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to perform or to be used in performing the method in question. Brief description of the figures
[0027] Other characteristics and advantages of the present invention will emerge from the description of the particular and non-limiting exemplary embodiments of the present invention below, with reference to the appended figures 1 to 4, in which:
[0028] [Fig. 1] schematically illustrates a stereoscopic vision system on board a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention;
[0029] [Fig. 2] illustrates a flowchart of the different steps of a method for determining a depth by a stereoscopic vision system on board the vehicle of FIG. 1, according to a particular and non-limiting exemplary embodiment of the present invention;
[0030] [Fig. 3] schematically illustrates images acquired by the stereoscopic vision system on board the vehicle of FIG. 1, according to a particular and non-limiting exemplary embodiment of the present invention;
[0031] [Fig. 4] schematically illustrates a device configured for the determination of a depth by a stereoscopic vision system on board the vehicle of figure 1, according to a particular and non-limiting exemplary embodiment of the present invention. Description of examples of implementation
[0032] A method and a device for determining a depth by a stereoscopic vision system on board a vehicle will now be described in the following with joint reference to Figures 1 to 4. The same elements are identified with the same reference signs throughout the description which follows.
[0033] The terms "first(s)", "second(s)" (or "first(s)", "second(s)"), etc. are used in this document by arbitrary convention to identify and distinguish different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0034] According to a particular and non-limiting example of embodiment of the present invention, a method for determining a depth by a stereoscopic vision system on board a vehicle is for example implemented by a processor or a computer of the on-board system of the vehicle controlling this stereoscopic vision system.
[0035] The stereoscopic vision system comprises a set of cameras of at least two cameras arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the optical axes representative of an orientation of the field of vision of each camera being included in two parallel planes, for example oriented in a non-parallel manner. The stereoscopic vision system is thus said to be non-parallel.
[0036] For this purpose, the method for determining a depth by a stereoscopic vision system on board a vehicle comprises receiving data representative of a first image and a second image acquired respectively by a first and second camera at the same acquisition time instant. Coordinates of a second pixel in the second image corresponding to a pixel in the first image are determined via the reprojection of the first pixel into the three-dimensional scene to obtain a projected point in the second image.
[0037] An optical flow associated with the first pixel and representative of a displacement vector between the first pixel and a third pixel of the second image is determined.
[0038] Second depths associated with the first pixel are then determined from the coordinates and the optical flow, the second and third pixels being the same.
[0039] Figure 1 schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.
[0040] Such an environment 1 corresponds, for example, to a road environment formed of a network of roads accessible to the vehicle 10.
[0041] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, an electric motor(s) or a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, a motorcycle. Finally, the vehicle 10 corresponds to an autonomous vehicle or not, that is to say a vehicle traveling according to a determined level of autonomy or under the total supervision of the driver.
[0042] The vehicle 10 advantageously comprises several on-board cameras 11, 12, each configured to acquire images of a scene in the environment of the vehicle 10. This set of cameras 11, 12 forms the stereoscopic vision system. Two cameras 11 and 12 are illustrated in FIG. 1. The present invention is however not limited to a stereoscopic vision system comprising two cameras but extends to any stereoscopic vision system comprising 2 or more cameras, for example 2, 3, 4 or 5 cameras.
[0043] Both cameras 11, 12 have known intrinsic parameters. These parameters include: - the focal length of the first camera 11; - the focal length of the second camera 12; - distortions which are due to imperfections in the optical system of each camera; - the direction C1 of the optical axis of the first camera 11; - the direction C2 of the optical axis of the second camera 12; and - the respective resolutions of cameras 11, 12.
[0044] Intrinsic parameters characterize the transformation that associates, for an image point, the camera coordinates with the pixel coordinates, in each camera. These parameters do not change if the camera is moved.
[0045] The intrinsic matrix of the first camera 11 is defined by:
[0046] [Math 1] fi ^x
[0047] Ki = fi c y .0 0 1.
[0048] With : • K the intrinsic matrix of the first camera 11, • if the focal length of the first camera 11, • c x l the abscissa of the optical center of an image acquired by the first camera 11, and • c the ordinate of the optical center of an image acquired by the first camera 11.
[0049] Similarly, the intrinsic matrix of the second camera 12 is defined by:
[0050] [Math 2]
[0052] With : • K the intrinsic matrix of the second camera 12, • f r the focal length of the second camera 12, • cj the abscissa of the optical center of an image acquired by the second camera 12, and • Cy the ordinate of the optical center of an image acquired by the second camera 12.
[0053] The optical axes of the first and second cameras are arranged horizontally, that is, in two horizontal parallel planes. These two planes are separated by a height H, H having a zero value if the planes are coincident.
[0054] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, that is to say that the distortions are not taken into account or that their correction is processed at the time of image acquisition.
[0055] These two cameras 11, 12 are arranged so as to each acquire an image of a scene from a different point of view, the first point of view is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10, the second point of view is for example located on or in the right rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. In the case where the two cameras are located at the top of the windshield of the vehicle, they are then placed at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, the second camera 12 is located in the right rearview mirror of the vehicle 10.
[0056] A first marker is associated with the first camera 11: - the direction of the x axis is defined horizontal and normal to the optical axis of the first camera 11. The distance B separating the optical center of the first camera 11 from the projection of the optical center of the second camera 12 onto the horizontal plane passing through the optical center of the first camera 11 is called the reference base (in English “baseline”); - the direction of the y axis is defined vertical and normal to the optical axis of the first camera 11; - the direction of the z axis is defined orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal reference frame.
[0057] The extrinsic parameters related to the position of the cameras 11, 12 are the following parameters: - three translations in the x, y and z directions: Tx, Ty and Tz constituting the translation vector T with Tx = B, Ty = H and Tz = 0; and - a rotation around the y axis by an angle 0 called the yaw angle.
[0058] A matrix for moving from the reference frame of the first camera 11 to the reference frame of the second camera 12 is thus defined by:
[0059] [Math 3]
[0061] With : • [ / ?, T] the matrix to move from the frame of reference of the first camera 11 to the frame of reference of the second camera 12, • 0 yaw angle, • B the reference base, and • H the height difference between the first 11 and second 12 cameras.
[0062] Extrinsic parameters are determined, for example, during a calibration phase of the stereoscopic vision system.
[0063] A major constraint of the stereoscopic vision system used in automobiles is, for example, the large distance between the two cameras. Indeed, for To be able to cover a measuring range of 200 meters, the reference base must reach 60cm for cameras commonly used in this field.
[0064] The two cameras 11, 12 acquire images of a scene located in front of the vehicle 10, the first camera 11 covering only a first acquisition field 13, the second camera 12 covering only a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic vision of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic vision of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic vision of the scene by the stereoscopic vision system composed of the two cameras 11, 12.
[0065] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.
[0066] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system composed of the second camera 12.
[0067] Field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable using the monoscopic vision system composed of the first camera 11.
[0068] Finally, field 19 is not visible from any of the cameras. The part of the scene present in this field 19 is therefore not observable.
[0069] It is obvious that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with differently placed and oriented cameras.
[0070] The images acquired by the cameras 11, 12 at an acquisition time instant are presented in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colors and brightness of objects in the observed scene in the form, for example, of RGB colorimetric coordinates (from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness).
[0071] The images acquired by the cameras 11, 12 represent views of the same scene taken from different viewpoints, the positions of the cameras being distinct. On this scene are found for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0072] These images are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.
[0073] A process for determining a depth by a stereoscopic vision system on board the vehicle 10 is advantageously implemented by the vehicle 10, that is to say by a processor, a computer or a combination of computers of the on-board system of the vehicle 10, for example by the computer(s) in charge of the stereoscopic vision system of the vehicle 10.
[0074] In a first operation, the calculator receives initial data representative of: • a first image 31 acquired by the first camera 11 at a first acquisition time instant, and • a second image 32 acquired by the second camera 12 at the first acquisition time instant.
[0075] The first and second images 31, 32 form a pair of stereoscopic images, that is to say images acquired by two separate cameras at the same time instant and representative of the same three-dimensional scene.
[0076] As illustrated in Figure 3, the first image 31 comprises a set of pixels including a first pixel 311, each pixel being representative of an object of the three-dimensional scene taking place in the environment of the vehicle 10 and present in the field of vision of the first camera 11. Indeed, a pixel of the first image 31 or of the second image 32 is the smallest visible unit and corresponds to a luminous point resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When the light strikes an object, photons are emitted or reflected, captured by the first camera 11 or the second camera 12, each being equipped with a photosensitive sensor. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details.The combination of millions of pixels creates a first or second image faithfully representing the physical object recorded by the first or second camera.
[0077] In a second operation, a first pixel 311 of the first image 31 is reprojected into the three-dimensional scene to obtain a point 312 corresponding to the first pixel as a function of first coordinates of the first pixel 311 in the first image 31, of a first depth associated with the first pixel and of intrinsic parameters of the first camera 11. This second operation is equivalent to determining the spatial coordinates of the point thus corresponding to those of the object having emitted or reflected the light captured by the first camera 11 and corresponding to the first pixel 311.
[0078] The first pixel 311 of the first image 31 is for example reprojected into the three-dimensional scene by the following function:
[0079] [Math 4]
[0081] With: coordinates of the point of the three-dimensional scene in the reference frame of the first camera 11, • K t 1 the inverse matrix of the intrinsic matrix of the first camera 11, Xl- yi a homogeneous vector comprising the coordinates of the first pixel 311 of the LiJ first image 31 having x t for abscissa and y, for ordinate, • Zw a second depth associated with the pixel of the first image, • c x the abscissa of the optical center of an image acquired by the first camera 11, and • c the ordinate of the optical center of an image acquired by the first camera 11.
[0082] It should be noted that the abscissa x t is the component of the coordinates of pixel 311 in the first image 31 along a horizontal axis and the y ordinate tis the component of the coordinates of pixel 311 in the first image 31 along a vertical axis.
[0083] Second coordinates of point 312 are thus defined in a reference system associated with the first camera 11.
[0084] It should also be noted that if the origin of a reference frame associated with the first image 31 is located at the top left of the first image 31, then the coordinates of the optical center of the first image 31 are defined by:
[0085] [Math 5]
[0087] With : • c x the abscissa of the optical center of the first image 31, • Cy the ordinate of the optical center of the first image 31, • wi the width of the first image 31 , and • hi the height of the first image 31 .
[0088] Similarly, for the second image 32, if the origin of a reference frame associated with the second image 32 is located at the top left of the second image 32, then the coordinates of the optical center of the second image 32 are defined by:
[0089] [Math 6]
[0091] With : • c x r the abscissa of the optical center of the second image 32, the ordinate of the optical center of the second image 32, • w r the width of the second image 32, and • h r the height of the second image 32.
[0092] In a third operation, third coordinates of the point 312 are determined in a reference system associated with the second camera 12 as a function of the second coordinates and extrinsic parameters of the stereoscopic vision system.
[0093] The transition from second coordinates expressed in the reference frame of the first camera 11 to third coordinates expressed in the reference frame of the second camera 12 is obtained, for example, by the following function:
[0094] [Math 7] Y1
[0096] With: coordinates of point 312 of the three-dimensional scene in the reference frame of the second camera 12, • [R, T] the matrix to move from the reference frame of the first camera 11 to the reference frame of the second camera 12, a homogeneous vector comprising coordinates of point 312 of the scene three-dimensional in the frame of reference of the first camera 11, • 0 yaw angle, • B the reference base, and • H the height difference between the first 11 and second 12 cameras.
[0097] In a fourth operation, fourth coordinates of a second pixel 321 are determined in the second image 32, the second pixel 321 corresponding to the projection of the point 312 in the second image 32. The fourth coordinates are determined as a function of the third coordinates and intrinsic parameters of the second camera 12.
[0098] The fourth coordinates are, for example, obtained by the following function:
[0099] [Math 8]
[0101] With : - x w- • yÿ' a homogeneous vector comprising the homogeneous coordinates of the second pixel 321 of the second image 32 having x as homogeneous abscissa, y,- 17 for homogeneous ordinate and z r for uniform depth, • K r the intrinsic matrix of the second camera 12, coordinates of a point of the three-dimensional scene in the reference frame of the second camera 12, • f r the focal length of the second camera 12, • c x l the abscissa of the optical center of the first image 31, • c the ordinate of the optical center of the first image 31, • c x r the abscissa of the optical center of the second image 32, • c y r the ordinate of the optical center of the second image 32, • 0 yaw angle, • B the reference base, • H the height difference between the first 11 and second 12 cameras, and • Z w l a depth of point 312 in the reference frame of the first camera 11.
[0102] The following function converts homogeneous coordinates to coordinates in the second image 32:
[0103] [Math 9]
[0105] With: second pixel 321 in second image 32 having xr For abscissa and y r for ordinate, • a homogeneous abscissa, y a homogeneous ordinate and z r a homogeneous depth, • f r the focal length of the second camera 12, • c x l the abscissa of the optical center of the first image 31, • Cy the ordinate of the optical center of the first image 31, • the abscissa of the optical center of the second image 32, the ordinate of the optical center of the second image 32, • 0 yaw angle, • B the reference base, • H the height difference between the first 11 and second 12 cameras, and • Z^ a depth of point 312 in the reference frame of the first camera 11.
[0106] In a fifth operation, an optical flow F associated with the first pixel 311 is determined, the optical flow being representative of a displacement vector between the first pixel 311 and a third pixel of the second image 32, the first and third pixels corresponding to the same object of the three-dimensional scene.
[0107] The fifth operation consists in particular of determining the third pixel in the second image 32 corresponding to the first pixel 311 of the first image 31, this operation is called "stereo matching" (in English "feature matching"). Stereo matching or disparity estimation is the process of searching for pixels in the stereoscopic views which correspond to the same object in the three-dimensional scene.
[0108] The stereo matching operation is performed by implementing a method called optical flow computing. Such a method is notably described in "PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume" by Deqing Sun, Xiaodong Yang, Ming-Yu Liu and Jan Kautz from September 2017.
[0109] The optical flow calculation method is performed, for example, by a convolutional neural network (CNN). This type of tool is commonly used in image processing.
[0110] The output data of this fifth operation is a vector representative of a displacement between the first pixel 311 of the first image 31 and the third pixel corresponding to the first pixel 311 in the second image 32.
[0111] According to a particular exemplary embodiment, the optical flow F is assimilated to a vector having two components, a first component Fx along a horizontal axis of the second image corresponding to the abscissa axis, and a second component Fy along a vertical axis of the second image corresponding to the ordinate axis. Thus, the coordinates of the third pixel in the second image are defined by:
[0112] [Math 10]
[0114] With: the coordinates of the third pixel in the second image 32 having x' r For abscissa and y' r for ordinate, • c x the abscissa of the optical center of the first image 31, • c the ordinate of the optical center of the first image 31, • c x r the abscissa of the optical center of the second image 32, the ordinate of the optical center of the second image 32, • F xthe first component of optical flow, and • F y the second component of optical flow.
[0115] Since the first and second images 31, 32 are not rectified, it is not possible to directly predict a depth associated with the first pixel 311.
[0116] Thus, in a sixth operation, second depths associated with the first pixel 311 are determined from the fourth coordinates and the optical flow, the second and third pixels being merged. Indeed, the second 321 and third pixels of the second image 32 correspond to the projection of the same object in the second image 32 itself merged with the point 312. Thus, in the absence of an anomaly, the coordinates of the second pixel 321 and the third pixel are equal. The following equations are then verified:
[0117] [Math 11]
[0118] x' r = x r and there's r = y r
[0119] With : • y X the coordinates of the third pixel in the second image 32 having x' r for abscissa and y' r for ordinate, and • [ r ] the coordinates of the second pixel 321 in the second image 32 having x r for abscissa and y r for ordinate.
[0120] From the equations determined in [Math 9] and [Math 11], we then obtain the following equations making it possible to determine different depths of point 312 in the reference frame of the first camera 11, a third depth along the abscissa axis of the second image 32 and a fourth depth along the ordinate axis of the second image 32.
[0121] The third depth is obtained by the following functions:
[0122] [Math 12]
[0124] With: x r an abscissa of the second pixel 321, f rthe focal length of the second camera 12, c x the abscissa of the optical center of the first image 31, c the ordinate of the optical center of the first image 31, c x r the abscissa of the optical center of the second image 32, the ordinate of the optical center of the second image 32, 0 the yaw angle, B the reference base, H the height difference between the first 11 and second 12 cameras, F x the first component of optical flow, and Z w l x the third depth of point 312 in the reference frame of the first camera 11.
[0125] The fourth depth is obtained by the following functions:
[0126] [Math 13]
[0128] With: y r an ordinate of the second pixel 321, f r the focal length of the second camera 12, c xthe abscissa of the optical center of the first image 31, c y the ordinate of the optical center of the first image 31, c x r the abscissa of the optical center of the second image 32, c y r the ordinate of the optical center of the second image 32, 0 the yaw angle, B the reference base, H the height difference between the first 11 and second 12 cameras, F y the second component of optical flow, and Zw y the fourth depth of point 312 in the reference frame of the first camera 11.
[0129] Thus it is possible to determine the second depth associated with the first pixel 311 from at least one of the two preceding equations.
[0130] According to a first particular embodiment, the second depth is equal to the third depth.
[0131] According to a second particular embodiment, the second depth is equal to the fourth depth.
[0132] According to a third particular embodiment, the second depth is determined from the third and fourth depths, for example by the following function:
[0133] [Math 14]
[0135] With : • Z w l the second depth of point 312 in the reference frame of the first camera 11, • has a weighting coefficient between 0 and 1, • Zw x the third depth of point 312 in the reference frame of the first camera 11, and • Z w l y the fourth depth of point 312 in the reference frame of the first camera 11.
[0136] According to a particular embodiment of the third particular embodiment, the coefficient a is equal to 0.5, the second depth is then equal to the average of the third and fourth depths.
[0137] It should be noted that the first particular embodiment corresponds to a value of a equal to 1 and the second particular embodiment corresponds to a value of a equal to 0.
[0138] Thus, the process of determining a depth by a stereoscopic vision system on board the vehicle 10 makes it possible to determine a depth associated with a pixel of an image of a pair of stereoscopic images quickly and accurately. Indeed, the images do not have to be processed, which allows a gain of time, image rectification requiring significant machine capacity and a long processing time.
[0139] An ADAS receiving the depth data predicted by this depth prediction process associated with the stereoscopic vision system on board the vehicle 10 then has rapidly updated information and is therefore more responsive. Its operation is then safer and more precise.
[0140] Figure 2 illustrates a flowchart of the different steps of a method 2 for determining a depth by a stereoscopic vision system on board a vehicle, for example in the vehicle 10 of Figure 1, according to a particular and non-limiting exemplary embodiment of the present invention.
[0141] The method 2 is for example implemented by one or more processors of one or more computers on board the vehicle 10, for example by a computer controlling the stereoscopic vision system.
[0142] In a first step 21, data representative of: • a first image 31 acquired by the first camera 11 at an acquisition time instant, and • a second image 32 acquired by the second camera 12 at the same acquisition time instant is received.
[0143] In a second step 22, a first pixel 311 of the first image 31 is reprojected into the three-dimensional scene to obtain a point 312 corresponding to the first pixel as a function of first coordinates of the first pixel 311 in the first image 31, of a first depth associated with the first pixel and of intrinsic parameters of the first camera 11. Second coordinates of the point 312 are defined in a reference system associated with the first camera 11.
[0144] In a third step 23, third coordinates of the point 312 are determined in a reference system associated with the second camera 12 as a function of the second coordinates and extrinsic parameters of the stereoscopic vision system.
[0145] In a fourth step 24, fourth coordinates of a second pixel 321 are determined in the second image 32, the second pixel 321 corresponding to the projection of the point 312 in the second image 32. The fourth coordinates are determined as a function of the third coordinates and intrinsic parameters of the second camera 12.
[0146] In a fifth step 25, an optical flow F associated with the first pixel 311 is determined, the optical flow being representative of a displacement vector between the first pixel 311 and a third pixel of the second image 32, the first and third pixels corresponding to the same object of the three-dimensional scene.
[0147] In a sixth step 26, second depths associated with the first pixel 311 are determined from the fourth coordinates and the optical flow, the second and third pixels being the same.
[0148] Figure 4 schematically illustrates a device 4 configured for determining a depth by a stereoscopic vision system on board a vehicle, for example in the vehicle 10 of Figure 1, according to a particular and non-limiting exemplary embodiment of the present invention. The device 4 corresponds for example to a device on board the vehicle 10, for example a computer.
[0149] The device 4 is for example configured for the implementation of the operations described with regard to figures 1 and 3 and / or the steps described with regard to figure 2. Examples of such a device 4 include, but are not limited to, on-board electronic equipment such as an on-board computer of a vehicle, an electronic calculator such as an ECU (“Electronic Control Unit”), a smartphone, a tablet, a laptop. The elements of the device 4, individually or in combination, can be integrated in a single integrated circuit, in several integrated circuits, and / or in discrete components. The device 4 can be produced in the form of electronic circuits or software (or computer) modules or even a combination of electronic circuits and software modules.
[0150] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the method and / or for executing the instructions of the software(s) embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41 corresponding for example to a volatile and / or non-volatile memory and / or comprises a memory storage device which may comprise volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.
[0151] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored in the memory 41.
[0152] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Telematic Control Unit”), for example via a communication bus or through dedicated input / output ports.
[0153] According to a particular and non-limiting exemplary embodiment, the device 4 comprises a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 comprise one or more of the following interfaces: - RF radio frequency interface, for example Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English “Universal Serial Bus” or “Universal Serial Bus” in French); HDMI interface (from the English “High Definition Multimedia Interface” or “High Definition Multimedia Interface” in French); - LIN interface (from the English “Local Interconnect Network”).
[0154] According to another particular and non-limiting exemplary embodiment, the device 4 comprises a communication interface 43 which makes it possible to establish communication with other devices (such as other computers of the on-board system) via a communication channel 430. The communication interface 43 corresponds for example to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds for example to a wired network of the CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by the ISO 17458 standard) or Ethernet (standardized by the ISO / IEC 802-3 standard).
[0155] According to a particular and non-limiting exemplary embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch-sensitive or not, one or more speakers 450 and / or other peripherals 460 via the output interfaces 44, 45, 46 respectively. According to a variant, one or other of the external devices is integrated into the device 4.
[0156] Of course, the present invention is not limited to the exemplary embodiments described above but extends to a method for measuring a distance separating an object from a vehicle carrying a stereoscopic vision system, which would include secondary steps without thereby departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0157] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-based motor vehicle, comprising the device 4 of FIG. 4.
Claims
CLAIMS 1. Method for determining a depth by a stereoscopic vision system on board a vehicle (10), said method being implemented by a processor, the stereoscopic vision system comprising a set of cameras of at least two cameras (11, 12) arranged so as to each acquire an image of a three-dimensional scene from a different point of view, said method being characterized in that it comprises the following steps: - reception (21) of data representative of a first image (31) acquired by the first camera (11) at an acquisition time instant and of a second image (32) acquired by the second camera (12) at said acquisition time instant - reprojection (22) of a first pixel (311) of the first image (31) into the three-dimensional scene to obtain a point (312) corresponding to said first pixel as a function of first coordinates of said first pixel in the first image, of a first depth associated with said first pixel and of intrinsic parameters of the first camera (11), second coordinates of the point (312) being defined in a reference system associated with the first camera (11); - determination (23) of third coordinates of the point (312) in a reference system associated with the second camera (12) as a function of the second coordinates and extrinsic parameters of the stereoscopic vision system; - determination (24) of fourth coordinates of a second pixel (321) in the second image (32) corresponding to the projection of the point (312) in the second image (32), the fourth coordinates being determined as a function of the third coordinates and intrinsic parameters of the second camera (12); - determination (25) of an optical flow (F) associated with said first pixel (311), said optical flow being representative of a displacement vector between the first pixel (311) and a third pixel of the second image, the first and third pixels corresponding to the same object of the three-dimensional scene; - determination (26) of second depths associated with the first pixel (311) from the fourth coordinates and the optical flow, the second and third pixels being the same.
2. Method according to claim 1, further comprising a step of determining a third depth associated with the first pixel from components along a horizontal axis of the second image of said fourth coordinates and said optical flow, the second depth being determined from the third depth.
3. Method according to claim 1 or 2, further comprising a step of determining a fourth depth associated with the first pixel from components along a vertical axis of the second image of said fourth coordinates and said optical flow, the second depth being determined from the fourth depth.
4. Method according to claim 3 as dependent on claim 2, for which the second depth is determined from the third and fourth depths.
5. The method of claim 4, wherein the second depth is equal to an average of the third and fourth depths.
6. Method according to one of claims 1 to 5, for which said optical flow is predicted by a convolutional neural network.
7. Computer program comprising instructions for implementing the method according to any one of the preceding claims, when these instructions are executed by a processor.
8. Computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to claims 1 to 6.
9. Device (4) for determining a depth by a stereoscopic vision system on board a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for implementing the steps of the method according to any one of claims 1 to 6.
10. Vehicle (10) comprising the device (4) according to claim 9.
Citation Information
Patent Citations
Bidirectional interface utilizing read-only memory, decoder and multiplexer
FR2400235A1
Stereo optical module and stereo camera
US20060082879A1