Method and device for rapid depth determination using a stereoscopic vision system mounted in a vehicle.
The stereoscopic vision system with image rectification and disparity analysis enhances ADAS responsiveness by rapidly determining depth, improving safety through precise distance estimation for vehicle navigation.
Patent Information
- Application Number
- FR2023012277
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2026-02-20
- Estimated Expiration
- 2043-11-10
AI Technical Summary
Existing vehicle ADAS systems face challenges in rapidly determining depth at high speeds due to limitations in image processing time, affecting their responsiveness and reliability.
A stereoscopic vision system using two cameras to acquire images from different viewpoints, followed by image rectification, resolution reduction, and disparity analysis to determine precise and rapid depth estimation, with bounding box pairing and resizing for enhanced accuracy.
Improves the responsiveness and reliability of ADAS systems by providing accurate and timely depth information for vehicle navigation, enhancing safety by quickly updating distance data for objects near the vehicle.
Abstract
Description
Title of the invention: Method and device for rapid determination of depth by means of a stereoscopic vision system mounted in a vehicle. technical field
[0001] The present invention relates to methods and devices for determining depth by stereoscopic vision system mounted in a vehicle, for example in a motor vehicle.
[0002] The present invention also relates to a method and a device for controlling one or more systems embedded in a vehicle from depth data determined by a stereoscopic vision system embedded in a vehicle. Technological background
[0003] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). These ADAS are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS use advanced technologies to assist the driver while driving and thus improve performance. ADAS use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.
[0004] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.
[0005] The AD AS systems embedded in a vehicle are powered by data obtained one or more onboard sensors, such as cameras. These cameras make it possible to detect and locate other road users or potential obstacles around a vehicle in order to, for example: - to adapt the vehicle's lighting according to the presence of other road users; - to automatically regulate the vehicle's speed; - to act on the braking system in case of risk of impact with an object.
[0006] The quality of the data emitted by a vision system therefore determines the proper functioning of the driving assistance devices using this data.
[0007] A vision system installed in a vehicle makes it possible to determine a depth or a distance separating the vehicle from an object present in the vehicle's environment. However, when a vehicle is traveling at high speed, for example on the highway, depth prediction time is critical. Indeed, the reaction time of an ADAS depends on image processing time and the depth prediction time of the predictive model associated with the vision system. This latter time must therefore be reduced to allow the vehicle, via the ADAS controls, to react quickly and effectively. Summary of the present invention
[0008] One object of the present invention is to solve at least one of the problems of the technological background described above.
[0009] Another object of the present invention is to improve the responsiveness of a depth prediction model associated with a vision system embedded in a vehicle.
[0010] Another object of the present invention is to improve road safety, in particular by improving the reliability of ADAS systems powered by data obtained from a vision system on board a vehicle.
[0011] According to a first aspect, the present invention relates to a method for determining depth using a stereoscopic vision system mounted in a vehicle traveling on a road, the stereoscopic vision system comprising a first camera and a second camera, each arranged to acquire an image of a three-dimensional scene from a different viewpoint with the same definition, the method being characterized in that it comprises the following steps: - reception of representative data from a first and second image acquired by the first and second cameras respectively at the same time of acquisition; - rectification of the first and second images according to extrinsic parameters of the stereoscopic vision system, so as to render horizontal epi-polar lines in the first and second rectified images; - generation of a third and a fourth image from the first and second images respectively by decreasing the definition of the first and second images; - determination of initial depths associated with pixels of the third image from initial disparities determined from the third and fourth images; - determination of a first set of pixels representative of the traffic lane and a first set of bounding boxes in the first image by processing the first image and a second set of pixels representative of the traffic lane and a second set of bounding boxes in the second image by processing the second image, each bounding box of the first and second sets of bounding boxes being associated with an object of the three-dimensional scene; - selection of a third set of bounding boxes comprising bounding boxes from the first set of bounding boxes, each associated with an object present on the traffic lane, and a fourth set of bounding boxes comprising bounding boxes from the second set of bounding boxes, each associated with an object present on the traffic lane; - pairing of bounding boxes from the third and fourth sets of bounding boxes to form a set of bounding box pairs associated with the same object, each pair of bounding boxes comprising one bounding box from the third set of bounding boxes and one bounding box from the fourth set of bounding boxes and resizing of the bounding boxes of each pair of bounding boxes according to the respective positions and sizes of the bounding boxes of the same pair of bounding boxes; - determination of second depths associated with pixels of the third set of bounding boxes from second disparities determined from the pixels of the bounding boxes of each pair of bounding boxes; and - association of the first depths to the pixels of the first image not included in a bounding box of the third set of bounding boxes.
[0012] Initial depth values are assigned to pixels of the first image outside the bounding boxes, allowing a depth value to be obtained for each pixel of the first image with greater precision within the bounding boxes. Thus, the distance between an object in the vehicle's lane is determined more accurately and precisely, while the distance to objects outside the vehicle's lane is determined more roughly and more quickly.
[0013] An AD AS using input data such as depths determined by the The stereoscopic vision system, which determines the distance between a part of the vehicle and another road user in its vicinity, then has access to precise and rapidly updated data. This increases the responsiveness of the ADAS, thus improving its operation.
[0014] According to a variant of the method, the pixel coordinates constituting the first and second images being defined in a coordinate system having as its origin a top left corner of the first and second images respectively, each pair of bounding boxes comprising a first bounding box in the first image and a second bounding box in the second image, the resizing comprises the following steps: - determining the coordinates of the upper left corner, its width and height the first bounding box and coordinates of the upper left corner, width and height of the second bounding box, - determination : • with a minimum abscissa equal to the minimum value among the abscissas of the upper left corners of the first and second bounding boxes, • with a minimum ordinate equal to the minimum value among the ordinates of the upper left corners of the first and second bounding boxes, • with a maximum width calculated from the widths of the first and second bounding boxes, • of a maximum height measured from the heights of the first and second encompassing boxes, and - assignment: • from the minimum x-coordinate and minimum y-coordinate to the coordinates of the upper left corners of the first and second bounding boxes, and • of the maximum width and maximum height at the first and second encompassing boxes.
[0015] The bounding boxes of the same pair of bounding boxes are then of equal size and placed in the same relative position on each of the first and second images. The content of each bounding box of the same pair of bounding boxes allows for an analysis equivalent to that of two images, each image being contained within each bounding box respectively.
[0016] According to yet another variant, the process includes cropping: - of the first image so as to obtain a first image whose width and height are each divisible by 32, when the height and / or width of the first image before cropping is not divisible by 32, - of the second image so as to obtain a second image whose width and height are each divisible by 32, when the height and / or width of the second image before cropping is not divisible by 32.
[0017] According to a further variant of the method, a height and a width of the first bounding box and of the second bounding box after resizing are divisible by 32.
[0018] According to another variant of the method, a first depth and a second depth are determined respectively from the first and second disparities by the following formula: D(p) = fx B / d(p) With : • D(p) the first depth of a pixel p of the third image, respectively the second depth of a pixel p of the first image, • f a focal length of the first camera, • B, a distance separating the first camera from the second camera, and • d(p) the first disparity determined for a pixel p of the third image, respectively the second disparity determined for a pixel p of the first image.
[0019] According to a variant of the method, the first and second disparities are determined by implementing a convolutional neural network.
[0020] According to a second aspect, the present invention relates to a device for determining depth by means of a stereoscopic vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.
[0021] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0022] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0023] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0024] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.
[0025] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.
[0026] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.
[0027] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures
[0028] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 4, in which:
[0029] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a a particular and non-limiting example of a realization of the present invention;
[0030] [Fig.2] illustrates a flowchart of the different stages of a process for determining depth by a vision system on board the vehicle of [Fig.1], according to a particular and non-limiting example of the present invention;
[0031] [Fig.3] schematically illustrates a device configured for determining depth by a vision system onboard in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention;
[0032] [Fig.4] schematically illustrates a first and second image acquired by a vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements
[0033] A method and device for determining depth by means of a vision system mounted in a vehicle will now be described in what follows with joint reference to Figures 1 to 4. The same elements are identified with the same reference signs throughout the description that follows.
[0034] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0035] According to a particular and non-limiting example of an embodiment of the present invention, a method for determining depth by a stereoscopic vision system embedded in a vehicle is for example implemented by a computer of the vehicle's embedded system controlling this vision system.
[0036] The vision system comprises at least two cameras, each arranged to acquire an image of a scene from a different point of view.
[0037] For this purpose, the method of determining a depth by a vision system on board a vehicle includes the reception of data representative of a first and second image acquired by respectively a first and second camera of the set of cameras at the same time instant of acquisition.
[0038] The method also includes the generation of a third and a fourth image from the first and second images respectively by reducing their definition and the determination of first depths associated with the pixels of the third image from the third and fourth images.
[0039] The method further includes determining bounding boxes associated with objects present on the vehicle's traffic lane and determining second depths associated with the pixels of the first image included in these bounding boxes.
[0040] The first depths are assigned to the pixels of the first image outside the bounding boxes, making it possible to obtain a depth for any pixel of the first image with greater precision in the bounding boxes.
[0041] Fig. 1 schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.
[0042] An environment 1 corresponds, for example, to a road environment consisting of a network of roads accessible to vehicle 10.
[0043] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.
[0044] The vehicle 10 advantageously comprises a camera array including a first camera 11 and a second camera 12, each configured to acquire images of a three-dimensional scene in the environment 1 of the vehicle 10. This camera array forms the vision system. Two cameras are illustrated in [Fig. 1]. The present invention is not limited, however, to a vision system comprising two cameras but extends to any vision system comprising two or more cameras, for example, two, three, four, or five cameras.
[0045] The first 11 and second 12 cameras have known intrinsic parameters. These parameters include, in particular: - the focal length fl of the first camera 11; - the focal length f2 of the second camera 12; - distortions which are due to imperfections in the optical system of each camera; - the direction Cl of the optical axis of the first camera 11; - the C2 direction of the optical axis of the second camera 12; and - the respective resolutions of cameras 11, 12.
[0046] The intrinsic parameters characterize the transformation which associates, for a point image, camera coordinates to pixel coordinates, in each camera. These parameters do not change if the camera is moved.
[0047] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in lens curvature, parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are either not taken into account or their correction is addressed during image acquisition.
[0048] These first 11 and second 12 cameras are arranged so that each acquires an image of a three-dimensional scene from a different viewpoint. The first viewpoint is, for example, located on or in the left-hand rearview mirror of vehicle 10 or at the top of the windshield of vehicle 10. The second viewpoint is, for example, located on or in the right-hand rearview mirror of vehicle 10 or at the top of the windshield of vehicle 10. If both cameras are located at the top of the windshield, they are positioned at a certain distance. In this example, the first camera 11 is located at the top of the windshield of vehicle 10, and the second camera 12 is located in the right-hand rearview mirror of vehicle 10.
[0049] A first marker is associated with the first camera 11: - the direction of the y-axis is defined by the position of the second camera 12, so as to place the second camera 12 on the y-axis of the first camera 11. The distance B separating the two cameras 11,12 is called the reference base (in English "baseline") and the direction separating the two cameras 11,12 is that of the y-axis; - the direction of the x axis is defined orthogonal to that of the y axis and orthogonal to that of the optical axis Cl of the first camera 11; - The direction of the z-axis is defined as orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal coordinate system.
[0050] The extrinsic parameters related to the position of cameras 11, 12 are the following parameters: - 3 translations in the x, y, and z directions: Tx, Ty, and Tz forming the translation vector T; and - 3 rotations around the x, y and z axes: Rx, Ry and Rz, constituting the rotation matrix R.
[0051] The determination of extrinsic parameters is carried out for example during the calibration of the vision system.
[0052] A stereoscopic vision system is a vision system comprising a A plurality of cameras, for example, the first camera 11 and the second camera 12. A key constraint of a stereoscopic vision system used in the automotive industry is, for example, the large distance between the two cameras. Indeed, to cover a measurement range of 200 meters, the reference base must be 60 cm for cameras commonly used in this field.
[0053] The first 11 and second 12 cameras acquire images of a three-dimensional scene located in front of the vehicle 10, the first camera 11 alone covering a first acquisition field 13, the second camera 12 alone covering a second acquisition field 14 and the first 11 and second 12 cameras both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic view of the three-dimensional scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic view of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic view of the scene by the stereoscopic vision system composed of the first 11 and second 12 cameras.
[0054] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.
[0055] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system composed of the second camera 12.
[0056] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable with the monoscopic vision system composed of the first camera 11.
[0057] Finally, field 19 is not visible from any of the cameras. The part of the scene present in this field 19 is therefore not observable.
[0058] When the vehicle 10 is in motion, each of the first 11 and second 12 cameras forms a monoscopic vision system.
[0059] According to one embodiment, the directions Cl, C2 of the optical axes representing an orientation of the field of vision of each camera are oriented non-parallel so as to obtain the third acquisition field 15 of the environment 1 as wide as possible.
[0060] According to another embodiment, the directions Cl, C2 of the optical axes representing an orientation of the field of vision of each camera are oriented parallel.
[0061] It is evident that it is possible to use such a vision system, stereoscopic or monoscopic, to take images of scenes located to the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently.
[0062] The images acquired by the first 11 and second 12 cameras at a given acquisition time are in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric coordinates (from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).
[0063] The images acquired by the first 11 and second 12 cameras at the same time instant represent views of the same scene taken from different viewpoints, the camera positions being distinct. This scene contains, for example, objects, each object belonging to a type of object such as: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0064] These images are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.
[0065] A process for determining depth by the stereoscopic vision system on board the vehicle 10 is advantageously implemented by the vehicle 10, i.e. by a computer or a combination of computers of the vehicle 10's on-board system, for example by the computer or computers in charge of the vehicle 10's vision system.
[0066] The vehicle 10 travels, in particular, on a traffic lane. For example, the method is implemented when the vehicle 10 travels on a multi-lane road at high speed, such as a national highway or motorway. In this type of situation, it is important to predict the distances of objects in the vicinity of the vehicle 10 quickly. Indeed, since the vehicle 10 is traveling at high speed, it is important that the data relating to its environment be updated rapidly. In particular, it is important to know the presence and position of objects located on or near the traffic lane in which the vehicle 10 is traveling in order to feed, for example, an ADAS (Advanced Driver Assistance System).
[0067] In a first operation, representative data from a first 41 and second 42 image acquired by the first 11 and second 12 cameras respectively at the same acquisition time are received, for example by a computer. These images have, for example, the same resolution and pixel coordinates for each image are for example expressed in a reference frame whose origin is the upper left corner of each image.
[0068] According to a particular embodiment, the first image 41 is cropped when its width and / or height is not divisible by 32 so as to obtain a first image 41 whose width and height are divisible by 32. Similarly, the second image 42 is cropped when its width and / or height is not divisible by 32 so as to obtain a second image 42 whose width and height are divisible by 32.
[0069] In a second operation, the first 41 and second 42 images are rectified according to extrinsic parameters of the stereoscopic vision system, so as to render horizontal epipolar lines in the first 41 and second 42 rectified images.
[0070] The first and second images are rectified using a method known to those skilled in the art. Such a method is described, for example, in "Projective Rectification of Uncalibrated Infrared Stereo Images with Global Consideration of Distortion Minimization" by Benoit Ducarouge, Thierry Sentenac, Florian Bugarin and Michel Devy of July 16, 2009.
[0071] The rectification method consists of reorienting the epipolar lines so that they are parallel with the horizontal axis of the image. This method is described by a transformation that projects the epipoles to infinity and whose corresponding points are necessarily on the same ordinate.
[0072] A rectification algorithm consists, for example, of 4 steps: - Rotate (virtually) the first camera 11 so that the epipole goes to infinity along the horizontal axis of the frame associated with it; - Apply the same rotation to the second camera to return to the initial geometric configuration; - Rotate the second camera 12 by the rotation associated with the rotation matrix 'R', corresponding to the extrinsic parameter of the starting stereoscopic vision system; - Adjust the scale in both camera reference points.
[0073] It should be noted that rectification simplifies the matching of pixels in stereoscopic images, i.e., images obtained by a stereoscopic vision system. The pixel corresponding in the second image 42 to a pixel in the first image 41 (and vice versa) is positioned on the same line. Based on knowledge of the epipolar geometry and thus of a fundamental matrix of the stereoscopic vision system, the objective is then to determine a pair of projective transformations, called homographies, which reorient the epipolar projections parallel to the image lines, and therefore to the horizontal axis of the rectified cameras.
[0074] In the continuation of the description of this depth determination process, the first 41 and second 42 images refer to the first and second rectified images.
[0075] In a third operation, a third and fourth image are generated from the first 41 and second 42 images respectively by reducing the resolution of the first and second images. The reduction in width and height of the first 41 and second 42 images is done with the same scaling factor so as not to generate distortion in the third and fourth images.
[0076] Thus, by working on lower definition images, the number of images per second (in English "Frame Per Second" or "FPS") processed can be increased, for example to reach an image processing speed of around 50 images per second.
[0077] In a fourth operation, a fourth set of pixels of the fourth image corresponding to a third set of pixels of the third image is determined and first disparities associating the pixels of the third set of pixels with the pixels of the fourth set of pixels are obtained, for example by implementing a convolutional neural network, this operation is called "stereo matching" (in English "feature matching").
[0078] Stereo matching, or disparity estimation, is the operation of finding pixels in stereoscopic views that correspond to the same 3D point in the scene. Rectified epipolar geometry simplifies this matching operation along the same epipolar line. It is not necessary to calculate the coordinates of the 3D point to find the corresponding pixel on the same line in the other image. The disparity of a pixel is the distance d between a pixel and its corresponding pixel in the other image. The first disparities are horizontal here, with the third and fourth images being obtained from the first 41 and second 42 images previously rectified.
[0079] The output data of this operation is a disparity representing a displacement between each pixel of the third set of pixels of the third image and a fourth pixel corresponding to each third pixel in the fourth image.
[0080] In a fifth operation, first depths associated with pixels of the third image are determined from first disparities determined from the third and fourth images.
[0081] According to a particular embodiment, the first depths are determined from the first disparities by the following function: D(p) = fx B / d(p) With : • D(p) the first depth of a pixel p of the third image, • f a focal length of the first camera 11, • B a distance separating the first camera 11 from the second camera 12, and • d(p) the first disparity determined for a pixel p of the third image.
[0082] Thus, a first disparity is determined for each pixel of the third image, thereby determining a first depth associated with each pixel of the third image. A distance separating the objects of the three-dimensional scene present in the field of vision of the first camera 11 of the vehicle 10 is then determined.
[0083] The first depth is then called the panoptic depth. Indeed, the first depth is determined very quickly and for the entirety of the third image.
[0084] As illustrated in [Fig.4], in a sixth operation, a first set of pixels 410 representing the traffic lane and a first set of bounding boxes 411, 412, 413, 414 are determined in the first image 41 by processing the first image 41.
[0085] Similarly, in a seventh operation, a second set of pixels 420 representing the traffic lane and a second set of bounding boxes 421, 422, 423, 424 is determined in the second image 42 by processing the second image 42.
[0086] Each bounding box of the first and second sets of bounding boxes is then associated with an object of the three-dimensional scene.
[0087] These determinations of bounding boxes and traffic lanes from images are known to the person skilled in the art; the sixth and seventh operations are for example implemented using the YOLOPv2® algorithm applied to the first 41 and second 42 stereographic images.
[0088] In an eighth operation, a third set of bounding boxes 411, 414 comprising bounding boxes from the first set of bounding boxes 411, 412, 413, 414 each associated with an object present on the traffic lane is selected.
[0089] Similarly, in a ninth operation, a fourth set of bounding boxes 421, 424 comprising bounding boxes from the second set of bounding boxes 421, 422, 423, 424 each associated with an object present on the traffic lane is selected.
[0090] A bounding box of the first image 41 is, for example, associated with an object present on the traffic lane when pixels included in a bounding box of the first image 41 correspond to pixels included in the first set of pixels 410, and a bounding box of the second image 42 is, for example, associated with an object present on the traffic lane when pixels included in a bounding box of the second image 42 correspond to pixels included in the second pixel set 420.
[0091] In a tenth operation, bounding boxes from the third and fourth sets of bounding boxes are paired to form a set of bounding box pairs associated with the same object, each bounding box pair comprising one bounding box from the third set of bounding boxes and one bounding box from the fourth set of bounding boxes. Thus, each bounding box pair comprises a first bounding box in the first image 41 and a second bounding box in the second image 42, each associated with the same object in the three-dimensional scene.
[0092] In an eleventh operation, the bounding boxes of each pair of bounding boxes are resized according to the respective positions and sizes of the bounding boxes of the same pair of bounding boxes.
[0093] According to a particular embodiment in which the pixel coordinates constituting the first 41 and second 42 images are defined in a coordinate system having as its origin a top left corner of the first 41 and second 42 images respectively, each pair of bounding boxes comprising a first bounding box 411,414 in the first image and a second bounding box 421,424 in the second image, the resizing comprises the following twelfth, thirteenth and fourteenth operations.
[0094] In a twelfth operation, and according to this particular embodiment, the coordinates of the upper left corner, the width and height of the first bounding box 411, 414, and the coordinates of the upper left corner, the width and height of the second bounding box 421, 424 are determined. In a thirteenth operation, a minimum abscissa equal to the minimum value among the abscissas of the upper left corners of the first and second bounding boxes, a minimum ordinate equal to the minimum value among the ordinates of the upper left corners of the first and second bounding boxes, a maximum width from the widths of the first and second bounding boxes, and a maximum height from the heights of the first and second bounding boxes are determined.In a fourteenth operation, the minimum abscissa and minimum ordinate are then assigned to the coordinates of the upper left corners of the first 411, 414 and second 421, 424 bounding boxes, and the maximum width and maximum height to the first and second bounding boxes.
[0095] For example, for a first pair of bounding boxes comprising the first bounding box 411 and the second bounding box 421, the coordinates of the upper left corner (xp;}'p), the width wp and the height of the first bounding box 411 and the coordinates of the upper left corner (xp; yp) are determined. width wp and height of the second bounding box 421. Next, a minimum abscissa xmin is determined, equal to the minimum value among the x-coordinates of the upper left corners of the first and second bounding boxes, i.e., xmin = min(xp; X'P), a minimum ordinate ymin is equal to the minimum value among the ordinates of the upper left corners of the first and second bounding boxes, i.e., ymin = minyÿ; yp, a maximum width wmax is derived from the widths of the first and second bounding boxes, i.e., wmax = max(H'P; H'P), and a maximum height hmax is derived from the heights of the first and second bounding boxes, i.e., hmax = maxQ^j1; ^). The minimum abscissa xmin and minimum ordinate ymin are then assigned to the coordinates of the upper left corners of the first 411, 414 and second 421, 424 bounding boxes, and the maximum width wmax and the maximum height hmax at the first and second encompassing boxes.
[0096] According to another particular embodiment, in order for the width and height of the bounding boxes to be divisible by 32 (see above the particular embodiment in which the number of pixels of the first and second images is divided by 32 during the operation of generating the third and fourth images), the maximum width wmax and the maximum height hmax are obtained by the following functions: w max = max(wp; M^) + 32 - max(wp; h?)%32 and hmax = maxQ^; - max^O; ^)%32, %32 representing the operation of retrieving the remainder of a division by 32. Thus, according to this particular embodiment, each resized bounding box has a width and a height divisible by 32.
[0097] In a fifteenth operation, a sixth set of pixels from the second image 42 corresponding to a fifth set of pixels from the first image 41 is determined, and second disparities associating the pixels of the fifth set of pixels with the pixels of the sixth set of pixels are obtained, for example, as before, by implementing a convolutional neural network. It should be noted that the fifth set of pixels is included in the third set of bounding boxes, and that the fourth set of bounding boxes is included, and that the pixel correspondences are guaranteed by the pairing and resizing of the bounding boxes.Indeed, the bounding boxes of the same pair of bounding boxes function as two stereoscopic images extracted from the first 41 and second 42 images, their pairing allowing to have pixels associated with the same object of the three-dimensional scene and the resizing allowing to retain all the pixels associated with this same object.
[0098] The second disparities are here horizontal, the first 41 and second 42 images having been previously rectified (second operation).
[0099] In a sixteenth operation, second depths associated with pixels of the third set of bounding boxes are determined from second disparities determined from the pixels of the bounding boxes of each pair of bounding boxes.
[0100] According to a particular embodiment, the second depths are determined from the second disparities by the following function: D(p) = fx B / d(p) With : • D(p) the second depth of a pixel p of the first image, • f a focal length of the first camera 11, • B a distance separating the first camera 11 from the second camera 12, and • d(p) the second disparity determined for a pixel p of the first image.
[0101] Thus, a second disparity is determined for each pixel of the first image included in a bounding box associated with an object present on or around the traffic lane of vehicle 10. A distance separating the objects present on the traffic lane of vehicle 10 present in the field of vision of the first camera 11 of vehicle 10 is then determined.
[0102] In a seventeenth operation, the first depths are associated with the pixels of the first image 41 not included in a bounding box of the third set of bounding boxes 411, 414.
[0103] The determination of the second depths is performed on a first image 41 whose resolution is higher than that of the third image. The second depth determined is therefore more precise but also takes longer to determine. The advantage of precisely defining this second depth only in the third set of bounding boxes, that is to say only for objects present on or near the vehicle's lane 10, is that no second depth is predicted for the rest of the first image, thus saving time.
[0104] Depending on the reduction in resolution of the first image to generate the third image, the same first depth associated with a pixel of the third image is, for example, associated with several pixels of the first image, these same pixels of the first image 41 being joined into a single pixel in the third image.
[0105] Thus, this process makes it possible to determine a depth for each pixel of the first image.
[0106] Figure [Fig.2] illustrates a flowchart of the different steps of a method 2 for determining depth by a stereoscopic vision system mounted in a vehicle travelling on a traffic lane, for example the vehicle of [Fig.1].
[0107] The vision system comprises at least two cameras 11, 12, each arranged to acquire an image of a scene from a different viewpoint at the same time instant, according to a particular and non-limiting embodiment of the present invention.
[0108] The method 2 is for example implemented by one or more processors of one or more computers embedded in the vehicle 10, for example by a computer controlling the stereoscopic vision system.
[0109] In a step 21, the computer receives data representative of a first image acquired by the first camera 11 at an acquisition time instant and of a second image acquired by the second camera 12 at the same acquisition time instant.
[0110] The two images received correspond to two views of the same scene taking place around vehicle 10.
[0111] In a step 22, the first and second images are rectified according to extrinsic parameters of the stereoscopic vision system, so as to render horizontal epipolar lines in the first and second rectified images.
[0112] In a step 23, a third and a fourth image are generated from the first and second images respectively by decreasing the definition of the first and second images.
[0113] In a step 24, first depths associated with pixels of the third image are determined from first disparities determined from the third and fourth images.
[0114] In a step 25, a first set of pixels (410) representing the traffic lane and a first set of bounding boxes (411, 412, 413, 414) are determined in the first image (41) by processing the first image (41) and a second set of pixels (420) representing the traffic lane and a second set of bounding boxes (421, 422, 423, 424) are determined in the second image (42) by processing the second image (42), each bounding box of the first and second sets of bounding boxes being associated with an object of the three-dimensional scene.
[0115] In a step 26, a third set of bounding boxes comprising bounding boxes from the first set of bounding boxes, each associated with an object present on the traffic lane, and a fourth set of bounding boxes comprising bounding boxes from the second set of bounding boxes, each associated with an object present on the traffic lane, are determined.
[0116] In a step 27, bounding boxes from the third and fourth sets of bounding boxes are paired to form a set of bounding box pairs associated with the same object, each pair of bounding boxes comprising one bounding box from the third set of bounding boxes and one bounding box from the fourth set of bounding boxes, and the bounding boxes of each pair of bounding boxes are resized according to positions and respective sizes of the bounding boxes of the same pair of bounding boxes.
[0117] In a step 28, second depths associated with pixels of the third set of bounding boxes are determined from second disparities determined from the pixels of the bounding boxes of each pair of bounding boxes.
[0118] In a step 29, first depths are associated with the pixels of the first image not included in a bounding box of the third set of bounding boxes.
[0119] Each pixel of the first image (41) then has a depth, a first depth determined quickly and roughly for a pixel not included in a bounding box of the third set of bounding boxes, and a second depth determined more finely and more precisely for a pixel included in a bounding box of the third set of bounding boxes.
[0120] The set of depths associated with the pixels of the first image is thus determined more quickly than if it had been determined entirely from a first and second image with their initial definition.
[0121] The distance of an object present on the vehicle's traffic lane, determined by the depth of a pixel associated with that same object, is determined precisely while minimizing the time to determine the depth of a pixel of the first image.
[0122] An ADAS using input data such as depths determined by the stereoscopic vision system to determine the distance between a part of the vehicle, for example the front bumper, and another road user, then has accurate and rapidly updated data. The responsiveness of the ADAS is thus increased, thereby improving the operation of the ADAS.
[0123] Figure 3 schematically illustrates a device 3 configured for determining depth by a vision system mounted in a vehicle 10, according to a particular and non-limiting embodiment of the present invention. The device 3 corresponds, for example, to a device mounted in the first vehicle 10, for example a computer.
[0124] Device 3 is, for example, configured to carry out the operations described opposite Figures 1 and 4 and / or steps described opposite [Fig. 2]. Examples of such a device 3 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 3, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 3 may be implemented in the form of electronic circuits or software (or computer) modules or a combination of electronic circuits and software modules.
[0125] The device 3 comprises one (or more) processor(s) 30 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 3. The processor 30 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 3 further comprises at least one memory 31, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.
[0126] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 31.
[0127] According to various particular and non-limiting embodiments, the device 3 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.
[0128] According to a particular and non-limiting embodiment, the device 3 includes a block 32 of interface elements for communicating with external devices. The interface elements of the block 32 include one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HDMI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").
[0129] According to another particular and non-limiting embodiment, device 3 includes a communication interface 33 which allows communication with other devices (such as other computers in the embedded system) via a communication channel 330. The communication interface 33 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 330. The communication interface 33 corresponds, for example, to a wired network of the type CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3).
[0130] According to a particular and non-limiting embodiment, the device 3 can provide output signals to one or more external devices, such as a display screen 340, touch or not, one or more speakers 350 and / or other peripherals 360 (projection system) via the output interfaces 34, 35, 36 respectively. According to a variant, one or more of the external devices is integrated into the device 3.
[0131] Of course, the present invention is not limited to the embodiments described above but extends to a method for measuring distance using a stereoscopic vision system mounted in a vehicle, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0132] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising device 3 of [Fig.3].
Claims
1. Demands Method for determining depth by means of a stereoscopic vision system mounted in a vehicle (10) travelling on a traffic lane, the stereoscopic vision system comprising a first camera (11) and a second camera (12), each arranged so as to acquire an image of a three-dimensional scene from a different point of view with the same definition, said method being characterized in that it comprises the following steps: - reception (21) of data representative of a first (41) and second (42) images acquired by respectively the first (11) and second (12) cameras at the same time instant of acquisition; - rectification (22) of the first and second images according to extrinsic parameters of said stereoscopic vision system, so as to render horizontal epipolar lines in the first and second rectified images; - generation (23) of a third and a fourth image from respectively the first and second images by decreasing the definition of the first and second images; - determination (24) of first depths associated with pixels of the third image from first disparities determined from the third and fourth images; - determination (25) of a first set of pixels (410) representative of said traffic lane and of a first set of bounding boxes (411,412, 413,414) in the first image (41) by processing the first image (41) and of a second set of pixels (420) representative of said traffic lane and of a second set of bounding boxes (421, 422, 423, 424) in the second image (42) by processing the second image (42), each bounding box of said first and second sets of bounding boxes being associated with an object of the three-dimensional scene; - selection (26) of a third set of bounding boxes comprising bounding boxes from the first set of bounding boxes each associated with an object present on said traffic lane and of a fourth set of bounding boxes comprising bounding boxes from the second set of bounding boxes each associated with an object present on said traffic lane;
2. - pairing (27) of the bounding boxes of the third and fourth sets of bounding boxes to form a set of bounding box pairs associated with the same object, each pair of bounding boxes comprising a bounding box of said third set of bounding boxes and a bounding box of said fourth set of bounding boxes and resizing of the bounding boxes of each pair of bounding boxes according to the respective positions and sizes of the bounding boxes of the same pair of bounding boxes; - determination (28) of second depths associated with pixels of the third set of bounding boxes from second disparities determined from the pixels of the bounding boxes of each pair of bounding boxes; and - association (29) of the first depths to the pixels of the first image not included in a bounding box of the third set of bounding boxes. A method according to claim 1, the pixel coordinates constituting the first (41) and second (42) images being defined in a coordinate system having as its origin a top left corner of the first and second images respectively, each pair of bounding boxes comprising a first bounding box (411,414) in the first image and a second bounding box (421, 424) in the second image, said resizing comprising the following steps: - Determination of the coordinates of the upper left corner, the width and height of the first bounding box (411, 414) and the coordinates of the upper left corner, the width and height of the second bounding box (421, 424), - determination: • with a minimum abscissa equal to the minimum value among the abscissas of the upper left corners of the first and second bounding boxes, • with a minimum ordinate equal to the minimum value among the ordinates of the upper left corners of the first and second bounding boxes, • with a maximum width calculated from the widths of the first and second bounding boxes, • of a maximum height measured from the heights of the first and second encompassing boxes, and - assignment: • of the minimum abscissa and the minimum ordinate to said coordinates of the upper left corners of the first and second bounding boxes, and • of the maximum width and the maximum height to the first and second bounding boxes.
3. A method according to any one of claims 1 to 2, comprising cropping: - the first image (41) so as to obtain a first image (41) whose width and height are each divisible by 32, when the height and / or width of the first image (41) before cropping is not divisible by 32, - the second image (42) so as to obtain a second image (42) whose width and height are each divisible by 32, when the height and / or width of the second image (42) before cropping is not divisible by 32.
4. A method according to claim 3 depending on claim 2, wherein a height and a width of the first bounding box and of the second bounding box after said resizing are divisible by 32.
5. A method according to any one of claims 1 to 4, wherein a first depth and a second depth are determined respectively from the first and second disparities by the following formula: D(p) = fx B / d(p) With: • D(p) the first depth of a pixel p of the third image, respectively the second depth of a pixel p of the first image (41), • f a focal length of the first camera (11), • B a distance separating the first camera (11) from the second camera (12), and • d(p) the first disparity determined for a pixel p of the third image, respectively the second disparity determined for a pixel p of the first image (41).
6. A method according to any one of claims 1 to 5, wherein said first and second disparities are determined by implementing a convolutional neural network.
7. Computer program containing instructions for putting into
8.
9.
10. operation of the process according to any one of the preceding claims, when these instructions are executed by a processor. Computer-readable recording medium on which a computer program according to claim 7 is stored. Device (4) for determining depth by means of a stereoscopic vision system mounted in a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for carrying out the steps of the method according to any one of claims 1 to 6. Vehicle (10) comprising the device (4) according to claim 9.