Method and device for determining the distance between an object and a vehicle using a locally learned depth prediction model.

A locally learned depth prediction model for ADAS systems processes a portion of images from a stereoscopic vision system, addressing high-speed depth prediction challenges and enhancing ADAS reliability and safety.

FR3163759A1Pending Publication Date: 2025-12-26STELLANTIS AUTO SAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
FR2024006670
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing Advanced Driver-Assistance Systems (ADAS) in vehicles face challenges in accurately determining depth and reacting quickly due to high image processing times, especially at high speeds, which can compromise safety.

Method used

A method and device using a locally learned depth prediction model that processes only a portion of an image acquired by a stereoscopic vision system, comprising two cameras, to predict depth by training on thumbnail pairs of images, reducing processing time and improving accuracy, especially for highly distorted images.

Benefits of technology

The method significantly reduces processing time for depth prediction and enhances the reliability of ADAS systems by providing accurate depth information even in distorted images, thereby improving road safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method or device for determining the depth of a pixel in an image acquired by a camera in a two-camera vision system, based on a portion of the image containing the pixel. The depth is predicted by a depth prediction model learned in a training phase comprising the reception (31) of pairs of training images representative of at least one training object present in an observed three-dimensional scene, the detection (33) of a first set of pixels and a second set of pixels in the two images corresponding to a training object, the generation (36) of thumbnail pairs comprising pixels from both sets of pixels for each training object, and training (37) the depth prediction model from the thumbnail pairs. Figure 3 (for the abstract)
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for determining the distance between an object and a vehicle by means of a locally learned depth prediction model. technical field

[0001] The present invention relates to methods and devices for determining the distance separating an object from a vision system mounted in a vehicle, for example in a motor vehicle.

[0002] The present invention also relates to a method and device for predicting the depth associated with a pixel of an image acquired by a vision system embedded in a vehicle. Technological background

[0003] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). These ADAS are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS use advanced technologies to assist the driver while driving and thus improve performance. ADAS use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.

[0004] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.

[0005] The AD AS systems embedded in a vehicle are powered by data obtained one or more on-board sensors such as, for example, cameras. These cameras make it possible to detect and locate other road users or any obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle's speed; - to act on the braking system in case of risk of impact with an object.

[0006] The quality of the data emitted by a vision system therefore determines the proper functioning of the driving assistance devices using this data.

[0007] A vision system installed in a vehicle makes it possible to determine the depth or distance separating the vehicle from an object in the vehicle's environment. However, when a vehicle is moving at high speed, For example, on the highway, depth prediction time is critical. Indeed, the reaction time of an ADAS (Advanced Driver Assistance System) depends on image processing time and the depth prediction time of the predictive model associated with the vision system. This latter time must therefore be reduced to allow the vehicle, via the ADAS controls, to react quickly and effectively. Summary of the present invention

[0008] One object of the present invention is to solve at least one of the problems of the technological background described above.

[0009] Another object of the present invention is to improve the quality of data from the processing of an image acquired by a vision system, in particular by a depth prediction model implemented by a neural network associated with a stereoscopic vision system, which is learned locally in images so as to be accurate even for highly distorted images.

[0010] Another object of the present invention is to improve road safety, in particular by improving the reliability of ADAS systems powered by data obtained from a vision system.

[0011] According to a first aspect, the present invention relates to a method for determining the depth of a pixel of an image acquired by a first camera of a vision system embedded in a vehicle comprising the first camera and a second camera arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the method being implemented by at least one processor, and being characterized in that the depth is determined from a part of the image comprising the pixel by a depth prediction model learned in a learning phase comprising the following steps: - reception of representative data of pairs of training images, each pair of images comprising a first image and a second image acquired respectively by the first camera and the second camera at the same time instant of acquisition and representative of at least one training object present in an observed three-dimensional scene; - first training of the depth prediction model from the first and second images of each pair of training images; - for each at least one learning object represented in the first and second images of a pair of training images: • detection of a first set of pixels in the first image and a second set of pixels in the second image corresponding to at least one training object; • determination of a first bounding box including the first set of pixels and a second bounding box including the second set of pixels; • determining the height of a window from the heights of the first and second bounding boxes and determining the width of said window from the widths of the first and second bounding boxes; • generation of a first thumbnail comprising pixels of the first image contained within a rectangle whose height and width are equal to respectively the height and width of the window and whose center corresponds to the center of the first bounding box, and generation of a second thumbnail comprising pixels of the second image contained within a rectangle whose height and width are equal respectively to the height and width of the window and whose center corresponds to the center of the second bounding box; - a second training of the depth prediction model from a set of thumbnail pairs, each thumbnail pair in the set of thumbnail pairs comprising first and second thumbnails associated with at least a portion of a set of training objects comprising the training objects.

[0012] Such a method makes it possible to predict the depth of a pixel by processing only a portion of a received image acquired by a camera of the vehicle's onboard vision system. The processing time for a portion of an image is therefore considerably reduced compared to the time required to process the entire image, as the depth prediction model is used on only a limited number of pixels.

[0013] The first training session matures the depth prediction model, and the second training session improves accuracy when depth is predicted locally, i.e., over a portion of an image, and is particularly well-suited to images with high distortion. Indeed, the same object located at the same distance from the vehicle has a different appearance depending on the location of its corresponding pixels in a highly distorted image.

[0014] According to a variant of the method, the height of the window is between the height of the first encompassing box and the height of the second encompassing box, and the width of the window is between the width of the first encompassing box and the width of the second encompassing box.

[0015] According to another variant of the method, coordinates of the pixels of the first and second thumbnails are encoded in a reference frame associated with the first and second images.

[0016] According to a further variant of the method, the first training step of the depth prediction model is carried out in a self-supervised manner and further comprises the following steps: • generation of a third image from the first image and depths predicted by the depth prediction model for pixels of the first image, and of a fourth image from the second image and depths predicted by the depth prediction model for pixels of the second image, and • minimization of a first error determined by comparing the third image and the second image and by comparing the fourth image and the first image.

[0017] According to yet another variant, the second learning step of the depth prediction model is carried out in a self-supervised manner and further comprises the following steps: • generation of a third thumbnail from the first thumbnail and depths predicted by the depth prediction model for pixels of the first thumbnail, and of a fourth thumbnail from the second thumbnail and depths predicted by the depth prediction model for pixels of the second thumbnail, and • minimization of a second error determined by comparing the third thumbnail and the second thumbnail and by comparing the fourth thumbnail and the first thumbnail.

[0018] According to yet another variant of the method, when thumbnails from at least two sets of thumbnail pairs include the same pixels from a first or second image, then the set of thumbnail pairs does not include the at least two sets of thumbnail pairs.

[0019] According to a second aspect, the present invention relates to a device for determining the depth of a pixel of an image acquired by a camera of a vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the method according to the first aspect of the present invention.

[0020] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0021] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0022] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0023] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.

[0024] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.

[0025] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.

[0026] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures

[0027] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 5, in which:

[0028] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a a particular and non-limiting example of a realization of the present invention;

[0029] [Fig.2] illustrates a flowchart of the different steps of a process for determining the depth of a pixel of an image acquired by a camera of a vision system on board the vehicle of the [Fig.1] by a depth prediction model, according to a particular and non-limiting embodiment of the present invention;

[0030] [Fig.3] illustrates a flowchart of the different stages of a method for learning the depth prediction model used in the method illustrated in [Fig.2], according to a particular and non-limiting embodiment of the present invention;

[0031] [Fig.4] schematically illustrates bounding boxes and thumbnails generated for sets of pixels corresponding to an object in a three-dimensional scene observed by a camera of the vehicle's vision system in [Fig. 1], according to a particular and non-limiting embodiment of the present invention; and

[0032] [Fig.5] schematically illustrates a device configured to determine the depth of a pixel of an image acquired by a vision system onboard in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements

[0033] A method and device for determining the distance separating a common object from a vehicle carrying a vision system will now be described in what follows with joint reference to Figures 1 to 5. The same elements are identified with the same reference signs throughout the description that follows.

[0034] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0035] The present invention relates to a method and device for determining the depth of a pixel of an image acquired by a camera of a vision system comprising two cameras, the depth being predicted from a part of the image comprising the pixel.

[0036] Indeed, depth is predicted by a depth prediction model learned in a learning phase comprising the reception of pairs of training images representative of training objects present in three-dimensional scenes observed by the cameras of the on-board vision system, the detection of a first set of pixels and a second set of pixels in the two images of a pair of images and corresponding to a training object present in the observed three-dimensional scene, the generation of pairs of thumbnails comprising pixels from the two sets of pixels for each training object and a learning of the depth prediction model from the pairs of thumbnails.

[0037] Fig. 1 schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.

[0038] Such an environment 1 corresponds, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.

[0039] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, with electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a Land vehicle such as a car, truck, bus, motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say a vehicle operating according to a determined level of autonomy or under the total supervision of the driver.

[0040] The vehicle 10 advantageously comprises at least two on-board cameras, a first camera 11 and a second camera 12, configured to acquire images of a three-dimensional scene unfolding in the environment of the vehicle 10 from distinct observation positions. The first camera 11 and the second camera 12 form a stereoscopic vision system when used together, as illustrated in [Fig. 1]. The first camera 11 forms a monoscopic vision system when used alone, and similarly, the second camera 12 forms another monoscopic vision system when used alone. The present invention, however, extends to any vision system comprising at least two cameras, for example, 2, 3, or 5 cameras.

[0041] The intrinsic parameters of the first camera 11 characterize the transformation which associates, for an image point, hereafter called a "point", its three-dimensional coordinates in the reference frame of the first camera 11 with the pixel coordinates in an image acquired by the first camera 11. These parameters do not change if the first camera 11 is moved. The intrinsic parameters of the first camera 11 include in particular a first focal length fl associated with the first camera 11.

[0042] The intrinsic parameters of the second camera 12 characterize, for their part, the transformation which associates, for an image point, its three-dimensional coordinates in the reference frame of the second camera 12 with the pixel coordinates in an image acquired by the second camera 12. These parameters do not change if the second camera 12 is moved. The intrinsic parameters of the second camera 12 include in particular a second focal length f2 associated with the second camera 12.

[0043] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in the curvature and parallelism of the lenses and the coaxiality of the optical axes.

[0044] According to a particular embodiment, the first camera 11 and / or the second camera 12 is of the "wide-angle" type, a wide-angle camera being, for example, equipped with a lens designed to acquire an image representative of a three-dimensional scene perceived over a wider field of view than that of a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens allows you to capture a larger portion of the three-dimensional scene unfolding in front of or around the camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by that camera. The angle of view of the first camera 11 is, for example, equal to 120°, 145°, 180°, or 360°, whereas a standard camera offers, for example, a field of view open at an angle of 45° or less. Such a first camera 11 corresponds, for example, to a camera equipped with mirrors or to a "fisheye" camera.Wide-angle lenses have a shorter focal length compared to standard lenses, making them suitable for capturing images of landscapes, architecture, road intersections, or any other subject requiring a wide perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with an extended depth of field.

[0045] These two cameras 11, 12 are arranged so that each acquires an image of a scene from a different viewpoint. The first viewpoint is, for example, located on or in the left-hand rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. The second viewpoint is, for example, located on or in the right-hand rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. If both cameras are located at the top of the windshield of the vehicle, they are then positioned at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, and the second camera 12 is located in the right-hand rearview mirror of the vehicle 10.

[0046] A first marker is associated with the first camera 11: - the direction of the x-axis is defined as horizontal and normal to the optical axis of the first camera 11. The distance B separating the optical center of the first camera 11 from the projection of the optical center of the second camera 12 onto the horizontal plane passing through the optical center of the first camera 11 is called the reference basis (in English "baseline"); - the direction of the y-axis is defined as vertical and normal to the optical axis of the first camera 11; - The direction of the z-axis is defined as orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal coordinate system.

[0047] The extrinsic parameters related to the position of cameras 11, 12 are the following parameters: - three translations in the x, y, and z directions: Tx, Ty, and Tz, constituting the translation vector T; and - three rotations in the x, y and z directions: 0x, 0y and 0z.

[0048] An extrinsic matrix of the vision system then includes the extrinsic parameters previously defined.

[0049] The extrinsic parameters are determined, for example, during a calibration phase of the stereoscopic vision system comprising the first camera 11 and the second camera 12.

[0050] A key constraint of stereoscopic vision systems used in the automotive industry is, for example, the large distance between the two cameras. Indeed, to cover a measurement range of 200 meters, the reference base must be 60 cm for cameras commonly used in this field.

[0051] The two cameras 11, 12 acquire images of a scene located in front of the vehicle 10, the first camera 11 alone covering a first acquisition field 13, the second camera 12 alone covering a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic view of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic view of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic view of the scene by the stereoscopic vision system composed of the two cameras 11, 12.

[0052] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.

[0053] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system comprising the second camera 12.

[0054] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable using the monoscopic vision system comprising the first camera 11.

[0055] Finally, field 19 is not visible to any of the cameras. The part of the scene present in this field 19 is therefore not observable.

[0056] According to one particular embodiment, the field of view of the second camera 12 covers at least half of the field of view of the first camera 11.

[0057] It is evident that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently.

[0058] The images acquired by the cameras 11, 12 at a given acquisition time are in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric coordinates (from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).

[0059] According to a particular embodiment, an image acquired by the first camera 11 and / or by the second camera 12 includes a distortion equal to 0.5%, 0.8% or greater than 1%. The measurement of such distortion corresponds to the determination of a ratio between: - the maximum spacing of a pixel in the image of a straight line in the first three-dimensional scene whose image is a line touching the longest edge of the first image, either at the center of the image edge or at the corners of the image edge, and - the length of this edge.

[0060] In the world of photography, distortion is commonly considered to be: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • bothersome if it is greater than or equal to 1% or more.

[0061] Barrel distortion is characterized by a positive percentage, while crescent distortion is characterized by a negative percentage.

[0062] Each pixel of the acquired image represents an object in the three-dimensional scene present in the camera's field of view. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a point of light resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, which are captured by a photosensitive sensor of the camera after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image that faithfully represents the physical object observed by the camera. An image point described above is therefore a point on the surface of an object in the three-dimensional scene.

[0063] The images acquired by cameras 11, 12 represent views of the same scene taken from different viewpoints, the camera positions being distinct. This scene includes, for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.

[0064] According to a particular embodiment, a field of view of the first camera 11 covers at least half of a field of view of the second camera 12, and a field of view of the second camera 12 covers at least half of a field of view of the first camera 11. In other words, more than half of the pixels of an image acquired by the first camera 11 correspond to an object in the three-dimensional scene seen by the second camera 12, and pixels of an image acquired by the second camera 12 also correspond to this object in the three-dimensional scene. Similarly, more than half of the pixels of an image acquired by the second camera 12 correspond to an object in the three-dimensional scene seen by the first camera 11, and pixels of an image acquired by the first camera 11 also correspond to this object in the three-dimensional scene.

[0065] The images acquired by the first camera 11 and by the second camera 12 are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.

[0066] A method for determining a depth associated with a pixel of an image acquired by one of the first and second cameras and / or for determining a distance separating an object from a vehicle carrying the vision system is advantageously implemented by the vehicle 10, i.e. by a processor, a computer or a combination of computers of the vehicle 10's embedded system, for example by the computer or computers in charge of the vehicle 10's vision system.

[0067] Figure 2 illustrates a flowchart of the different steps of a method 2 for determining the depth of a pixel in an image acquired by the first camera of a vehicle-mounted vision system, for example, by the first camera 11 of the vehicle-mounted vision system 10 of Figure 1. The vehicle-mounted vision system 10 thus comprises a first and a second camera arranged so as to each acquire an image of a three-dimensional scene, i.e., a scene taking place around the vehicle 10. Such a three-dimensional scene includes, in particular, an object located in the field of vision of the first and second cameras, i.e., in the third acquisition field 15 shown opposite Figure 1. The method 2 is implemented, for example, by a device of the vehicle-mounted vision system 10 or by the device 5 of Figure 5.

[0068] According to a particular embodiment, process 2 comprises the following steps.

[0069] In a step 21, an image acquired by the first camera 11 and an image acquired by the second camera 12 are received. The reception of these images consists of receiving data representative of these images. These images are acquired respectively by the first camera 11 and by the second camera 12 at the same acquisition time. They therefore represent the same three-dimensional scene observed from two distinct viewpoints at the same time.

[0070] According to a particular embodiment, the received images are of the same definition, that is to say they have the same number of pixels, have the same number of pixels according to their height and the same number of pixels according to their width and a large part of each image represents the same part of the current three-dimensional scene.

[0071] According to another particular embodiment, the received images are not of the same resolution and / or only a small part of each of them represents the same part of the current three-dimensional scene. An additional step then consists of resizing or cropping them to obtain two images of the same resolution and whose pixels correspond mainly to the same part of the three-dimensional scene observed at the time of acquisition, for example more than 80% of the pixels of each image are representative of the same part of the current three-dimensional scene seen by the two cameras 11, 12.

[0072] If the cameras are calibrated, for example by following a calibration method known to those skilled in the art such as that presented in the document "Single View Point Omnidirectional Camera Calibration from Planar Grids", written by Christopher Mei and Patrick Rives and published in April 2007, then the intrinsic parameters of the first and second cameras 11,12 are known. Edges of the common field of view can be identified in the images received by a simulation via a back-projection and projection program, which are described below.

[0073] The rear projection program performs the following operations: • The pixel coordinates are corrected to remove distortion defined by distortion parameters: the parameter of the model presented in the cited document, called the Mei model, and distortion coefficients kh k2, k3, pi and p2, the definition of which can be found in the free library OpenCV® for example. • The coordinates are projected onto the unit sphere according to the Mei model. • Projection lines are calculated, always according to the Mei model. • Projection lines are multiplied by a depth.

[0074] The projection program then proceeds by performing the following operations: • The coordinates of the points in space are transformed into the coordinate system of the other image. • The coordinates of the points in space are modified via the distortion function defined by the same parameters as those mentioned previously. • Multiplication with the inverse of the intrinsic matrix similar to the pinhole camera lens.

[0075] By associating a depth corresponding to the midpoint of a measurement range, for example 100 m (one hundred meters) for a measurement range from 10 m (ten meters) to 200 m (two hundred meters), with the pixels of the first current image, and by backprojecting and projecting one of the received images onto the other received image with this depth, the simulation makes it possible to determine the average edge of the shared field of view on the other received image. The received image is to be cropped with the same resolution as the other received and cropped image, and on the opposite side along a diagonal to the cropped side of the other received image.

[0076] If the first and second cameras are not calibrated, the projection and reprojection models are defined by a neural network such as the one described in the paper "Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion" by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich, and Adrien Gaidon, published in August 2020. This particular embodiment is well-suited to images acquired by the vision system cameras and exhibiting significant distortion. It is not possible to determine the exact cropping before training. The common field of view in the received images can be found using a trial-and-error method, starting with cropping by observing current images and training the projection and reprojection models to improve their effectiveness.If the pixels at the edge of the received images do not have a reasonable prediction, for example with a value that is too high or too low, or when a depth map does not match the texture of an original image, the size of the cropped image must be reduced. If some of the pixels, preferably half, have reasonable predictions while the others do not, the correct cropping is achieved.

[0077] In a step 22, pixels corresponding to the object of the observed three-dimensional scene are detected in one of the received images.

[0078] In a step 23, depths associated with a set of pixels in the received image are predicted. The pixel set is, for example, contained within a rectangular window corresponding to a bounding box comprising all or part of the pixels corresponding to the observed object. The depths are notably predicted by a depth prediction model capable of predicting depths for the pixel set of the received image without requiring image processing. entire. Indeed, this depth prediction model is learned in a learning phase as described in relation to [Fig.3] so as to allow depth prediction from a part of a received image.

[0079] This depth prediction model is notably learned from training images comprising pixels corresponding to training objects similar to the current object and positioned in a region of an image close to the region in which the observed object is located.

[0080] The depths are then predicted by the depth prediction model from the set of pixels and pixels of the other received image corresponding to the pixels of the set of pixels, i.e. the pixels of the received image taken into consideration.

[0081] In an optional step 24, the distance separating the observed object from the vehicle 10 is determined from the depths associated with the pixels of the pixel set.

[0082] The distance separating the observed object from the vehicle 10 is, for example, equal to an average of the depths associated with the pixels in the pixel set, for example the average of the depths determined for all the pixels in the pixel set or the average of a portion of the pixels in the pixel set. The portion of pixels in the pixel set represents, for example, one-quarter of the pixels in the pixel set.

[0083] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another road user, the ADAS is then able to determine this distance precisely and act accordingly. For example, if the ADAS is designed to activate the braking system of the vehicle 10 in the event of a risk of collision with another road user, and the distance between the vehicle 10 and that road user decreases sharply, then the ADAS is able to detect this sudden closeness and activate the braking system of the vehicle 10 to avoid a possible accident.

[0084] Figure 3 illustrates a flowchart of the different stages of a method for learning the local depth prediction model used in a method for determining the depth of a pixel of an image and / or in a method for determining the distance separating a vehicle from an object, for example in method 2 of Figure 2, according to a particular and non-limiting embodiment of the present invention.

[0085] The learning process 3 is for example implemented by the device on board the vehicle 10 implementing the process of determining a distance separating an object from the vehicle 10 or by the device 5 of the [Fig.5].

[0086] In a step 31, pairs of training images are received, each pair of training images comprising a first image 41 acquired by the first camera 11 and a second image acquired by the second camera 12 at the same time instant of acquisition.

[0087] The pairs of images are for example acquired previously by the vision system of the vehicle 10 and stored in a memory of an embedded system of the vehicle 10 or in a memory accessible to an embedded system of the vehicle 10.

[0088] According to one particular embodiment, the images of the training image pairs were pre-filtered to be relevant for the execution of the learning process 3.

[0089] As with method 2, a resizing or cropping step of the images acquired by the first and second cameras is carried out, if necessary, so as to retain parts of the images corresponding to a part of the three-dimensional scene observed by both the first camera 11 and the second camera 12.

[0090] In step 32, the depth prediction model is learned from the first and second images of each pair of training images. This training constitutes a first training of the depth prediction model performed from whole images.

[0091] According to a first particular embodiment, the training of the depth prediction model is done by supervised machine learning as done by the model called Unimatch® presented in the document "Unifying Flow, Stereo and Depth Estimation" written by Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao and Andreas Geiger, published in July 2023. Annotated data, for example obtained from a LiDAR®, are used for example.

[0092] According to a second particular embodiment, the first learning step 32 of the depth prediction model is carried out in a self-supervised manner and further comprises the following steps: • generation of a third image from the first image 41 and depths predicted by the depth prediction model for pixels of the first image and generation of a fourth image from the second image 42 and depths predicted by the depth prediction model for pixels of the second image, and • minimization of a first error determined by comparing the third image and the second image and by comparing the fourth image and the first image.

[0093] According to this second particular embodiment, the depth prediction model is trained by self-supervision by changing the loss function. Self-supervision is based on image reconstruction. Image reconstruction is performed using the equation below:

[0094] [Math.l] And [Math.l]

[0095] With: • P] the coordinates of a pixel in the first image 41, • P2 the coordinates of a pixel in the second image 42, • 71 a function to convert from homogeneous coordinates to pixel coordinates in removing one dimension from a vector, • K, a projection model associated with the first camera 11, • a reprojection model associated with the first camera 11, • K' a projection model associated with the second camera 12, • a reprojection model associated with the second camera 12, • T and p1 of matrices including extrinsic parameters, • 0 a projection function in the three-dimensional scene of a pixel P\ in function of its coordinates in the first image, and of the depth who is associated, and • 0' a projection function in the three-dimensional scene of a pixel P 2 in function of its coordinates in the second image, and of the depth D^p^ who is associated.

[0096] Depending on the type of cameras, K and K' represent intrinsic matrices of the first and second cameras when these cameras have pinhole lenses.

[0097] Generating an image from an image acquired by a camera of the vision system and resized consists of reprojecting a pixel of the resized image into the three-dimensional scene as a point, and then projecting this point into the image plane of another camera of the vision system, so as to obtain an image corresponding to a view of the three-dimensional scene from the viewpoint of the other camera. The image plane of a camera corresponds to a plane defined in the camera's frame of reference, normal to the camera's optical axis and located at the camera's first focal length. Thus, the third image generated from the primary image is comparable to the secondary image. Similarly, the fourth image generated from the secondary image is comparable to the primary image. Since objects can obscure other objects in the scene, the generated images are not identical to acquired images.Furthermore, depth prediction, just like the models used to generate the images, are not. error-free. Thus, comparing a generated image to an acquired and resized image allows us to evaluate the relevance of the different models used.

[0098] The first error is for example determined from a main error associated with each first pixel by comparing pixels of the first training image and the fourth generated image and a secondary error associated with each second pixel by comparing pixels of the second training image and the third image.

[0099] According to a first particular embodiment, the primary and secondary errors are photometric errors as presented in the document "Digging Into Self-Supervised Monocular Depth Estimation" by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow published in August 2019 and are determined by the following function: [Math.2] L^P) = ' 1^)\+a\\-^SIM(l(p\ ï{p) ) ) ]

[0100] With: • L*(p) the principal photometric error denoted L^p), respectively the secondary photometric error denoted L2(p), p being a pixel defined by its coordinates in an image, • l(p) a value of pixelp in the first image 41, respectively second image 42, * l(p) a value of pixel p in the fourth image, respectively third image, • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of road environment in which the vehicle travels 10.

[0101] According to a second particular embodiment, the primary and secondary errors include a determination of a reconstruction error of a generated image determined, furthermore, by the following function:

[0102] [Math.3] IV o) = W( Pl )

[0103] With: • Lgmooth( D(pFF, o) is a main error for a pixel pt of the third image, respectively a secondary error for a pixel pt of the fourth image, • D (pf ) is a depth associated with a pixel pt, • W is a parameter matrix, • 0 is the order of a smoothing gradient, • an L1 norm of the second-order depth gradients is calculated with W = 1, et0 = 2, • x and y are the dimensions of the third image, and the fourth image respectively, • P is a hyperparameter dependent on the road environment in which vehicle 10 is traveling, and • is a value of the pixel Pt in the third image, respectively fourth image.

[0104] This second function is generally used to deal with discontinuity at the boundary of objects.

[0105] The first error is thus defined, for example, from the photometric errors and reconstruction errors previously defined.

[0106] According to a particular embodiment, the first error is determined by the following function: [Math.4] £ = ^1111(^( / 7),^( / 7) )

[0107] With: • The first mistake, • ^l(p) the principal error for a pixel p of the first image 41, and • L2(p) the secondary error for a pixelp of the second image 42 corresponding to the pixel P of the first image 41.

[0108] The learning of the depth prediction model then consists of minimizing this first error and adjusting the input parameters of the convolutional neural network used by the depth prediction model.

[0109] The following steps 33, 34, 35 and 36 are implemented successively for a learning object represented in the first and second images of the same pair of training images. They are repeated for each learning object represented in the first and second images of a pair of training images.

[0110] In a step 33, as illustrated in [Fig.4], for a training object represented in the first and second images of a pair of training images, a first set of pixels El is detected in the first image 41 and a second set of pixels E2 is detected in the second image 42, the pixels of the first and second sets of pixels El, E2 corresponding to the training object.

[0111] The first and second sets of pixels are likely to be of different shapes and sizes. Indeed, since the training object is not viewed from the same point of view, it does not necessarily have the same appearance in each of the images. Moreover, if the images exhibit strong distortion, the appearance of the object across the images differs markedly.

[0112] In a step 34, for the learning object represented in the first and second images of the pair of training images, a first bounding box B4i is determined in the first image 41, the first bounding box B4[ comprising the first set of pixels El, and a second bounding box B42 is determined in the second image 42, the second bounding box B42 comprising the second set of pixels E2.

[0113] A width L4i and a height H4i of the first bounding box B4[ are determined as well as a position of its center C4i whose coordinates in the first image 41 are defined by the pair (x4i, y4i).

[0114] Similarly, a width L42 and a height H42 of the second bounding box B42 are determined as well as a position of its center C42 whose coordinates in the second image 42 are defined by the pair (x42, y42).

[0115] Note that the first and second images 41, 42 have a defined width L and height H, the coordinates of the pixels of these images being expressed in a coordinate system having as its origin the upper left corner of these images.

[0116] In a step 35, for the learning object represented in the first and second images of the pair of training images, a height HF of a window is determined from the heights of the first and second bounding boxes Hb4i, Hb42 and a width LF of the window is determined from the widths of the first and second bounding boxes Lb4i, LB42.

[0117] According to a particular embodiment, the height HF of the window is between the height HB4 of the first bounding box B4 and the height HB42 of the second bounding box B42, and the width LF of the window is between the width Lb4i of the first bounding box B41 and the width LB42 of the second bounding box L42, for example the height HF of the window is the average of the heights Hb4i, Hb42 of the first and second bounding boxes B4H B42, and the width LF of the window is equal to the average of the widths Lb4i, LB42 of the first and second bounding boxes B4B B42

[0118] In a step 36, for the training object represented in the first and second images of the pair of training images, a first thumbnail Fl and a second thumbnail F2 are generated. The first thumbnail Fl comprises pixels from the first image 41 contained within a rectangle whose height and width are equal to the height HF and width LF of the window, respectively, and whose center corresponds to the center C4i of the first bounding box B4. The second thumbnail F2 comprises pixels from the second image 42 contained within a rectangle whose height and width are equal to respectively the height HF and the width LF of the window and whose center corresponds to the center C42 of the second bounding box B42.

[0119] According to a particular embodiment, coordinates of the pixels of the first and second thumbnails Fl, F2 are encoded in a reference frame associated with the first and second images 41, 42.

[0120] Position coding is performed by transforming the position into a combined function using a sinusoidal curve and a cosine curve, following the operations below: • In a first operation, the abscissa x in the horizontal direction and the ordinate y in the vertical direction in pixel coordinates (0,1,2,3,4,5,6...) are normalized in a range 0 to 2Æ. • In a second operation, a denominator that increases in pairs is calculated, the denominator thus describes the following series: 1, 1, 2, 2, 3, 3... with the magnitudes greater than 1 to decrease the magnitude of the coordinate value, and preferably with an exponential increase. • In a third operation, the values ​​of the abscissa x and ordinate y are divided by the denominator. • In a fourth operation, a sine function is applied to the digits in odd positions and a cosine function to the digits in even positions of the x-coordinate and y-coordinate. The result forms two matrices with the height and width of the characteristic map respectively for x and y. • In a fifth operation, the two position coding matrices are added to all channels of the two feature maps.

[0121] The adaptation of the position encoding proposed here makes it possible to retain the position encoding of the original image with the initial resolution for the image cropped by the previously defined window. The first window Fl is represented by the coordinates of the upper left corner, the height and the width (xFb yFb LF, HF), while the second window F2 is represented by the coordinates of the upper left corner, the height and the width (xF2, yF2, LF, HF).

[0122] The x-coordinate coding of a pixel belonging to the first window Fl is therefore included, in the x-coordinate coding matrix of the feature map of the model trained with the entire image, between the values:

[0123] [Math.5] r L Lp And [Math.5] xF}+LF L

[0124]

[0125]

[0126] With : • xfi the x-coordinate of the upper left corner of the first window Fl, • L the width of the first image 41, and • Lf the width of the feature map. Similarly, the y-coordinate encoding of a pixel belonging to the first window Fl is therefore included, in the y-coordinate encoding matrix of the feature map of the model trained with the entire image, between the values: [Math.6] H aF And [Math.6] yF.+HF Fl rj H

[0127] With: • y fi the ordinate of the upper left corner of the first window Fl, • H the height of the first image 41, and • HF the height of the characteristic map.

[0128] Similarly, the x-coordinate coding of a pixel belonging to the first window F2 is therefore included, in the x-coordinate coding matrix of the feature map of the model trained with the entire image, between the values:

[0129] [Math.7] XF2 j L And [Math.7] Xp^Fp “T •>

[0130] With: • xfi the x-coordinate of the upper left corner of the second window F2, • L the width of the second image 42, and • Lf the width of the feature map.

[0131] Similarly, the y-coordinate encoding of a pixel belonging to the second window F2 is therefore included, in the y-coordinate encoding matrix of the feature map of the model trained with the entire image, between the values:

[0132] [Math.8] And [Math.8] H •>

[0133] With: • ^2 the ordinate of the upper left corner of the second window F2, • H the height of the second image 42, and • HF the height of the characteristic map.

[0134] In a step 37, the depth prediction model is learned from a set of pairs of thumbnails, each pair of thumbnails in the set of thumbnail pairs comprising first and second thumbnails associated with at least a part of a set of training objects comprising the learning objects. In this second learning step 37, the depth prediction model is thus learned from pairs of two thumbnails, the thumbnails of the same pair of thumbnails having the same definition, this definition being less than or equal to that of the first and second images 41, 42.

[0135] To avoid any interference between several objects, thumbnails from at least two sets of thumbnail pairs containing the same pixels from a first or second image are not taken into account in this second learning process. Thus, the set of thumbnail pairs does not include the at least two sets of thumbnail pairs; that is, it does not include thumbnails having pixels in common.

[0136] On the other hand, if several objects are represented distinctly in the images of the same pair of images, then several thumbnails are generated from the same pair of images.

[0137] This second learning is carried out in a self-supervised manner, for example according to the method presented opposite the first learning step 32. The first and second images 41, 42 are then replaced by the first and second thumbnails Fl, F2. According to this particular embodiment, the second learning step 37 further comprises the following steps: • generation of a third thumbnail from the first thumbnail Fl and depths predicted by the depth prediction model for pixels of the first thumbnail, and generation of a fourth thumbnail from the second thumbnail F2 and depths predicted by the depth prediction model for pixels of the second thumbnail, and • minimization of a second error determined by comparing the third thumbnail and the second thumbnail and by comparing the fourth thumbnail and the first thumbnail.

[0138] Thus, the depth prediction model used to predict the depth of a pixel within a set of pixels defined around pixels corresponding to an object detected in an image acquired by the first camera 11 or the second camera 12 is made more reliable by this learning process. Furthermore, the data used for this learning are obtained from the stereoscopic vision system itself; they therefore correspond to data that are perfectly representative of the use of the vision system mounted on the vehicle 10.

[0139] Figure 5 schematically illustrates a device 5 configured for determining the depth of a pixel in an image acquired by a camera of a vision system embedded in a vehicle 10, for determining the distance separating an object from the vehicle 10, and / or for training a depth prediction model associated with a vision system embedded in the vehicle 10, according to a particular and non-limiting embodiment of the present invention. The device 5 corresponds, for example, to a device embedded in the first vehicle 10, for example, a computer associated with the stereoscopic vision system.

[0140] The device 5 is, for example, configured to carry out the operations described opposite Figures 1 and 4 and / or the steps described opposite Figures 2 and 3. Examples of such a device 5 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of the device 5, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. The device 5 may be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.

[0141] The device 5 comprises one (or more) processor(s) 50 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 5. The processor 50 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 5 further comprises at least one memory 51, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.

[0142] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 51.

[0143] According to various particular and non-limiting embodiments, the device 5 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.

[0144] According to a particular and non-limiting embodiment, the device 5 includes a block 52 of interface elements for communicating with external devices. The interface elements of the block 52 include one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HD MI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").

[0145] According to another particular and non-limiting embodiment, the device 5 includes a communication interface 53 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 530. The communication interface 53 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 530. The communication interface 53 corresponds, for example, to a wired network of the CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3) type.

[0146] According to a particular and non-limiting embodiment, the device 5 can provide output signals to one or more external devices, such as a display screen 540, touch or not, one or more speakers 550 and / or other peripherals 560 via output interfaces 54, 55, 56 respectively. In one variant, one or more of the external devices is integrated into device 5.

[0147] Of course, the present invention is not limited to the embodiments described above but extends to a method for measuring the distance between an object and a vehicle equipped with a vision system, which would include secondary steps without falling outside the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0148] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 5 of [Fig.5].

Claims

1. Demands Method for determining the depth of a pixel of an image acquired by a first camera (11) of a vision system embedded in a vehicle (10) comprising the first camera (11) and a second camera (12) arranged so as to each acquire an image of a three-dimensional scene from a different point of view, said method being implemented by at least one processor, and being characterized in that said depth is determined from a part of said image comprising said pixel by a depth prediction model learned in a learning phase comprising the following steps: - reception (31) of representative data of pairs of training images, each pair of images comprising a first image (41) and a second image (42) acquired respectively by the first camera (11) and the second camera (12) at the same time instant of acquisition and representative of at least one training object present in an observed three-dimensional scene; - first training (32) of the depth prediction model from the first and second images of each pair of training images; - for each at least one learning object represented in the first and second images of a pair of training images: • detection (33) of a first set of pixels (E1) in the first image (41) and a second set of pixels (E2) in the second image (42) corresponding to said at least one learning object; • determination (34) of a first bounding box (B4i) comprising the first set of pixels (El) and of a second bounding box (B42) comprising the second set of pixels (E2); • determination (35) of a height of a window from the heights of the first and second bounding boxes and determination of a width of said window from the widths of the first and second bounding boxes; • generation (36) of a first thumbnail (Fl) comprising pixels of the first image (41) contained within a rectangle of which a height and width are equal to respectively the height and width of the window and whose center corresponds to the center (C4i) of the first bounding box, and generation of a second thumbnail (F2) comprising pixels of the second image (42) contained in a rectangle whose height and width are equal to respectively the height and width of the window and whose center corresponds to the center (C42) of the second bounding box; - second learning (37) of the depth prediction model from a set of thumbnail pairs, each thumbnail pair of said set of thumbnail pairs comprising first and second thumbnails associated with at least a part of a set of training objects comprising the training objects.

2. A method according to claim 1, wherein the height (HF) of the window is between the height (HB4i) of the first encompassing box and the height (Hb42) of the second encompassing box, and the width (LF) of the window is between the width (Lb4i) of the first encompassing box and the width (Lb42) of the second encompassing box.

3. A method according to claim 1 or 2, wherein coordinates of the pixels of the first and second thumbnails (F1, F2) are encoded in a reference frame associated with the first and second images (41, 42).

4. A method according to any one of claims 1 to 3, wherein the first learning step (32) of the depth prediction model is carried out in a self-supervised manner and further comprises the following steps: • generation of a third image from the first image (41) and depths predicted by the depth prediction model for pixels of the first image and of a fourth image from the second image (42) and depths predicted by the depth prediction model for pixels of the second image, and • minimization of a first error determined by comparison of the third image and the second image and by comparison of the fourth image and the first image.

5. A method according to any one of claims 1 to 4, wherein the second learning step (37) of the depth prediction model is carried out in a self-supervised manner and further comprises the following steps: • generation of a third thumbnail from the first thumbnail (F1) and depths predicted by the depth prediction model for pixels of the first thumbnail and of a fourth thumbnail from the second thumbnail (F2) and depths predicted by the depth prediction model for pixels of the second thumbnail, and • minimization of a second error determined by comparison of the third thumbnail and the second thumbnail and by comparison of the fourth thumbnail and the first thumbnail.

6. A method according to any one of claims 1 to 5, wherein when thumbnails of at least two sets of thumbnail pairs include the same pixels from a first or second image, then the set of thumbnail pairs does not include said at least two sets of thumbnail pairs.

7. Computer program comprising instructions for carrying out the method according to any one of claims 1 to 6, when such instructions are executed by at least one processor.

8. A computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to any one of claims 1 to 6

9. 1 d O. Device (5) for determining the depth of a pixel of an image acquired by a camera of a vision system on board a vehicle, said device (5) comprising a memory (51) associated with at least one processor (50) configured for carrying out the steps of the method according to any one of claims 1 to 6.

10. Vehicle (10) comprising the device (5) according to claim 9.

Citation Information

Patent Citations

  • Self-supervised training of a depth estimation model using depth hints

    US20200351489A1