Method and device for learning motion prediction and depth prediction models.

The method and device enhance ADAS systems by using convolutional neural networks to learn depth and motion prediction models, addressing the challenge of dynamic objects in monocular vision systems, resulting in improved accuracy and safety.

FR3168050A1Pending Publication Date: 2026-05-01STELLANTIS AUTO SAS
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
STELLANTIS AUTO SAS
Filing Date
2024-10-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Monocular vision systems in vehicles struggle to accurately predict the depth and motion of dynamic objects due to pixel displacement caused by both the movement of the vehicle and the objects within the three-dimensional scene, leading to biased depth predictions that affect the reliability of Advanced Driver-Assistance Systems (ADAS).

Method used

A method and device using convolutional neural networks to learn depth and motion prediction models by generating error-free images through a training process that minimizes loss errors, distinguishing between static and dynamic objects, and incorporating camera movement, enabling accurate depth and motion prediction even in scenes with moving objects.

Benefits of technology

Improves the reliability of ADAS systems by enhancing the accuracy of depth and motion prediction models, ensuring precise distance measurements and vehicle motion estimation, thereby improving road safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method or device for learning depth prediction models associated with a vision system. Specifically, the method comprises receiving images (21) acquired by a camera, generating (23) first and second three-dimensional point clouds from the received images and depths predicted (22) by the depth prediction model, and generating point clouds from the received images based on vehicle motion predicted by the motion prediction model and predicted depths. Points are identified (25) as associated with dynamic objects detected by a segmentation model, and images are generated (26) from points based on their correspondence to a dynamic object. The depth and motion prediction models are then learned (27) by comparing the generated images with the received images. Figure for the abstract: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for learning motion prediction and depth prediction models. technical field

[0001] The present invention relates to methods and devices for learning a depth prediction model associated with a vision system embedded in a vehicle, for example, in a motor vehicle. The present invention also relates to a method and device for determining depth and / or measuring the distance separating a static or dynamic object from a vehicle equipped with a vision system. Technological background

[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS systems use advanced technologies to assist the driver while driving and thus improve performance. ADAS systems use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.

[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.

[0004] The ADAS systems installed in a vehicle are powered by data obtained from one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or any obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle's speed; - to act on the braking system in case of risk of impact with an object.

[0005] The position of another user or an obstacle in a three-dimensional scene is, for example, determined by a vision system comprising a prediction model for a depth associated with a pixel or a distance separating the vision system from an object in a three-dimensional scene. However, the accuracy of a predicted depth or distance for a dynamic object, that is, an object in movement in the three-dimensional scene is generally not satisfactory for a vision system comprising only a single camera, also called a monocular vision system.

[0006] Indeed, a monocular vision system uses several consecutive images to predict depths associated with pixels in one of the images. This monocular vision system has the drawback of not being able to accurately predict depths associated with pixels in an image when objects in the observed three-dimensional scene are moving within that same scene, because the depth is obtained by comparing the positions of pixels corresponding to the same object in the three-dimensional scene in different images. However, if the observed object is moving, then the pixels associated with that object are displaced between two images acquired at different times, and this displacement is due both to the movement of the monocular vision system mounted in the vehicle and to the inherent movement of the dynamic object within the three-dimensional scene.A depth prediction model is therefore unable to accurately determine the position of this dynamic object.

[0007] Furthermore, a monocular vision system also predicts the camera's own motion between image acquisition times, i.e., the motion of the vehicle carrying the vision system's camera. Image processing then requires accurate prediction of this motion to avoid biasing, for example, the depth prediction described above.

[0008] The quality of the training of the depth or distance prediction model is, however, very important. Indeed, the predicted depths or distances, as well as the movement predicted by the model, make it possible to determine, for example, the distances to other road users or obstacles present in the road environment of the vehicle equipped with the vision system, whose output data feeds ADAS (Advanced Driver Assistance Systems). The proper functioning of the driver assistance devices using this data therefore depends on the quality of the data emitted by the vision system. Summary of the present invention

[0009] One object of the present invention is to solve at least one of the problems of the technological background described above.

[0010] Another object of the present invention is to improve the learning phase of a depth prediction model and a motion prediction model from images acquired by a vision system and representing objects in a three-dimensional scene moving within that same scene, the depth prediction and motion prediction models being in particular put into operated by neural networks and associated with a vision system embedded in a vehicle.

[0011] Another object of the present invention is to improve road safety, in particular by improving the reliability of AD AS systems powered by data obtained from a camera of a vision system.

[0012] According to a first aspect, the present invention relates to a method for learning depth prediction and motion prediction models implemented by convolutional neural networks associated with a vision system embedded in a vehicle, the vision system comprising a camera arranged to acquire an image of a three-dimensional scene taking place near the vehicle, the method being implemented by at least one processor, and comprising the following steps: - reception of representative data from a first image and a second image acquired by the camera at two distinct acquisition times; - prediction • first depths associated with pixels of the first image, called first pixels, and second depths associated with pixels of the second image, called second pixels, the first and second depths being predicted by the depth prediction model, and • of a camera movement between the two time instants of acquisition by the motion prediction model; - generation : • initial three-dimensional points from the first image, the spatial position of a first point being a function of the coordinates of a first pixel in the first image associated with the first point and the first depth associated with the first pixel, • of second three-dimensional points from the second image, a spatial position of a second point, being a function of the coordinates of a second pixel in the second image associated with the second point and of the second depth associated with the second pixel; • of three-dimensional third points from the first image, the spatial position of a third point being a function of the coordinates of a first pixel in the first image associated with the third point, the first depth associated with the first pixel, and the camera movement between the two temporal instants of acquisition, and • of fourth three-dimensional points from the second image, the spatial position of a fourth point being a function of the coordinates of a second pixel in the second image associated with the fourth point, of the second associated depth to the second pixel and the camera movement between the two temporal moments of acquisition; - determination of a scene flow for each first and second point associated with the same object in the three-dimensional scene as a function of a difference in position between the first and second points; - determination of a first set of points comprising first points corresponding to pixels of the first image associated with at least one dynamic object and of a second set of points comprising second points corresponding to pixels of the second image associated with at least one dynamic object, dynamic objects being detected by a segmentation model in the first and second images; - generation : • of a third image, pixels of the third image being defined by projection of second points associated with the same object of the three-dimensional scene as a first point included in the first set of points, and of third points associated with the same object of the three-dimensional scene as a first point not included in the first set of points, and • of a fourth image, the pixels of the fourth image being defined by projection of first points associated with the same object of the three-dimensional scene as a second point included in the second set of points and of fourth points associated with the same object of the three-dimensional scene as a second point not included in the second set of points, - learning depth prediction and motion prediction models by minimizing a loss error determined from: • the result of a comparison of the third image to the second image and the result of a comparison of the fourth image to the first image, and • a consistency error representative of an error in predicting camera movement between the two time moments of acquisition.

[0013] The training phase thus allows the depth prediction model to be trained simultaneously with the motion prediction model using error-free generated images that are free from the presence of dynamic objects in the three-dimensional scene. The presence of dynamic objects in the three-dimensional scene observed during the training phase therefore does not disrupt this learning process, which can be carried out using any pair of images containing pixels corresponding to dynamic objects.

[0014] According to a variant of the method, a spatial position of a first point, respectively second point, is determined by the following function: P3d=^^\dm) With : • P^D the spatial position of a first point, respectively second point, • K is a camera-associated direction prediction model, and • 0 a projection function in a three-dimensional scene of a pixelp as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, Dint which is associated with it.

[0015] According to another variant of the method, a spatial position of a third point, respectively fourth point, is determined by the following function: With : • P'^d the spatial position of a third point, respectively fourth point, • K is a camera-associated direction prediction model, • 0 a projection function in a three-dimensional scene of a pixel p as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, Dml associated with it, and • T a movement of the vehicle between the two time instants.

[0016] According to yet another variant of the process, a position of a pixel of the third image, respectively fourth image, is determined by the following function: Pm=^KP\D) With : • P ms the position of a pixel in the third image, respectively fourth image, • A function to convert from homogeneous coordinates to pixel coordinates by removing one dimension from a vector, • K is a camera-associated direction prediction model, and • P " the spatial position of a second or third point, respectively first or fourth point.

[0017] According to a further embodiment of the method, the loss error is determined by the following function: L' = +Lcons With : • The loss error, • L^p) a result of a comparison of a pixel p from the first image to a pixel p from the fourth image, • L2(p) a result of a comparison of a pixel p of the second image to a pixel p of the third image, and • lessons-consistency error.

[0018] According to yet another variant of the method, the consistency error is determined by the following function: PcOUS — 5Z P ) ( p ) ] With : • LcmK. The consistency error, • p GV a pixelp of the first image that is not associated with a dynamic object, • S^ip) a scene stream associated with pixel p, and • Tt^s(p) a displacement between the two time instants associated with pixelp.

[0019] According to another variant of the method, the segmentation model determines bounding boxes associated with dynamic objects in the first and second images.

[0020] According to a second aspect, the present invention relates to a device configured to learn a depth prediction model and a motion prediction model associated with a vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.

[0021] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0022] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0023] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0024] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.

[0025] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium may include a storage means, such as a ROM, a CD-ROM, or a type ROM microelectronic circuit, or even a magnetic recording device or a hard drive.

[0026] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.

[0027] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures

[0028] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 4, in which:

[0029] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a a particular and non-limiting example of a realization of the present invention;

[0030] [Fig.2] illustrates a flowchart of the different stages of a method for learning a depth prediction model associated with the vision system of [Fig.1], according to a particular and non-limiting embodiment of the present invention; and

[0031] [Fig.3] schematically illustrates a device configured to learn a depth prediction model and a motion prediction model associated with the vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention; and

[0032] [Fig.4] schematically illustrates images and point clouds used during the learning process of [Fig.2], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements

[0033] A method and device for learning a depth prediction model and a motion prediction model implemented by convolutional neural networks associated with a vision system embedded in a vehicle will now be described in what follows with joint reference to Figures 1 to 4. The same elements are identified with the same reference signs throughout the description that follows.

[0034] The terms "first(s)", "second(s)" (or "first(s)", "second(s)"), etc. are used in this document by arbitrary convention to allow identification and distinction of different elements (such as operations, means, etc.) put into work in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0035] Fig. 1 schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.

[0036] The vehicle 10 is located in an environment 1 corresponding, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.

[0037] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.

[0038] The vehicle 10 advantageously comprises at least one onboard camera 11, configured to acquire images of a three-dimensional scene unfolding in the environment of the vehicle 10 from a common viewing position. The camera 11 forms a monocular vision system if used alone, as illustrated in [Fig. 1]. However, the present invention is not limited to a monocular vision system comprising a single camera but extends to any vision system comprising at least one camera, for example, 1, 2, 3, or 5 cameras.

[0039] The camera 11 has intrinsic parameters, including: - a focal length, - distortions which are due to imperfections in the optical system of camera 11, - a direction of the optical axis of camera 11, and - a resolution.

[0040] The intrinsic parameters characterize the transformation which associates, for an image point, hereafter called "point", its three-dimensional coordinates in the camera 11 reference frame with the pixel coordinates in an image acquired by the camera 11. These parameters do not change if the camera is moved.

[0041] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in curvature, lens parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, that is to say, the distortions are not not taken into account, whether their correction is processed at the time of image acquisition or at the time of calibration.

[0042] The camera 11 is arranged to acquire an image of a three-dimensional scene from a defined viewpoint, the viewpoint is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10 as illustrated in [Fig.1].

[0043] The camera 11, for example, acquires images of a three-dimensional scene located in front of the vehicle 10, the camera 11 covering an acquisition field 12. An object 13 is placed in the acquisition field 12 of the camera 11, defining an occlusion field 14 for the vision system, an object present in the occlusion field 14 not being observable by the camera 11 from its current observation position.

[0044] It is evident that it is possible to use such a vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently, the invention not being limited to the observation of a three-dimensional scene taking place in front of the vehicle 10 carrying the camera 11.

[0045] According to one particular embodiment, the camera 11 is a wide-angle camera, a wide-angle camera being, for example, equipped with a lens designed to acquire a representative image of a three-dimensional scene seen over a wider field of view than a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens makes it possible to capture a larger portion of the three-dimensional scene unfolding in front of or around the camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by this camera. The angle α of the field of view of the camera 11 is, for example, equal to 120°, 145°, 180°, or 360°, whereas a standard camera offers, for example, a field of view open at an angle of 45° or less.Such a camera 11 corresponds, for example, to a camera equipped with mirrors or a "fisheye" camera. Wide-angle lenses have a shorter focal length compared to standard lenses, making them suitable for capturing images of landscapes, architecture, road intersections, or any other subject requiring a wide perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with an extended depth of field.

[0046] An image acquired by the camera 11 at a given acquisition time is in the form of data representing pixels characterized by: - ​​coordinates in the image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric values ​​(from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).

[0047] Each pixel of the acquired image represents an object in the three-dimensional scene present in the field of view of the camera 11. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a point of light resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, which are captured by a photosensitive sensor of the camera 11 after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image that faithfully represents the physical object observed by the camera 11. An image point described above is therefore a point on the surface of an object in the three-dimensional scene.

[0048] When the vehicle 10 is in motion, then two images acquired by the camera 11 at two distinct time instants represent views of the same three-dimensional scene taken from different viewpoints or observation positions, the observation positions of the camera 11 being distinct. On this three-dimensional scene there are, for example: - buildings; - road infrastructure; - other users or stationary objects, for example a parked vehicle; and / or - other users or moving objects, for example another vehicle, a cyclist or a moving pedestrian.

[0049] According to a particular embodiment, an image acquired by the camera 11 includes a distortion equal to 0.5%, 0.8%, or greater than 1%. The measurement of such distortion corresponds to determining a ratio between: - the maximum spacing of a pixel in the image of a straight line in the first three-dimensional scene whose image is a line touching the longest edge of the first image, either at the center of the image edge or at the corners of the image edge, and - the length of this edge.

[0050] In the world of photography, distortion is commonly considered to be: • negligible if it is less than 0.3%, • slightly sensitive if it is between 0.3% and 0.4%, • significant if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • bothersome if it is greater than or equal to 1% or more.

[0051] Barrel distortion is characterized by a positive percentage, while crescent distortion is characterized by a negative percentage.

[0052] Each image acquired by the camera 11 is for example sent to a processor, for example a computer of a device equipping the vehicle 10, or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10. It is then used during the implementation of a method for determining the depth of one of its pixels and / or during the implementation of a method for learning a depth prediction model associated with this camera of the vision system.

[0053] Figure 2 illustrates a flowchart of the different steps in a method for learning a depth prediction model and a motion prediction model. Such depth prediction and motion prediction models are, for example, associated with the vision system embedded in the vehicle 10 and comprising the camera 11, and are also used in methods for determining the depth of a pixel in an image and / or for determining camera motion between two time instants of image acquisition, according to a particular and non-limiting embodiment of the present invention.

[0054] The learning process 2 is for example implemented by a processor of a device embedded in the vehicle 10 or of the device 3 of the [Fig.3].

[0055] In a first step 21, representative data of a first image and a second image acquired by the camera 11 at two distinct acquisition times are received; that is, a first image acquired by the camera 11 at a first acquisition time and a second image acquired by the camera 11 at a second acquisition time distinct from the first acquisition time are received. Image reception consists of receiving representative image data, which is, for example, recorded in a device embedded in the vehicle 10 and comprising a memory accessible by a computer implementing the learning process 2. The first and second acquisition times are sufficiently close in time so that the observed three-dimensional scene is the same but observed from two different viewpoints when the vehicle 10 is in motion.The first and second images are acquired at time intervals on the order of one or more milliseconds.

[0056] In a second step 22, first depths associated with pixels of the first image, called first pixels, are predicted by the depth prediction model and, similarly, second depths associated with pixels of the second image, called second pixels, are predicted by the depth prediction model.

[0057] Such a depth prediction model is known to those skilled in the art, and is described for example in the following documents: - “Digging Into Self-Supervised Monocular Depth Estimation” by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow published in August 2019, and - “HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation” by Xiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong, Lina Liu, Yong Liu, Xinxin Chen and Yi Yuan published in December 2020.

[0058] It should be noted that the first and second depths become increasingly accurate for static objects in the observed three-dimensional scene; that is, these depths become more and more accurate as the learning process is repeated with new images. However, the first and second depths predicted for dynamic objects in the observed three-dimensional scene are less reliable. It is therefore necessary to consider the nature of the object corresponding to a first or second pixel to determine the accuracy of the predicted first or second depth, which is addressed in the following steps.

[0059] In the second step 22, a movement of the camera 11 between the two acquisition times is also predicted by the motion prediction model from the first and second images. Such a motion prediction model is known to those skilled in the art, and is described, for example, in the document "HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation" cited above.

[0060] It should be noted that, as with the depth prediction model, the motion predicted by the motion prediction model becomes more and more accurate when the learning process is repeated with new images.

[0061] In a third step 23, four point clouds are generated, these clouds comprising three-dimensional points positioned in the observed scene: • a first point cloud containing first points, • a second point cloud containing second points, • a third point cloud containing third points, and • a fourth point cloud containing fourth points.

[0062] The first points are generated from the first image, a spatial position of a first point being a function of the coordinates of a first pixel in the first image associated with this first point and of the first depth associated with the first pixel, and the second points are generated similarly from the second image, a spatial position of a second point being a function of the coordinates of a second pixel in the second image associated with this second point and of the second depth associated with the second pixel.

[0063] As illustrated in [Fig.4], the generation of the first point cloud consists of the projection of the pixels of the first image 41 into the three-dimensional scene and into the camera frame 11 at the first time instant of acquisition.

[0064] Similarly, the generation of the second point cloud consists of the projections of the pixels of the second image 42 into the three-dimensional scene and into the frame of the camera 11 at the second time instant of acquisition, which is different from the frame of the camera 11 at the first time instant of acquisition when the vehicle 10 is in motion, the camera 11 having moved in the three-dimensional scene between the first and second time instants of acquisition.

[0065] According to a particular embodiment, a spatial position of a first point, respectively second point, is determined by the following function: [Math.l] p}B=

[0066] With: • P^d the spatial position of a first point, respectively second point, • K is a direction prediction model associated with camera 11, and • 0 a projection function in a three-dimensional scene of a pixel p as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, Dmt which is associated with it.

[0067] According to a first variant, the direction prediction model corresponds to an intrinsic matrix of the camera 11. This first variant is particularly applicable to a pinhole camera and calibrated.

[0068] According to a second variant, the direction prediction model corresponds to a learned model implementing a different neural network. Such a direction prediction model is known to those skilled in the art; it is notably presented in the paper "Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion" written by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich, and Adrien Gaidon and published in August 2020. This second variant is suitable for wide-angle or fisheye cameras, or even for uncalibrated cameras. According to this second variant, a direction prediction step associated with the pixels of the first image and the pixels of the second image is necessary.

[0069] Thus, as illustrated in [Fig.4], a first point pl and a second point p2 are respectively generated from the first image 41 and the second image 42. The first and second points are notably represented in the form of black dots.

[0070] The third points are, for their part, generated from the first image, a spatial position of a third point being a function of the coordinates of a first pixel in the first image associated with the third point, of the first depth associated with the first pixel and of the movement of the vehicle 10 between the two time instants of acquisition, and the fourth points are generated from the second image, a spatial position of a fourth point being a function of the coordinates of a second pixel in the second image associated with the fourth point, of the second depth associated with the second pixel and of the movement of the vehicle 10 between the two time instants of acquisition.

[0071] As illustrated in [Fig. 4], the generation of the third point cloud consists of projecting the pixels of the first image 41 into the three-dimensional scene and into the frame of the camera 11 at the second acquisition time instant. Alternatively, considering the first point cloud obtained from the first image 41, the generation of the third point cloud consists of projecting the first points of the first point cloud into the frame of the camera 11 at the second acquisition time instant, taking into account the movement of the vehicle 10 between the two acquisition times. The third point p3 is then the image of a pixel from the first image 41 or of the first point pl, according to the definition above.

[0072] Similarly, the generation of the fourth point cloud consists of projecting the pixels of the second image 42 into the three-dimensional scene and into the frame of the camera 11 at the first time instant of acquisition. Or, considering the second point cloud obtained from the second image 41, the generation of the fourth point cloud consists of projecting the second points of the second point cloud into the frame of the camera 11 at the first time instant of acquisition, taking into account the movement of the vehicle 10 between the two time instants of acquisition. The fourth point p4 is then the image of a pixel of the second image 42 or of the second point p2, according to the definition above.

[0073] According to a particular embodiment, a spatial position of a third point, respectively a fourth point, is determined by the following function: [Math.2] p3 d =t^\d m )

[0074] With: • P'zi) the spatial position of a third point, respectively fourth point, • K a direction prediction model associated with camera 11, • 0 a projection function in a three-dimensional scene of a pixelp as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, Dint which is associated with it, and • T a movement of the vehicle 10 between the two time instants.

[0075] It is noted here that part of this function corresponds to the previous function determining the spatial position of a first or second point.

[0076] Note that the movement of the vehicle between the two time instants does not have the same value when determining the position of a third point and a fourth point. Indeed, the movement considered for determining the position of a fourth pixel is opposite to the movement considered for determining the position of a third pixel.

[0077] The movement is notably presented in the form of a matrix comprising extrinsic parameters of the moving camera 11, which is determined during the second step 22.

[0078] Thus, as illustrated in [Fig.4], a third point p3 and a fourth point p4 are respectively generated from the first image 41 and the second image 42. The third and fourth points are notably represented as white dots.

[0079] In a fourth step 24, a scene flow is determined for each first point and for each second point.

[0080] A scene flow is determined for each first and second point associated with the same object in the three-dimensional scene based on a difference in position between the first and second points. A model for determining scene flow is known to those skilled in the art and is presented in particular in the document "DeFlow: Decoder of Scene Flow Network in Autonomous Driving" written by Qingwen Zhang, Yi Yang, Heng Fang, Ruoyu Geng and Patrie Jensfelt and published in January 2024.

[0081] In [Fig.4], a scene flow F is determined for the first point pl and a scene flow opposite to the scene flow F is determined for the second point p2.

[0082] The scene flow allows us to determine, for a first point, the corresponding second point, that is, the second point corresponding to the same object in the three-dimensional scene. Similarly, the scene flow allows us to determine, for a second point, the corresponding first point, that is, the second point corresponding to the same object in the three-dimensional scene. Note that the pixel in the first image 41 that yields the first point corresponds to the pixel in the second image 42 that yields the second point, the determination of corresponding pixels in the first and second images being known to those skilled in the art.

[0083] Thus, each first point corresponds to: • a second point positioned relative to the first point according to a scene flow, and • a third point, the second and third points corresponding to the first point thus correspond to the same object of the three-dimensional scene.

[0084] In [Fig. 4], it is thus distinguished that the first point pl corresponds to: • the second point p2 positioned relative to the first point pl according to the scene flow F, and • the third point p3 positioned relative to the first point pl according to a displacement T (not determined).

[0085] Similarly, each second point corresponds to: • a first point positioned relative to the second point according to a scene flow, and • a fourth point positioned.

[0086] In a fifth step 25, a first set of points comprising first points corresponding to pixels of the first image associated with at least one dynamic object and a second set of points comprising second points corresponding to pixels of the second image associated with at least one dynamic object are determined. The dynamic objects are detected in particular by a segmentation model in the first and second images, such a segmentation model being known to those skilled in the art and presented, for example, in the paper "Masked Autoencoders for Point Cloud Self-supervised Learning" written by Yatian Pang, Wenxiao Wang, Francis EH Tay, Wei Liu, Yonghong Tian and Li Yuan and published in March 2022.

[0087] According to a particular embodiment, the dynamic objects belong to a set of dynamic objects comprising: - automobiles, - trucks, - coaches, - motorcycles, - bicycles, and - pedestrians.

[0088] Such objects are sometimes in motion within the observed three-dimensional scene, and their presence then disrupts the learning of depth prediction and motion prediction models. Indeed, their own motion within the three-dimensional scene is difficult to determine by a vision system comprising only a single camera, as their own motion within the three-dimensional scene combines with the camera's motion within the three-dimensional scene when the vehicle carrying it is moving. In order not to bias the learning of depth prediction and motion prediction models movement, it is then necessary to identify the pixels of the images and therefore the points of the previously generated point clouds corresponding to such dynamic objects.

[0089] Detecting and taking into account dynamic objects then amounts to generating a dynamic object mask or, conversely, a static object mask. The static object mask thus includes the points in the first and second point clouds corresponding to dynamic objects, that is, points determined by projecting the pixels of the first and second images included in the dynamic object mask applied to the first and second images, these points forming the first and second point sets. Conversely, the static object mask, also called the environment mask, includes the points in the first and second point clouds not corresponding to dynamic objects, that is, points determined by projecting the pixels of the first and second images included in the static object mask applied to the first and second images.

[0090] In a sixth step 26, a third image and a fourth image are generated.

[0091] The third image is generated from points in the second and third point clouds. For each pair of second and third points corresponding to the same object in the three-dimensional scene, only one of these points is projected, depending on whether the first point to which they correspond belongs to the dynamic object mask. Thus, the pixels of the third image are defined by projection: • of second points associated with the same object in the three-dimensional scene as a first point corresponding to a dynamic object, and therefore included in the first set of points, and • third points associated with the same object in the three-dimensional scene as a first point not corresponding to a dynamic object, that is to say corresponding to a static object and not included in the first set of points.

[0092] Similarly, the fourth image is generated from points in the first and fourth point clouds. For each pair of first and fourth points corresponding to the same object in the three-dimensional scene, only one of these points is projected according to the membership of the second point to which they correspond in the dynamic object mask. Thus, the pixels of the fourth image are defined by projection: • first points associated with the same object in the three-dimensional scene as a second point corresponding to a dynamic object, therefore included in the second set of points, and • fourth points associated with the same object in the three-dimensional scene as a second point not corresponding to a dynamic object, that is to say corresponding to a static object and not included in the second set of points.

[0093] According to a particular embodiment, the position of a pixel in the third image, or fourth image respectively, is determined by the following function: [Math.4] P^^P".^

[0094] With: • P ms the position of a pixel in the third image, respectively fourth image, •77 a function to convert from homogeneous coordinates to pixel coordinates by removing one dimension from a vector, • K is a direction prediction model associated with camera 11, and • P " 3" the spatial position of a second or third point, respectively first or fourth point.

[0095] Thus, in [Fig.4], the third image 43 is generated by projecting the second pixels included in a second set of pixels E2 and the third pixels included in a third set of pixels E3. Indeed, the second set of pixels E2 includes the second pixels associated with a dynamic object, while the third set of pixels E3 includes the third pixels associated with a static object.

[0096] Similarly, the fourth image 44 is generated by projecting the first pixels included in a first set of pixels El and the fourth pixels included in a fourth set of pixels E4. Indeed, the first set of pixels El includes the first pixels associated with a dynamic object, while the fourth set of pixels E4 includes the fourth pixels associated with a static object.

[0097] In a seventh step 27, the depth prediction model and the motion prediction model are learned by minimizing a loss error determined from: • the result of a comparison of the third image to the second image and the result of a comparison of the fourth image to the first image, and • a consistency error representative of a prediction error of said movement.

[0098] According to a particular embodiment, these comparison results include a photometric error determined by the following function: [Math.5] U{p} = EJ ( 1-a) • \l{p)-î(p) | +a■ ( 1 -±SSIM(l(p), ï(p) ) ) ]

[0099] With: • L*{p} is the photometric error associated with a pixel p defined by its coordinates in the first image, and respectively in the second image. • l(p) a value of the pixelp in the first image, respectively second image, • J(p) a value of pixel p in the fourth image, respectively third image, • SSIM is a function that takes into account a local structure, and • has a weighting factor that depends in particular on the type of road environment.

[0100] It should be noted that a photometric error is lower the more accurate the depth and motion prediction models are, i.e., the more effective their training. Conversely, if the depth and motion prediction models lack accuracy, for example because their training is incomplete, then a photometric error is significant.

[0101] According to another specific embodiment, the loss error includes reconstruction errors determined by the following function: [Math.6] / « / ,( / >) =^(^,)

[0102] With: • ^smoothiP ') a reconstruction error for a pixelp of the third image, respectively fourth image, • a first depth, respectively a second depth, associated with pixel p, • VF is a parameter matrix, • ° the order of a smoothing gradient, • an L1 norm of the second-order depth gradients is calculated with VF = 1, et0 = 2, • x and the dimensions of the third and fourth images, • P is a hyperparameter dependent on the road environment in which vehicle 10 is traveling, and • It(p) a value of pixel p in the third image, respectively fourth image.

[0103] This function is generally used to deal with the discontinuity at the edge of objects (in English, "edge aware smoothness").

[0104] According to a particular embodiment, the loss error is determined by the following function: [Math.7] L ' = E^min (L^p^L^p')) + Lcons

[0105] With: • The loss error, • L^p) a result of a comparison of a pixel p from the first image to a pixel p from the fourth image, • L2(p) is the result of a comparison of a pixel p from the second image to a pixel p from the third image, and • Lessons: the consistency error.

[0106] According to another particular embodiment, the loss error is determined by the following function:

[0107] [Math.8] = E^mîll (L^p), L2 {p) ) + E^smoo / fô (P) + ^p^smooth^ (P) + Lcons

[0108] With: • The loss error, • L^p) a result of a comparison of a pixel p from the first image to a pixel p from the fourth image, • L2(p) is the result of a comparison of a pixel p from the second image to a pixel p from the third image. • Lsmootit2(p) a reconstruction error for a pixelp of the third image, • Lsmaot^{p) a reconstruction error for a pixelp of the fourth image, and • Lcons.the consistency error.

[0109] According to a particular embodiment, the consistency error is determined by the following function: [Math.9] LCons E„eV [^t-^s(.P^ ~ (P ) ] 2 ew

[0110] With: • LcmK. The consistency error, • P^ V a pixelp of the first image that is not associated with a dynamic object, • Sf-»s(p) a scene stream associated with pixel p, and • Tt~.s(p} a displacement between the two time instants associated with pixelp.

[0111] Learning the depth prediction model consists of adjusting the input parameters of the convolutional neural network implementing the depth prediction model in order to minimize the previously calculated loss error and similarly learning the motion prediction model consists of adjusting the input parameters of the convolutional neural network implementing the motion prediction model in order to minimize the previously calculated loss error.

[0112] In the case where representative data from several pairs of images are received, an aggregate loss error is then determined, corresponding to the average of the loss errors determined for each pair of images.

[0113] Once the training of the depth prediction and motion prediction models has been completed, for example when the loss error is less than a threshold value defining an acceptable level of accuracy, it is then possible to use the depth prediction model to predict depths associated with pixels of images acquired by the camera 11 and to use the motion prediction model to predict the motion of the camera between images acquired consecutively by the camera 11.

[0114] Each determined depth then corresponds to a distance separating the vehicle 10 or a part of the vehicle 10 from an object in the three-dimensional scene to which a pixel is associated, the determination of a depth of a pixel then corresponding to a measurement of a distance separating an object from the vehicle carrying the vision system.

[0115] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another road user, the ADAS is then able to determine this distance precisely. For example, if the ADAS is designed to activate the braking system of the vehicle 10 in the event of a risk of collision with another road user, and the distance between the vehicle 10 and that road user decreases sharply, then the ADAS is able to detect this sudden closeness and activate the braking system of the vehicle 10 to avoid a possible accident.

[0116] Figure 3 schematically illustrates a device configured to learn a a depth prediction model and a motion prediction model associated with a vehicle-mounted vision system and / or for predicting the depth associated with a pixel of an image acquired by a camera of the vehicle-mounted vision system, according to a particular and non-limiting embodiment of the present invention. Device 3 corresponds, for example, to a device embedded in the vehicle 10, for example a computer associated with the monocular vision system.

[0117] Device 3 is, for example, configured to carry out the steps described opposite Figures 1, 2, and 4. Examples of such a device 3 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 3, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 3 may be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.

[0118] The device 3 comprises one (or more) processor(s) 30 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 3. The processor 30 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 3 further comprises at least one memory 31, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.

[0119] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 31.

[0120] According to various particular and non-limiting embodiments, the device 3 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.

[0121] According to a particular and non-limiting embodiment, the device 3 includes a block 32 of interface elements for communicating with external devices. The interface elements of the block 32 include one or more of the following interfaces: - Radio frequency (RF) interface, for example, Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long- Long-term Evolution” or in French “Long-term Evolution”), LTE-Advanced (or in French LTE-advanced); - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HD MI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").

[0122] According to another particular and non-limiting embodiment, the device 3 includes a communication interface 33 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 330. The communication interface 33 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 330. The communication interface 33 corresponds, for example, to a wired network of the CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3) type.

[0123] According to a particular and non-limiting embodiment, the device 3 can provide output signals to one or more external devices, such as a display screen 340, touch or not, one or more speakers 350 and / or other peripherals 360 via the output interfaces 34, 35, 36 respectively. According to a variant, one or more of the external devices is integrated into the device 3.

[0124] Of course, the present invention is not limited to the embodiments described above but extends to a method for determining the movement of a camera between two acquisition times, and / or for determining the depth of a pixel in an image acquired by a vision system, and / or for measuring the distance between an object and a vehicle equipped with a vision system, the depth and / or distance being predicted and / or measured via a depth prediction model learned according to the learning method described above, which would include secondary steps without falling outside the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0125] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising device 3 of [Fig.3].

Claims

1. Demands A method for learning depth prediction and motion prediction models implemented by convolutional neural networks associated with a vision system embedded in a vehicle (10), the vision system comprising a camera (11) arranged to acquire an image of a three-dimensional scene taking place near the vehicle (10), said method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception (21) of data representative of a first image and a second image acquired by the camera (11) at two distinct time moments of acquisition; - prediction (22) • first depths associated with pixels of the first image, called first pixels, and second depths associated with pixels of the second image, called second pixels, the first and second depths being predicted by said depth prediction model, and • of a camera movement (11) between the two time instants of acquisition by the motion prediction model; - generation (23): • first three-dimensional points from the first image, the spatial position of a first point being a function of the coordinates of a first pixel in the first image associated with said first point and of said first depth associated with said first pixel, • of second three-dimensional points from the second image, a spatial position of a second point, being a function of the coordinates of a second pixel in the second image associated with said second point and of said second depth associated with said second pixel; • of three-dimensional third points from the first image, the spatial position of a third point being a function of the coordinates of a first pixel in the first image associated with said third point, of said first depth associated with said first pixel, and of said movement, and • of fourth three-dimensional points from the second image, a spatial position of a fourth point being a function of the coordinates of a second pixel in the second image associated with said fourth point, of said second depth associated with said second pixel and of said movement; - determination (24) of a scene flow for each first and second points associated with the same object of the three-dimensional scene as a function of a difference in position between the first and second points; - determination (25) of a first set of points comprising first points corresponding to pixels of the first image associated with at least one dynamic object and of a second set of points comprising second points corresponding to pixels of the second image associated with at least one dynamic object, dynamic objects being detected by a segmentation model in the first and second images; - generation (26): • of a third image, pixels of the third image being defined by projection of second points associated with the same object of the three-dimensional scene as a first point included in the first set of points, and of third points associated with the same object of the three-dimensional scene as a first point not included in the first set of points, and • of a fourth image, the pixels of the fourth image being defined by projection of first points associated with the same object of the three-dimensional scene as a second point included in the second set of points and of fourth points associated with the same object of the three-dimensional scene as a second point not included in the second set of points, - learning (27) depth prediction and motion prediction models by minimizing a loss error determined from: • the result of a comparison of the third image to the second image and the result of a comparison of the fourth image to the first image, and • a consistency error representative of a prediction error of said movement.

2. A method according to claim 1, wherein a spatial position of a first point, respectively second point, is determined by the following function: With: • P^d the spatial position of a first point, respectively second point, • K a direction prediction model associated with the camera (11), and • 0 a projection function in a three-dimensional scene of a pixelp as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, Dmt associated with it.

3. A method according to claim 1 or 2, wherein a spatial position of a third point, respectively fourth point, is determined by the following function: With: • P'^d the spatial position of a third point, respectively fourth point, • a direction prediction model associated with the camera (11), • 0 a projection function in a three-dimensional scene of a pixelp as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, Dmt which is associated with it, and • T a movement of the vehicle (10) between the two time instants.

4. A method according to any one of claims 1 to 3, wherein a position of a pixel of the third image, respectively fourth image, is determined by the following function: pm=n{KP'\D) With: • P'n the position of a pixel of the third image, respectively fourth image, • 7' a function for going from homogeneous coordinates to pixel coordinates by removing one dimension of a vector, • K a direction prediction model associated with the camera (11), and • P'n the spatial position of a second or third point, respectively first or fourth point.

5. A method according to any one of claims 1 to 4, wherein the loss error is determined by the following function: L' = E^minCLj(p), £2(p) ) + £CWIS With: • L' is the loss error, • L^p) is the result of a comparison of a pixelp from the first image to a pixel p from the fourth image, • L2(p) is the result of a comparison of a pixelp from the second image to a pixel p from the third image, and • LconsX is the consistency error.

6. A method according to any one of claims 1 to 5, wherein the consistency error is determined by the following function: ^cons E / pey [ P^ ~ ( P ) ] With: • Lcons.V consistency error, •pE V a pixelp of the first image that is not associated with a dynamic object, • S^p) a scene flow associated with pixelp, and • Tt^s(p) a displacement between the two time instants associated with pixel p.

7. A method according to any one of claims 1 to 5, wherein the segmentation model determines bounding boxes associated with dynamic objects in the first and second images.

8. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by a processor.

9. Device (3) configured to learn a depth prediction model and a motion prediction model associated with a vision system embedded in a vehicle (10), said device (3) comprising a memory (31) associated with at least one processor (30) configured for carrying out the steps of the method according to any one of claims 1 to 7.

10. Vehicle (10) comprising the device (3) according to claim 9.

Citation Information

Patent Citations

  • Joint learning of geometry and motion with three-dimensional holistic understanding

    US20200211206A1