Method and device for learning a depth prediction model assisted by a scene flow model.

The method improves ADAS reliability by using a convolutional neural network to train a depth prediction model that excludes dynamic objects, ensuring accurate depth predictions for both static and dynamic scenes, thus enhancing vehicle safety systems.

FR3167468A1Pending Publication Date: 2026-04-17STELLANTIS AUTO SAS
0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
STELLANTIS AUTO SAS
Filing Date
2024-10-14
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Monocular vision systems struggle to accurately predict depths for dynamic objects in a three-dimensional scene due to pixel displacement caused by both the movement of the vehicle and the objects, which degrades the learning of the depth prediction model and affects the reliability of Advanced Driver-Assistance Systems (ADAS).

Method used

A method using a convolutional neural network to learn a depth prediction model by receiving images from a vehicle-mounted camera, generating three-dimensional points from these images, determining scene flow and displacement, and minimizing loss errors to exclude dynamic objects during training, thereby improving the model's accuracy.

Benefits of technology

The method enhances the reliability of ADAS systems by accurately predicting depths for both static and dynamic objects, ensuring precise distance measurements for improved road safety.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

Method or device for learning a depth prediction model associated with a vision system. Specifically, the method comprises receiving images (21) acquired by a camera, generating (23) first and second three-dimensional point clouds from the received images and depths predicted (22) by the depth prediction model, and generating point clouds from the received images based on vehicle movement and predicted depths. The point positions are compared (25) to identify points relative to dynamic objects, and images are generated (26) from points based on their correspondence with a dynamic object. The depth prediction model is then learned (27) by comparing the generated images with the received images. Figure for the abstract: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for learning a depth prediction model assisted by a scene flow model. technical field

[0001] The present invention relates to methods and devices for learning a depth prediction model associated with a vision system embedded in a vehicle, for example, in a motor vehicle. The present invention also relates to a method and device for determining depth and / or measuring the distance separating a static or dynamic object from a vehicle equipped with a vision system. Technological background

[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS systems use advanced technologies to assist the driver while driving and thus improve performance. ADAS systems use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.

[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.

[0004] The ADAS systems installed in a vehicle are powered by data obtained from one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or any obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle's speed; - to act on the braking system in case of risk of impact with an object.

[0005] The position of another user or an obstacle in a three-dimensional scene is, for example, determined by a vision system comprising a prediction model for a depth associated with a pixel or a distance separating the vision system from an object in a three-dimensional scene. However, the accuracy of a predicted depth or distance for a dynamic object, that is, an object in movement in the three-dimensional scene is generally not satisfactory for a vision system comprising only a single camera, also called a monocular vision system.

[0006] Indeed, a monocular vision system uses several consecutive images to predict depths associated with pixels in one of the images. This monocular vision system has the drawback of not being able to accurately predict depths associated with pixels in an image when objects in the observed three-dimensional scene are moving within that same scene, because the depth is obtained by comparing the positions of pixels corresponding to the same object in the three-dimensional scene in different images. However, if the observed object is moving, then the pixels associated with that object are displaced between two images acquired at different times, and this displacement is due both to the movement of the monocular vision system mounted in the vehicle and to the inherent movement of the dynamic object within the three-dimensional scene.A depth prediction model is therefore unable to accurately determine the position of this dynamic object.

[0007] Moreover, the presence of a dynamic object in a learning phase then disrupts the learning of the depth prediction model, that is to say that the quality of the learning of the depth prediction model is degraded when pixels corresponding to dynamic objects are taken into consideration during the learning phase.

[0008] The quality of the training of the depth or distance prediction model is, however, very important. Indeed, the depths or distances predicted by the model represent, for example, the distances to other road users or obstacles present in the road environment of the vehicle equipped with the vision system, whose output data feeds ADAS (Advanced Driver Assistance Systems). The proper functioning of the driver assistance devices using this data therefore depends on the quality of the data emitted by the vision system. Summary of the present invention

[0009] One object of the present invention is to solve at least one of the problems of the technological background described above.

[0010] Another object of the present invention is to improve the learning phase of a depth prediction model from images acquired by a vision system and representing objects of a three-dimensional scene moving in that same scene, the depth prediction model being implemented in particular by a neural network and associated with a vision system embedded in a vehicle.

[0011] Another object of the present invention is to improve road safety, in particular by improving the reliability of AD AS systems powered by data obtained from a camera of a vision system.

[0012] According to a first aspect, the present invention relates to a method for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, the vision system comprising a camera arranged to acquire an image of a three-dimensional scene taking place near the vehicle, the method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception of representative data from a first image and a second image acquired by the camera at two distinct acquisition times; - prediction of first depths associated with pixels of the first image, called first pixels, and of second depths associated with pixels of the second image, called second pixels, the first and second depths being predicted by said depth prediction model; - generation : • initial three-dimensional points from the first image, the spatial position of a first point being a function of the coordinates of a first pixel in the first image associated with the first point and the first depth associated with the first pixel, and • of second three-dimensional points from the second image, a spatial position of a second point, being a function of the coordinates of a second pixel in the second image associated with the second point and of the second depth associated with the second pixel; • of three-dimensional third points from the first image, the spatial position of a third point being a function of the coordinates of a first pixel in the first image associated with the third point, the first depth associated with the first pixel, and a movement of the vehicle between the two temporal instants of acquisition, and • of fourth three-dimensional points from the second image, a spatial position of a fourth point being a function of the coordinates of a second pixel in the second image associated with the fourth point, of the second depth associated with the second pixel and of said movement; - determination : • a scene flow for each first and second point associated with the same object in the three-dimensional scene, based on a difference in position between the first and second points, and • a displacement for each first point based on a difference in position between the first and third points associated with the same object in the three-dimensional scene and a displacement for each second point based on a difference in position between the second and fourth points associated with the same object in the three-dimensional scene; - determination of an offset associated with each first point as a function of the scene flows and displacement associated with each first point and of an offset associated with each second point as a function of the scene flows and displacement associated with each second point; - generation : • of a third image, pixels of the third image being defined by projection of second points associated with the same object of the three-dimensional scene as a first point whose offset is greater than a threshold value and of third points associated with the same object of the three-dimensional scene as a first point whose offset is less than the threshold value, and • of a fourth image, pixels of the fourth image being defined by projection of first points associated with the same object of the three-dimensional scene as a second point whose offset is greater than the threshold and of fourth points associated with the same object of the three-dimensional scene as a second point whose offset is less than the threshold value, - learning the depth prediction model by minimizing a loss error determined from a result of a comparison of the third image to the second image and a result of a comparison of the fourth image to the first image.

[0013] The training phase thus allows the depth prediction model to be trained using generated images free from errors related to the presence of dynamic objects in the three-dimensional scene. The presence of dynamic objects in the three-dimensional scene observed during the training phase therefore does not disrupt this learning process, which can be carried out using any pair of images containing pixels corresponding to dynamic objects.

[0014] According to a variant of the method, a spatial position of a first point, respectively second point, is determined by the following function: P}„ = With: • P^D the spatial position of a first point, respectively second point, • K a direction prediction model associated with the camera, and • 0 a projection function in a three-dimensional scene of a pixelp as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, Dtnt which is associated with it.

[0015] According to another variant of the method, a spatial position of a third point, respectively a fourth point, is determined by the following function: P'w = T^\Dm) With: • P'zi) the spatial position of a third point, respectively fourth point, • K is a camera-associated direction prediction model, • 0 a projection function in a three-dimensional scene of a pixelp as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, associated with it, and • T a movement of the vehicle between the two time instants.

[0016] According to a further variant of the method, an offset associated with a first point, respectively second point, corresponds to a distance between a second and third point, respectively between a first and fourth point.

[0017] According to yet another variant of the method, an offset associated with a first point, respectively a second point, is determined by the following function: d- yr-PH With: • d the offset associated with a first point, respectively second point, • F a scene stream associated with a first point, respectively second point, and • P a displacement associated with a first point, respectively second point.

[0018] According to yet another variant of the method, the position of a pixel of the third image, or fourth image respectively, is determined by the following function: With : • Pms the position of a pixel in the third image, respectively fourth image, • A function to convert from homogeneous coordinates to pixel coordinates by removing one dimension from a vector, • F is a camera-associated direction prediction model, and • P " 3D the spatial position of a second or third point, respectively first or fourth point.

[0019] According to an additional variant of the method, the loss error is determined by the following function: L' = p), L^p)) With : • The loss error, • L^p) a result of a comparison of a pixel p from the first image to a pixel p from the fourth image, and • L^p} a result of a comparison of a pixel p of the second image to a pixel p of the third image.

[0020] According to a second aspect, the present invention relates to a device configured to learn a depth prediction model by a vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.

[0021] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0022] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0023] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0024] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.

[0025] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.

[0026] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.

[0027] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures

[0028] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 4, in which:

[0029] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a a particular and non-limiting example of a realization of the present invention;

[0030] [Fig.2] illustrates a flowchart of the different stages of a method for learning a depth prediction model associated with the vision system of [Fig.1], according to a particular and non-limiting embodiment of the present invention; and

[0031] [Fig.3] schematically illustrates a device configured to learn a depth prediction model associated with the vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention; and

[0032] [Fig.4] schematically illustrates images and point clouds used during the learning process of [Fig.2], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements

[0033] A method and device for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle will now be described in what follows with joint reference to Figures 1 to 4. The same elements are identified with the same reference signs throughout the description that follows.

[0034] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0035] Fig. 1 schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.

[0036] The vehicle 10 is located in an environment 1 corresponding, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.

[0037] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, with electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a Land vehicles such as cars, trucks, coaches, and motorcycles. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.

[0038] The vehicle 10 advantageously comprises at least one onboard camera 11, configured to acquire images of a three-dimensional scene unfolding in the environment of the vehicle 10 from a common viewing position. The camera 11 forms a monocular vision system if used alone, as illustrated in [Fig. 1]. However, the present invention is not limited to a monocular vision system comprising a single camera but extends to any vision system comprising at least one camera, for example, 1, 2, 3, or 5 cameras.

[0039] The camera 11 has intrinsic parameters, including: - a focal length, - distortions which are due to imperfections in the optical system of camera 11, - a direction of the optical axis of camera 11, and - a resolution.

[0040] The intrinsic parameters characterize the transformation which associates, for an image point, hereafter called "point", its three-dimensional coordinates in the camera 11 reference frame with the pixel coordinates in an image acquired by the camera 11. These parameters do not change if the camera is moved.

[0041] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in lens curvature, parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are not taken into account, and their correction is addressed during image acquisition or calibration.

[0042] The camera 11 is arranged so as to acquire an image of a three-dimensional scene from a defined viewpoint, the viewpoint is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10 as illustrated in [Fig.1].

[0043] The camera 11, for example, acquires images of a three-dimensional scene located in front of the vehicle 10, the camera 11 covering an acquisition field 12. An object 13 is placed in the acquisition field 12 of the camera 11, defining an occlusion field 14 for the vision system, an object present in the occlusion field 14 not being observable by camera 11 from its current observation position.

[0044] It is evident that it is possible to use such a vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently, the invention not being limited to the observation of a three-dimensional scene taking place in front of the vehicle 10 carrying the camera 11.

[0045] According to one particular embodiment, the camera 11 is a wide-angle camera, a wide-angle camera being, for example, equipped with a lens designed to acquire a representative image of a three-dimensional scene seen over a wider field of view than a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens makes it possible to capture a larger portion of the three-dimensional scene unfolding in front of or around the camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by this camera. The angle α of the field of view of the camera 11 is, for example, equal to 120°, 145°, 180°, or 360°, whereas a standard camera offers, for example, a field of view open at an angle of 45° or less.Such a camera 11 corresponds, for example, to a camera equipped with mirrors or a "fisheye" camera. Wide-angle lenses have a shorter focal length compared to standard lenses, making them suitable for capturing images of landscapes, architecture, road intersections, or any other subject requiring a wide perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with an extended depth of field.

[0046] An image acquired by the camera 11 at a given acquisition time is in the form of data representing pixels characterized by: - ​​coordinates in the image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric values ​​(from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).

[0047] Each pixel of the acquired image represents an object in the three-dimensional scene present in the field of view of the camera 11. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a point of light resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, captured by a photosensitive sensor of the camera 11 after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image that faithfully represents the physical object observed by the camera. An image point, as previously described, is therefore a point on the surface of an object in the three-dimensional scene.

[0048] When the vehicle 10 is in motion, then two images acquired by the camera 11 at two distinct time instants represent views of the same three-dimensional scene taken from different viewpoints or observation positions, the observation positions of the camera 11 being distinct. On this three-dimensional scene there are, for example: - buildings; - road infrastructure; - other users or stationary objects, for example a parked vehicle; and / or - other users or moving objects, for example another vehicle, a cyclist or a moving pedestrian.

[0049] According to a particular embodiment, an image acquired by the camera 11 includes a distortion equal to 0.5%, 0.8%, or greater than 1%. The measurement of such distortion corresponds to determining a ratio between: - the maximum spacing of a pixel in the image of a straight line in the first three-dimensional scene whose image is a line touching the longest edge of the first image, either at the center of the image edge or at the corners of the image edge, and - the length of this edge.

[0050] In the world of photography, distortion is commonly considered to be: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • problematic if it is greater than or equal to 1% or more.

[0051] Barrel distortion is characterized by a positive percentage, while crescent distortion is characterized by a negative percentage.

[0052] Each image acquired by the camera 11 is, for example, sent to a processor, for example a computer of a device equipping the vehicle 10, or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10. It is then used during the implementation of a method for determining the depth of one of its pixels and / or during the implementation of a method learning a depth prediction model associated with this camera of the vision system.

[0053] Figure 2 illustrates a flowchart of the different stages of a method for learning a depth prediction model, such a depth prediction model being for example associated with the vision system on board the vehicle 10 and including the camera 11 and also being used in a method for determining the depth of a pixel of an image, according to a particular and non-limiting embodiment of the present invention.

[0054] The learning process 2 is for example implemented by a processor of a device embedded in the vehicle 10 or of the device 3 of the [Fig.3].

[0055] In a first step 21, representative data of a first image and a second image acquired by the camera 11 at two distinct acquisition times are received; that is, a first image acquired by the camera 11 at a first acquisition time and a second image acquired by the camera 11 at a second acquisition time distinct from the first acquisition time are received. Image reception consists of receiving representative image data, which are, for example, recorded in a device embedded in the vehicle 10 and comprising a memory accessible by a computer implementing the learning process 2. The first and second acquisition times are sufficiently close in time so that the observed three-dimensional scene is the same but observed from two different viewpoints when the vehicle 10 is in motion.The first and second images are acquired at time intervals on the order of one or more milliseconds.

[0056] In a second step 22, first depths associated with pixels of the first image, called first pixels, are predicted by the depth prediction model and, similarly, second depths associated with pixels of the second image, called second pixels, are predicted by the depth prediction model.

[0057] Such a depth prediction model is known to those skilled in the art, and is described for example in the following documents: - “Digging Into Self-Supervised Monocular Depth Estimation” by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow published in August 2019, and - “HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation” by Xiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong, Lina Liu, Yong Liu, Xinxin Chen and Yi Yuan published in December 2020.

[0058] It should be noted that the first and second depths are increasingly precise for static objects in the observed three-dimensional scene, that is to say These depths become increasingly accurate as the learning process is repeated with new images. However, the first and second depths predicted for dynamic objects in the observed three-dimensional scene are less reliable; therefore, it is necessary to consider the nature of the object corresponding to a first or second pixel to determine the accuracy of the first or second predicted depth, which is what the following steps allow.

[0059] In a third step 23, four point clouds are generated, these clouds comprising three-dimensional points positioned in the observed scene: • a first point cloud containing first points, • a second point cloud containing second points, • a third point cloud containing third points, and • a fourth point cloud containing fourth points.

[0060] The first points are generated from the first image, a spatial position of a first point being a function of the coordinates of a first pixel in the first image associated with this first point and of the first depth associated with the first pixel, and the second points are generated similarly from the second image, a spatial position of a second point being a function of the coordinates of a second pixel in the second image associated with this second point and of the second depth associated with the second pixel.

[0061] As illustrated in [Fig.4], the generation of the first point cloud consists of the projection of the pixels of the first image 41 into the three-dimensional scene and into the camera frame 11 at the first time instant of acquisition.

[0062] Similarly, the generation of the second point cloud consists of the projection of the pixels of the second image 42 into the three-dimensional scene and into the frame of the camera 11 at the second time instant of acquisition, which is different from the frame of the camera 11 at the first time instant of acquisition when the vehicle 10 is in motion, the camera 11 having moved in the three-dimensional scene between the first and second time instants of acquisition.

[0063] According to a particular embodiment, a spatial position of a first point, respectively second point, is determined by the following function: [Math.l] ^30 = 0(^-

[0064] With: • P3D the spatial position of a first point, respectively second point, • K is a direction prediction model associated with camera 11, and • 0 a projection function in a three-dimensional scene of a pixel as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, Dint which is associated with it.

[0065] According to a first variant, the direction prediction model corresponds to an intrinsic matrix of the camera 11. This first variant is particularly applicable to a pinhole camera and calibrated.

[0066] According to a second variant, the direction prediction model corresponds to a learned model implementing a different neural network. Such a direction prediction model is known to those skilled in the art; it is notably presented in the paper "Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion" written by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich, and Adrien Gaidon and published in August 2020. This second variant is suitable for wide-angle or fisheye cameras, or even for uncalibrated cameras. According to this second variant, a direction prediction step associated with the pixels of the first image and the pixels of the second image is necessary.

[0067] Thus, as illustrated in [Fig.4], a first point pl and a second point p2 are respectively generated from the first image 41 and the second image 42. The first and second points are notably represented in the form of black dots.

[0068] The third points are, for their part, generated from the first image, a spatial position of a third point being a function of the coordinates of a first pixel in the first image associated with the third point, of the first depth associated with the first pixel and of a movement of the vehicle 10 between the two time instants of acquisition, and the fourth points are generated from the second image, a spatial position of a fourth point being a function of the coordinates of a second pixel in the second image associated with the fourth point, of the second depth associated with the second pixel and of the movement of the vehicle 10 between the two time instants of acquisition.

[0069] As illustrated in [Fig. 4], the generation of the third point cloud consists of projecting the pixels of the first image 41 into the three-dimensional scene and into the frame of the camera 11 at the second acquisition time instant. Alternatively, considering the first point cloud obtained from the first image 41, the generation of the third point cloud consists of projecting the first points of the first point cloud into the frame of the camera 11 at the second acquisition time instant, taking into account the movement of the vehicle 10 between the two acquisition times. The third point p3 is then the image of a pixel from the first image 41 or of the first point pl, according to the definition above.

[0070] Similarly, the generation of the fourth point cloud consists of projecting the pixels of the second image 42 into the three-dimensional scene and into the frame of the camera 11 at the first time instant of acquisition. Or, considering the second point cloud obtained from the second image 41, the generation of the fourth point cloud consists of projecting the second points of the second point cloud into the frame of the camera 11 at the first time instant of acquisition, taking into account the movement of the vehicle 10 between the two time instants of acquisition. The fourth point p4 is then the image of a pixel of the second image 42 or of the second point p2, according to the definition above.

[0071] According to a particular embodiment, a spatial position of a third point, respectively fourth point, is determined by the following function: [Math.2] P'w

[0072] With: • P'^d the spatial position of a third point, respectively fourth point, • K a direction prediction model associated with camera 11, • 0 a projection function in a three-dimensional scene of a pixel p as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, Dmt which is associated with it, and • T a movement of the vehicle 10 between the two time instants.

[0073] It is noted here that part of this function corresponds to the previous function determining the spatial position of a first or second point.

[0074] Note that the movement of the vehicle between the two time instants does not have the same value when determining the position of a third point and a fourth point. Indeed, the movement considered for determining the position of a fourth pixel is opposite to the movement considered for determining the position of a third pixel.

[0075] The motion is notably presented in the form of a matrix comprising extrinsic parameters of the moving camera 11, which can be obtained in several ways. This matrix is, for example, obtained from data emitted by a driver assistance system of the vehicle 10, called the ADAS system, or from sensors on board the vehicle 10. According to another example, the extrinsic parameters corresponding to the motion of the camera 11 between the first and second time instants are determined in an additional step by a motion prediction model, for example by the one presented in the document “Neural Ray Surfaces for Self-Supervised Leaming of Depth and Ego-motion”.

[0076] Thus, as illustrated in [Fig.4], a third point p3 and a fourth point p4 are respectively generated from the first image 41 and the second image 42. The third and fourth points are notably represented as white dots.

[0077] In a fourth step 24, a scene flow and a displacement are determined for each first point and for each second point.

[0078] A scene flow is determined for each first and second point associated with the same object in the three-dimensional scene based on a difference in position between the first and second points. A model for determining scene flow is known to those skilled in the art and is presented in particular in the document "DeFlow: Decoder of Scene Flow Network in Autonomous Driving" written by Qingwen Zhang, Yi Yang, Heng Fang, Ruoyu Geng and Patrie Jensfelt and published in January 2024.

[0079] In [Fig.4], a scene flow F is determined for the first point pl and a scene flow opposite to the scene flow F is determined for the second point p2.

[0080] The scene flow allows us to determine, for a first point, the corresponding second point, that is, the second point corresponding to the same object in the three-dimensional scene. Similarly, the scene flow allows us to determine, for a second point, the corresponding first point, that is, the second point corresponding to the same object in the three-dimensional scene. Note that the pixel in the first image 41 that yields the first point corresponds to the pixel in the second image 42 that yields the second point, the determination of corresponding pixels in the first and second images being known to those skilled in the art.

[0081] A displacement for each first point is also determined as a function of a difference in position between the first and third points associated with the same object in the three-dimensional scene and, similarly, a displacement for each second point is determined as a function of a difference in position between the second and fourth points associated with the same object in the three-dimensional scene.

[0082] The displacement allows us to determine, for a first point, the corresponding third point, that is to say, corresponding to the same object in the three-dimensional scene. Similarly, the displacement allows us to determine, for a second point, the corresponding fourth point, that is to say, corresponding to the same object in the three-dimensional scene.

[0083] Thus, each first point corresponds to: • a second point positioned relative to the first point according to a scene flow, and • a third point positioned relative to the first point according to a displacement.

[0084] In [Fig.4], it is thus distinguished that at the first point pl correspond: • the second point p2 positioned relative to the first point pl according to the scene flow F, and • the third point p3 positioned relative to the first point pl according to the displacement T.

[0085] Similarly, each second point corresponds to: • a first point positioned relative to the second point according to a scene flow, and • a fourth point positioned relative to the second point according to a displacement.

[0086] In a fifth step 25, an offset associated with each first point is determined based on the scene flows and displacement associated with that first point and an offset associated with each second point is determined based on the scene flows and displacement associated with that second point.

[0087] According to a particular embodiment, an offset associated with a first point, respectively second point, corresponds to a distance between a second and third point, respectively between a first and fourth point.

[0088] According to one variant, the offset is determined by a distance calculation from the coordinates of the first and fourth points, respectively second and third points.

[0089] According to another variant, an offset associated with a first point, respectively second point, is determined from the scene and displacement flows by the following function associated with this first point, respectively with this second point: [Math.3] d = ||FD||

[0090] With: • d the offset associated with a first point, respectively second point, • F the scene flow associated with a first point, respectively second point, and • the displacement associated with a first point, respectively second point.

[0091] In a sixth step 26, a third image and a fourth image are generated.

[0092] The third image is generated from points in the second and third point clouds. For each pair of second and third points corresponding to the same object in the three-dimensional scene, only one of these points is projected according to of the offset determined for a first corresponding point. Thus, the pixels of the third image are defined by projection: • second points associated with the same object in the three-dimensional scene as a first point whose offset is greater than a threshold value, and • third points associated with the same object in the three-dimensional scene as a first point whose offset is less than the threshold value.

[0093] Similarly, the fourth image is generated from points in the first and fourth point clouds. For each pair of first and fourth points corresponding to the same object in the three-dimensional scene, only one of these points is projected according to the offset determined for a second corresponding point. Thus, the pixels of the fourth image are defined by projection: • first points associated with the same object in the three-dimensional scene as a second point whose offset is greater than a threshold value, and • fourth points associated with the same object in the three-dimensional scene as a second point whose offset is less than the threshold value.

[0094] Comparing an offset to the threshold value makes it possible to identify the first and second points corresponding to a dynamic object of the three-dimensional scene when the offset is greater than the threshold value and the first and second points corresponding to a static object of the three-dimensional scene when the offset is less than the threshold value.

[0095] According to a particular embodiment, the position of a pixel in the third image, or fourth image respectively, is determined by the following function: [Math.4] Pms = ^(KP^D)

[0096] With: • Pms the position of a pixel in the third image, respectively the fourth image, • 71 a function to go from homogeneous coordinates to pixel coordinates by removing one dimension of a vector, • K is a direction prediction model associated with camera 11, and • P " 3D the spatial position of a second or third point, respectively first or fourth point.

[0097] Thus, in [Fig. 4], the third image 43 is generated by projecting the second pixels included in a second set of pixels E2 and the third pixels included in a third set of pixels E3. Indeed, the second set of pixels E2 comprises the second pixels associated with a dynamic object, that is, corresponding to a first pixel whose offset is greater than the threshold value, while the third set of pixels E3 comprises the third pixels associated to a static object, that is to say corresponding to a first pixel whose offset is less than the threshold value.

[0098] Similarly, the fourth image 44 is generated by projecting the first pixels included in a first set of pixels El and the fourth pixels included in a fourth set of pixels E4. Indeed, the first set of pixels El includes the first pixels associated with a dynamic object, that is to say corresponding to a second pixel whose offset is greater than the threshold value while the fourth set of pixels E4 includes the fourth pixels associated with a static object, that is to say corresponding to a second pixel whose offset is less than the threshold value.

[0099] In a seventh step 27, the depth prediction model is learned by minimizing a loss error determined from: • a result of comparing the third image to the second image, and • a result of comparing the fourth image to the first image.

[0100] According to a particular embodiment, these comparison results include a photometric error determined by the following function: [Math.5] U{p} = EJ ( 1-a) • \l{p)-î(p) | +a■ ( 1 -±SSIM(l(p), ï(p) ) ) ]

[0101] With: • L*{p} is the photometric error associated with a pixel p defined by its coordinates in the first image, and respectively in the second image. • l(p) a value of pixel p in the first image, respectively second image, • J(p) a value of pixel p in the fourth image, respectively third image, • SSIM is a function that takes into account a local structure, and • has a weighting factor that depends in particular on the type of road environment.

[0102] It should be noted that a photometric error is lower the more accurate the depth prediction model, i.e., the more effective its training. Conversely, if the depth prediction model lacks accuracy, for example because its training is incomplete, then a photometric error is significant.

[0103] According to another particular embodiment, the loss error includes reconstruction errors determined by the following function: [Math.6] l^Ip) =L Jsv (W(p i )

[0104] With: • ^smooth^P^ a reconstruction error for a pixel p of the third image, respectively fourth image, • a first depth, respectively a second depth, associated with pixel p, • W is a parameter matrix, • 0 the order of a smoothing gradient, • an L1 norm of the second-order depth gradients is calculated with W = 1, et0 = 2, • x and y are the dimensions of the third and fourth images, • P is a hyperparameter dependent on the road environment in which vehicle 10 is traveling, and • It(p) a value of pixel p in the third image, respectively fourth image.

[0105] This function is generally used to deal with the discontinuity at the edge of objects (in English, "edge aware smoothness").

[0106] According to a particular embodiment, the loss error is determined by the following function: [Math.7] L^p} )

[0107] With: • The loss error, • L(p) is the result of a comparison of a pixel p from the first image to a pixel p from the fourth image, and • L^ip) a result of a comparison of a pixel p of the second image to a pixel p of the third image.

[0108] According to another particular embodiment, the loss error is determined by the following function:

[0109] [Math.8] £' = ^min(L^pYL^p') ) + [Math.8] '—'pf^smooth'i ( P ) + [Math.8] ^^p^smoothA ( P )

[0110] With: • The loss error, •L^p) a result of comparing a pixel p from the first image to a pixel p from the fourth image, • L2(p) is the result of a comparison of a pixel p from the second image to a pixel p from the third image. • Lsmooth2(p) a reconstruction error for a pixelp of the third image, and • ^smooth^P} a reconstruction error for a pixelp of the fourth image.

[0111] The training of the depth prediction model consists of adjusting the input parameters of the convolutional neural network implementing the depth prediction model in order to minimize the previously calculated loss error.

[0112] Once the training of the depth model is complete, for example when the loss error is less than a threshold value defining an acceptable level of accuracy, it is then possible to use the depth prediction model to predict depths associated with pixels of images acquired by the camera 11.

[0113] Each determined depth then corresponds to a distance separating the vehicle 10 or a part of the vehicle 10 from an object in the three-dimensional scene to which a pixel is associated, the determination of a depth of a pixel then corresponding to a measurement of a distance separating an object from the vehicle carrying the vision system.

[0114] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another road user, the ADAS is then able to determine this distance precisely. For example, if the ADAS is designed to activate the braking system of the vehicle 10 in the event of a risk of collision with another road user, and the distance between the vehicle 10 and that road user decreases sharply, then the ADAS is able to detect this sudden closeness and activate the braking system of the vehicle 10 to avoid a possible accident.

[0115] Figure 3 schematically illustrates a device configured to learn a depth prediction model associated with a vehicle-mounted vision system and / or to predict a depth associated with a pixel of an image acquired by a camera of the vehicle-mounted vision system, according to a particular and non-limiting embodiment of the present invention. The device corresponds for example to a device embedded in the vehicle 10, for example a computer associated with the monocular vision system.

[0116] Device 3 is, for example, configured to carry out the steps described opposite Figures 1, 2, and 4. Examples of such a device 3 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 3, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 3 may be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.

[0117] The device 3 comprises one (or more) processor(s) 30 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 3. The processor 30 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 3 further comprises at least one memory 31, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.

[0118] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 31.

[0119] According to various particular and non-limiting embodiments, the device 3 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.

[0120] According to a particular and non-limiting embodiment, the device 3 comprises a block 32 of interface elements for communicating with external devices. The interface elements of the block 32 comprise one or more of the following interfaces: - Radio frequency (RF) interface, for example, Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long- Long-term Evolution” or in French “Long-term Evolution”), LTE-Advanced (or in French LTE-advanced); - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HD MI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").

[0121] According to another particular and non-limiting embodiment, the device 3 includes a communication interface 33 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 330. The communication interface 33 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 330. The communication interface 33 corresponds, for example, to a wired network of the CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3) type.

[0122] According to a particular and non-limiting embodiment, the device 3 can provide output signals to one or more external devices, such as a display screen 340, touch or not, one or more speakers 350 and / or other peripherals 360 via the output interfaces 34, 35, 36 respectively. According to a variant, one or more of the external devices is integrated into the device 3.

[0123] Of course, the present invention is not limited to the embodiments described above but extends to a method for determining the depth of a pixel in an image acquired by a vision system, and / or for measuring the distance between an object and a vehicle equipped with a vision system, the depth and / or distance being predicted and / or measured via a depth prediction model learned according to the learning method described above, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0124] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 3 of [Fig.3].

Claims

1. Demands Method for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle (10), the vision system comprising a camera (11) arranged to acquire an image of a three-dimensional scene taking place near the vehicle (10), said method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception (21) of data representative of a first image and a second image acquired by the camera (11) at two distinct time moments of acquisition; - prediction (22) of first depths associated with pixels of the first image, called first pixels, and of second depths associated with pixels of the second image, called second pixels, the first and second depths being predicted by said depth prediction model; - generation (23): • first three-dimensional points from the first image, the spatial position of a first point being a function of the coordinates of a first pixel in the first image associated with said first point and of said first depth associated with said first pixel, and • of second three-dimensional points from the second image, a spatial position of a second point, being a function of the coordinates of a second pixel in the second image associated with said second point and of said second depth associated with said second pixel; • of three-dimensional third points from the first image, a spatial position of a third point being a function of the coordinates of a first pixel in the first image associated with said third point, of said first depth associated with said first pixel and of a movement of the vehicle (10) between the two temporal instants of acquisition, and • of fourth three-dimensional points from the second image, the spatial position of a fourth point being a function of the coordinates of a second pixel in the associated second image

2. said fourth point, of the said second depth associated with said second pixel and of said movement; - determination (24): • a scene flow for each first and second point associated with the same object in the three-dimensional scene, based on a difference in position between the first and second points, and • a displacement for each first point based on a difference in position between the first and third points associated with the same object in the three-dimensional scene and a displacement for each second point based on a difference in position between the second and fourth points associated with the same object in the three-dimensional scene; - determination (25) of an offset associated with each first point as a function of the scene flows and displacement associated with said first point and of an offset associated with each second point as a function of the scene flows and displacement associated with said second point; - generation (26): • of a third image, pixels of the third image being defined by projection of second points associated with the same object of the three-dimensional scene as a first point whose offset is greater than a threshold value and of third points associated with the same object of the three-dimensional scene as a first point whose offset is less than the threshold value, and • of a fourth image, the pixels of the fourth image being defined by projection of first points associated with the same object of the three-dimensional scene as a second point whose offset is greater than the threshold and of fourth points associated with the same object of the three-dimensional scene as a second point whose offset is less than the threshold value, - learning (27) of the depth prediction model by minimizing a loss error determined from a result of a comparison of the third image to the second image and a result of a comparison of the fourth image to the first image. A method according to claim 1, wherein a spatial position of a first point, respectively a second point, is determined by the following function: With: • P}d the spatial position of a first point, respectively second point, • K a direction prediction model associated with the camera (11), and • 0 a projection function in a three-dimensional scene of a pixelp as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, Dm( which is associated with it.

3. A method according to claim 1 or 2, wherein a spatial position of a third point, respectively fourth point, is determined by the following function: With: • P'sd the spatial position of a third point, respectively fourth point, • K a direction prediction model associated with the camera (11), • 0 a projection function in a three-dimensional scene of a pixelp as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, Dmt associated with it, and • T a movement of the vehicle (10) between the two time instants.

4. A method according to any one of claims 1 to 3, wherein an offset associated with a first point, respectively second point, corresponds to a distance between a second and third point, respectively between a first and fourth point.

5. Method according to claim 4, wherein an offset associated with a first point, respectively second point, is determined by the following function: d= With: • d the offset associated with a first point, respectively second point, • F a scene flow associated with a first point, respectively second point, and • D a displacement associated with a first point, respectively second point.

6. A method according to any one of claims 1 to 4, wherein a position of a pixel of the third image, respectively fourth image, is determined by the following function: Ptm = ^(KP"3D) With: • Pnu the position of a pixel of the third image, respectively fourth image, •57 a function for going from homogeneous coordinates to pixel coordinates by removing a dimension of a vector, • K a direction prediction model associated with the camera (11), and • P"3D the spatial position of a second or third point, respectively first or fourth point.

7. A method according to any one of claims 1 to 6, wherein the loss error is determined by the following function: L' = Epmin(L^p), L2(p)) With: • L' the loss error, • a result of a comparison of a pixel p of the first image to a pixel p of the fourth image, and • L2(p) a result of a comparison of a pixel p of the second image to a pixel p of the third image.

8. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by a processor.

9. Device (3) configured to learn a depth prediction model by a vision system embedded in a vehicle (10), said device (3) comprising a memory (31) associated with at least one processor (30) configured to implement the steps of the method according to any one of claims 1 to 7.

10. Vehicle (10) comprising the device (3) according to claim 9.