Method and device for learning a model for predicting the positions of three-dimensional points associated with a vision system.

A convolutional neural network-based method for predicting three-dimensional point positions from camera images addresses the inefficiencies of existing depth learning and LiDAR, enhancing ADAS systems' accuracy and responsiveness.

FR3167233A1Pending Publication Date: 2026-04-10STELLANTIS AUTO SAS
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
STELLANTIS AUTO SAS
Filing Date
2024-10-08
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Learning a depth prediction model for vehicle-mounted vision systems is tedious and requires significant time and high-quality training data, while using LiDAR sensors is expensive and has long data acquisition and processing times, hindering the improvement of ADAS reliability and accuracy.

Method used

A method using a convolutional neural network to predict three-dimensional point positions from images acquired by a vehicle-mounted camera, involving image reception, depth prediction, point cloud generation, and error minimization to learn a model that forecasts future object positions, enhancing the accuracy and responsiveness of ADAS systems.

Benefits of technology

The method provides precise and rapid data for ADAS systems, improving road safety by accurately predicting the future positions of objects, outperforming LiDAR in efficiency and reliability without the need for expensive hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method or device for learning a three-dimensional point position prediction model associated with a vehicle-mounted vision system and including a camera. Images acquired by the camera at different times are received (21), and depths associated with pixels are predicted (22) by a depth prediction model. Initial three-dimensional point clouds are generated (23) from the images based on their coordinates and the predicted depths, and the three-dimensional point position prediction model generates (24) second past or future three-dimensional point clouds from the first point clouds. Images are then generated (25) from the second point clouds and compared (26) to the acquired images to learn (27) the three-dimensional point position prediction model by minimizing a loss error. Figure for the abstract: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for learning a model for predicting the positions of three-dimensional points associated with a vision system. technical field

[0001] The present invention relates to methods and devices for learning a model for predicting the positions of three-dimensional points associated with a vision system embedded in a vehicle, for example, in a motor vehicle. The present invention also relates to a method for determining depth using a vision system embedded in a vehicle. Technological background

[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS systems use advanced technologies to assist the driver while driving and thus improve performance. ADAS systems use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.

[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.

[0004] Vehicle-mounted ADAS systems are powered by data obtained from one or more on-board sensors, such as cameras. These cameras make it possible, in particular, to detect and locate other road users or any obstacles present around a vehicle in order to, for example: • to adapt the vehicle's lighting according to the presence of other road users; • to automatically regulate the vehicle's speed; • to act on the braking system in case of risk of impact with an object.

[0005] Learning a depth prediction model associated with these cameras is tedious and requires both a large amount of high-quality training data and time. In order to learn the depth prediction model from data acquired by the vision system itself, the training phase should be optimized to reduce its duration while improving its reliability.

[0006] To improve the accuracy of knowledge of the vehicle's environment, it is preferable to directly use points arranged in space to feed the ADAS, which is done, for example, using data obtained by a three-dimensional environmental sensor such as a LiDAR®. However, this type of sensor is particularly expensive, its implementation in a vehicle is complex, and the acquisition and processing time of the data received from the LiDAR® is significantly longer than the image processing time, since the rotation speed of a LiDAR® sensor is on the order of 10 Hz (ten Hertz), generating an exposure time of a three-dimensional scene on the order of 30 to 100 ms (thirty to one hundred milliseconds), while this same scene must be exposed in front of the camera for 1 ms (one millisecond). Summary of the present invention

[0007] One object of the present invention is to solve at least one of the problems of the technological background described above.

[0008] Another object of the present invention is to improve the quality of data from the processing of an image acquired by a vision system, in particular by a depth prediction model implemented by a neural network associated with a stereoscopic vision system.

[0009] Another object of the present invention is to improve road safety, in particular by improving the reliability of AD AS systems powered by data obtained from a vision system.

[0010] According to a first aspect, the present invention relates to a method for learning a model for predicting the positions of three-dimensional points implemented by a convolutional neural network associated with a vision system embedded in a vehicle, the vision system comprising a camera arranged to acquire an image of a three-dimensional scene taking place near the vehicle, the method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception of representative data of a first image acquired by the camera at a first time instant of acquisition and of a second image acquired by the camera at a second time instant of acquisition distinct from the first time instant of acquisition; - prediction of first depths associated with pixels of the first image and of second depths associated with pixels of the second image, the first and second depths being predicted by a depth prediction model; - generation of a first and a second three-dimensional point clouds from the first and second images, the spatial position of a point in the first point cloud being a function of the coordinates of a pixel in the first image and a depth associated with the pixel of the first image and a spatial position of a point of the second point cloud being a function of the coordinates of a pixel of the second image and a depth associated with the pixel of the second image; - prediction, using the three-dimensional point position prediction model, of a third three-dimensional point cloud from the first three-dimensional point cloud and of the vehicle's movement between the first and second time instants, and • a fourth three-dimensional point cloud from the second three-dimensional point cloud and the movement of the vehicle between the first and second time instants; - generation : • of a third image by projection of points from the third point cloud into an image plane associated with the camera and defined at the second time instant of acquisition, and • of a fourth image by projection of points from the fourth point cloud into an image plane associated with the camera and defined at the first time instant of acquisition; - determination of a first error by comparing the first image to the fourth image and of a second error by comparing the fourth image to the second image; - learning the model for predicting the positions of three-dimensional points by minimizing a loss error determined from the first and second errors.

[0011] Such a method thus makes it possible to use the three-dimensional point position prediction model to forecast the future position of objects in an observed three-dimensional scene in order to provide a driver assistance system with precise and rapidly determined data. Such a three-dimensional point position prediction model is therefore faster than a device incorporating a LiDAR® and allows for better anticipation of the behavior of surrounding objects by predicting their future position. Driver assistance systems using data from this three-dimensional point position prediction model are thus more accurate and more responsive.

[0012] According to a variant of the method, a spatial position of a point in the first and second point clouds is determined by the following function: P}„ = With: • P^D the spatial position of a point in the first point cloud, respectively in the second point cloud, • K is a camera-associated direction prediction model, and • 0 a projection function in a three-dimensional scene of a pixelp as a function of its coordinates in the first image, respectively in the second image, and of the first, respectively second, depth Dmt which is associated with it.

[0013] According to another variant of the method, the coordinates of a pixel in the third image, and respectively the fourth image, are determined by the following function: p'-7r(KP'3D) With: • p' the coordinates of a pixel in the third image, respectively fourth image, • K is a camera-associated direction prediction model, • A function to convert from homogeneous coordinates to pixel coordinates by removing one dimension from a vector, and • P'^d the spatial position of a point in the third point cloud, respectively in the fourth point cloud.

[0014] According to yet another variant of the method, a first depth associated with a pixel of the first image or a second depth associated with a pixel of the second image is determined from depth candidates, a probability being determined for each depth candidate by the depth prediction model, the first or second depth being equal to a sum of products of the probabilities and depth candidates associated with said pixel of the first or second image.

[0015] According to a further embodiment of the method, the first and second errors comprise a photometric error determined by the following function: With : • L*(p) the photometric error associated with a pixel p defined by its coordinates in the first image, respectively second image, • lip) a value of pixel p in the first image, respectively second image, * l(p) a value of pixel p in the fourth image, respectively third image, • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of road environment.

[0016] According to yet another variant of the process, the first and second errors include a reconstruction error determined by the following function: L^Çp) = ^,.(^(^,) With : • Lsmooth{p) is a reconstruction error for a pixelp of the third image, or fourth image respectively, • ^mt is a depth associated with the pixelp, • VF is a parameter matrix, • 0 is the order of a smoothing gradient, • an L1 norm of the second-order depth gradients is calculated with VF = 1, et0 = 2, • x and y are dimensions of the third image, and fourth image respectively, • P is a hyperparameter dependent on the road environment in which vehicle 10 is traveling, and • It(p) is a value of pixel p in the third image, respectively fourth image.

[0017] According to a further embodiment of the method, the loss error is determined by the following function: L'^Ypmm(Li(p),L2(p)) With : • The loss error, • L} (p) the first error for a pixelp of the first image, and • L2(p) is the second error for a pixelp of the second image corresponding to the pixel p of the first image.

[0018] According to a second aspect, the present invention relates to a learning device for a model predicting the positions of three-dimensional points associated with a vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.

[0019] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0020] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0021] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0022] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.

[0023] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.

[0024] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.

[0025] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures

[0026] Other features and advantages of the present invention will become apparent from the description of the specific and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 3, in which:

[0027] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a a particular and non-limiting example of a realization of the present invention;

[0028] [Fig.2] illustrates a flowchart of the different stages of a learning process for a model predicting the positions of three-dimensional points associated with a vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention; and

[0029] [Fig.3] schematically illustrates a device configured to determine the depth of a pixel of an image acquired by a vision system on board the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements

[0030] A method and device for determining the depth of a pixel of an image by a neural network associated with a vision system embedded in a vehicle will now be described in what follows with joint reference to Figures 1 to 3. The same elements are identified with the same reference signs throughout the description that follows.

[0031] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0032] Fig. 1 schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.

[0033] The vehicle 10 is located in an environment 1 corresponding, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.

[0034] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.

[0035] The vehicle 10 advantageously comprises at least one onboard camera 11, configured to acquire images of a three-dimensional scene occurring near the vehicle 10, i.e., in the immediate vicinity of the vehicle 10, from a typical viewing position. The camera 11 forms a monocular vision system if used alone, as illustrated in [Fig. 1]. However, the present invention is not limited to a monocular vision system comprising a single camera but extends to any vision system comprising at least one camera, for example, 1, 2, 3, or 5 cameras.

[0036] The camera 11 has intrinsic parameters, including: - a focal length, - distortions which are due to imperfections in the optical system of camera 11, - a direction of the optical axis of camera 11, and - a resolution.

[0037] The intrinsic parameters characterize the transformation which associates, for an image point, hereafter called "point", its three-dimensional coordinates in the camera 11 reference frame with the pixel coordinates in an image acquired by the camera 11. These parameters do not change if the camera is moved.

[0038] According to one particular embodiment, the camera 11 is calibrated, for example, by following a calibration method known to those skilled in the art such as that presented in the document "Single View Point Omnidirectional Camera Calibration from Planar Grids", written by Christopher Mei and Patrick Rives and published in April 2007, hereafter referred to as the Mei model, the intrinsic parameters of the camera 11 being then known.

[0039] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in lens curvature, parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are not taken into account, and their correction is addressed during image acquisition or calibration.

[0040] The camera 11 is arranged to acquire an image of a three-dimensional scene from a defined viewpoint, the viewpoint is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10 as illustrated in [Fig.1].

[0041] The camera 11, for example, acquires images of a three-dimensional scene located in front of the vehicle 10, the camera 11 covering an acquisition field 12. An object 13 is placed in the acquisition field 12 of the camera 11, defining an occlusion field 14 for the vision system, an object present in the occlusion field 14 not being observable by the camera 11 from its current observation position.

[0042] It is evident that it is possible to use such a vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently, the invention not being limited to the observation of a three-dimensional scene taking place in front of the vehicle 10 carrying the camera 11.

[0043] According to one particular embodiment, the camera 11 is a wide-angle camera, a wide-angle camera being, for example, equipped with a lens designed to acquire a representative image of a three-dimensional scene seen over a wider field of view than a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens makes it possible to capture a larger portion of the three-dimensional scene unfolding in front of or around the camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by this camera. The angle α of the field of view of the camera 11 is, for example, equal to 120°, 145°, 180°, or 360°, whereas a standard camera offers, for example, a field of view open at an angle of 45° or less.Such a camera 11 corresponds, for example, to a camera equipped with mirrors or to a "fisheye" camera. Wide-angle lenses have a shorter focal length. Compared to standard lenses, this makes them suitable for capturing images of landscapes, architecture, road intersections, or any other subject requiring a wide perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with an extended depth of field.

[0044] An image acquired by the camera 11 at a given acquisition time is in the form of data representing pixels characterized by: - ​​coordinates in the image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric values ​​(from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).

[0045] Each pixel of the acquired image represents an object in the three-dimensional scene present in the field of view of the camera 11. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a point of light resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, which are captured by a photosensitive sensor of the camera 11 after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image that faithfully represents the physical object observed by the camera 11. An image point described above is therefore a point on the surface of an object in the three-dimensional scene.

[0046] When the vehicle 10 is in motion, then two images acquired by the camera 11 at two distinct time instants represent views of the same three-dimensional scene taken from different viewpoints or observation positions, the observation positions of the camera 11 being distinct. On this three-dimensional scene there are, for example: - buildings; - road infrastructure; - other users or stationary objects, for example a parked vehicle; and / or - other users or moving objects, for example another vehicle, a cyclist or a moving pedestrian.

[0047] According to a particular embodiment, an image acquired by the camera 11 includes a distortion equal to 0.5%, 0.8%, or greater than 1%. The measurement of such distortion corresponds to determining a ratio between: - the maximum spacing of a pixel in the image of a straight line in the first three-dimensional scene whose image is a line touching the longest edge of the first image, either in the center of the image edge, or at the corners of the image edge, and - the length of this edge.

[0048] In the world of photography, distortion is commonly considered to be: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • problematic if it is greater than or equal to 1%.

[0049] Barrel distortion is characterized by a positive percentage, while crescent distortion is characterized by a negative percentage.

[0050] Each image acquired by the camera 11 is, for example, sent to a processor, for example a computer of a device equipping the vehicle 10, or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10. It is then used during the implementation of one of the following processes: • a process for determining the depth of at least one of its pixels, and / or • a process for learning a depth prediction model associated with this camera of the vision system, and / or • a method for predicting the positions of three-dimensional points, and / or • a method for learning a model for predicting the positions of three-dimensional points.

[0051] Figure [Fig. 2] illustrates a flowchart of the different stages of a learning process for a model predicting the positions of three-dimensional points according to a particular and non-limiting embodiment of the present invention.

[0052] A method for learning a model for predicting the positions of three-dimensional points is advantageously implemented by the vehicle 10, i.e. by a computer or a combination of computers of the vehicle 10's on-board system, for example by the computer(s) in charge of the vision system of the vehicle 10 or by the device 3 of [Fig.3].

[0053] The model for predicting the positions of three-dimensional points is implemented by a convolutional neural network associated with the vision system embedded in the vehicle 10 described above. Thus, the model for predicting the positions of three-dimensional points is associated with the vision system comprising a camera 11 arranged to acquire an image of a three-dimensional scene unfolding around the vehicle 10.

[0054] In a first step 21, a first image acquired by the camera 11 at a first time instant of acquisition and a second image acquired by the Images are received from camera 11 at a second acquisition time, distinct from the first acquisition time. Image reception consists of receiving data representing images, which are, for example, recorded in a device embedded in the vehicle 10 and comprising memory accessible by a computer implementing the process. The first and second acquisition times are sufficiently close in time so that the observed three-dimensional scene is the same but viewed from two different perspectives while the vehicle 10 is in motion. The first and second images are acquired at time intervals on the order of one or more milliseconds.

[0055] According to a particular embodiment, the positions of the pixels in the first and second images are not expressed in Cartesian coordinates but are encoded in another reference frame associated with these images.

[0056] Position coding is performed by transforming the position into a combined function using a sinusoidal curve and a cosine curve, following the sub-operations below: • In a first operation, the abscissa x in the horizontal direction and the ordinate y in the vertical direction in pixel coordinates (0,1,2,3,4,5,6...) are normalized in a range 0 to 2Æ. • In a second operation, a denominator that increases in pairs is calculated, the denominator thus describes the following series: 1, 1, 2, 2, 3, 3... with the magnitudes greater than 1 to decrease the magnitude of the coordinate value, and preferably with an exponential increase. • In a third operation, the values ​​of the abscissa x and ordinate y are divided by the denominator. • In a fourth operation, a sine function is applied to the digits in odd positions and a cosine function to the digits in even positions of the x-coordinate and y-coordinate. The result forms two matrices with the height and width of the characteristic map respectively for x and y. • In a fifth operation, the two position coding matrices are added to all channels of the two feature maps.

[0057] Such position coding subsequently makes it possible to preserve the position of the pixels in the original image and to evaluate a level of distortion associated with them, which is especially important when the images exhibit strong distortion, such as when the images are acquired by a wide-angle camera.

[0058] In a second step 22, first depths associated with pixels of the first image and second depths associated with pixels of the second image are predicted by a depth prediction model associated with the system of vehicle-mounted vision 10. Such a depth prediction model is known to the person skilled in the art, and is described for example in the following documents: - “Digging Into Self-Supervised Monocular Depth Estimation” by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow published in August 2019, and - “HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation” by Xiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong, Lina Liu, Yong Liu and Xinxin Chen, Yi Yuan published in December 2020.

[0059] According to a particular embodiment, a first depth associated with a pixel in the first image or a second depth associated with a pixel in the second image is determined from depth candidates, a probability being determined by the depth prediction model for each depth candidate, the first or second depth being equal to a sum of the products of the probabilities and the depth candidates associated with said pixel in the first or second image. The determination of depth candidates and probabilities is used, for example, in the Unimatch® model presented in the paper "Unifying Flow, Stereo and Depth Estimation" by Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger, published in July 2023.A series of depth candidates covers, for example, the entire measurement range determined for camera 11, for example from 10 to 100m (from ten to one hundred meters), with an interval of 1m (one meter) for example.

[0060] In a third step 23, a first and a second three-dimensional point cloud are generated from the first and second images. A spatial position of a point in the first point cloud is then a function of the coordinates of a pixel in the first image and a depth associated with that pixel in the first image and, similarly, a spatial position of a point in the second point cloud is a function of the coordinates of a pixel in the second image and a depth associated with that pixel in the second image.

[0061] According to a particular embodiment, a spatial position of a point in the first and second point clouds is determined by the following function: [Math.l]

[0062] With: • P3D the spatial position of a point in the first point cloud, respectively in the second point cloud, • K is a direction prediction model associated with camera 11, and • 0 a projection function in a three-dimensional scene of a pixel as a function of its coordinates in the first image, respectively in the second image, and of the first, respectively second, depth Dtnt which is associated with it.

[0063] The direction prediction model associated with the camera 11 is also known to those skilled in the art, such a prediction model being notably presented in the document "Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion" written by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich and Adrien Gaidon published in August 2020.

[0064] Note that for a camera 11 with a pinhole lens, the direction prediction model is defined by an analytical function and thus corresponds to the intrinsic matrix of the camera 11, also called the projection matrix of the camera 11 and comprising four parameters representing horizontal focal lengths (fx) and vertical focal lengths (fy) in pixels and the position (cx,cy) in pixels of the optical center in the whole of the image plane.

[0065] Each pixel of an image acquired by the camera 11 corresponds to an object point in the three-dimensional scene. A point cloud generated from an image theoretically reconstructs the object points of the observed three-dimensional scene. However, the accuracy of their position depends on the quality of the depth prediction and direction prediction models.

[0066] It should be noted that the first point cloud is defined in the frame of reference associated with the camera 11 in its position at the first time instant of acquisition, and that the second point cloud is defined in the frame of reference associated with the camera 11 in its position at the second time instant of acquisition. The link between these frames of reference is defined by the movement of the vehicle 10 between these two time instants of acquisition, which is known by analyzing the images or by analyzing telemetry data obtained from sensors onboard the vehicle 10.

[0067] In a fourth step 24, the three-dimensional point position prediction model predicts a third three-dimensional point cloud from the first three-dimensional point cloud and the movement of the vehicle 10 between the first and second time instants of acquisition, and a fourth three-dimensional point cloud from the second three-dimensional point cloud and the movement of the vehicle 10 between the first and second time instants of acquisition.

[0068] Such a model for predicting the positions of three-dimensional points is known to the person skilled in the art, and is presented in the document "Visual Point Cloud Forecasting enables Scalable Autonomous Driving" written by Zetong Yang, Li Chen, Yanan Sun and Hongyang Li and published in December 2023.

[0069] The third point cloud thus corresponds to a theoretical position of the first point cloud in the reference frame associated with camera 11 at the second time instant during acquisition, and certain points associated with dynamic objects also have their positions modified according to the predicted proper displacement of the dynamic object between the two acquisition times. Similarly, the fourth point cloud corresponds to a theoretical position of the second point cloud in the frame of reference associated with camera 11 at the first acquisition time, and as before, certain points associated with dynamic objects also have their positions modified according to the predicted proper displacement of the dynamic object between the two acquisition times.

[0070] It should be noted that if the models for predicting depth, direction, and the positions of three-dimensional points were perfect, the first and fourth point clouds would be identical, and the second and third point clouds would also be identical, which is not the case due to the imprecision of each of the models. Thus, the first and fourth point clouds are comparable, and the second and third point clouds are also comparable.

[0071] In a fifth step 25, a third image and a fourth image are generated. The third image is generated by projecting points from the third point cloud into an image plane associated with the camera 11 and defined at the second time instant of acquisition, and the fourth image is generated by projecting points from the fourth point cloud into an image plane associated with the camera 11 and defined at the first time instant of acquisition.

[0072] According to a particular embodiment, the coordinates of a pixel in the third image, respectively fourth image, are determined by the following function: [Math.2] p'^(KP3D)

[0073] With: • p' the coordinates of a pixel in the third image, respectively fourth image, • K is a direction prediction model associated with camera 11, • A function to convert from homogeneous coordinates to pixel coordinates by removing one dimension from a vector, and • P?,d the spatial position of a point in the third point cloud, respectively in the fourth point cloud.

[0074] As previously explained, the first and fourth point clouds, just like the second and third point clouds, being comparable, the fourth image is comparable to the first image and the third image is comparable to the second image.

[0075] In a sixth step 26, a first error is determined by comparing the first image to the fourth image and a second error is determined by comparing the fourth image to the second image.

[0076] According to a particular embodiment, the first and second errors include photometric errors as presented in the document "Digging Into Self-Supervised Monocular Depth Estimation" by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow published in August 2019 and are determined by the following function: [Math.3] U{p} = EJ ( 1-a) • \l{p)-î(p) | +a■ ( 1 -±SSIM(l(p), ï(p) ) ) ]

[0077] With: • L*{p} is the photometric error associated with a pixel p defined by its coordinates in the first image, respectively the second image, • Ip is a pixel value in the first image, respectively second image, • / p a value of pixel p in the fourth image, respectively third image, • SSIM is a function that takes into account a local structure, and • has a weighting factor that depends in particular on the type of road environment.

[0078] According to another particular embodiment, the first and second errors include a reconstruction error of a generated image determined by the following function: [Math.4]

[0079] With: • L^gg^p} a reconstruction error for a pixel p of the third image, respectively fourth image, • Dmt a depth associated with the pixelp, • VF is a parameter matrix, • 0 is the order of a smoothing gradient, • an L1 norm of the second-order depth gradients is calculated with VF = 1, et0 = 2, • x and y are dimensions of the third image, and fourth image respectively, • P is a hyperparameter dependent on the road environment in which vehicle 10 is traveling, and * h(P) is a value of pixel p in the third image, respectively fourth image.

[0080] This function is generally used to deal with discontinuity at the boundary of objects.

[0081] The first and second errors are thus defined, for example, based on the photometric and reconstruction errors defined previously. Note that a photometric error is determined for the first and second images respectively, while a reconstruction error is determined for the fourth and third images respectively. They can nevertheless be combined, for example, added together, since these four images have the same resolution and the errors are determined in pixel coordinates, which are homogeneous for all four images.

[0082] In a seventh step 27, the model for predicting the positions of three-dimensional points is learned by minimizing a loss error determined from the first and second errors.

[0083] According to a particular embodiment, the loss error is determined by the following function: [Math.5] L' = Epmin(£]( / ?), )

[0084] With: • The loss error, • L^p) the first error for a pixelp of the first image, and • L2{p) the second error for a pixelp of the second image corresponding to the pixel p of the first image.

[0085] The use of the minimum (min) function is preferred here to a sum or average function in order to avoid taking into account errors due to occlusions in the acquired images. According to other variants, however, these sum or average functions remain possible for determining the loss error.

[0086] Learning the model for predicting the positions of three-dimensional points consists of adjusting the input parameters of the convolutional neural network in order to minimize the previously calculated loss error.

[0087] Thus, the model for predicting the positions of three-dimensional points is made more reliable through this learning process. It is therefore able to accurately predict the future position of objects in the three-dimensional scene associated with three-dimensional points, whether the object is static or dynamic, and regardless of the vehicle's movement. Furthermore, the data used for this learning is obtained from the vision system itself; it therefore corresponds to data that perfectly represents the use of the vision system onboard the vehicle 10 and thus to the real-world environments in which the vehicle 10 operates.

[0088] It should be noted that it is also possible to learn the depth prediction model by further adjusting the parameters of the convolutional neural network used by the depth prediction model.

[0089] The predicted depths and / or the positions of the points in the predicted point clouds then make it possible to determine a distance separating the vehicle from an object to which a pixel or a three-dimensional point corresponds.

[0090] If the ADAS uses predicted depths and / or predicted point clouds as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another road user, the ADAS is then able to determine this distance precisely. For example, if the ADAS is designed to activate the braking system of the vehicle 10 in the event of a risk of collision with another road user, and the distance between the vehicle 10 and that road user decreases sharply, then the ADAS is able to detect this sudden closeness and activate the braking system of the vehicle 10 to avoid a potential accident.

[0091] Figure 3 schematically illustrates a device 3 configured for predicting the positions of three-dimensional points and / or for learning a model for predicting the positions of three-dimensional points associated with a vision system embedded in a vehicle 10, according to a particular and non-limiting embodiment of the present invention. The device 3 corresponds, for example, to a device embedded in the first vehicle 10, for example, a computer associated with the stereoscopic vision system.

[0092] Device 3 is, for example, configured to carry out the operations described opposite Figures 1 and 4 and / or the steps described opposite Figures 2 and 3. Examples of such a device 3 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 3, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 3 may be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.

[0093] The device 3 comprises one (or more) processor(s) 30 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 3. The processor 30 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 3 further comprises at least one memory 31 corresponding for example to volatile and / or non-volatile memory and / or includes a memory storage device which may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.

[0094] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 31.

[0095] According to various particular and non-limiting embodiments, the device 3 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Unité de Contrôle Télématique”), for example via a communication bus or through dedicated input / output ports.

[0096] According to a particular and non-limiting embodiment, the device 3 comprises a block 32 of interface elements for communicating with external devices. The interface elements of the block 32 comprise one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HD MI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").

[0097] According to another particular and non-limiting embodiment, the device 3 includes a communication interface 33 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 330. The communication interface 33 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 330. The communication interface 33 corresponds, for example, to a wired CAN (Controller Area Network) or CAN FD (Controller Area Network Flexible Data-Rate) type network. flexible data rate”), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3).

[0098] According to a particular and non-limiting embodiment, the device 3 can provide output signals to one or more external devices, such as a display screen 340, touch or not, one or more speakers 350 and / or other peripherals 360 via the output interfaces 34, 35, 36 respectively. According to a variant, one or more of the external devices is integrated into the device 3.

[0099] Of course, the present invention is not limited to the embodiments described above but extends to a method for determining the depth of a pixel in an image using the depth prediction model, or to a method for measuring the distance between an object and a vehicle equipped with a vision system, which would include secondary steps without falling outside the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0100] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 3 of [Fig.3].

Claims

1. Demands Method for learning a model for predicting the positions of three-dimensional points implemented by a convolutional neural network associated with a vision system embedded in a vehicle (10), the vision system comprising a camera (11) arranged to acquire an image of a three-dimensional scene taking place near the vehicle (10), said method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception (21) of data representing a first image acquired by the camera (11) at a first time instant of acquisition and a second image acquired by the camera (11) at a second time instant of acquisition distinct from said first time instant; - prediction (22) of first depths associated with pixels of the first image and of second depths associated with pixels of the second image, the first and second depths being predicted by a depth prediction model; - generation (23) of a first and a second three-dimensional point cloud from the first and second images, a spatial position of a point of the first point cloud being a function of the coordinates of a pixel of the first image and of a depth associated with said pixel of the first image and a spatial position of a point of the second point cloud being a function of the coordinates of a pixel of the second image and of a depth associated with said pixel of the second image; - prediction (24), by said model for predicting the positions of three-dimensional points, • of a third three-dimensional point cloud from the first three-dimensional point cloud and of a movement of the vehicle (10) between said first and second time instants, and • of a fourth three-dimensional point cloud from the second three-dimensional point cloud and said movement; - generation (25): • of a third image by projection of points from the third point cloud onto an image plane associated with the camera (11) and defined at said second time instant, and • of a fourth image by projection of points from the fourth point cloud into an image plane associated with the camera (11) and defined at said first time instant; - determination (26) of a first error by comparison of the first image to the fourth image and of a second error by comparison of the fourth image to the second image; - learning (27) of the model for predicting positions of three-dimensional points by minimizing a loss error determined from the first and second errors.

2. A method according to claim 1, wherein a spatial position of a point in the first and second point clouds is determined by the following function: P}D = With: • Pjd the spatial position of a point in the first point cloud, respectively in the second point cloud, • K a direction prediction model associated with the camera (11), and • 0 a projection function in a three-dimensional scene of a pixel as a function of its coordinates in the first image, respectively in the second image, and of the first, respectively second, depth Dmt associated with it.

3. A method according to claim 1 or 2, wherein the coordinates of a pixel in the third image, respectively fourth image, are determined by the following function: p'^7T{KP\ï}) With: • p' the coordinates of a pixel in the third image, respectively fourth image, • K a direction prediction model associated with the camera (11), • Æ a function for going from homogeneous coordinates to pixel coordinates by removing one dimension of a vector, and • P'^d the spatial position of a point in the third point cloud, respectively fourth point cloud.

4. A method according to any one of claims 1 to 3, wherein a first depth associated with a pixel of the first image or a second depth associated with a pixel of the second image is determined from depth candidates, a probability being determined by the depth prediction model for each depth candidate, the first or second depth being equal to a sum of products of the probabilities and depth candidates associated with said pixel of the first or second image.

5. A method according to any one of claims 1 to 4, wherein said first and second errors include a photometric error determined by the following function: L,(p) = Ep[(la) • \Hp)-î(p)\+a (14sSIM{l(p).ï(p)'})] With: • L*(p) the photometric error associated with a pixel defined by its coordinates in the first image, respectively second image, • l(p) a value of pixel p in the first image, respectively second image, * l(p) a value of pixel p in the fourth image, respectively third image, • SSIM a function which takes into account a local structure, and • a a weighting factor depending in particular on a type of road environment.

6. A method according to any one of claims 1 to 5, wherein said first and second errors comprise a reconstruction error determined by the following function: L^^p) (P,) With: • ^smoath^ P ) a reconstruction error for a pixel p of the third image, respectively fourth image, • a depth associated with pixel p, • W is a parameter matrix, • ° is the order of a smoothing gradient, • an L1 norm of second-order depth gradients is calculated with VF = 1, and ° = 2, • x and are dimensions of the third image, respectively fourth image, • / 1 is a hyperparameter dependent on the road environment in which the vehicle is traveling (10), and • 1t(p) is a value of pixel p in the third image, respectively fourth image.

7. A method according to any one of claims 1 to 6, wherein the loss error is determined by the following function: L' = p), L2(p) ) With: • L' the loss error, • Lj( p) the first error for a pixel p of the first image, and • L2( p) the second error for a pixel p of the second image corresponding to the pixel p of the first image.

8. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by a processor.

9. Device (3) for learning a model for predicting the positions of three-dimensional points for a vision system embedded in a vehicle (10), said device (3) comprising a memory (31) associated with at least one processor (30) configured for carrying out the steps of the method according to any one of claims 1 to 7.

10. Vehicle (10) comprising the device (3) according to claim 9.

Citation Information

Patent Citations

  • Unsupervised learning of image depth and ego-motion prediction neural networks

    US10810752B2

  • Computer-implemented method to improve scale consistency and / or scale awareness in a model of self-supervised depth and ego-motion prediction neural networks

    US20220156882A1

  • Depth detection method, method for training depth estimation branch network, electronic device, and storage medium

    US20220351398A1

  • Method for obtaining depth images for improved driving safety and electronic device

    US20230394690A1