Method and device for learning a scene flow prediction model.
The scene flow prediction model using a convolutional neural network addresses inaccuracies in ADAS systems by improving depth and distance predictions, enhancing the reliability and safety of vehicle operations.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- STELLANTIS AUTO SAS
- Filing Date
- 2024-11-04
- Publication Date
- 2026-05-08
AI Technical Summary
Existing ADAS systems in vehicles face inaccuracies in predicting depth and distance from acquired images, leading to inefficiencies in processing and decision-making, which can be improved by enhancing the scene flow prediction model using a convolutional neural network.
A method for learning a scene flow prediction model using a convolutional neural network that processes images from a vehicle's camera to predict depths and scene flows, minimizing loss errors through image comparisons and displacement determinations to enhance accuracy.
Improves the reliability of ADAS systems by providing precise depth and distance predictions, enabling safer vehicle operations by enhancing the accuracy of ADAS decision-making processes.
Abstract
Description
Title of the invention: Method and device for learning a scene flow prediction model. technical field
[0001] The present invention relates to methods and devices for learning a scene flow prediction model associated with a vision system embedded in a vehicle, for example, in a motor vehicle. The present invention also relates to a method and device for determining a scene flow using a scene flow prediction model. Technological background
[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS systems use advanced technologies to assist the driver while driving and thus improve performance. ADAS systems use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.
[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.
[0004] The ADAS systems installed in a vehicle are powered by data obtained from one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or any obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle's speed; - to act on the braking system in case of risk of impact with an object.
[0005] The position of another user or an obstacle in a three-dimensional scene is, for example, determined by a vision system comprising a model for predicting the depth associated with a pixel or the distance separating the vision system from an object in a three-dimensional scene. However, the accuracy of a depth or distance predicted from acquired images is not satisfactory, and the execution time of such a process is long; it is therefore preferable to process points in space in order to determine the position of objects more precisely and more quickly. in the three-dimensional scene and predict their future position, which is achievable using a scene flow prediction model.
[0006] The quality of the training of the scene flow prediction model is, however, very important and must be adapted to the vision system with which it is associated. Indeed, the predicted spatial positions of points associated with objects in the three-dimensional scene represent, for example, the distances to other road users or obstacles present in the road environment of the vehicle equipped with the vision system, whose output data feeds ADAS. The proper functioning of the driver assistance devices using this data therefore depends on the quality of the data emitted by the vision system. Summary of the present invention
[0007] One object of the present invention is to solve at least one of the problems of the technological background described above.
[0008] Another object of the present invention is to improve the learning phase of a scene flow prediction model from images acquired by a vision system and representing objects of a three-dimensional scene, the scene flow prediction model being implemented in particular by a neural network and associated with a vision system embedded in a vehicle.
[0009] Another object of the present invention is to improve road safety, in particular by improving the reliability of ADAS systems powered by data obtained from a camera of a vision system.
[0010] According to a first aspect, the present invention relates to a method for learning a scene flow prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, the vision system comprising a camera arranged to acquire an image of a three-dimensional scene taking place near the vehicle, the method being implemented by at least one processor and comprising the following steps: - reception of data representative of a first image and a second image respectively acquired by the camera at a first time instant of acquisition and at a second time instant of acquisition distinct from the first time instant of acquisition; - prediction of first depths associated with pixels of the first image, called first pixels, and of second depths associated with pixels of the second image, called second pixels, the first and second depths being predicted by a depth prediction model; - generation : • initial points from the first image, the spatial position of a first point being a function of the coordinates of a first pixel in the first image associated with the first point and the first depth associated with the first pixel, and • of second points from the second image, a spatial position of a second point, being a function of the coordinates of a second pixel in the second image associated with the second point and of the second depth associated with the second pixel; - prediction of a scene flow associated with the first points and a scene flow associated with the second points, the scene flows being predicted by the scene flow prediction model; - determination of spatial positions of third points and fourth points, a spatial position of a third point being determined from a spatial position of a first point associated with it and a scene flow associated with the first point and a spatial position of a fourth point being determined from a spatial position of a second point associated with it and a scene flow associated with the second point; - learning the scene flow prediction model by minimizing a loss error determined by comparison: • from the second image to a third image generated by projecting the third points onto an image plane associated with the camera at the second time instant of acquisition, and from the first image to a fourth image generated by projecting the fourth points onto an image plane associated with the camera at the first time instant of acquisition, or • scene flow associated with first and second points to displacements determined for first and second points, a displacement being determined for each first point and for each second point associated with the same object of the three-dimensional scene by comparison of the spatial positions of each first point and each second point.
[0011] The learning phase thus makes it possible to train the scene flow prediction model from images acquired by the camera of the vision system embedded in the vehicle, these images being particularly relevant because they are representative of real three-dimensional scenes including objects commonly encountered by this on-board vision system.
[0012] According to a variant of the method, a spatial position of a first point, respectively second point, is determined by the following function: With : • P3D the spatial position of a first point, respectively second point, • K is a camera-associated direction prediction model, and • 0 a projection function in a three-dimensional scene of a pixelp as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, Dmt which is associated with it.
[0013] According to another variant of the method, a spatial position of a third point, respectively a fourth point, is determined by the following function: P'id- Pid+F With : • P'3D the spatial position of a third point, respectively fourth point, • P3D the spatial position of a first point associated with the third point, respectively a second point associated with the fourth, and • F a scene flow determined for a first point, respectively second point.
[0014] According to yet another variant of the method, a position of a pixel of the third image, respectively fourth image, is determined by the following function: Pms = ^(KP'3D) With: • Pms the position of a pixel in the third image, respectively fourth image, • 77 a function to convert from homogeneous coordinates to pixel coordinates by removing one dimension from a vector, • Has a camera-associated direction prediction model, and • P'3D the spatial position of a third point, respectively fourth point.
[0015] According to a further embodiment of the method, the loss error is determined by the following function: L = L2(,P) ) With : • The loss error, • -^1(P) a result of a comparison of a pixel p from the first image to a pixel p from the fourth image, and • ^2(^) a result of a comparison of a pixel p of the second image to a pixel p of the third image.
[0016] According to yet another variant of the method, the loss error is determined by the following function: 0=^,11^,0)-0(^0)11 With : • The loss error, * F(P^ a "scene ux associated with a first or second point P^d, and • a displacement associated with a first or second point P^d.
[0017] According to a further embodiment of the method, the loss error is determined by the following function: With : • The loss error, • ?3D a third, respectively fourth point P^d, and • PiD a second, respectively first point P^u.
[0018] According to a second aspect, the present invention relates to a device configured to learn a scene flow prediction model by a vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.
[0019] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0020] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0021] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0022] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.
[0023] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.
[0024] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency or by beam. self-steering laser or by other means. The computer program according to the present invention can in particular be downloaded onto an Internet-type network.
[0025] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures
[0026] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 3, in which:
[0027] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a a particular and non-limiting example of a realization of the present invention;
[0028] [Fig.2] illustrates a flowchart of the different stages of a method for learning a scene flow prediction model associated with the vision system of [Fig.1], according to a particular and non-limiting embodiment of the present invention; and
[0029] [Fig.3] schematically illustrates a device configured to learn a scene flow prediction model associated with the vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements
[0030] A method and device for learning a scene flow prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle will now be described in what follows with joint reference to Figures 1 to 3. The same elements are identified with the same reference signs throughout the description that follows.
[0031] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0032] Fig. 1 schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.
[0033] The vehicle 10 is located in an environment 1 corresponding, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.
[0034] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, with electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, Vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say a vehicle operating according to a determined level of autonomy or under the total supervision of the driver.
[0035] The vehicle 10 advantageously comprises at least one onboard camera 11, configured to acquire images of a three-dimensional scene unfolding in the environment of the vehicle 10 from a common viewing position. The camera 11 forms a monocular vision system if used alone, as illustrated in [Fig. 1]. However, the present invention is not limited to a monocular vision system comprising a single camera but extends to any vision system comprising at least one camera, for example, 1, 2, 3, or 5 cameras.
[0036] The camera 11 has intrinsic parameters, including: - a focal length, - distortions which are due to imperfections in the optical system of camera 11, - a direction of the optical axis of camera 11, and - a resolution.
[0037] The intrinsic parameters characterize the transformation which associates, for an image point, hereafter called "point", its three-dimensional coordinates in the camera 11 reference frame with the pixel coordinates in an image acquired by the camera 11. These parameters do not change if the camera is moved.
[0038] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in lens curvature, parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are not taken into account, and their correction is addressed during image acquisition or calibration.
[0039] The camera 11 is arranged to acquire an image of a three-dimensional scene from a defined viewpoint, the viewpoint is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10 as illustrated in [Fig.1].
[0040] The camera 11, for example, acquires images of a three-dimensional scene located in front of the vehicle 10, the camera 11 covering an acquisition field 12. An object 13 is placed in the acquisition field 12 of the camera 11, defining an occlusion field 14 for the vision system, an object present in the occlusion field 14 not being observable by camera 11 from its current observation position.
[0041] It is evident that it is possible to use such a vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently, the invention not being limited to the observation of a three-dimensional scene taking place in front of the vehicle 10 carrying the camera 11.
[0042] According to one particular embodiment, the camera 11 is a wide-angle camera, a wide-angle camera being, for example, equipped with a lens designed to acquire a representative image of a three-dimensional scene seen over a wider field of view than a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens makes it possible to capture a larger portion of the three-dimensional scene unfolding in front of or around the camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by this camera. The angle α of the field of view of the camera 11 is, for example, equal to 120°, 145°, 180°, or 360°, whereas a standard camera offers, for example, a field of view open at an angle of 45° or less.Such a camera 11 corresponds, for example, to a camera equipped with mirrors or a "fisheye" camera. Wide-angle lenses have a shorter focal length compared to standard lenses, making them suitable for capturing images of landscapes, architecture, road intersections, or any other subject requiring a wide perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with an extended depth of field.
[0043] An image acquired by the camera 11 at a given acquisition time is in the form of data representing pixels characterized by: - coordinates in the image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric values (from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).
[0044] Each pixel of the acquired image represents an object in the three-dimensional scene present in the field of view of the camera 11. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a point of light resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, captured by a photosensitive sensor of the camera 11 after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image that faithfully represents the physical object observed by the camera. An image point, as previously described, is therefore a point on the surface of an object in the three-dimensional scene.
[0045] When the vehicle 10 is in motion, then two images acquired by the camera 11 at two distinct time instants represent views of the same three-dimensional scene taken from different viewpoints or observation positions, the observation positions of the camera 11 being distinct. On this three-dimensional scene there are, for example: - buildings; - road infrastructure; - other users or stationary objects, for example a parked vehicle; and / or - other users or moving objects, for example another vehicle, a cyclist or a moving pedestrian.
[0046] According to a particular embodiment, an image acquired by the camera 11 includes a distortion equal to 0.5%, 0.8%, or greater than 1%. The measurement of such distortion corresponds to determining a ratio between: - the maximum spacing of a pixel in the image of a straight line in the first three-dimensional scene whose image is a line touching the longest edge of the first image, either at the center of the image edge or at the corners of the image edge, and - the length of this edge.
[0047] In the world of photography, distortion is commonly considered to be: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • problematic if it is greater than or equal to 1% or more.
[0048] Barrel distortion is characterized by a positive percentage, while crescent distortion is characterized by a negative percentage.
[0049] Each image acquired by the camera 11 is, for example, sent to a processor, for example a computer of a device equipping the vehicle 10, or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10. It is then used during the implementation of a method for determining the depth of one of its pixels and / or for determining a scene stream associated with points whose position is determined from the pixels of the images and the depths. predicted and / or during the implementation of a learning process of a scene flow prediction model associated with this camera of the vision system.
[0050] Figure 2 illustrates a flowchart of the different stages of a method for learning a scene flow prediction model, such a scene flow prediction model being for example associated with the vision system on board the vehicle 10 and including the camera 11 and also being used in a method for determining a scene flow associated with a point whose spatial position is determined as a function of a pixel of an image acquired by a camera and a depth associated with that pixel, according to a particular and non-limiting embodiment of the present invention.
[0051] The learning process 2 is for example implemented by a processor of a device embedded in the vehicle 10 or of the device 3 of the [Fig.3].
[0052] In a first step 21, representative data of a first image and a second image acquired by the camera 11 at two distinct acquisition times are received; that is, a first image acquired by the camera 11 at a first acquisition time and a second image acquired by the camera 11 at a second acquisition time distinct from the first acquisition time are received. Image reception consists of receiving representative image data, which are, for example, recorded in a device embedded in the vehicle 10 and comprising a memory accessible by a computer implementing the learning process 2. The first and second acquisition times are sufficiently close in time so that the observed three-dimensional scene is the same but observed from two different viewpoints when the vehicle 10 is in motion.The first and second images are acquired at time intervals on the order of one or more milliseconds.
[0053] In a second step 22, first depths associated with pixels of the first image, called first pixels, are predicted by the depth prediction model and, similarly, second depths associated with pixels of the second image, called second pixels, are predicted by the depth prediction model.
[0054] Such a depth prediction model is known to those skilled in the art, is implemented by a convolutional neural network, and is described, for example, in the following documents: - “Digging Into Self-Supervised Monocular Depth Estimation” by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow, published in August 2019, and - “HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation” by Xiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong, Lina Liu, Yong Liu, Xinxin Chen and Yi Yuan published in December 2020.
[0055] It should be noted that the accuracy of the prediction of the first and second depths depends on the maturity of the learning of the depth prediction model.
[0056] In a third step 23, two point clouds are generated, these clouds comprising points positioned in the three-dimensional scene observed at different time instants of acquisition: • a first point cloud containing initial points, • a second point cloud containing second points.
[0057] The first points are generated from the first image, a spatial position of a first point being a function of the coordinates of a first pixel in the first image associated with this first point and of the first depth associated with the first pixel, and the second points are generated similarly from the second image, a spatial position of a second point being a function of the coordinates of a second pixel in the second image associated with this second point and of the second depth associated with the second pixel.
[0058] Similarly, the generation of the second point cloud consists of projecting the pixels of the second image into the three-dimensional scene and into the frame of the camera 11 at the second time instant of acquisition, which is different from the frame of the camera 11 at the first time instant of acquisition when the vehicle 10 is in motion, the camera 11 having moved in the three-dimensional scene between the first and second time instants of acquisition.
[0059] According to a particular embodiment, a spatial position of a first point, respectively second point, is determined by the following function: [Math.l]
[0060] With: • P^D the spatial position of a first point, respectively second point, • K is a direction prediction model associated with camera 11, and • 0 a projection function in a three-dimensional scene of a pixel as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, which is associated with it.
[0061] According to a first variant, the direction prediction model corresponds to an intrinsic matrix of the camera 11. This first variant is particularly applicable to a pinhole camera and calibrated.
[0062] According to a second variant, the direction prediction model corresponds to a learned model implementing a different neural network. Such a direction prediction model is known to those skilled in the art; it is notably presented in the paper "Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion" written by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich, and Adrien Gaidon and published in August 2020. This second variant is suitable for wide-angle or fisheye cameras, or even for uncalibrated cameras. According to this second variant, a direction prediction step associated with the pixels of the first image and the pixels of the second image is necessary.
[0063] In a fourth step 24, a scene flow is predicted by the scene flow prediction model for each first point and for each second point.
[0064] A scene flow prediction model is known to those skilled in the art and is presented in particular in the document "SeFlow: A Self-Supervised Scene Flow Method in Autonomous Driving" written by Qingwen Zhang, Yi Yang, Peizheng Li, Olov Andersson and Patrie Jensfelt and published in July 2024.
[0065] The scene flow associated with a first point then allows the position of an image point of the first point to be predicted at the second time instant of acquisition, and the scene flow associated with a second point then allows the position of an image point of the second point to be predicted at the first time instant of acquisition. The accuracy of the predicted scene flow increases with its learning level; that is, the accuracy of the scene flow increases after several iterations of this learning process.
[0066] In a fifth step 25, spatial positions of third and fourth points are determined. A spatial position of a third point is determined from a spatial position of a first point associated with it and a scene flow associated with the first point; that is, the third point is the image point of the first point whose spatial position is determined using the scene flow associated with the first point. Similarly, a spatial position of a fourth point is determined from a spatial position of a second point associated with it and a scene flow associated with the second point; that is, the fourth point is the image point of the second point whose spatial position is determined using the scene flow associated with the second point.
[0067] In other words, a third point cloud comprising the third points is the image of the first point cloud predicted by the flow prediction model of scene and a fourth point cloud including the fourth points is the image of the second point cloud predicted by the scene flow prediction model.
[0068] According to a particular embodiment, a spatial position of a third point, respectively fourth point, is determined by the following function: [Math.2]
[0069] With: • P'zi) the spatial position of a third point, respectively a fourth point, • P3D the spatial position of a first point associated with the third point, respectively a second point associated with the fourth, and • F a scene flow determined for a first point, respectively second point.
[0070] In a sixth step 26, the scene flow prediction model is learned according to one of the following variants.
[0071] According to a first variant, the scene flow prediction model is learned by minimizing a loss error determined by comparison: • from the second image to a third image, and • from the first image to a fourth image.
[0072] In a substep 261a, a third image and a fourth image are generated from the third and fourth points respectively.
[0073] The third image is generated by projecting the third points onto an image plane associated with camera 11 at the second acquisition time instant, and the fourth image is generated by projecting the fourth points onto an image plane associated with camera 11 at the first acquisition time instant. In theory, if the first and second depths are perfectly predicted, and if the scene flows are also perfectly predicted, the first and fourth images are identical, as are the second and third images. However, since the predictions are imperfect, comparisons of the first image with the fourth image and of the second image with the third image allow for the evaluation of a prediction error.
[0074] According to a particular embodiment, the position of a pixel in the third image, or fourth image respectively, is determined by the following function: [Math.3] Pms = ^(KP^D)
[0075] With: • Pms the position of a pixel in the third image, respectively the fourth image, • 71 a function to go from homogeneous coordinates to pixel coordinates by removing one dimension of a vector, • K is a direction prediction model associated with camera 11, and • P " yo the spatial position of a third point, respectively fourth point.
[0076] In a substep 262a, the scene flow prediction model is learned by minimizing a loss error determined from: • a comparison of the third image to the second image, and • a comparison of the fourth image to the first image.
[0077] According to one particular embodiment, the depth prediction model is also learned during this substep 262a.
[0078] According to a particular embodiment, the loss error includes photometric errors determined by the following function: [Math.4] H(P) = E4 (!-«)• \Hp)-HP) 1 +« • ( i-iSSIM(Hp). ï(p) ) ) ]
[0079] With: • L*(p) the photometric error associated with a pixel p defined by its coordinates in the first image, respectively second image • l(p) a pixel value in the first image, respectively second image, * I(p) is a value of pixel p in the fourth image, respectively the third image, • SSIM is a function that takes into account a local structure, and • has a weighting factor that depends in particular on the type of road environment.
[0080] It should be noted that a photometric error is lower the more accurate the depth prediction and scene flow prediction models are, i.e., the more advanced or even successful their training is. Conversely, if one of the depth prediction or scene flow prediction models lacks accuracy, for example because its training is insufficient, then a photometric error is not negligible.
[0081] According to another particular embodiment, the loss error includes reconstruction errors determined by the following function: [Math.5] = ^( WJ
[0082] With: • ^smooth^P) a reconstruction error for a pixel p of the third image, respectively fourth image, • a first depth, respectively second depth, associated with pixel p.
[0083]
[0084]
[0085]
[0086]
[0087]
[0088] • VF is a parameter matrix, • 0 the order of a smoothing gradient, • an L1 norm of the second-order depth gradients is calculated with VF = 1, et° = 2, • x and y are the dimensions of the third and fourth images, • P is a hyperparameter dependent on the road environment in which vehicle 10 is traveling, and • l^P) a value of pixel p in the third image, respectively fourth image. This function is generally used to handle discontinuity at the edge of objects (in English, "edge aware smoothness"). According to a particular implementation example, the loss error is determined by the following function: [Math.6] p), L2(p) ) With : • The loss error, • L^p) a result of a comparison of a pixel p from the first image to a pixel p from the fourth image, and • L2{p) a result of a comparison of a pixel p of the second image to a pixel p of the third image. According to another specific implementation example, the loss error is determined by the following function: [Math.7] [Math.7] ^—ipL'smoMî ( P ) [Math.7] (P) With : • The loss error, • L^p) a result of a comparison of a pixel p from the first image to a pixel p from the fourth image, • L2(p) is the result of a comparison of a pixel p from the second image to a pixel ? from the third image. • a reconstruction error for one pixel of the third image, and • ^smooth^P) a reconstruction error for one pixelp of the fourth image.
[0089] Learning the scene flow prediction model and, where applicable, the depth prediction model consists of adjusting the input parameters of the convolutional neural networks implementing the scene flow prediction model and the depth prediction model in order to minimize the previously calculated loss error.
[0090] According to a second variant, the scene flow prediction model is learned by comparing the first and second point clouds to the fourth and third point clouds.
[0091] In a substep 261b, a displacement is determined for each first point and for each second point associated with the same object of the three-dimensional scene by comparing the spatial positions of these each first point and each second point.
[0092] Indeed, when the depth prediction model is accurate, a second point corresponds to the theoretical position of a third point. Since the position of the second point is obtained from the second image, which represents the scene as it actually is at the second time instant of acquisition and not as it is predicted, the second point corresponds to an annotated value for the scene flow prediction model. Similarly, a first point corresponds to the theoretical position of a fourth point. Since the position of the first point is obtained from the first image, which represents the scene as it actually is at the first time instant of acquisition and not as it is predicted, the first point corresponds to an annotated value for the scene flow prediction model.
[0093] In a substep 262b, by minimizing a loss error determined by comparing the predicted scene flows, associated with the first and second points, to displacements determined for the first and second points.
[0094] The scene flow prediction model is thus learned by minimizing a loss error determined by one of the following functions: [Math.8]
[0095] With: • The loss error, * F(P3D) a "ux scene" associated with a first or second point Pw, and * I)(P3D} a displacement associated with the first or second point P31), or [Math.9]
[0096] With: • The loss error, • ?3D a third, respectively fourth point P^, and • P3D a second, respectively first point P>d.
[0097] It is noted here that these two functions are equivalent, the positions of the third and fourth points integrating the scene flow and the position of a second point or a first point in the last equation amounting to determining the displacement for a first point, respectively second point.
[0098] Once the learning of the scene flow prediction model and, where applicable, of the depth prediction model has been completed, for example when the loss error is less than a threshold value defining an acceptable level of accuracy, it is then possible to use the depth prediction model to predict depths associated with pixels of images acquired by the camera 11 and the scene flow prediction model to predict the position of points corresponding to objects of the three-dimensional scene.
[0099] Each determined depth then corresponds to a distance separating the vehicle 10, or a part of the vehicle 10, from an object in the three-dimensional scene to which a pixel is associated. Determining the depth of a pixel then corresponds to measuring the distance separating an object from the vehicle carrying the vision system. Similarly, the position of a point predicted by the scene flow prediction model also makes it possible to determine a distance separating the vehicle 10 or the camera 11 from an object in the three-dimensional scene to which that point corresponds.
[0100] If the ADAS uses predicted point positions, depths, or distances as input data to determine the distance between a part of the vehicle 10, for example, the front bumper, and another road user, the ADAS is then able to determine this distance precisely. For example, if the ADAS is designed to activate the braking system of the vehicle 10 in the event of a risk of collision with another road user, and the distance between the vehicle 10 and that road user decreases sharply, then the ADAS is able to detect this sudden closeness and activate the braking system of the vehicle 10 to avoid a potential accident.
[0101] Figure 3 schematically illustrates a device configured to learn a scene flow prediction model and / or to learn a depth prediction model associated with a vision system embedded in a vehicle and / or to predict a point position or a depth associated with a pixel of an image. acquired by a camera of the vehicle's on-board vision system, according to a particular and non-limiting embodiment of the present invention. Device 3 corresponds, for example, to a device on-board in the vehicle 10, for example, a computer associated with the monocular vision system.
[0102] Device 3 is, for example, configured to carry out the steps described opposite Figures 1 and 2. Examples of such a device 3 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 3, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 3 may be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.
[0103] The device 3 comprises one (or more) processor(s) 30 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 3. The processor 30 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 3 further comprises at least one memory 31, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.
[0104] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 31.
[0105] According to various particular and non-limiting embodiments, the device 3 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.
[0106] According to a particular and non-limiting embodiment, the device 3 comprises a block 32 of interface elements for communicating with external devices. The interface elements of the block 32 comprise one or more of the following interfaces: - Radio frequency (RF) interface, for example Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HD MI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").
[0107] According to another particular and non-limiting embodiment, the device 3 includes a communication interface 33 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 330. The communication interface 33 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 330. The communication interface 33 corresponds, for example, to a wired network of the CAN (Controller Area Network) type, CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3).
[0108] According to a particular and non-limiting embodiment, the device 3 can provide output signals to one or more external devices, such as a display screen 340, touch or not, one or more speakers 350 and / or other peripherals 360 via the output interfaces 34, 35, 36 respectively. According to a variant, one or more of the external devices is integrated into the device 3.
[0109] Of course, the present invention is not limited to the embodiments described above but extends to a method for predicting the position of a point in a three-dimensional scene from an image acquired by a vision system, the position of the point being predicted by a scene flow prediction model learned according to the learning method described above, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0110] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 3 of [Fig.3].
Claims
1. Demands Method for learning a scene flow prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle (10), the vision system comprising a camera (11) arranged to acquire an image of a three-dimensional scene taking place near the vehicle (10), said method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception (21) of data representative of a first image and a second image respectively acquired by the camera (11) at a first time instant of acquisition and at a second time instant of acquisition distinct from the first time instant of acquisition; - prediction (22) of first depths associated with pixels of the first image, called first pixels, and of second depths associated with pixels of the second image, called second pixels, the first and second depths being predicted by a depth prediction model; - generation (23): • first points from the first image, a spatial position of a first point being a function of the coordinates of a first pixel in the first image associated with said first point and of said first depth associated with said first pixel, and • second points from the second image, a spatial position of a second point, being a function of the coordinates of a second pixel in the second image associated with said second point and of said second depth associated with said second pixel; - prediction (24) of a scene flow associated with the first points and of a scene flow associated with the second points, the scene flows being predicted by the scene flow prediction model; - determination (25) of spatial positions of third points and fourth points, a spatial position of a third point being determined from a spatial position of a first point associated with it and a scene flow associated with said first point, and a spatial position of a fourth point being determined from a spatial position of a second point associated with it and of a scene flow associated with said second point; - learning (26) of the scene flow prediction model by minimizing a loss error determined by comparison: • of the second image to a third image generated by projection of the third points into an image plane associated with the camera (11) at the second time instant of acquisition, and of the first image to a fourth image generated by projection of the fourth points into an image plane associated with the camera (11) at the first time instant of acquisition, or • of scene flows associated with the first and second points to displacements determined for the first and second points, a displacement being determined for each first point and for each second point associated with the same object of the three-dimensional scene by comparison of the spatial positions of said each first point and said each second point.
2. A method according to claim 1, wherein a spatial position of a first point, respectively second point, is determined by the following function: With: • P^d the spatial position of a first point, respectively second point, • K a direction prediction model associated with the camera (11), and • 0 a projection function in a three-dimensional scene of a pixelp as a function of its coordinates in the first image, respectively second image and of the first depth, respectively second depth, Dmt associated with it.
3. A method according to claim 1 or 2, wherein a spatial position of a third point, respectively a fourth point, is determined by the following function: P^P^d + F With: • P'^d the spatial position of a third point, respectively a fourth point, • P^d the spatial position of a first point associated with the third point, respectively a second point associated with the fourth, and • F a scene flow determined for a first point, respectively second point.
4. A method according to any one of claims 1 to 3, wherein a position of a pixel of the third image, respectively fourth image, is determined by the following function: With: • Pms the position of a pixel of the third image, respectively fourth image, • Æ a function for going from homogeneous coordinates to pixel coordinates by removing one dimension of a vector, • K a direction prediction model associated with the camera (11), and • P'^d the spatial position of a third point, respectively fourth point.
5. A method according to any one of claims 1 to 4, wherein the loss error is determined by the following function: L^pMintL^ ) With: • L the loss error, • ^i(P) a result of a comparison of a pixel of the first image to a pixel p of the fourth image, and • L2(jP) a result of a comparison of a pixel p of the second image to a pixel p of the third image.
6. A method according to any one of claims 1 to 4, wherein the loss error is determined by the following function: With: • L the loss error, * F(Pd) a scene variable associated with a first or second point P3D, and * a displacement associated with a first or second point P3D-
7. A method according to any one of claims 1 to 4, wherein the loss error is determined by the following function: 4- = 1.11^^-^11 Where: • £ the loss error, • ^3D a third, respectively fourth point P,d, and • ?3D a second, respectively first point P^.
8. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by a processor.
9. Device (3) configured to learn a scene flow prediction model and a depth prediction model by a vision system embedded in a vehicle (10), said device (3) comprising a memory (31) associated with at least one processor (30) configured for the implementation of the steps of the method according to any one of claims 1 to 7.
10. Vehicle (10) comprising the device (3) according to claim 9.