Method and device for learning a depth prediction model associated with a stereoscopic vision system by comparing positions of points in a three-dimensional scene.

The method improves depth prediction in vehicle vision systems by using a convolutional neural network with stereoscopic vision, addressing data requirements and occlusion issues to enhance ADAS precision and safety.

FR3160789A1Active Publication Date: 2025-10-03STELLANTIS AUTO SAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
FR2024003030
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-26
Publication Date
2025-10-03
Estimated Expiration
2044-03-26

AI Technical Summary

Technical Problem

Existing depth prediction models for vehicle vision systems face challenges due to the need for massive data accumulation and the presence of occluded objects in stereoscopic vision, leading to errors and unsuitability for various configurations, particularly affecting the precision of ADAS systems.

Method used

A method for learning a depth prediction model using a convolutional neural network with a stereoscopic vision system, involving image and depth map processing, visibility masking, and minimizing loss errors based on consistency and photometric comparisons to improve depth estimation accuracy.

Benefits of technology

Enhances the quality of depth prediction data for ADAS systems, improving road safety by providing precise depth measurements for vehicle operations, especially in environments with occluded objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method or device for learning a depth prediction model implemented by a neural network associated with a vision system. Indeed, the method comprises the generation (32) of depth maps associated with images acquired by each camera of the vision system and the generation (34) of other images and other depth maps from the received images and the depth maps by change of reference frame. A visibility mask is determined (35) from the generated images and a zero value is assigned (36) to the pixels of the generated images not included in the visibility mask. The depth prediction model is then learned (38) by minimizing a loss error determined by comparing positions of projected points (33, 37) in the three-dimensional scene from the different depth maps, the points corresponding to pixels. Figure for abstract: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for learning a depth prediction model associated with a stereoscopic vision system by comparing positions of points in a three-dimensional scene. Technical field

[0001] The present invention relates to methods and devices for determining a depth by a vision system on board a vehicle, for example in a motor vehicle. The present invention also relates to a method and a device for measuring a distance separating an object from a vehicle carrying a vision system. Technological background

[0002] Many modern vehicles are equipped with so-called AD AS (Advanced Driver-Assistance System). Such AD AS systems are passive and active safety systems designed to eliminate the element of human error in the driving of vehicles of all types. AD AS use advanced technologies to assist the driver while driving and thus improve their performance. AD AS use a combination of sensor technologies to perceive the environment around a vehicle, then provide information to the driver or act on certain vehicle systems.

[0003] There are several levels of ADAS, such as rearview cameras and blind spot sensors, lane departure warning systems, adaptive cruise control and automatic parking systems.

[0004] The AD AS embedded in a vehicle are supplied with data obtained one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or possible obstacles present around a vehicle in order, for example: • to adapt the vehicle's lighting according to the presence of other users; • automatically regulate the vehicle speed; • to act on the braking system in the event of a risk of impact with an object.

[0005] The prediction of a distance separating an object from a vehicle, that is to say the depth associated with a pixel of an image acquired by a camera, the pixel representing the object, requires great precision in order to feed the AD AS. The prediction of a depth associated with a pixel notably calls upon a model of depth prediction, the depth prediction model being adapted to the vehicle's onboard vision system. Such a prediction model is commonly implemented by a neural network, which is notably learned in a training phase in order to predict accurate depths in relation to the onboard vision system and the type of environment in which the vehicle operates. A training phase then requires data, for example annotated data from libraries or annotated via other more precise onboard systems such as a LIDAR®. However, this type of learning requires a massive accumulation of data or the presence of this other onboard system which the vehicle does not always have.It is then possible to carry out a learning phase of the depth prediction model from data acquired by the on-board vision system, but existing solutions are not suitable for the configuration of any on-board vision system in a vehicle.

[0006] Furthermore, objects present in a three-dimensional scene observed by a stereoscopic vision system, i.e. a vision system comprising at least two cameras, may not be visible on certain images acquired by the stereoscopic vision system, these objects then being occluded from the point of view of a camera of the stereoscopic vision system. The presence of occluded objects is then likely to generate errors and distort the learning of a depth prediction model. Summary of the present invention

[0007] An object of the present invention is to solve at least one of the problems of the technological background described above.

[0008] Another object of the present invention is to improve the quality of the data resulting from the processing of an image acquired by a vision system, in particular by a depth prediction model implemented by a neural network associated with a stereoscopic vision system.

[0009] Another object of the present invention is to improve road safety, in particular by improving the operational safety of AD AS systems supplied by data obtained from a camera, in particular a wide-angle camera.

[0010] According to a first aspect, the present invention relates to a method for learning a depth prediction model implemented by a convolutional neural network associated with a stereoscopic vision system on board a vehicle, the stereoscopic vision system comprising a first camera and a second camera arranged so as to each acquire an image of a three-dimensional scene from a different point of view, said method being implemented by at least one processor, and being characterized in which includes the following steps: - reception of a first image and a second image acquired by the first camera and the second camera respectively at the same acquisition time instant; - generating a first depth map comprising depths associated with a set of pixels of the first image and a second depth map comprising depths associated with a set of pixels of the second image, the depths being predicted with the depth prediction model from the first and second images; - determining three-dimensional positions of points of a first set of points from the first depth map, each point of the first set of points corresponding to a pixel of the first depth map, and determining three-dimensional positions of points of a second set of points from the second depth map, each point of the second set of points corresponding to a pixel of the second depth map; - generating a third image and a third depth map from the first image, the first depth map and extrinsic parameters of the stereoscopic vision system, the third depth map comprising depths associated with a set of pixels of the third image, and generating a fourth image and a fourth depth map from the second image, the second depth map and extrinsic parameters of the stereoscopic vision system, the fourth depth map comprising depths associated with a set of pixels of the fourth image; - determining a visibility mask as the union of a first visibility mask and a second visibility mask, each pixel of the third image resulting from at least one pixel of the first image belonging to the first visibility mask and each pixel of the fourth image resulting from at least one pixel of the second image belonging to the second visibility mask; - assigning a zero value to the pixels of the third and fourth images not included in the visibility mask; - determining three-dimensional positions of points of a third set of points from the third depth map, each point of the third set of points corresponding to a pixel of the third depth map, and determining three-dimensional positions of points of a fourth set of points from the fourth depth map, each point of the fourth set of points corresponding to a pixel of the fourth depth map; - learning the depth prediction model by minimizing a loss error, the loss error being determined from: • a first consistency error determined by comparing three-dimensional positions of points of the first and fourth sets of points, • a second consistency error determined by comparing three-dimensional positions of points of the second and third sets of points, and • a photometric error determined by comparing the first and fourth images and by comparing the second and third images.

[0011] According to a variant, the first consistency error and the second consistency error are determined respectively by the following function: l^p)=^(**(p)-^(p) )2+ (y^p) )2+ Mp)r With : • Lc^p) corresponding to the first consistency error Lc^p) for a pixel P of the first depth map, respectively the second consistency error For a pixel of the second depth map, a pixel P being defined by coordinates in two dimensions; • p) a component along a first axis of a three-dimensional position of a point of the first set of points corresponding to the pixel P of the first depth map noted xf ( p ), respectively of the second set of points corresponding to the pixel P of the second depth map noted x2 ( p ); • x'4 p) a component along a first axis of a three-dimensional position of a point of the fourth set of points corresponding to the pixel P of the fourth depth map noted x\(p), respectively of the third set of points corresponding to the pixel P of the third depth map noted x \ ( p ); • y^p) a component along a second axis of a three-dimensional position of a point of the first set of points corresponding to pixel P of the first map of depths noted y ( p ), respectively of the second set of points cor corresponding to pixel P of the second depth map noted y (p}; * y*(p) a component along a second axis of a three-dimensional position of a point from the fourth set of points corresponding to the pixel P of the fourth depth map noted y' (p), respectively from the third set of points corresponding to the pixel P of the third depth map noted y ( p ); • zUp) a component along a third axis of a three-dimensional position of a point of the first set of points corresponding to the pixel P of the first depth map noted z^p), respectively of the second set of points corresponding to the pixel P of the second depth map noted z7(p); and * Z«(p) a component along a third axis of a three-dimensional position of a point of the fourth set of points corresponding to the pixel P of the fourth depth map noted z\(p), respectively of the third set of points corresponding to the pixel P of the third depth map noted z'2(p) ■

[0012] According to another variant, the third set of pixels comprises only pixels of the third depth map included in the visibility mask and the fourth set of pixels comprises only pixels of the fourth depth map included in the visibility mask.

[0013] According to yet another variant, the loss error is determined by the following function: L = L^min(Lcï(p), Lc2(p))+ min(Lpl(p), Lp2(P) )) With : • The loss error, • Lpi(jp) the first photometric error for a pixel P defined by its coordinates in two dimensions, • Lpi the second photometric error for pixel P, • Lcl the first consistency error for pixel P, and • Lc2 the second consistency error for pixel P.

[0014] According to an additional variant, the determination of a three-dimensional position of a point corresponding to a pixel is obtained by the following function: P3o=^(pF3-d(p,)) With : • Three-dimensional position of a point from the first, second, third and respectively fourth set of points, • K' a direction prediction model associated with the first camera, respectively with the second camera, • 0 a projection function in the three-dimensional scene of a pixel P, according to its coordinates and the depth p^p ) which is associated with it in the first, second, third and respectively fourth depth map.

[0015] According to another variant, the direction prediction model is implemented by a neural network.

[0016] According to yet another variant, the third and fourth depth maps are generated by the following function: Ps = T^p^K^, D(pf) ) ] ) With : • Ps the coordinates of a pixel of the third depth map, respectively of the fourth depth map, • n a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • K a direction prediction model associated with the second camera, respectively with the first camera, • K' a direction prediction model associated with the first camera, respectively with the second camera, • T an extrinsic matrix of the stereoscopic vision system, and • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates in the first image, respectively second image, and of the depth j^pj which is associated with it in the first depth map, respectively second depth map.

[0017] According to a second aspect, the present invention relates to a device configured to learn a depth prediction model by a vision system on board a vehicle, the device comprising a memory associated with at least one processor configured to implement the steps of the method according to the first aspect of the present invention.

[0018] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0019] According to a fourth aspect, the present invention relates to a computer program which comprises instructions adapted for executing the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0020] Such a computer program may use any programming language and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0021] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to the first aspect of the present invention.

[0022] On the one hand, the recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM memory, a CD-ROM or a microelectronic circuit type ROM memory, or a magnetic recording means or a hard disk.

[0023] On the other hand, this recording medium can also be a trans medium missible such as an electrical or optical signal, such a signal being able to be conveyed via an electrical or optical cable, by conventional or hertzian radio or by self-directed laser beam or by other means. The computer program according to the present invention can in particular be downloaded from an Internet-type network.

[0024] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to perform or to be used in performing the method in question. Brief description of the figures

[0025] Other characteristics and advantages of the present invention will emerge from the description of the particular and non-limiting exemplary embodiments of the present invention below, with reference to the appended figures 1 to 4, in which:

[0026] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting example of embodiment of the present invention;

[0027] [Fig.2] illustrates a flowchart of the different stages of a determination process mination of a depth of a pixel of an image by a depth prediction model associated with a vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;

[0028] [Fig.3] illustrates a flowchart of the different stages of a learning process of the depth prediction model used in the method of [Fig.2], according to a particular and non-limiting exemplary embodiment of the present invention; and

[0029] [Fig.4] schematically illustrates a device configured to learn a model depth prediction by a vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention. Description of examples of implementation

[0030] A method and a device for learning a depth prediction model implemented by a convolutional neural network associated with a stereoscopic vision system on board a vehicle will now be described in what follows with joint reference to Figures 1 to 4. The same elements are identified with the same reference signs throughout the description which follows.

[0031] The terms "first(s)", "second(s)" (or "first(s)", "second(s)"), etc. are used in this document by arbitrary convention to enable different elements (such as operations, means, etc.) implemented in the embodiments described below to be identified and distinguished. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0032] For the entire description, reception of an image means the reception of data representative of an image. Similarly, for the generation of an image, we mean the generation of data representative of an image and for the generation of a depth map, the generation of data representative of a depth map. These shortcuts are only intended to simplify the description, however, since the processes are implemented by one or more processors, it is obvious that the input and output data of the different stages of a process are computer data.

[0033] According to a particular and non-limiting example of embodiment of the present invention, a depth prediction model associated with a stereoscopic vision system comprising two cameras is learned in a learning phase.

[0034] Indeed, the method relating to the learning phase comprises the generation of depth maps associated with images acquired by each camera of the vision system and the generation of other images and other depth maps from the received images and the depth maps by change of reference frame.

[0035] A visibility mask is determined from the generated images and a zero value is assigned to the pixels of the generated images not included in the visibility mask.

[0036] The depth prediction model is then learned by minimizing a loss error determined by comparing positions of points projected into the three-dimensional scene from the different depth maps, the points corresponding to pixels.

[0037] [Fig. 1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.

[0038] Such an environment 1 corresponds, for example, to a road environment formed of a network of roads accessible to the vehicle 10.

[0039] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, with an electric motor(s) or even a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as an automobile, a truck, a bus, a motorcycle. Finally, the vehicle 10 corresponds to an autonomous vehicle or not, that is to say a vehicle traveling according to a determined level of autonomy or under the total supervision of the driver.

[0040] The vehicle 10 advantageously comprises at least two on-board cameras, a first camera 11 and a second camera 12, configured to acquire images of a three-dimensional scene taking place in the environment of the vehicle 10 from separate observation positions. The first camera 11 and the second camera 12 form a stereoscopic vision system when used together as illustrated in [Fig.l]. The first camera 11 forms a monoscopic vision system when used alone, likewise the second camera 12 forms another monoscopic vision system when used alone. The present invention, however, extends to any vision system comprising at least two cameras, for example 2, 3 or 5 cameras.

[0041] The intrinsic parameters of the first camera 11 characterize the transformation which associates, for an image point, subsequently called “point”, its three-dimensional coordinates in the reference frame of the first camera 11 with the pixel coordinates in an image acquired by the first camera 11. These parameters do not change if the first camera 11 is moved. The intrinsic parameters of the first camera 11 include in particular a first focal distance fl associated with the first camera 11.

[0042] The intrinsic parameters of the second camera 12 characterize, for their part, the transformation which associates, for an image point, its three-dimensional coordinates in the reference frame of the second camera 12 with the pixel coordinates in an image acquired by the second camera 12. These parameters do not change if the second camera 12 is moved. The intrinsic parameters of the second camera 12 include in particular a second focal length f2 associated with the second camera 12.

[0043] The distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions which generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, that is to say that the distortions are not taken into account, that their correction is processed at the time of image acquisition or at the time of calibration.

[0044] These two cameras 11, 12 are arranged so as to each acquire an image of a scene from a different point of view, the first point of view is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10, the second point of view is for example located on or in the right rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. In the case where the two cameras are located at the top of the windshield of the vehicle, they are then placed at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, the second camera 12 is located in the right rearview mirror of the vehicle 10.

[0045] A first marker is associated with the first camera 11: - the direction of the x axis is defined horizontal and normal to the optical axis Cl of the first camera 11. The distance B separating the optical center of the first camera 11 the projection of the optical center of the second camera 12 onto the horizontal plane passing through the optical center of the first camera 11 is called the reference base (in English “baseline”); - the direction of the y axis is defined vertical and normal to the optical axis Cl of the first camera 11; - the direction of the z axis is defined orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal reference frame.

[0046] The optical axis C1 of the first camera 11 and the optical axis C2 of the second camera 12 are not necessarily parallel or even included in the same plane.

[0047] According to a variant, optical axes of the first and second cameras are not coplanar.

[0048] The extrinsic parameters linked to the position of the cameras 11, 12 are the following parameters: - three translations in the x, y and z directions: Tx, Ty and Tz constituting the translation vector T; and - three rotations in the x, y and z directions: 0x, 0y and 0z.

[0049] An extrinsic matrix of the vision system then includes the previously defined extrinsic parameters.

[0050] The extrinsic parameters are determined, for example, during a calibration phase of the stereoscopic vision system comprising the first camera 11 and the second camera 12.

[0051] A main constraint of the stereoscopic vision system used in automobiles is, for example, the large distance between the two cameras. Indeed, to be able to cover a measurement range of 200 meters, the reference base must reach 60cm for the cameras commonly used in this field.

[0052] The two cameras 11, 12 acquire images of a scene located in front of the vehicle 10, the first camera 11 covering only a first acquisition field 13, the second camera 12 covering only a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic vision of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic vision of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic vision of the scene by the stereoscopic vision system composed of the two cameras 11, 12.

[0053] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.

[0054] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system comprising the second camera 12.

[0055] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable using the monoscopic vision system comprising the first camera 11.

[0056] Finally, field 19 is not visible to any of the cameras. The part of the scene present in this field 19 is therefore not observable.

[0057] According to a particular exemplary embodiment, the field of vision of the second camera 12 covers at least half of the field of vision of the first camera 11.

[0058] It is obvious that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with differently placed and oriented cameras.

[0059] The images acquired by the cameras 11, 12 at an acquisition time instant are presented in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colors and brightness of objects in the observed scene in the form, for example, of RGB colorimetric coordinates (from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness).

[0060] Each pixel of the acquired image is representative of an object in the three-dimensional scene present in the camera's field of vision. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a luminous point resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When the light strikes this object, photons are emitted or reflected, captured by a photosensitive sensor of the camera after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image faithfully representing the physical object observed by the camera. An image point previously presented is thus a point on a surface of an object in the three-dimensional scene.

[0061] The images acquired by the cameras 11, 12 represent views of the same scene taken from different viewpoints, the positions of the cameras being distinct. On this scene are found for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.

[0062] According to a particular embodiment, the first camera 11 and / or the second camera 12 is of the “wide-angle” type, a wide-angle camera being for example equipped with a lens designed to acquire an image representative of a three-dimensional scene perceived according to a wider field of vision than that of a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens makes it possible to capture a larger portion of the three-dimensional scene taking place in front of or around the wide-angle camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by this camera. The angle a of the field of vision of the wide-angle camera is for example equal to 120°, 145°, 180° or 360°, whereas a standard camera offers, for example, an open field of vision following an angle of 45° or less.Such a wide-angle camera is, for example, a camera equipped with mirrors or a "fisheye" camera. Wide-angle lenses have a shorter focal length than standard lenses, which makes them suitable for capturing images of landscapes, architecture, road intersections or any other subject requiring a wide perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with an extended depth of field.

[0063] According to a particular embodiment, an image acquired by the first camera 11 and / or an image acquired by the second camera 12 comprises a distortion equal to 0.5%, 0.8% or greater than 1%. The measurement of such a distortion corresponds to the determination of a ratio between: - the maximum spacing of a pixel of the image from a straight line of the first three-dimensional scene whose image is a line touching the longest edge of the first image, either at the center of the edge of the image, or at the corners of the edge of the image, and - the length of this edge.

[0064] Commonly, distortion is considered, in the world of photography, as: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • bothersome if it is greater than or equal to 1% or more.

[0065] A barrel distortion is characterized by a positive percentage, while a crescent distortion is characterized by a negative percentage.

[0066] According to a particular embodiment, a field of vision of the first camera 11 covers at least half of a field of view of the second camera 12 and a field of view of the second camera 12 covers at least half of a field of view of the first camera 11. In other words, more than half of the pixels of an image acquired by the first camera 11 correspond to an object of the three-dimensional scene seen by the second camera 12, pixels of an image acquired by the second camera 12 also corresponding to this object of the three-dimensional scene. Similarly, more than half of the pixels of an image acquired by the second camera 12 correspond to an object of the three-dimensional scene seen by the first camera 11, pixels of an image acquired by the first camera 11 also corresponding to this object of the three-dimensional scene.

[0067] The images acquired by the first camera 11 and by the second camera 12 are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.

[0068] A method for determining a depth by a vision system on board the vehicle 10 is advantageously implemented by the vehicle 10, that is to say by a processor, a computer or a combination of computers of the on-board system of the vehicle 10, for example by the computer(s) in charge of the vision system of the vehicle 10.

[0069] [Fig. 2] illustrates a flowchart of the different steps of a method 2 for determining a depth of a pixel of an image by depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, for example in the vehicle 10 of [Fig. 1], according to a particular and non-limiting exemplary embodiment of the present invention. The method 2 is for example implemented by a device of the vision system embedded in the vehicle 10 or by the device 4 of [Fig. 4].

[0070] In a step 21, data representative of an image acquired by the first camera 11 and of an image acquired by the second camera 12 are received.

[0071] In a step 22, depths associated with a set of pixels of one of the received images are determined by the depth prediction model from the two received images.

[0072] Each determined depth then corresponds to a distance separating the vehicle 10 or a part of the vehicle 10 from an object of the three-dimensional scene with which a pixel is associated, the determination of a depth of a pixel then corresponding to a measurement of a distance separating an object from the vehicle carrying the vision system.

[0073] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another user present on the road, the ADAS is then able to precisely determine this distance. For example, if the ADAS has the function of acting on a braking system of the vehicle 10 in the event of a risk of collision with another road user and the distance separating the vehicle 10 from this same road user decreases significantly, then the ADAS is able to detect this sudden approach and act on the braking system of the vehicle 10 to avoid a possible accident.

[0074] [Fig. 3] illustrates a flowchart of the different steps of a method for learning the depth prediction model used in a method for determining a depth of a pixel of an image, for example in method 2 of [Fig. 2], according to a particular and non-limiting exemplary embodiment of the present invention.

[0075] The learning method 3 is for example implemented by the device on board the vehicle 10 implementing the method for determining a depth by the vision system on board a vehicle or by the device 4 of [Fig.4].

[0076] In a step 31, a first image and a second image are received, the first image being acquired by the first camera 11 at an acquisition time instant and the second image being acquired by the second camera at the same acquisition time instant.

[0077] According to a particular exemplary embodiment, the first image and second image have the same definition, that is to say they comprise the same number of pixels, have the same number of pixels according to their height and the same number of pixels according to their width.

[0078] According to another particular exemplary embodiment, the first image and second image are not of the same definition. An additional step then consists of resizing or cropping them to obtain a first image and a second image of the same definition.

[0079] In a step 32, a first depth map comprising depths associated with a set of pixels of the first image is determined. The depths associated with the pixels of the first image are predicted with the depth prediction model from the first and second images.

[0080] Similarly, a second depth map comprising depths associated with a set of pixels of the second image is determined. The depths associated with the pixels of the second image are also predicted with the depth prediction model from the first and second images.

[0081] The depth prediction model, implemented by a convolutional neural network, is known to those skilled in the art and is for example presented in the document “Unifying Flow, Stereo and Depth Estimation” written by Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao and Andréas Geiger, published in July 2023 suitable for any stereoscopic vision system including those comprising cameras whose optical axes are not included in the same plane.

[0082] Thus, a predicted depth is associated with each pixel of a set of pixels of the first image, the first depth map comprising the coordinates of each pixel of the set of pixels of the first image and a depth associated with this each pixel recorded for example in a correspondence table, for example in a memory accessible to the processor implementing this learning method 3. This correspondence table then contains pairs (coordinates of a pixel of image 1; predicted depth for this pixel).

[0083] According to another particular exemplary embodiment, the depths are recorded in an additional channel associated with each image, thus a depth map and an image are the same object, each pixel having coordinates in two dimensions, one or more values ​​(image) and a depth (depth map).

[0084] Similarly, a predicted depth is associated with each pixel of a set of pixels of the second image, the second depth map comprising the coordinates of each pixel of the set of pixels of the second image and a depth associated with each pixel recorded for example in a correspondence table, for example in a memory accessible to the processor implementing this learning method 3. This correspondence table then contains pairs (coordinates of a pixel of the image 2; predicted depth for this pixel).

[0085] In a step 33, three-dimensional positions of points of a first set of points, which is a virtual point, are determined from the first depth map, each point of the first set of points corresponding to a pixel of the first depth map, and three-dimensional positions of points of a second set of points are determined from the second depth map, each point of the second set of points corresponding to a pixel of the second depth map.

[0086] The three-dimensional position of a point of the first set of points defines spatial coordinates in the three-dimensional scene, in a reference frame associated with the first camera 11, of a point corresponding to a first pixel of the first depth map from the coordinates of the first pixel in the first depth map, i.e. in the first image, and the depth which is predicted for this first pixel, its coordinates and depth being recorded in the first depth map. It should be noted that the coordinates of the first pixel in the first depth map as well as in the first image comprise two components while the spatial coordinates of the point associated with the first pixel comprise three components.

[0087] This step 33 amounts to reprojecting into the three-dimensional scene a pixel of the first image and / or of the first depth map.

[0088] According to a particular exemplary embodiment, the determination of a three-dimensional position of a point of the first set of points corresponding to a pixel of the first depth map is obtained by the following function:

[0089] [Math.l] P3d=^.p^PdM)

[0090] With: • P3D the three-dimensional position of a point in the first set of points, • K' a direction prediction model associated with the first camera 11, • 0 a projection function in the three-dimensional scene of a pixel Pt according to its coordinates and the depth [^p ) associated with it in the first depth map.

[0091] If the depths predicted by the depth prediction model are exact and the direction prediction model is also exact, then the position in the scene of this virtual point is confused with that of a point of an object in the scene, this point of the object in the scene then corresponding to the pixel of the first image of which the virtual point is the image.

[0092] According to a first variant, the direction prediction model associated with the first camera 11 is an intrinsic matrix of the first camera 11. Such a variant is particularly applicable in the case of pinhole cameras generating images with little distortion. An intrinsic matrix is ​​particularly determined during an upstream phase of calibration of the first camera 11 or of the cameras 11, 12 of the stereoscopic vision system.

[0093] According to a second variant, the direction prediction model is implemented by a neural network, this direction prediction model being more suited to wide-angle cameras generating highly distorted images. Such a direction prediction model is known to those skilled in the art; it is notably presented in the document “Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion” written by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich and Adrien Gaidon, published in August 2020. The projections and reprojections are then inverse functions obtained from the direction prediction model and are a function of a depth and pixel coordinates.

[0094] Similarly, the three-dimensional position of a point of the second set of points defines spatial coordinates in the three-dimensional scene, in a reference frame associated with the second camera 12, of a point corresponding to a second pixel of the second depth map from the coordinates of the second pixel in the second depth map, that is to say in the second image, and the depth which is predicted for this second pixel, its coordinates and depth being recorded in the second depth map. It should be noted that the coordinates of the second pixel in the second depth map as well as in the second image comprise two components while the spatial coordinates of the point associated with the second pixel comprise three components.

[0095] This step 33 amounts to reprojecting into the three-dimensional scene a pixel of the second image and / or of the second depth map.

[0096] According to a particular exemplary embodiment, the determination of a three-dimensional position of a point of the second set of points corresponding to a pixel of the second depth map is obtained by the following function:

[0097] [Math.2] P3D=^P^DM)

[0098] With: • P^d the three-dimensional position of a point in the second set of points, • K' a direction prediction model associated with the second camera 12, • 0 a projection function in the three-dimensional scene of a pixel Pt according to its coordinates and the depth associated with it in the second depth map.

[0099] If the depths predicted by the depth prediction model are exact and the direction prediction model is also exact, then the position in the scene of this virtual point is confused with that of a point of an object in the scene, this point of the object in the scene then corresponding to the pixel of the second image of which the virtual point is the image.

[0100] As previously, according to the first variant, the direction prediction model associated with the second camera 12 is an intrinsic matrix of the second camera 12 or, according to a second variant, the direction prediction model is implemented by a neural network.

[0101] In a step 34, a third image and a third depth map are generated from the first image, the first depth map and extrinsic parameters of the stereoscopic vision system, the third depth map comprising depths associated with a set of pixels of the third image.

[0102] The generation of the third depth map from the first image, The first depth and extrinsic parameter map of the stereoscopic vision system consists of: • determining spatial coordinates in the three-dimensional scene, in a reference system associated with the first camera 11, of a point corresponding to a first pixel of the first depth map from the coordinates of the first pixel in the first image and the depth which is associated with this first pixel, its coordinates and depth being recorded in the first depth map. It should be noted that the coordinates of the first pixel in the first image comprise two components while the spatial coordinates of the point associated with the first pixel comprise three components; • the determination of spatial coordinates in the three-dimensional scene of the point associated with the first pixel in a reference frame associated with the second camera 12, in other words a change of reference frame from that associated with the first camera 11 to that associated with the second camera 12, using the extrinsic parameters of the stereoscopic vision system; and • the projection of the point associated with the pixel in the image plane of the second camera 12, this image plane corresponding to that of the second image, making it possible to determine the arrival coordinates of the first pixel of the first image in an image such as the second camera 12 would have acquired it. The image plane of a camera corresponds to a plane defined in the frame of reference of the camera, normal to the optical axis of the camera and located at the first focal length of the camera. The third depth map then includes the arrival coordinates of the first pixels and the depth associated with the first pixel.

[0103] The arrival coordinates of a pixel in the third depth map are thus determined and are the same as the arrival coordinates of a pixel in the third image. In the third depth map, the depth of a pixel corresponds to a depth determined from pixel depths of the first depth map, while in the third image, the value of a pixel corresponds to a value determined from pixel values ​​of the first image.

[0104] Similarly, a fourth image and a fourth depth map are generated from the second image, the second depth map and the extrinsic parameters of the stereoscopic vision system, the fourth depth map comprising depths associated with a set of pixels of the fourth image. The fourth depth map then comprises the arrival coordinates of second pixels of the second depth map and the depth associated with the second pixel.

[0105] According to a particular exemplary embodiment, the third depth map is generated by the following function: [Math.3] ps = 7r(K[T^(p)iK'-\ D(pf) ) ] )

[0106] With: • Ps the coordinates of a pixel of the third depth map, •77 a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • K' a direction prediction model associated with the first camera 11, • K a direction prediction model associated with the second camera 12, • T an extrinsic matrix of the stereoscopic vision system, and • 0 a projection function in the three-dimensional scene of a first pixel Pt as a function of its coordinates in the first image and the depth L^p^ associated with it in the first depth map. This function thus determines the arrival coordinates of a pixel of an image acquired by the first camera 11 in an image as it would be acquired by the second camera 12.

[0107] The determination of the fourth depth map is similar to the determination of the third depth map, the cameras being in particular inverted. The arrival coordinates of a pixel in the fourth depth map are thus determined and are the same as the arrival coordinates of a pixel in the fourth image. In the fourth depth map, the depth of a pixel corresponds to a depth determined from pixel depths of the second depth map, while in the fourth image, the value of a pixel corresponds to a value of a pixel determined from pixel values ​​of the second image.

[0108] Thus, the previous function for determining the arrival coordinates of a first pixel of the first depth map is applicable for determining the arrival coordinates of a second pixel of the second depth map, the function being adapted as follows: [Math.4] p =PK[T^p^~\D(Pl) ) ])

[0109] With: • Ps the coordinates of a pixel of the fourth depth map, • 77 a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • K' a direction prediction model associated with the second camera 12, • K a direction prediction model associated with the first camera 11, • T an extrinsic matrix of the stereoscopic vision system, and • 0 a projection function in the three-dimensional scene of a second pixel Pt as a function of its coordinates in the second image and of the depth which is associated with it in the second depth map. This function thus determines the arrival coordinates of a pixel of an image acquired by the second camera 12 in an image as it would be acquired by the first camera 11.

[0110] According to an alternative embodiment, values ​​associated with pixels of the third and fourth images, that is to say colorimetric values ​​of the pixels of the third and fourth images, are obtained by interpolation of the values ​​of the pixels of the first image and respectively of the second image using a function such as torch.nn.functional.grid_sample() in Python® language which requires as arguments the arrival coordinates of the pixels of the third and fourth images and the values ​​of the preceding pixels in the first image, respectively second image.

[0111] In a step 35, a visibility mask is determined. This visibility mask is the union of a first visibility mask and a second visibility mask, each pixel of the third image resulting from at least one pixel of the first image belonging to the first visibility mask and each pixel of the fourth image resulting from at least one pixel of the second image belonging to the second visibility mask. In other words, the first visibility mask includes all the pixels of the third image which have an antecedent in the first image during the step of generating the third image, that is to say the set of pixels of the third image whose coordinates are arrival coordinates of a pixel of the first image.Similarly, the second visibility mask includes the set of pixels of the fourth image that have an antecedent in the second image during the step of generating the fourth image, that is, the set of pixels of the fourth image whose coordinates are arrival coordinates of a pixel of the second image.

[0112] Conversely, the pixels not included in the first visibility mask are pixels which have no antecedent in the first image acquired by the first camera 11, they then correspond to objects occluded from the point of view of the second camera 12 and the pixels not included in the second visibility mask are pixels which have no antecedent in the second image acquired by the second camera 12, they then correspond to objects occluded from the point of view of the first camera 11.

[0113] The determination of the visibility mask is then the union of the first and second visibility masks, that is to say that coordinates corresponding to those of a pixel of the third image that has no antecedent in the first image and also corresponding to those of a pixel of the fourth image that has no antecedent in the second image define coordinates of a pixel not included in the visibility mask. Conversely, if coordinates correspond to those of a pixel of the third image that has a antecedent in the first image and / or to those of a pixel of the fourth image that has a antecedent in the second image then these coordinates define coordinates of a pixel included in the visibility mask.

[0114] The determination of the visibility mask is for example obtained by using the torch.nn.functional.grid_sample() function in Python® language which requires as arguments the coordinates of the pixels of the third and fourth depth maps and the depths determined in the second and first depth maps.

[0115] In a step 36, a zero value is assigned to the pixels of the third and fourth images not included in the visibility mask. These pixels of the third and fourth images have no antecedent in the first image, respectively in the second image. A pixel with a zero value corresponds, for example, to a black pixel. In the case of several channels, each channel associated with the pixel is assigned a zero value.

[0116] In a step 37, three-dimensional positions of points of a third set of points are determined from the third depth map, each point of the third set of points corresponding to a pixel of the third depth map, and three-dimensional positions of points of a fourth set of points are determined from the fourth depth map, each point of the fourth set of points corresponding to a pixel of the fourth depth map.

[0117] According to a variant, similarly to step 33, the determination of a three-dimensional position of a point of the third set of points corresponding to a pixel of the third depth map is obtained by the following function:

[0118] [Math.5] P3d = ^P^PD(p,))

[0119] With: • the three-dimensional position of a point in the third set of points, • K' a direction prediction model associated with the second camera 12, • 0 a projection function in the three-dimensional scene of a pixel Pt according to its coordinates and the depth associated with it in the third depth map.

[0120] Similarly, determining a three-dimensional position of a point of the fourth set of points corresponding to a pixel of the fourth pro- founders is obtained by the following function:

[0121] [Math.6] p}d=^p^\d(pi))

[0122] With: • Pio the three-dimensional position of a point in the fourth set of points, • K' a direction prediction model associated with the first camera 11, • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates and the depth p^pj associated with it in the fourth depth map.

[0123] In a step 38, the depth prediction model is learned by minimizing a loss error, the loss error being determined from: • a first consistency error determined by comparing three-dimensional positions of points of the first and fourth sets of points, • a second consistency error determined by comparing three-dimensional positions of points of the second and third sets of points, and • a photometric error determined by comparing the first and fourth images and by comparing the second and third images.

[0124] According to a particular exemplary embodiment, the first consistency error is determined by the following function: [Math.7]

[0125] With: • LcÂp) corresponding to the first consistency error Lci(p) for a pixel P of the first depth map, a pixel P being defined by two-dimensional coordinates; • X*(p) a component along a first axis of a three-dimensional position of a point of the first set of points corresponding to the pixel P of the first depth map noted X^p); * x'*(p) a component along a first axis of a three-dimensional position of a point of the fourth set of points corresponding to the pixel P of the fourth depth map noted x\(p); * (P) a component along a second axis of a three-dimensional position of a point of the first set of points corresponding to the pixel P of the first depth map noted y ( p ); * y ' -(p) a component along a second axis of a three-dimensional position of a point of the fourth set of points corresponding to the pixel P of the fourth depth map noted y' (p); * Z*(p} a component along a third axis of a three-dimensional position of a point of the first set of points corresponding to the pixel P of the first depth map noted z^(p); and • z'*(p) a component along a third axis of a three-dimensional position of a point of the fourth set of points corresponding to the pixel P of the fourth depth map noted z\( p) ■

[0126] It should be noted that the first consistency error is based on the distance separating a point of the three-dimensional scene obtained from the first depth map from a point of the three-dimensional scene obtained from the fourth depth map.

[0127] According to a variant, the visibility mask is used to determine the first consistency error. Indeed, according to this variant, only the points associated with visible pixels, i.e. included in the visibility mask, of the fourth depth map are taken into account in this calculation. In other words, the fourth set of pixels only comprises pixels of the fourth depth map included in said visibility mask.

[0128] Similarly, the second consistency error is determined by the following function: [Math. 8] Lc{p) =

[0129] With: • L corresponding to the second consistency error L^Çp) for a pixel P of the second depth map, a pixel P being defined by coordinates in two dimensions; • X*(p) a component along a first axis of a three-dimensional position of a point of the second set of points corresponding to the pixel P of the second depth map noted x^(p); • x'*( p) a component along a first axis of a three-dimensional position of a point of the third set of points corresponding to the pixel P of the third depth map noted x'2 ( / ?); • y J p) a component along a second axis of a three-dimensional position of a point of the second set of points corresponding to the pixel P of the second depth map noted y^ ( p ); * y'Ap) a component along a second axis of a three-dimensional position of a point of the third set of points corresponding to the pixel P of the third depth map noted y' (p); * Z*(p} a component along a third axis of a three-dimensional position of a point of the second set of points corresponding to the pixel P of the second depth map noted z2(p); and • z'*(p) a component along a third axis of a three-dimensional position of a point of the third set of points corresponding to the pixel P of the third depth map noted z'2(p)-

[0130] It should be noted that the second consistency error is based on the distance separating a point of the three-dimensional scene obtained from the second depth map from a point of the three-dimensional scene obtained from the third depth map.

[0131] According to a variant, the visibility mask is used to determine the second consistency error. Indeed, according to this variant, only the points associated with visible pixels, i.e. included in the visibility mask, of the third depth map are taken into account in this calculation. In other words, the third set of pixels only comprises pixels of the third depth map included in the visibility mask.

[0132] According to a first particular exemplary embodiment, the photometric error is determined from a first photometric error by comparing pixel values ​​of the first and fourth images and from a second photometric error determined by comparing pixel values ​​of the second and third images, the loss error being determined from the first and second photometric errors.

[0133] A photometric error is, for example, presented in the document “Digging Into Self-Supervised Monocular Depth Estimation” by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow published in August 2019 and is determined by the following function:

[0134] [Math.9] LX / 7 ) ~ ( 1- ^) ■ U(p)-Hp) I +«' ) )

[0135] With: • Lp«(p) the first photometric error noted Lpi(p), respectively the second photometric error noted L^p), P being a pixel defined by its coordinates in an image, • Kp) a value of pixel P in the first image, respectively second image, • a value of pixel P in the fourth image, respectively third picture, • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.

[0136] According to this particular embodiment, the loss error is determined by the following function: [Math. 10] L = Yp(min(^ Lc2(p)) + min(Lpl(p), Lp2(p)))

[0137] With: • The loss error, • Lp^p) the first photometric error for a pixel P, P being a pixel defined by its coordinates in an image, • Lp2 the second photometric error for pixel P, • Lcl the first consistency error for pixel P, and • ^c2 the second consistency error for pixel P.

[0138] According to a second particular exemplary embodiment, the loss error further comprises a construction error, each of the components of which is determined by the following function: [Math. 11]

[0139] With: • L ( p) the component for a pixel p of the third image, respec tively the component L 4( p) ​​for a pixel p of the fourth image, • D(p) is a depth of a pixel P obtained from the third depth map, respectively obtained from the fourth depth map; • W is a parameter matrix; • 0 is the order of a smoothing gradient; • an L1 norm of second-order depth gradients is calculated with W =1, and0 =2; • x and y are the dimensions of the images; • / i is a hyperparameter dependent on the environment in which the vehicle operates; and • It{p ) is a value of pixel P in the third image, respectively fourth image.

[0140] This function is generally used to deal with discontinuity at the edge of objects (in English “edge aware smoothness”).

[0141] The loss error includes, for example, the consistency errors, photometric errors and construction error determined above: [Math. 12] L = £p(min(Lcl(p), Lc2(p) ) + min(Lp}(p), Lp2(p) ) + Ls3(p) + Ls4(p))

[0142] With: • The loss error, • Lp](p) the first photometric error for a pixel P, P being a pixel defined by its coordinates in an image, • Lp2 the second photometric error for pixel P, • Lc} the first consistency error for pixel P, • Lc2 the second consistency error for pixel P, • Ls3 the component in the third image for pixel P, • Lx4 the component in the fourth image for pixel P.

[0143] It should be noted that the pixels compared are those with similar or equal coordinates in the acquired images as in the generated images.

[0144] It should be noted that taking into account the consistency error also makes it possible to learn the depth prediction model if necessary.

[0145] Furthermore, the visibility mask used for determining the first and second consistency errors (depending on the variants) may be imprecise at the start of learning. Also, in order to make it larger as learning progresses, for example as a function of the number of iterations of the learning method 3, also called the number of epochs, the first and second consistency errors are for example weighted by a factor [3 increasing as a function of the number of iterations. Thus, according to another particular exemplary embodiment, the loss error comprises, for example, the weighted consistency errors, the photometric errors and the construction error determined above: [Math. 13] L^^min^ Up)) + mm(Lpï(p), Lp2(p))+Ls3(p)+Ls4(p))

[0146] With: • The loss error, • Lpx[p) the first photometric error for a pixel P, P being a pixel defined by its coordinates in an image, • Lp2 the second photometric error for pixel P, • a weighting factor between 0 and 1, • Lci the first consistency error for pixel P, • Lc2 the second consistency error for pixel P, • Ls3 the component in the third image for pixel P, • the component in the fourth image for pixel P.

[0147] Training the depth prediction model consists of adjusting input parameters of the convolutional neural network in order to minimize the previously calculated loss error.

[0148] In addition, an occluded or non-visible object in the field of vision of one camera and masked in the field of vision of the other camera does not impact the loss error, the use of a visibility mask making this learning method insensitive to occlusions. Thus, the depth prediction model used for the depth prediction of a pixel of an image acquired by one of the cameras of the stereoscopic vision system is made reliable thanks to this learning method.

[0149] This learning is carried out from data acquired by the on-board vision system and therefore does not require data annotated by another on-board system or storage of a library of learning images. In addition, the learning data is representative of the data received when the system is in operation or in production, in fact the learning data is representative of real environments in which the vehicle carrying the stereoscopic vision system evolves or moves, this learning data is therefore particularly relevant.

[0150] [Fig.4] schematically illustrates a device 4 configured to learn a depth prediction model by a vision system embedded in a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention. The device 4 corresponds for example to a device embedded in the first vehicle 10, for example a computer associated with the stereoscopic vision system.

[0151] The device 4 is for example configured for the implementation of the operations described with regard to figures 1 and 4 and / or steps described with regard to figures 2 and 3. Examples of such a device 4 include, but are not limited to, on-board electronic equipment such as an on-board computer of a vehicle, an electronic calculator such as an ECU (“Electronic Control Unit”), a smartphone, a tablet, a laptop. The elements of the device 4, individually or in combination, can be integrated in a single integrated circuit, in several integrated circuits, and / or in discrete components. The device 4 can be produced in the form of electronic circuits or software (or computer) modules or even a combination of electronic circuits and software modules.

[0152] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the method and / or for executing the instructions of the software(s) embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41 corresponding for example to a volatile and / or non-volatile memory and / or comprises a memory storage device which may comprise volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.

[0153] The computer code of the embedded software(s) comprising the instructions to be loaded and executed by the processor is for example stored in the memory 41.

[0154] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Telematic Control Unit”), for example via a communication bus or through dedicated input / output ports.

[0155] According to a particular and non-limiting exemplary embodiment, the device 4 comprises a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 comprise one or more of the following interfaces: - RF radio frequency interface, for example Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English “Universal Serial Bus” or “Universal Serial Bus” in French); HDMI interface (from the English “High Definition Multimedia Interface” or “High Definition Multimedia Interface” in French); - LIN interface (from the English “Local Interconnect Network”).

[0156] According to another particular and non-limiting embodiment, the device 4 comprises a communication interface 43 which makes it possible to establish communication with other devices (such as other computers of the on-board system) via a communication channel 430. The communication interface 43 corresponds for example to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds for example to a wired network of the CAN type (from the English “Controller Area Network” or in French “Réseau de contrôles”), CAN FD (from the English “Controller Area Network Flexible Data-Rate” or in French “Réseau de flexible data rate controllers”), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3).

[0157] According to a particular and non-limiting exemplary embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch-sensitive or not, one or more speakers 450 and / or other peripherals 460 via the output interfaces 44, 45, 46 respectively. According to a variant, one or other of the external devices is integrated into the device 4.

[0158] Of course, the present invention is not limited to the exemplary embodiments described above but extends to a method for determining the depth of a pixel of an image acquired by a vision system, and / or for measuring a distance separating an object from a vehicle carrying a vision system, the depth and / or the distance being predicted and / or measured via a depth prediction model learned according to the learning method described above, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0159] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-based motor vehicle, comprising the device 4 of [Fig.4].

Claims

1. Claims Method for learning a depth prediction model implemented by a convolutional neural network associated with a stereoscopic vision system embedded in a vehicle (10), the stereoscopic vision system comprising a first camera (11) and a second camera (12) arranged so as to each acquire an image of a three-dimensional scene from a different point of view, said method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception (31) of a first image and a second image acquired by the first camera (11) and the second camera (12) respectively at the same acquisition time instant; - generating (32) a first depth map comprising depths associated with a set of pixels of the first image and a second depth map comprising depths associated with a set of pixels of the second image, the depths being predicted with the depth prediction model from the first and second images; - determining (33) three-dimensional positions of points of a first set of points from the first depth map, each point of the first set of points corresponding to a pixel of the first depth map, and determining three-dimensional positions of points of a second set of points from the second depth map, each point of the second set of points corresponding to a pixel of the second depth map; - generation (34) of a third image and a third depth map from the first image, the first depth map and extrinsic parameters of the stereoscopic vision system, the third depth map comprising depths associated with a set of pixels of the third image, and generation of a fourth image and a fourth depth map from the second image, the second depth map and extrinsic parameters of the stereoscopic vision system, the fourth depth map comprising depths associated with a set of pixels of the fourth image; - determination (35) of a visibility mask as the union of a first visibility mask and a second visibility mask,

2. each pixel of the third image resulting from at least one pixel of the first image belonging to the first visibility mask and each pixel of the fourth image resulting from at least one pixel of the second image belonging to the second visibility mask; - assignment (36) of a zero value to the pixels of the third and fourth images not included in the visibility mask; - determining (37) three-dimensional positions of points of a third set of points from the third depth map, each point of the third set of points corresponding to a pixel of the third depth map, and determining three-dimensional positions of points of a fourth set of points from the fourth depth map, each point of the fourth set of points corresponding to a pixel of the fourth depth map; - learning (38) the depth prediction model by minimizing a loss error, the loss error being determined from: • a first consistency error determined by comparing three-dimensional positions of points of the first and fourth sets of points, • a second consistency error determined by comparing three-dimensional positions of points of the second and third sets of points, and • a photometric error determined by comparing the first and fourth images and by comparing the second and third images. The method of claim 1, wherein the first consistency error and the second consistency error are determined respectively by the following function: With : • Lc^p) corresponding to the first consistency error Lc^p) for a pixel P of the first depth map, respectively the second consistency error Lc~\p) for a pixel P of the second depth map, a pixel P being defined by coordinates in two dimensions; • X*(p ) a component along a first axis of a three-dimensional position mental of a point of the first set of points corresponding to the pixel P of the first depth map noted Xj( p), respectively of the second set of points corresponding to the pixel P of the second depth map noted X2(p); • x'*( p) a component along a first axis of a three-dimensional position of a point of the fourth set of points corresponding to the pixel P of the fourth depth map noted p), respectively of the third set of points corresponding to the pixel P of the third depth map noted X^Çp); * yÀP) a component along a second axis of a three-dimensional position of a point of the first set of points corresponding to the pixel P of the first depth map noted y ( p ), respectively of the second set of points corresponding to the pixel P of the second depth map noted ( p );* P) a component along a second axis of a three-dimensional position of a point of the fourth set of points corresponding to the pixel P of the fourth depth map denoted y ' ( p ), respectively of the third set of points corresponding to the pixel P of the third depth map denoted p); • z*( p) a component along a third axis of a three-dimensional position of a point of the first set of points corresponding to the pixel P of the first depth map denoted zY( p), respectively of the second set of points corresponding to the pixel P of the second depth map denoted z2(p) ' ct • z'^p) a component along a third axis of a three-dimensional position of a point of the fourth set of points corresponding to the pixel P of the fourth depth map denoted z\( p), respectively of the third set of points corresponding to the pixel P of the third depth map denoted z'2( p) ■;

3. The method of claim 2, wherein the third set of pixels comprises only pixels of the third depth map included in said visibility mask and the fourth set of pixels comprises only pixels of the fourth depth map included in said visibility mask.

4. A method according to claim 2 or 3, wherein the loss error is determined by the following function: L = ^p(min(Lcl(p\ Lc2(p) )+ min(Lpl(p), Lp2(p))) With: • L the loss error, • Lp{p) the first photometric error for a pixel P defined by its coordinates in two dimensions, • LP2 the second photometric error for the pixel P, • Lc\ the first consistency error for the pixel P, and • Lc2 the second consistency error for the pixel P.

5. Method according to one of claims 1 to 4, for which the determination of a three-dimensional position of a point corresponding to a pixel is obtained by the following function: p3O=^(pF'D(p,)) With: • the three-dimensional position of a point of the first, second, third and respectively fourth set of points, • K' a direction prediction model associated with the first camera (11), respectively with the second camera (12), • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates and the depth [^pj which is associated with it in the first, second, third and respectively fourth depth map.

6. The method of claim 5, wherein the direction prediction model is implemented by a neural network.

7. Method according to one of claims 5 to 6, for which the third and fourth depth maps are generated by the following function: ^=^(^1^(^^(^))]) With: • Ps the coordinates of a pixel of the third depth map, respectively of the fourth depth map, • 71 a function for going from homogeneous coordinates to pixel coordinates by removing a dimension of a vector, • K a direction prediction model associated with the second camera (12), respectively with the first camera (11), • K' a direction prediction model associated with the first camera (11), respectively with the second camera (12), • T an extrinsic matrix of the stereoscopic vision system, and • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates in the first image, respectively second image, and of the depth L^p^ which is associated with it in the first depth map, respectively second depth map.

8. Computer program comprising instructions for implementing the method according to any one of the preceding claims, when these instructions are executed by a processor.

9. Device (4) configured to learn a depth prediction model by a vision system on board a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured to implement the steps of the method according to any one of claims 1 to 7.

10. Vehicle (10) comprising the device (4) according to claim 9.

Citation Information

Patent Citations

  • Shared median-scaling metric for multi-camera self-supervised depth evaluation

    US11688090B2

  • Predicting depth from image data using a statistical model

    US20190213481A1