Method and device for learning a depth prediction model to reduce the loss of consistency in a stereoscopic vision system.
The method enhances depth prediction models for vehicle-mounted stereoscopic vision systems by using a convolutional neural network with stereoscopic vision systems, improving reliability and reducing reliance on external data, thus enhancing ADAS systems' performance.
Patent Information
- Application Number
- FR2024002905
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2044-03-22
AI Technical Summary
Existing depth prediction models for vehicle-mounted stereoscopic vision systems require significant amounts of annotated data or additional onboard systems like LiDAR, which may not always be available, leading to unsuitable configurations and reduced reliability in Advanced Driver-Assistance Systems (ADAS).
A method for learning a depth prediction model using a convolutional neural network that utilizes a stereoscopic vision system with two cameras, involving image flipping, depth map generation, and error minimization techniques to enhance depth consistency and reliability without relying on external data sources.
Improves the learning phase of depth prediction models, enhancing the reliability of ADAS systems by utilizing onboard vision system data, reducing the need for external data and ensuring precise depth estimation for improved road safety.
Smart Images

Figure 00000026_0000 
Figure 00000026_0001 
Figure 00000026_0002
Abstract
Description
Title of the invention: Method and device for learning a depth prediction model to reduce the loss of consistency of a stereoscopic vision system. technical field
[0001] The present invention relates to methods and devices for determining depth using a vision system embedded in a vehicle, for example, a motor vehicle. The present invention particularly relates to a method and device for learning a depth prediction model associated with a stereoscopic vision system embedded in a vehicle. Technological background
[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS systems use advanced technologies to assist the driver while driving and thus improve performance. ADAS systems use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.
[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.
[0004] Vehicle-mounted ADAS systems are powered by data obtained from one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or any obstacles present around a vehicle in order, for example: • to adapt the vehicle's lighting according to the presence of other road users; • to automatically regulate the vehicle's speed; • to act on the braking system in case of risk of impact with an object.
[0005] Predicting the distance between an object and a vehicle, that is, the depth associated with a pixel in an image acquired by a camera, the pixel representing the object, requires high precision to feed ADAS. Predicting the depth associated with a pixel relies in particular on a depth prediction model, the depth prediction model being adapted to the vehicle's onboard vision system. Such a prediction model is This is commonly implemented using a neural network, which is trained in a specific phase to predict accurate depths based on the vehicle's onboard vision system and the type of environment in which it operates. This training phase requires data, such as annotated data from libraries or data acquired from more precise onboard systems like LiDAR®. However, this type of training requires a massive amount of data or the presence of this other onboard system, which the vehicle may not always have. It is possible to train the depth prediction model using data acquired by the onboard vision system, but existing solutions are not suitable for the configuration of every onboard vision system in a vehicle. Summary of the present invention
[0006] One object of the present invention is to solve at least one of the problems of the technological background described above.
[0007] Another object of the present invention is to improve the learning phase of a depth prediction model from data from image processing acquired by a vision system, in particular by a depth prediction model implemented by a neural network associated with a stereoscopic vision system.
[0008] Another object of the present invention is to improve road safety, in particular by improving the reliability of AD AS systems powered by data obtained from a camera of a vision system.
[0009] According to a first aspect, the present invention relates to a method for learning a depth prediction model implemented by a convolutional neural network associated with a stereoscopic vision system embedded in a vehicle, the vision system comprising a first camera and a second camera arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the process being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception of representative data of a first image and a second image acquired by respectively the first camera and the second camera at the same time instant of acquisition; - generation of a first depth map including depths associated with a set of pixels from the first image, the depths being predicted with the depth prediction model from the first and second images; - generation of a first image flipped by symmetry of the first image with respect to a central vertical axis of the first image and generation of a second image flipped by symmetry of the second image with respect to a central vertical axis of the second image; - generation of a second depth map including depths associated with a set of pixels from the second image, the depths being predicted with the depth prediction model from the second and first returned images; - generation of a third image and a third depth map from the first image, the first depth map and extrinsic parameters of the stereoscopic vision system, the third depth map including depths associated with a set of pixels of the third image, and generation of a fourth image and a fourth depth map from the second image, the second depth map and extrinsic parameters of the stereoscopic vision system, the fourth depth map including depths associated with a set of pixels of the fourth image; - determination of a first consistency error by comparing the first and fourth depth maps and a second consistency error by comparing the second and third depth maps; - learning the depth prediction model by minimizing a loss error determined from a result of a comparison of the first and second consistency errors.
[0010] According to a variant of the method, the first consistency error and the second consistency error are determined respectively by the following function: (Dt(p)-Dt(p)Y With : • corresponding to the first consistency error L^p), respectively the second consistency error p Y for a pixel P, • D(p) a depth of a pixel Pt obtained from the first depth map, respectively from the second depth map, for a pixel P, and • D'(p) a depth of one pixel Pt obtained from the fourth depth map, respectively from the third depth map, for the pixel P.
[0011] According to another variant, the method further comprises the determination of a first photometric error determined by comparing pixel values of the first and fourth images and a second photometric error determined by comparing pixel values of the second and third images, the loss error being determined from the first and second photometric errors.
[0012] According to a further embodiment of the method, the loss error is determined by the following function: L^^min^L^p), Lc2(p)) + min(Lpi(p), Lp2(p))) With : • The loss error, • Lp^p) the first photometric error for a pixel P, P being a pixel defined by its coordinates in an image, • Lp2, the second photometric error for pixel P. • Lcl the first consistency error for pixel P, and • Le2 the second consistency error for pixel P.
[0013] According to yet another variant of the method, a photometric error is determined by the following function: L p *(p) - 0-«)■ R(p)-Hp) 1 +a ' With : • Lp*(p) the first photometric error denoted Lpi( p), respectively the second photometric error denoted Lp2(p ), P being a pixel defined by its coordinates in an image, • I(p) a value of pixel P in the first image, respectively second image, • a value of pixel P in the fourth image, respectively third picture, • SSIM is a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.
[0014] According to yet another variant of the process, the third and fourth depth maps are generated by the following function: D(pt) ) ] ) With : • Ps the coordinates of a pixel from the third depth map, respectively from the fourth depth map, • A function to convert from homogeneous coordinates to pixel coordinates by removing one dimension from a vector, • K' is an intrinsic matrix of the second camera, respectively to the first camera, • K is an intrinsic matrix of the first camera, respectively of the second camera, • T is an extrinsic matrix of the stereoscopic vision system, and • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates in the second image, respectively first image, and of the depth associated with it in the second depth map, respectively first depth map.
[0015] According to a further variant of the method, the generation of the second depth map includes the generation of another depth map associated with the second returned image, the depth of a pixel of the second depth map having the value of the depth of another pixel of the other depth map, a correspondence of a pixel of the other depth map with a pixel of the second depth map being determined during the generation of the second returned image.
[0016] According to a second aspect, the present invention relates to a device configured to learn a depth prediction model by a vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.
[0017] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0018] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0019] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0020] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.
[0021] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.
[0022] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.
[0023] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures
[0024] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 4, in which:
[0025] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a a particular and non-limiting example of a realization of the present invention;
[0026] [Fig.2] illustrates a flowchart of the different steps of a process for determining the depth of a pixel of an image by a depth prediction model associated with a vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention;
[0027] [Fig. 3] illustrates a flowchart of the different stages of a method for learning the depth prediction model used in the method of [Fig. 2], according to a particular and non-limiting embodiment of the present invention; and
[0028] [Fig.4] schematically illustrates a device configured to learn a depth prediction model by a vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements
[0029] A method and device for learning a depth prediction model implemented by a convolutional neural network associated with a stereoscopic vision system embedded in a vehicle will now be described in what follows with joint reference to Figures 1 to 4. The same elements are identified with the same reference signs throughout the description that follows.
[0030] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0031] According to a particular and non-limiting embodiment of the present invention, the depth prediction model is learned in a learning phase comprising the generation of depth maps associated with two images acquired by two cameras, the generation of images and other depth maps from the acquired images, the previously generated depth maps, and parameters extrinsic to the stereoscopic vision system as well as the generation of depth maps associated with the generated images.
[0032] A first and a second consistency error are determined by comparing the depth maps associated with the acquired images to depth maps associated with the generated images and the depth prediction model is learned by minimizing a loss error determined from a result of a comparison of the first and second consistency errors.
[0033] Fig. 1 schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.
[0034] Such an environment 1 corresponds, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.
[0035] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.
[0036] The vehicle 10 advantageously comprises at least two on-board cameras, a first camera 11 and a second camera 12, configured to acquire images of a three-dimensional scene unfolding in the environment of the vehicle 10 from distinct observation positions. The first camera 11 and the second camera 12 form a stereoscopic vision system when used together, as illustrated in [Fig. 1]. The first camera 11 forms a monoscopic vision system when used alone, and similarly, the second camera 12 forms another monoscopic vision system when used alone. The present invention, however, extends to any vision system comprising at least two cameras, for example, 2, 3, or 5 cameras.
[0037] The intrinsic parameters of the first camera 11 characterize the transformation which associates, for an image point, hereafter called "point", its three-dimensional coordinates in the reference frame of the first camera 11 with the pixel coordinates in an image acquired by the first camera 11. These parameters do not change if the first camera 11 is moved. The intrinsic parameters of the first camera 11 include in particular a first focal length fl associated with the first camera 11.
[0038] The intrinsic parameters of the second camera 12 characterize, for their part, the transformation which associates, for an image point, its three-dimensional coordinates in the frame of reference of the second camera 12 with the pixel coordinates in an image acquired by the second camera 12. These parameters do not change if the second camera 12. The intrinsic parameters of the second camera 12 include in particular a second focal length f2 associated with the second camera 12.
[0039] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in lens curvature, parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are not taken into account, and their correction is addressed during image acquisition or calibration.
[0040] These two cameras 11, 12 are arranged so that each acquires an image of a scene from a different viewpoint. The first viewpoint is, for example, located on or in the left-hand rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. The second viewpoint is, for example, located on or in the right-hand rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. If both cameras are located at the top of the windshield of the vehicle, they are then positioned at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, and the second camera 12 is located in the right-hand rearview mirror of the vehicle 10.
[0041] A first marker is associated with the first camera 11: - the direction of the x-axis is defined as horizontal and normal to the optical axis Cl of the first camera 11. The distance B separating the optical center of the first camera 11 from the projection of the optical center of the second camera 12 onto the horizontal plane passing through the optical center of the first camera 11 is called the reference base (in English "baseline"); - the direction of the y-axis is defined as vertical and normal to the optical axis Cl of the first camera 11; - The direction of the z-axis is defined as orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal coordinate system.
[0042] According to a particular embodiment, the optical axes Cl, C2 of the first and second cameras are contained in the same plane but are not parallel.
[0043] The extrinsic parameters related to the position of cameras 11, 12 are the following parameters: - three translations in the x, y, and z directions: Tx, Ty, and Tz, constituting the translation vector T; and - three rotations in the x, y and z directions: 0x, 0y and 0z.
[0044] An extrinsic matrix of the vision system then includes the extrinsic parameters previously defined.
[0045] The extrinsic parameters are determined, for example, during a calibration phase of the stereoscopic vision system comprising the first camera 11 and the second camera 12.
[0046] A key constraint of stereoscopic vision systems used in automobiles is, for example, the large distance between the two cameras. Indeed, to cover a measurement range of 200 meters, the reference base must be 60 cm for cameras commonly used in this field.
[0047] The two cameras 11, 12 acquire images of a scene located in front of the vehicle 10, the first camera 11 alone covering a first acquisition field 13, the second camera 12 alone covering a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic view of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic view of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic view of the scene by the stereoscopic vision system composed of the two cameras 11, 12.
[0048] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.
[0049] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system comprising the second camera 12.
[0050] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable using the monoscopic vision system comprising the first camera 11.
[0051] Finally, field 19 is not visible to any of the cameras. The part of the scene present in this field 19 is therefore not observable.
[0052] According to one particular embodiment, the field of view of the second camera 12 covers at least half of the field of view of the first camera 11.
[0053] It is evident that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently.
[0054] The images acquired by cameras 11, 12 at a given acquisition time are in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric coordinates (from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).
[0055] Each pixel of the acquired image represents an object in the three-dimensional scene present in the camera's field of view. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a point of light resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, which are captured by a photosensitive sensor of the camera after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image that faithfully represents the physical object observed by the camera. An image point described above is therefore a point on the surface of an object in the three-dimensional scene.
[0056] The images acquired by cameras 11, 12 represent views of the same scene taken from different viewpoints, the camera positions being distinct. For example, this scene includes: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0057] According to a particular embodiment, a field of view of the first camera 11 covers at least half of a field of view of the second camera 12, and a field of view of the second camera 12 covers at least half of a field of view of the first camera 11. In other words, more than half of the pixels of an image acquired by the first camera 11 correspond to an object in the three-dimensional scene seen by the second camera 12, and pixels of an image acquired by the second camera 12 also correspond to this object in the three-dimensional scene. Similarly, more than half of the pixels of an image acquired by the second camera 12 correspond to an object in the three-dimensional scene seen by the first camera 11, and pixels of an image acquired by the first camera 11 also correspond to this object in the three-dimensional scene.
[0058] The images acquired by the first camera 11 and by the second camera 12 are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.
[0059] A method for determining depth by a vision system on board the vehicle 10 is advantageously implemented by the vehicle 10, i.e. by a processor, a computer or a combination of computers of the on-board system of the vehicle 10, for example by the computer or computers in charge of the vision system of the vehicle 10.
[0060] Figure 2 illustrates a flowchart of the different steps of a method 2 for determining the depth of a pixel in an image by means of a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, for example in vehicle 10 of Figure 1, according to a particular and non-limiting embodiment of the present invention. Method 2 is implemented, for example, by a device of the vision system embedded in vehicle 10 or by device 4 of Figure 4.
[0061] In a step 21, representative data of an image acquired by the first camera 11 and of an image acquired by the second camera 12 are received.
[0062] In a step 22, depths associated with a set of pixels of one of the received images are determined by the depth prediction model from the two received images.
[0063] Each determined depth then corresponds to a distance separating the vehicle 10 or a part of the vehicle 10 from an object in the three-dimensional scene to which a pixel is associated, the determination of a depth of a pixel then corresponding to a measurement of a distance separating an object from the vehicle carrying the vision system.
[0064] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another road user, the ADAS is then able to determine this distance precisely. For example, if the ADAS's function is to activate the braking system of the vehicle 10 in the event of a risk of collision with another road user, and the distance between the vehicle 10 and that same road user decreases sharply, then the ADAS is able to detect this sudden approach and activate the braking system of the vehicle 10 to avoid a possible accident.
[0065] Figure 3 illustrates a flowchart of the different stages of a method for learning the depth prediction model used in a method for determining the depth of a pixel of an image, for example in method 2 of Figure 2, according to a particular and non-limiting embodiment of the present invention.
[0066] The learning process 3 is, for example, implemented by the device embedded in the vehicle 10 implementing the method for determining depth by means of a vision system embedded in a vehicle or by the device 4 of the [Fig.4].
[0067] In a step 31, representative data of a first image and a second image are received, the first image being acquired by the first camera 11 at an acquisition time instant and the second image being acquired by the second camera at the same acquisition time instant.
[0068] According to a particular embodiment, the first image and second image are of the same definition, that is to say they have the same number of pixels, have the same number of pixels according to their height and the same number of pixels according to their width.
[0069] According to another particular embodiment, the first and second images are not of the same resolution. An additional step then consists of resizing or cropping them to obtain a first and second image of the same resolution.
[0070] In step 32, a first depth map comprising depths associated with a set of pixels from the first image is determined. The depths are predicted using the depth prediction model from the first and second images.
[0071] The depth prediction model, implemented by a convolutional neural network, is known to those skilled in the art and is presented for example in the document "UnOS: Unified Unsupervised Optical-flow and Stereo-depth Estimation by Watching Videos" by Yang Wang, Peng Wang, Zhenheng Yang, Chenxu Luo, Yi Yang and Wei Xu published in June 2019, adapted to a stereoscopic vision system including cameras whose optical axes are contained in the same plane, or in the document "Unifying Flow, Stereo and Depth Estimation" written by Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao and Andreas Geiger, published in July 2023, adapted for any stereoscopic vision system including those including cameras whose optical axes are not contained in the same plane.
[0072] Thus, a predicted depth is associated with each pixel of a set of pixels of the first image, the first depth map comprising the coordinates of each pixel of the set of pixels of the first image and an associated depth recorded for example in a lookup table, in a memory accessible to the processor implementing this learning process 3.
[0073] In a step 33, a first flipped image is generated by mirroring the first image with respect to a central vertical axis of the first image, and a second flipped image is generated by mirroring the second image with respect to a central vertical axis of the second image (such an operation being called "flipping"). The central vertical axis of an image is the axis that divides this image into two equal parts, for example, each comprising the same number of pixels.
[0074] The pair comprising the second returned image and the first returned image is equivalent to a pair of images acquired by the first 11 and second 12 cameras acquiring a virtual three-dimensional scene corresponding to the symmetry of a real scene with respect to a plane, the plane being normal to the axis separating the positions of the two cameras and equidistant from the positions of the two cameras.
[0075] The pair of returned images thus allows the depth prediction model to predict depths associated with the pixels of the second returned image, the latter having the appearance of a first image.
[0076] In a step 34, a second depth map is generated, the second depth map comprising depths associated with a set of pixels from the second image, the depths being predicted with the depth prediction model from the second and first returned images.
[0077] The generation of the second depth map includes, for example, the generation of another depth map associated with the second returned image, the depths being predicted by the depth prediction model. A depth of a pixel in the second depth map then has the value of the depth of another pixel in the other depth map, a correspondence between a pixel in the other depth map and a pixel in the second depth map being determined during the generation of the second returned image. In other words, the depth prediction model predicts depths associated with the pixels of the second returned image. These depths are then associated with the pixels of the second image, each pixel of the second returned image corresponding to a pixel in the second image of which it is the mirror image. The second depth map is then mirror image of the first depth map.
[0078] Similar to the first depth map, a predicted depth is associated with each pixel of a set of pixels of the second image, the second depth map comprising the coordinates of each pixel of the set of pixels of the second image and an associated depth recorded for example in the lookup table described above.
[0079] In a step 35, a third image and a third depth map are generated from the first image, the first depth map and extrinsic parameters of the stereoscopic vision system, the third depth map comprising depths associated with a set of pixels of the third image.
[0080] The generation of the third depth map from the first image, the first depth map, and extrinsic parameters of the stereoscopic vision system consists of: • the determination of spatial coordinates in the three-dimensional scene, in a reference frame associated with the first camera 11, of a point associated with a first pixel of the first depth map, from the coordinates of the first pixel in the first image and the depth associated with this first pixel, its coordinates and depth being recorded in the first depth map. It should be noted that the coordinates of the first pixel in the first image comprise two components while the spatial coordinates of the point associated with the first pixel comprise three components; • the determination of spatial coordinates in the three-dimensional scene of the point associated with the first pixel in a reference frame associated with the second camera 12, in other words a change of reference frame from that associated with the first camera 11 to that associated with the second camera 12, using the extrinsic parameters of the stereoscopic vision system; and • The projection of the point associated with the pixel onto the image plane of the second camera 12, this image plane corresponding to that of the second image, allows us to determine the arrival coordinates of the first pixel of the first image in an image as the second camera 12 would have acquired it. The image plane of a camera corresponds to a plane defined in the camera's frame of reference, normal to the optical axis of the camera and located at the first focal length of the camera. The third depth map then includes the arrival coordinates of the first pixels and the depth associated with the first pixel.
[0081] The arrival coordinates of a pixel in the third depth map are thus determined and are the same as the arrival coordinates of a pixel in the third image. In the third depth map, the depth of a pixel corresponds to a depth determined from the pixel depths of the first depth map, while in the third image, the value of a pixel corresponds to a pixel value determined from the pixel values of the first image.
[0082] Similarly, a fourth image and a fourth depth map are generated from the second image, the second depth map and the extrinsic parameters of the stereoscopic vision system, the fourth depth map comprising depths associated with a set of pixels from the fourth image.
[0083] According to a particular embodiment, the third depth map is generated using the following function: [Math.l] Px = ^(K[T^(piK,~^ D(PT) ) ] )
[0084] With : • Ps the coordinates of a pixel on the third depth map, • 77 a function to convert from homogeneous coordinates to pixel coordinates by removing one dimension from a vector, • K' the intrinsic matrix of the first camera 11, • K the intrinsic matrix of the second camera 12, • T is an extrinsic matrix of the stereoscopic vision system, and • 0 a projection function in the three-dimensional scene of a first pixel Pt as a function of its coordinates in the first image and the depth D^p) which is associated with it in the first depth map.
[0085] Determining the fourth depth map is similar to determining the third depth map, with the cameras being reversed. The arrival coordinates of a pixel in the fourth depth map are thus determined and are the same as the arrival coordinates of a pixel in the fourth image. In the fourth depth map, the depth of a pixel corresponds to a depth determined from the pixel depths of the second depth map, while in the fourth image, the value of a pixel corresponds to a pixel value determined from the pixel values of the second image.
[0086] Thus, the preceding function allows the arrival coordinates to be determined The function of a first pixel of the first depth map is applicable to determine the arrival coordinates of a second pixel of the second depth map, the function being adapted as follows: [Math.2] P, = T^PkK^ D(pr} ) ] )
[0087] With : • Ps the coordinates of a pixel on the fourth depth map, •77 a function to convert from homogeneous coordinates to pixel coordinates by removing one dimension from a vector, • K' the intrinsic matrix of the second camera 12, • K, the intrinsic matrix of the first camera 11, • T is an extrinsic matrix of the stereoscopic vision system, and • 0 a projection function in the three-dimensional scene of a second pixel Pt as a function of its coordinates in the second image and of the depth [^pj] associated with it in the second depth map.
[0088] According to one embodiment, depth values associated with pixels of the third and fourth images, i.e., depths of the third and fourth depth maps, are obtained by interpolating the depths using a function such as torch.nn.functional.grid_sample() in Python® language which requires as arguments the coordinates of the pixels of the third and fourth depth maps and the depths determined in the second and first depth maps.
[0089] In a step 36, a first consistency error is determined by comparing the first and fourth depth maps and a second consistency error is determined by comparing the second and third depth maps.
[0090] According to a particular embodiment, the first consistency error and the second consistency error are determined respectively by the following function: [Math.3] (Dt(p)-D't(p)Ÿ
[0091] With: • Llj)) corresponding to the first consistency error L^pY and respectively the second consistency error I^p), for a pixel P, • D(p) a depth of a pixel Pt obtained from the first depth map, respectively from the second depth map, for a pixel P, and • D'(p) a depth of one pixel Pt obtained from the fourth depth map, respectively from the third depth map, for the pixel P.
[0092] Other functions can be used to perform this comparison, for example by taking the absolute value of the difference in depths. Indeed, it is important that the result of the depth comparison be a positive value for each pixel.
[0093] In a step 37, the depth prediction model is learned by minimizing a loss error determined from a result of a comparison of the first and second consistency errors.
[0094] According to a first particular embodiment, the learning process 3 further includes the determination of a first photometric error determined by comparing pixel values of the first and fourth images and a second photometric error determined by comparing pixel values of the second and third images, the loss error being determined from the first and second photometric errors.
[0095] A photometric error is, for example, presented in the document "Digging Into Self-Supervised Monocular Depth Estimation" by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow published in August 2019 and is determined by the following function:
[0096] [Math.4] Lp*(p) = (l-«) ■ U(.P)-Hp)\+a-
[0097] With: • Lp*(p) the first photometric error denoted Lp](p), respectively the second photometric error denoted Lp2{p), P being a pixel defined by its coordinates in an image, • I(p) a value of pixel P in the first image, respectively second image, • a value of pixel P in the fourth image, respectively third picture, • SSIM is a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.
[0098] According to this particular embodiment, the loss error is determined by the following function: [Math.5] L=Hp(min(Lci(p), Lc2(p))+ min(Lpl(p), Lp2(p)))
[0099] With: • The loss error, • Lpi(p) is the first photometric error for a pixel P, where P is a pixel defined by its coordinates in an image. • Lp2, the second photometric error for pixel P. • L(:l the first consistency error for pixel P, and • Lc2 the second consistency error for pixel P.
[0100] According to a second particular embodiment, the loss error comprises components further determined by the following function: [Math.6]
[0101] With: • The nwoth(p) is the component p) for a pixel p of the third image, respectively the component for a pixel p of the fourth image, • D(p) is a depth of a pixel P obtained from the third depth map, respectively obtained from the fourth depth map; • W is a parameter matrix; • 0 is the order of a smoothing gradient; • an L1 norm of the second-order depth gradients is calculated with W = 1, and ° = 2; • x and y are the dimensions of the images; • / 3 is a hyperparameter dependent on the environment in which the vehicle operates; and • / ( nj is a value of pixel P in the third image, respectively fourth image.
[0102] This function is generally used to deal with the discontinuity at the edge of objects (in English, "edge aware smoothness").
[0103] The loss error includes, for example, consistency errors, photometric errors and the components determined above: [Math.7] £ = ^min(Lcl(p), Lc2(p) ) + min(Lpi(p), Lp2(p)) + Ls3(p) + Ls4(p))
[0104] With: • the loss error, • Lp-^p) the first photometric error for a pixel P, P being a pixel defined by its coordinates in an image, • Ap2, the second photometric error for pixel P, • Lcl the first consistency error for pixel P, • Lc2, the second consistency error for pixel P, • Ls3 is the component in the third image for pixel P, • Ls4 the component in the fourth image for pixel P.
[0105] It should be noted that the pixels compared are those with similar or equal coordinates in both the acquired and generated images.
[0106] The training of the depth prediction model consists of adjusting the input parameters of the convolutional neural network in order to minimize the previously calculated loss error.
[0107] Furthermore, an object visible in the field of view of one camera and obscured in the field of view of the other camera does not affect the loss error; therefore, this method is insensitive to occlusions. Thus, the depth prediction model used for predicting the depth of a pixel in an image acquired by the first camera 11 is made more reliable by this learning method.
[0108] This learning is performed using data acquired by the on-board vision system and therefore does not require data annotated by another on-board system or the storage of a training image library. Furthermore, the training data is representative of the data received when the system is in operation or in production; indeed, the training data is representative of real-world environments in which the vehicle carrying the system operates or moves. stereoscopic vision system, therefore this training data is particularly relevant.
[0109] Figure 4 schematically illustrates a device 4 configured to learn a depth prediction model from a vehicle-mounted vision system, according to a particular, non-limiting embodiment of the present invention. The device 4 corresponds, for example, to a device mounted in the first vehicle 10, for example, a computer associated with the stereoscopic vision system.
[0110] Device 4 is, for example, configured to carry out the operations described opposite Figures 1 and 4 and / or the steps described opposite Figures 2 and 3. Examples of such a device 4 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 4, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 4 may be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.
[0111] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.
[0112] The computer code of the embedded software(s), including the instructions to be loaded and executed by the processor, is for example stored on memory 4L
[0113] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.
[0114] According to a particular and non-limiting embodiment, the device 4 comprises a block 42 of interface elements for communicating with external devices. Interface elements of block 42 include one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HD MI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").
[0115] According to another particular and non-limiting embodiment, the device 4 includes a communication interface 43 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 430. The communication interface 43 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds, for example, to a wired network of the CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3) type.
[0116] According to a particular and non-limiting embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch or not, one or more speakers 450 and / or other peripherals 460 via the output interfaces 44, 45, 46 respectively. According to a variant, one or more of the external devices is integrated into the device 4.
[0117] Of course, the present invention is not limited to the embodiments described above but extends to a method for determining the depth of a pixel in an image acquired by a vision system, and / or for measuring the distance between an object and a vehicle equipped with a vision system, the depth and / or distance being predicted and / or measured via a depth prediction model learned according to the learning method described above, which would include secondary steps without departing from the scope of the present invention. It would be a case of even of a device configured for the implementation of such a process.
[0118] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 4 of [Fig.4].
Claims
1. Demands Method for learning a depth prediction model implemented by a convolutional neural network associated with a stereoscopic vision system embedded in a vehicle (10), the vision system comprising a first camera (11) and a second camera (12) arranged so as to each acquire an image of a three-dimensional scene from a different point of view, said method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception (31) of data representative of a first image and a second image acquired respectively by the first camera (11) and the second camera (12) at the same time instant of acquisition; - generation (32) of a first depth map including depths associated with a set of pixels from the first image, the depths being predicted with the depth prediction model from the first and second images; - generation (33) of a first image reversed by symmetry of the first image with respect to a central vertical axis of the first image and generation of a second image reversed by symmetry of the second image with respect to a central vertical axis of the second image; - generation (34) of a second depth map including depths associated with a set of pixels from the second image, the depths being predicted with the depth prediction model from the second and first returned images; - generation (35) of a third image and a third depth map from the first image, the first depth map and extrinsic parameters of the stereoscopic vision system, the third depth map comprising depths associated with a set of pixels of the third image, and generation of a fourth image and a fourth depth map from the second image, the second depth map and extrinsic parameters of the stereoscopic vision system, the fourth depth map comprising depths associated with a set of pixels of the fourth image; - determination (36) of a first consistency error by com- comparison of the first and fourth depth maps and a second consistency error by comparison of the second and third depth maps; - learning (37) of the depth prediction model by minimizing a loss error determined from a result of a comparison of the first and second consistency errors.
2. A method according to claim 1, wherein the first consistency error and the second consistency error are determined respectively by the following function: ^0= (A4» -Dt(p))2 With: • L / p) corresponding to the first consistency error L^p), respectively the second consistency error For a pixel P, • D(p) a depth of a pixel Pt obtained from the first depth map, respectively from the second depth map, for a pixel P, and • D'(p) a depth of a pixel Pt obtained from the fourth depth map, respectively from the third depth map, for the pixel P.
3. A method according to any one of claims 1 to 2, further comprising determining a first photometric error determined by comparing pixel values of the first and fourth images and a second photometric error determined by comparing pixel values of the second and third images, the loss error being determined from the first and second photometric errors.
4. A method according to claim 3, wherein the loss error is determined by the following function: L = Yp(min(Lcl(p), Lc2(p) ) + min (Lpi(p), Lp2(p) )) With: • P the loss error, • Lp^p) the first photometric error for a pixel P, P being a pixel defined by its coordinates in an image, • Lp2 the second photometric error for the pixel P, • Lci the first consistency error for the pixel P, and • Lc2 the second consistency error for the pixel P.
5. A method according to claim 4, wherein a photometric error is determined by the following function: L}Ap) = (l-«) ' \Ap)A(p)\+a- (l-jSSIM^I(p),l(p) ) ) With: • Lp*( p} the first photometric error denoted Lp\( p), respectively the second photometric error denoted Lpz(p)A being a pixel defined by its coordinates in an image, • Kp) a value of the pixel P in the first image, respectively second image, * l(p) a value of the pixel P in the fourth image, respectively third image, • SSIM a function which takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.
6. A method according to any one of claims 1 to 5, wherein the third and fourth depth maps are generated by the following function: Ps = 71(K[(PiKtD(Pt))]) With: • Ps the coordinates of a pixel in the third depth map, respectively in the fourth depth map, • 77 a function for converting homogeneous coordinates to pixel coordinates by removing one dimension from a vector, • K' an intrinsic matrix of the first camera (11), respectively in the second camera (12), • K an intrinsic matrix of the second camera (12), respectively in the first camera (11), • T an extrinsic matrix of the stereoscopic vision system, and • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates in the first image, respectively second image, and the depth I)(p) associated with it in the first depth map,respectively, the second depth map.
7. A method according to any one of claims 1 to 6, wherein the generation of the second depth map comprises the generation of another depth map associated with the second returned image, a depth of one pixel of the second depth map having value the depth of another pixel of the other depth map, a correspondence of a pixel of the other depth map with a pixel of the second depth map being determined during the generation of the second returned image.
8. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by a processor.
9. Device (4) configured to learn a depth prediction model by a vision system embedded in a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured to implement the steps of the method according to any one of claims 1 to 7.
10. Vehicle (10) comprising the device (4) according to claim 9.