Method and device for determining depth by a self-supervised stereoscopic vision system.
The method improves ADAS systems by using a stereoscopic vision system with multiple cameras and neural networks to correct disparities, enhancing depth determination and safety through accurate depth estimation.
Patent Information
- Application Number
- FR2023011623
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-10-26
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-10-26
AI Technical Summary
The quality of data obtained by vehicle cameras affects the performance and safety of Advanced Driver-Assistance Systems (ADAS), necessitating improved depth determination methods for stereoscopic vision systems to enhance reliability and accuracy.
A method utilizing a stereoscopic vision system with multiple cameras to determine depth by correcting disparities through convolutional neural networks, adjusting input parameters to minimize errors, and generating images from corrected disparities to improve depth estimation.
Enhances the reliability of ADAS systems by providing more accurate depth data, leading to improved operational safety and efficiency.
Smart Images

Figure 00000030_0000 
Figure 00000031_0000 
Figure 00000031_0001
Abstract
Description
Title of the invention: Method and device for determining depth by a self-supervised stereoscopic vision system. Technical field
[0001] The present invention relates to methods and devices for determining a depth by means of a stereoscopic vision system on board a vehicle, for example in a motor vehicle. The present invention also relates to a method and a device for measuring such a depth. The present invention also relates to a method and a device for controlling one or more AD AS systems on board a vehicle from the determined depth. Technological background
[0002] Many modern vehicles are equipped with so-called AD AS (Advanced Driver-Assistance System). Such AD AS systems are passive and active safety systems designed to eliminate the element of human error in the driving of vehicles of all types. AD AS use advanced technologies to assist the driver while driving and thus improve their performance. AD AS use a combination of sensor technologies to perceive the environment around a vehicle, then provide information to the driver or act on certain vehicle systems.
[0003] There are several levels of ADAS, such as rearview cameras and blind spot sensors, lane departure warning systems, adaptive cruise control and automatic parking systems.
[0004] The AD AS embedded in a vehicle are supplied with data obtained from one or more embedded sensors such as, for example, cameras. These cameras make it possible in particular to detect and locate other road users or possible obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle speed; - to act on the braking system in the event of a risk of impact with an object.
[0005] The proper functioning of the driving assistance peripherals using this data therefore depends on the quality of the data emitted by a vision system. Summary of the present invention
[0006] An object of the present invention is to solve at least one of the problems of the technological background described above.
[0007] Another object of the present invention is to improve the quality of the data obtained of these cameras.
[0008] Another object of the present invention is to improve road safety, in particular by improving the operational safety of AD AS systems supplied by data obtained from at least one camera.
[0009] According to a first aspect, the present invention relates to a method for determining a depth by a stereoscopic vision system on board a vehicle, the stereoscopic vision system comprising a set of cameras of at least three cameras arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the method being characterized in that it comprises the following steps: - reception of data representative of a set of three images comprising a first, a second and a third images acquired respectively by a first, a second and a third camera of the set of cameras at the same acquisition time instant; - determining first disparities associated with a first set of pixels of the first image from the first set of pixels and a second set of pixels of the second image corresponding to the first set of pixels and determining second disparities associated with a third set of pixels of the second image from the third set of pixels and a fourth set of pixels of the first image corresponding to the third set of pixels, the first and second disparities being predicted by an optical flow method implemented by a first convolutional neural network; - determination of first and second corrected disparities by correcting the first and second disparities as a function of a disparity induced by a relative position of the first and second cameras; - determining first and second depths associated respectively with the first and third sets of pixels from the first and second corrected disparities respectively; - generating a fourth image from the first image and the first depths and generating a fifth image from the second image and the second depths; - determining a first error by comparing the fourth and second images and determining a second error by comparing the fifth and first images; - determining third disparities associated with a fifth set of pixels of the first image from the fifth set of pixels and a sixth set of pixels of the third image corresponding to the fifth set of pixels and determining fourth disparities associated with a seventh set of pixels of the third image from the seventh set of pixels and an eighth set of pixels of the first image corresponding to the seventh set of pixels, the third and fourth disparities being predicted by an optical flow method implemented by the first convolutional neural network; - determination of third and fourth disparities corrected by correction of the third and fourth disparities as a function of a disparity induced by a relative position of the first and third cameras; - determination of third and fourth depths associated respectively with the fifth and seventh sets of pixels from respectively the third and fourth corrected disparities; - generation of a sixth image from the first image and the third depths and generation of a seventh image from the third image and the fourth depths; - determination of a third error by comparing the sixth and third images and determination of a fourth error by comparing the seventh and first images; - adjusting input parameters of the first convolutional neural network by minimizing a fifth error, the fifth error being determined from an average of the first, second, third and fourth errors.
[0010] Such a method thus makes it possible to carry out training of the first convolutional neural network, thus making it possible to make the disparities predicted by this convolutional neural network more reliable and thus to make the determined depths more reliable. An AD AS system using these depths is then more efficient because it has more reliable data.
[0011] This learning is further carried out using data obtained directly from the stereoscopic vision system, the data used for this learning thus reflect the conditions of use of this stereoscopic vision system and are easily accessible.
[0012] According to a variant of the method, the fifth error is determined by the following loss function: As = E^avg ( Lx ( p ), L2 ( p ), L3( p ), L4(p) ) With: • At the fifth error, • with an average of arguments, • L^p) the first error for a pixel P of the second image, • L^p) the second error for a pixel P of the first image, • L / p) the third error for a pixel P of the third image, and • L / p) the fourth error for a pixel P of the first image.
[0013] According to another variant of the method, an image is generated from an input image and depths associated with the pixels of the input image so as to reproduce an image seen by another camera from the following formula: \ Ds((p() ) ] ) With : • Ps one pixel of the generated image, • n a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • F an intrinsic matrix of the other camera, • F an extrinsic matrix of the stereoscopic vision system formed by the camera having acquired the input image and the other camera, • 0 a projection function in the three-dimensional scene of a pixel as a function of its depth, • K' an intrinsic matrix of the camera having acquired the input image, • is a depth of the pixel Pt associated with the input image.
[0014] According to a variant of the additional method, an error is determined by the following loss function: With : • L*(p) the error associated with a pixel P, • I(p) a value of the pixel P in an image acquired by a camera, • [{pj a value of the pixel P in a generated image' • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.
[0015] According to yet another variant of the method, an error is determined by the following loss function: Lsmooth(Dst(pt), W, =LpLdGx>y( w(pt)\^dDst(pt) pw,(p;)| ) With : •LsmOoth{0, W,o} the error, • p ) is a depth of one pixel Pt predicted by the stereoscopic vision system; • VF is a parameter matrix; • 0 is the order of a smoothing gradient; • an L1 norm of second-order depth gradients is calculated with W =1, and ° =2 ; • x and y are the dimensions of an image; • / 3 is an environment-dependent hyperparameter; and * ht P ) is the value of the pixel Pt in a generated image.
[0016] According to yet another variant of the method, a disparity induced by a relative position of two cameras of the set of cameras is determined by a second convolutional neural network.
[0017] According to a second aspect, the present invention relates to a device for determining a depth by a stereoscopic vision system on board a vehicle, the device comprising a memory associated with at least one processor configured for implementing the steps of the method according to the first aspect of the present invention.
[0018] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0019] According to a fourth aspect, the present invention relates to a computer program which comprises instructions adapted for executing the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0020] Such a computer program may use any programming language and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0021] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to the first aspect of the present invention.
[0022] On the one hand, the recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM memory, a CD-ROM or a microelectronic circuit type ROM memory, or a magnetic recording means or a hard disk.
[0023] Furthermore, this recording medium may also be a transmissible medium such as an electrical or optical signal, such a signal being able to be conveyed via an electrical or optical cable, by conventional or hertzian radio or by self-directed laser beam or by other means. The computer program according to the present invention may in particular be downloaded from an Internet-type network.
[0024] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to perform or to be used in performing the method in question. Brief description of the figures
[0025] Other characteristics and advantages of the present invention will emerge from the description of the particular and non-limiting exemplary embodiments of the present invention below, with reference to the appended figures 1 to 5, in which:
[0026] [Fig-1] schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention;
[0027] [Fig.2] illustrates a flowchart of the different steps of a method for determining a depth by a stereoscopic vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;
[0028] [Fig.3] schematically illustrates a first and a second image obtained from a stereoscopic vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;
[0029] [Fig.4] schematically illustrates a first and a third images obtained from a stereoscopic vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;
[0030] [Fig.5] schematically illustrates a device configured for determining a depth by a stereoscopic vision system on board the vehicle of [Fig.1], according to a particular and non-limiting exemplary embodiment of the present invention. Description of examples of implementation
[0031] A method and a device for determining a depth by means of a stereoscopic vision system on board a vehicle will now be described in the following with joint reference to FIGS. 1 to 5. The same elements are identified with the same reference signs throughout the description which follows.
[0032] The terms "first(s)", "second(s)" (or "first(s)", "second(s)"), etc. are used in this document by arbitrary convention to enable different elements (such as operations, means, etc.) implemented in the embodiments described below to be identified and distinguished. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0033] According to a particular and non-limiting example of embodiment of the present invention, a method for determining a depth by stereoscopic vision system embedded in a vehicle is for example implemented by a computer of the vehicle's on-board system controlling this stereoscopic vision system.
[0034] The stereoscopic vision system comprises a set of cameras of at least three cameras arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the optical axes representative of an orientation of the field of vision of each camera being oriented in a non-parallel manner. The stereoscopic vision system is thus said to be non-parallel.
[0035] For this purpose, the method for determining a depth by a stereoscopic vision system on board a vehicle comprises receiving data representative of a set of three images comprising a first, a second and a third images acquired respectively by a first, a second and a third camera of the set of cameras at the same acquisition time instant.
[0036] Disparities are determined for each image of a pair of images from these same images by a first convolutional neural network, a first pair of images comprising the first and second images and a second pair of images comprising the first and third images. Thus, first and second disparities are determined for pixels of the first and second images respectively and third and fourth disparities are determined for pixels of the first and third images respectively.
[0037] These disparities are then corrected as a function of the disparity induced by the relative position of the cameras having acquired the images of a pair of images and depths are then determined for the pixels of the images of the pairs of images.
[0038] Each image of a pair of images then makes it possible to generate a new image from the depths determined for this pair of images and the generated image is then compared to the second image of the pair of images in order to determine an error, subsequently called a reconstruction error.
[0039] Input parameters of the first convolutional neural network are then adjusted by minimizing a fifth error, the fifth error being determined from an average of the four previously determined reconstruction errors.
[0040] [Fig. 1] schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.
[0041] The stereoscopic vision system is then arranged in an environment 1 corresponding, for example, to a road environment formed of a network of roads accessible to the vehicle 10.
[0042] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, with an electric motor(s) or even a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as a car, truck, bus, motorcycle. Finally, vehicle 10 corresponds to an autonomous vehicle or not, that is to say a vehicle traveling according to a determined level of autonomy or under the total supervision of the driver.
[0043] The vehicle 10 advantageously comprises several on-board cameras 11, 12, 13, each configured to acquire images of a scene in the environment of the vehicle 10. This set of cameras 11, 12, 13 forms the stereoscopic vision system. Three cameras 11, 12 and 13 are illustrated in [Fig.l]. The present invention is however not limited to a stereoscopic vision system comprising three cameras but extends to any stereoscopic vision system comprising 3 or more cameras, for example 3, 4 or 5 cameras.
[0044] The three cameras 11, 12, 13 have known intrinsic parameters. These parameters consist in particular of: - the focal length fl of the first camera 11; - the focal length f2 of the second camera 12; - the focal length f3 of the second camera 13; - distortions which are due to imperfections in the optical system of each camera; - the direction Cl of the optical axis of the first camera 11; - the direction C2 of the optical axis of the second camera 12; - the direction C3 of the optical axis of the second camera 13; and - the respective resolutions of cameras 11, 12, 13.
[0045] The intrinsic parameters characterize the transformation which associates, for an image point, the camera coordinates with the pixel coordinates, in each camera. These parameters do not change if the camera is moved.
[0046] The distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions which generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, that is to say that the distortions are not taken into account or that their correction is processed at the time of image acquisition.
[0047] These three cameras 11, 12, 13 are arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the first point of view, associated for example with the first camera 11, is for example located at the top of the windshield of the vehicle 10, the second point of view, associated with the second camera 12, is for example located on or in the left rearview mirror of the vehicle 10 and the third point of view, associated with the third camera 13, is for example located on or in the right rearview mirror of the vehicle 10.
[0048] The example illustrated in [Fig.l] shows cameras placed at the front of the vehicle 10. The invention is however not limited to this particular embodiment, the cameras being able to be placed at the rear of the vehicle 10 or on the sides of the vehicle 10. Similarly, a camera placed at the front of the vehicle 10 can be placed at the level of a headlight or in a bumper for example.
[0049] A marker is associated with each camera. For example, a first marker (xb yb zj is associated with the first camera 11: - the Zi axis corresponds to the optical axis Cl of the first camera 11, - the axis Xi corresponds to a first transverse axis of the first camera 11 and corresponding to a centered horizontal axis of an image acquired by the first camera 11, and - the yi axis corresponds to a second transverse axis of the first camera 11 and corresponding to a centered vertical axis of an image acquired by the first camera 11. The three axes x, y and z form an orthonormal reference frame.
[0050] The optical axes C2 and C3 are for example coplanar, the plane comprising them being parallel to a plane comprising the axes xi and zb
[0051] The optical axis C2 of the second camera 12 is further included in a plane forming an angle 012 with the plane comprising the axes yl and zl of the first camera 11 and the optical axis C3 of the third camera 13 is further included in a plane forming an angle 013 with the plane comprising the axes yl and zl of the first camera 11.
[0052] Each camera is distinct from another camera, the distance separating the optical centers of two cameras being called the baseline distance.
[0053] The optical centers of the first 11 and second 12 cameras are located at a first distance B12 while the optical centers of the first 11 and third 13 cameras are located at a second distance B13.
[0054] A main constraint of the stereoscopic vision system used in automobiles is, for example, the large distance between the two cameras. Indeed, to be able to cover a measurement range of 200 meters, the reference base must reach 60cm for the cameras commonly used in this field. Thus, the first B12 and second B13 distances are for example equal to 50cm, 60cm or 80cm.
[0055] In this example, the first camera 11 is located at a height from the ground greater than the height of the second 12 and third 13 cameras, the second camera 12 is located to the left of the first camera 11 and the third camera 13 is located to the right of the first camera 11. The right and left directions are determined relative to the direction of movement of the vehicle when it is moving forward, that is to say along a horizontal axis. corresponding to a transverse axis of the vehicle 10.
[0056] The extrinsic parameters related to the relative position of the first 11 and second 12 cameras are the following parameters: - 3 translations in the x, y and z directions: Txi2, Tyn and Tzn constituting the translation vector Tn; and - 3 rotations around the x, y and z axes: Rx^, Ryi2 and Rzn, constituting the rotation matrix R[2.
[0057] The extrinsic parameters related to the relative position of the first 11 and third 13 cameras are the following parameters: - 3 translations in the x, y and z directions: Txn, Tyn and Tzn constituting the translation vector Tn; and - 3 rotations around the x, y and z axes: Rxn, Ryn and Rzn, constituting the rotation matrix Rn.
[0058] The extrinsic parameters are further determined during a calibration phase of the stereoscopic vision system.
[0059] The first camera 11 acquires images of a three-dimensional scene located in its field of vision 110, the second camera 12 acquires images of a three-dimensional scene located in its field of vision 120 and the third camera 13 acquires images of a three-dimensional scene located in its field of vision 130. These three fields of vision 110, 120, 130 are concurrent, thus several zones are defined such that: - a first acquisition zone 111 visible only by the first camera 11, - a second acquisition zone 121 visible only by the second camera 12, - a third acquisition zone 131 visible only by the third camera 13, - a fourth acquisition zone 112 visible by the first 11 and second 12 cameras, - a fifth acquisition zone 113 visible by the first 11 and third 13 cameras, and - a sixth acquisition zone 114 visible by the first 11, second 12 and third 13 cameras.
[0060] The first 111, second 121 and third 131 acquisition zones thus allow a monoscopic vision of the three-dimensional scene by the first 11, second 12 and third 13 cameras respectively.
[0061] The fourth 112 and fifth 113 acquisition zones allow a stereoscopic vision of the three-dimensional scene respectively by the first 11 and second 12 cameras and by the first 11 and third 13 cameras.
[0062] The sixth 114 acquisition zone allows a stereoscopic vision of the three-dimensional scene by both the first 11 and second 12 cameras and by the first 11 and third 13 cameras.
[0063] It is obvious that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with differently placed and oriented cameras.
[0064] The images acquired by the cameras 11, 12, 13 at a given acquisition time instant are presented in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colors and brightness of objects in the observed scene in the form, for example, of RGB colorimetric coordinates (from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness).
[0065] The images acquired by the cameras 11, 12, 13 represent views of the same scene taken from different viewpoints, the positions of the cameras being distinct. On this scene are for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0066] These images are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.
[0067] A process for determining a depth by a stereoscopic vision system on board the vehicle 10 is advantageously implemented by the vehicle 10, that is to say by a computer or a combination of computers of the on-board system of the vehicle 10, for example by the computer(s) in charge of the stereoscopic vision system of the vehicle 10.
[0068] In a first operation, the computer receives data representative of a set of three images comprising a first 31, a second 32 and a third 33 images acquired respectively by a first 11, a second 12 and a third 13 cameras of the set of cameras, each of the images of this set of three images being acquired at the same acquisition time instant.
[0069] Each of these images thus represents the three-dimensional scene seen from the point of view of each of the cameras at the same time instant.
[0070] A part 112, 114 of the field of vision 111 of the first camera 11 also being included in the field of vision 121 of the second camera 12, the first 31 and second 32 images observe the same part of the three-dimensional scene. ional. Part of the pixels of the first image and part of the pixels of the second image are thus associated with the same objects present in the three-dimensional scene.
[0071] As illustrated in [Fig. 3], the first image 31 has a first part 311 associated with objects located in the field of vision 121 of the second camera 12, representative of the objects located in the field of vision 112 common to the first 11 and second 12 cameras as well as a second part 312 associated with objects located outside the field of vision 121 of the second camera 12.
[0072] Similarly, the second image 32 has a first part 321 associated with objects located in the field of vision 111 of the first camera 11, representative of the objects located in the field of vision 112 common to the first 11 and second 12 cameras as well as a second part 322 associated with objects located outside the field of vision 111 of the first camera 11.
[0073] The respective positions of the first parts 311 and 321 of the first 31 and second 32 images are defined by the relative positions of the first 11 and second 12 cameras as well as the orientation of their respective fields of vision 111, 121. Thus, according to the example described previously, if the first camera 11 is located above and to the right of the second camera, then the first part 311 of the first image 31 is located at the bottom left of the first image 31 and the first part 321 of the second image 32 is located at the top right of the second image 32.
[0074] Similarly, a part 113, 114 of the field of vision 111 of the first camera 11 also being included in the field of vision 131 of the third camera 13, the first 31 and third 33 images observe the same part of the three-dimensional scene. A part of the pixels of the first image and a part of the pixels of the third image are thus associated with the same objects present in the three-dimensional scene.
[0075] As illustrated in [Fig.4], the first image 31 has a third part 313 associated with objects located in the field of vision 131 of the third camera 13, representative of the objects located in the field of vision 113 common to the first 11 and third 13 cameras as well as a fourth part 314 associated with objects located outside the field of vision 131 of the third camera 13.
[0076] Similarly, the third image 33 has a first part 331 associated with objects located in the field of vision 111 of the first camera 11, representative of the objects located in the field of vision 113 common to the first 11 and third 13 cameras as well as a second part 332 associated with objects located outside the field of vision 111 of the first camera 11.
[0077] The respective positions of the third part 313 of the first image 31 and of the first part 331 of the third 33 are defined by the relative positions of the first 11 and third 12 cameras as well as the orientation of their respective fields of vision 111, 131. Thus, according to the example described previously, if the first camera 11 is located above and to the left of the second camera, then the third part 313 of the first image 31 is located at the bottom right of the first image 31 and the first part 331 of the third image 33 is located at the top left of the third image 33.
[0078] According to a particular exemplary embodiment, the first part 311 of the first image 31, the third part 313 of the first image, the first part 321 of the second image 32 and the first part 331 of the third image 33 associated with one of the stereoscopic fields of vision 112, 113 or 114 represent at least 50%, 60% or 80% of the pixels of the respective complete image 11, 12, 13.
[0079] In a second operation, first disparities associated with a first set of pixels of the first image are determined from the first set of pixels and a second set of pixels of the second image corresponding to the first set of pixels.
[0080] Similarly, in a third operation, second disparities associated with a third set of pixels of the second image are determined from the third set of pixels and a fourth set of pixels of the first image corresponding to the third set of pixels.
[0081] The first and second disparities are predicted by an optical flow method implemented by a first convolutional neural network (in English "Convolutional Neural Networks", CNN), this type of tool is commonly used in image processing. Thus, the second and third operations are carried out by implementing a method called optical flow calculation. Such a method is notably described in "PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume" by Deqing Sun, Xiaodong Yang, Ming-Yu Liu and Jan Kautz of June 25, 2018 and in "RAFT: Recurrent Ail-Pairs Field Transforms for Optical Flow" by Zachary Teed and Jia Deng of August 25, 2020.
[0082] An intermediate data item of the second operation is an optical flow representative of a displacement between each pixel of the first set of pixels and a second pixel of the first image 31 corresponding to each first pixel in the second image 32, the set of second pixels corresponding to the second set of pixels.
[0083] Similarly, an intermediate data item of the third operation is an optical flow representative of a displacement between each pixel of the third set of pixels and a fourth pixel of the second image 32 corresponding to each third pixel in the first cropped image 31b, the set of fourth pixels corresponding to the fourth set of pixels.
[0084] Indeed, the optical flow calculation makes it possible to process the first 31 and second 32 images in both directions without additional operation.
[0085] In a fourth operation, first corrected disparities are determined by correcting the first disparities as a function of a disparity induced by a relative position of the first 11 and second 12 cameras and in a fifth operation second corrected disparities are determined by correcting second disparities as a function of a disparity induced by a relative position of the first 11 and second 12 cameras.
[0086] Indeed, the rotation and translation to pass from the optical axis of the first camera 11 to the second camera 12, and vice versa, create in particular a so-called “induced” disparity. This disparity is, moreover, included in the optical flow determined from the first 31 and second 32 images.
[0087] The rotation is for example due to the fact that the optical axes of the first 11 and second 12 cameras are not parallel. The translation is for example due to the fact that the first 11 and second 12 cameras are not placed at the same height on the vehicle 10.
[0088] According to a particular exemplary embodiment, a disparity induced by a relative position of two cameras of said set of cameras is determined by a second convolutional neural network, here by the relative position of the first 11 and second 12 cameras.
[0089] According to a particular exemplary embodiment, in order not to take into account this induced disparity likely to distort the prediction of first and second depths, a convolutional neural network of the FCN type (from the English “Fully Convo-lutional network”), subsequently called an FCN, is implemented to determine the induced disparity and filter the optical flow value before predicting the depth of a pixel.
[0090] Thus, according to this particular exemplary embodiment, in a sixth operation, an induced disparity is determined by the FCN from the first 31 and second 32 images making it possible to subtract the induced disparity from the optical flow to obtain first and second corrected disparities.
[0091] The input data of such an FCN is a one-dimensional matrix HxWxN with H the height of an image, W the width of an image and N a number of channels including the unnullified elements of the rotation matrix (maximum 9), the unnullified elements of the translation vector (maximum 3) as well as two optical flows (horizontal and vertical). The maximum number of channels is therefore 14. The number of intermediate layers and the number of channels of the matrix at each layer depends on the deviation of the parallel system.
[0092] The structure of this FCN is constructed by experimentation in such a way as to minimize the size of the FCN while providing an acceptable result.
[0093] In a seventh operation, first and second depths associated respectively with the first and third sets of pixels are determined from the first and second corrected disparities respectively.
[0094] A depth is determined from a disparity by the following function: [Math.l] Vf / “IPJ
[0095] With: • Ds(pj the first depth, respectively second depth, associated with a pixel Pt of the first image 31, respectively with a pixel Pt of the second image 32, • f the focal length fl of the first camera 11, respectively the focal length f2 of the second camera 12, • B is the distance B12 between the first 11 and second 12 cameras, and * d( p is the first corrected disparity associated with a pixel Pt of the first image 31, respectively the second corrected disparity associated with a pixel Pt of the second image 32.
[0096] In an eighth operation, a fourth image is generated from the first image and the first depths and a fifth image is generated from the second image and the second depths.
[0097] According to a particular embodiment, an image is generated from an input image and depths associated with the pixels of the input image so as to reproduce an image seen by another camera from the following formula: [Math.2] the D* ( Pt ) ) ] )
[0098] With: • Ps one pixel of the generated image, the fourth image, respectively the fifth image, • 77 a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • K an intrinsic matrix of the other camera, the second camera 12, respectively the first camera 11, • F an extrinsic matrix of the stereoscopic vision system formed by the camera having acquired the input image and the other camera, here the extrinsic matrix linked to the relative positions of the first 11 and second 12 cameras, • 0 a projection function in the three-dimensional scene of a pixel as a function of its depth, the first depth, respectively the second depth, • K' an intrinsic matrix of the camera having acquired the input image, the first camera 11, respectively the second camera 12, • D^p ) is a depth of the pixel Pt associated with the input image, the first image 31, respectively the second image 32.
[0099] Such a method is notably described in the document “Digging Into Self-Su-pervised Monocular Depth Estimation” by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow published in 2019. It should be noted, however, that the method used in the cited document is applied to a monoscopic vision system, i.e. comprising a single moving camera.
[0100] The principle of image generation consists of projecting a pixel of the first image 31 acquired by the first camera 11 towards a point in the three-dimensional scene as a function of its first depth, then reprojecting this point into an image such as it would be acquired from the point of view of the second camera 12 and, respectively, projecting a pixel of the second image 32 acquired by the second camera 12 towards a point in the three-dimensional scene as a function of its second depth, then reprojecting this point into an image such as it would be acquired from the point of view of the first camera 11.
[0101] In a ninth operation, a first error is determined by comparing the fourth and second images and a second error is determined by comparing the fifth and first images.
[0102] According to a first particular exemplary embodiment, the first and respectively second errors are determined by the following loss function: [Math.3] ^(p) = LP[(i-«) ■ Uip)-Hp)\+a'
[0103] With: • L*(p) the first, respectively second error, • Kp) a value of the pixel P in the second image 32, respectively the first image 31, • a value of the pixel P in the fourth image, respectively the fifth image' • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.
[0104] This loss function is based on the photometric error. Once a fourth image, or respectively fifth image, generated, this is compared with the real image acquired by the stereoscopic vision system, the second image 32 or respectively the first image 31.
[0105] According to a second particular exemplary embodiment, the first and respectively second errors are determined by the following loss function: [Math.4]
[0106] With: • LmMOth^0. W, o) the first, respectively second error, • D t( n ) is the second depth, respectively first depth Dtm t' of a pixel Pt predicted by the stereoscopic vision system; • VF is a parameter matrix; • 0 is the order of a smoothing gradient; • an L1 norm of the second-order depth gradients is calculated with VF =1, and ° =2; • and 3' are the dimensions of the second image, respectively first image; • is an environment-dependent hyperparameter; and • / ( p ) is the value of pixel Pt in the fourth image, respectively fifth image.
[0107] This second loss function is generally used to deal with discontinuity at the edge of objects (in English “edge aware smoothness”).
[0108] According to a third particular exemplary embodiment, the first and second errors are determined by a function combining the function presented in the first particular exemplary embodiment and by the function presented in the second particular exemplary embodiment, for example by performing the sum of the two functions or by applying one or the other of the functions according to a location of the pixels of the first 31 and second 32 images.
[0109] In a tenth operation, third disparities associated with a fifth set of pixels of the first image are determined from the fifth set of pixels and a sixth set of pixels of the third image corresponding to the fifth set of pixels and fourth disparities associated with a seventh set of pixels of the third image are determined from the seventh set of pixels and an eighth set of pixels of the first image corresponding to the seventh set of pixels.
[0110] As before, the third and fourth disparities are predicted by an optical flow method implemented by the first convolutional neural network.
[0111] In an eleventh operation, third and fourth corrected disparities are determined by correcting the third and fourth disparities based on a disparity induced by a relative position of the first 11 and third 13 cameras.
[0112] According to a particular exemplary embodiment, as previously, a disparity induced by a relative position of two cameras of the set of cameras is determined by the second convolutional neural network.
[0113] In a twelfth operation, third and fourth depths associated respectively with the fifth and seventh set of pixels are determined from respectively the third and fourth disparities corrected by the following function: [Math.5] DJp) = fx- / -J d(pt)
[0114] With: • D^p) the third depth, respectively fourth depth, associated with a pixel Pt of the first image 31, respectively third image 33, • / the focal length fl of the first camera 11, respectively the focal length f3 of the third camera 13, • B is the distance B13 between the first 11 and third 13 cameras, and • d( p ) is the third corrected disparity associated with a pixel Pt of the first image 31, respectively the fourth corrected disparity associated with a pixel Pt of the third image 33.
[0115] In a thirteenth operation, a sixth image is generated from the first image and the third depths and a seventh image is generated from the third image and the fourth depths.
[0116] According to a particular embodiment, an image is generated from an input image and depths associated with the pixels of the input image so as to reproduce an image seen by another camera from the following formula: [Math.6] Ps= M ) ] )
[0117] With: • Ps one pixel of the generated image, the sixth image, respectively the seventh image, •77 a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • K an intrinsic matrix of the other camera, the third camera 13, respectively the first camera 11, • P an extrinsic matrix of the stereoscopic vision system formed by the camera having acquired the input image and from the other camera, here the extrinsic matrix linked to the relative positions of the first 11 and third 13 cameras, • 0 a projection function in the three-dimensional scene of a pixel according to its depth, the third depth, respectively the fourth depth, • K' an intrinsic matrix of the camera having acquired the input image, the first camera 11, respectively the third camera 13, • D^p ) is a depth of the pixel Pt associated with the input image, the first image 31, respectively the third image 33.
[0118] In a fourteenth operation, a third error is determined by comparing the sixth and third images and a fourth error is determined by comparing the seventh and first images;
[0119] According to the first particular embodiment, the third and respectively fourth errors are determined by the following loss function: [Math.7] L*(p) = Ep[(la) ■ \I(p) -î(p)\+a- (l-±SSIM(l(p)^
[0120] With: • L*(p) read third, respectively fourth error, • Kp) a value of the pixel P in the third image 33, respectively the first image 31, * A pixel value in the sixth image, respectively the seventh image • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.
[0121] Once a sixth image, or respectively seventh image, has been generated, it is compared with the true image acquired by the stereoscopic vision system, the third image 33 or respectively the first image 31.
[0122] According to the second particular embodiment, the third and respectively fourth errors are determined by the following loss function: [Math. 8] l^PUp,)- w ' °) = w (p,) । )
[0123] With: • L smooih ( O, W. o) the third, respectively fourth error, • p) is the fourth depth, respectively third depth Dtm of a pixel Pt predicted by the stereoscopic vision system; • IV is a parameter matrix; • 0 is the order of a smoothing gradient; • an L1 norm of second-order depth gradients is calculated with W =1, and0 =2; • x and y are the dimensions of the third image, respectively first image; • is an environment-dependent hyperparameter; and • [ ( nj is the value of pixel Pt in the sixth image, respectively seventh image.
[0124] This second loss function is generally used to deal with discontinuity at the border of objects.
[0125] According to the third particular exemplary embodiment, the third and fourth errors are determined by a function combining the function presented in the first particular exemplary embodiment and by the function presented in the second particular exemplary embodiment, for example by performing the sum of the two functions or by applying one or the other of the functions according to a location of the pixels of the first 31 and third 33 images.
[0126] The operations for predicting the first, second, third and fourth depths are further implemented by a self-supervised convolutional neural network. This learning from the data obtained from the stereoscopic vision system makes it possible to refine the parameters of the convolutional neural network and thus make the intermediate data such as the disparities or optical flows or output such as the first, second, third and fourth depths more reliable.
[0127] Such a convolutional neural network is known to those skilled in the art, such as PWCNet (from the English “Pyramidal processing, Warping, and the use of a Cost volume Network”).
[0128] In a fifteenth operation, input parameters of the first convolutional neural network are adjusted by minimizing a fifth error, the fifth error being determined from an average of the first, second, third and fourth errors.
[0129] According to a particular exemplary embodiment, the fifth error is determined by the following loss function: [Math.9] Ls = Epavg(L](p\ L2(p), L3(p), L / p))
[0130] With: • £' the fifth error, • avS an average of arguments, • L^p) the first error for a pixel P of the second image, • the second error for a pixel P of the first image, • L^p) the third error for a pixel P of the third image, and • L4(p) the fourth error for a pixel P of the first image.
[0131] When generating an image, the reconstruction of pixels from an area not visible by another camera is impossible (monocular vision), generating an image reconstruction error in this area and not contributing to a good adjustment of the input parameters of the first convolutional neural network. The averaging of the first, second, third and fourth errors also makes it possible to take into consideration each of the first, second, third and fourth errors in an equivalent manner but also to limit the impact of a reconstruction error in an area seen by a single camera.
[0132] Furthermore, the averaging of the first, second, third and fourth errors makes it possible to consider the generation of an image in one direction and the other in an equivalent manner, that is to say that the averaging of the first and second errors makes it possible to take into consideration an error in the reconstruction of the fourth image and the fifth image equally. Similarly, the averaging of the third and fourth errors makes it possible to take into consideration an error in the reconstruction of the sixth image and the seventh image equally.
[0133] [Fig.2] illustrates a flowchart of the different steps of a method 2 for determining a depth by a stereoscopic vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention.
[0134] The method 2 is for example implemented by one or more processors of one or more computers on board the vehicle 10, for example by a computer controlling the stereoscopic vision system.
[0135] In a step 21, data representative of a set of three images comprising a first 31, a second 32 and a third 33 images acquired respectively by a first 11, a second 12 and a third 13 cameras of the set of cameras at the same acquisition time instant are received.
[0136] In a step 22, first disparities associated with a first set of pixels of the first image are determined from the first set of pixels and a second set of pixels of the second image corresponding to the first set of pixels and second disparities associated with a third set of pixels of the second image are determined from the third set of pixels and a fourth set of pixels of the first image corresponding to the third set of pixels.
[0137] The first and second disparities are predicted by an optical flow method implemented by a first convolutional neural network.
[0138] In a step 23a, first and second corrected disparities are determined by correcting the first and second disparities as a function of a disparity induced by a relative position of the first 11 and second 12 cameras.
[0139] In a step 24a, first and second depths associated respectively with the first and third sets of pixels are determined from the first and second corrected disparities respectively.
[0140] In a step 25a, a fourth image is generated from the first image and the first depths and a fifth image is generated from the second image and the second depths.
[0141] In a step 26a, a first error is determined by comparing the fourth and second images and a second error is determined by comparing the fifth and first images.
[0142] In a step 22b, third disparities associated with a fifth set of pixels of the first image are determined from the fifth set of pixels and a sixth set of pixels of the third image corresponding to the fifth set of pixels and fourth disparities associated with a seventh set of pixels of the third image are determined from the seventh set of pixels and an eighth set of pixels of the first image corresponding to the seventh set of pixels.
[0143] The third and fourth disparities are predicted by an optical flow method implemented by the first convolutional neural network.
[0144] In a step 23b, third and fourth corrected disparities are determined by correcting the third and fourth disparities as a function of a disparity induced by a relative position of the first 11 and third 13 cameras.
[0145] In a step 24b, third and fourth depths associated respectively with the fifth and seventh sets of pixels are determined from the third and fourth corrected disparities respectively.
[0146] In a step 25b, a sixth image is generated from the first image and the third depths and a seventh image is generated from the third image and the fourth depths.
[0147] In a step 26b, a third error is determined by comparing the sixth and third images and a fourth error is determined by comparing the seventh and first images.
[0148] In a step 27, input parameters of the first convolutional neural network are adjusted by minimizing a fifth error, the fifth error being determined from an average of the first, second, third and fourth errors.
[0149] Thus, the depth predictions by the vehicle-mounted stereoscopic vision system are refined, improving the reliability of the output data from the stereoscopic vision system.
[0150] An ADAS using depths as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another user present on the road, the ADAS is then able to determine this distance precisely. For example, if FADAS has the function of acting on a braking system of the vehicle 10 in the event of a risk of collision with another road user and the distance separating the vehicle 10 from this same road user decreases sharply, then the ADAS is able to detect this sudden approach and act on the braking system of the vehicle 10 to avoid a possible accident.
[0151] [Fig. 5] schematically illustrates a device 4 configured for the determination of a depth by a stereoscopic vision system on board a vehicle 10, according to a particular and non-limiting exemplary embodiment of the present invention. The device 4 corresponds for example to a device on board the first vehicle 10, for example a computer.
[0152] The device 4 is for example configured for the implementation of the operations described with regard to figures 1, 3 and 4 and / or the steps described with regard to [Fig.2]. Examples of such a device 4 include, but are not limited to, on-board electronic equipment such as an on-board computer of a vehicle, an electronic calculator such as an ECU (“Electronic Control Unit”), a smartphone, a tablet, a laptop. The elements of the device 4, individually or in combination, can be integrated in a single integrated circuit, in several integrated circuits, and / or in discrete components. The device 4 can be produced in the form of electronic circuits or software (or computer) modules or even a combination of electronic circuits and software modules.
[0153] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the method and / or for executing the instructions of the software(s) embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41 corresponding for example to a volatile and / or non-volatile memory and / or comprises a memory storage device which may comprise volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.
[0154] The computer code of the embedded software(s) comprising the instructions to be loaded and executed by the processor is for example stored in the 4L memory.
[0155] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Telematic Control Unit”), for example via a communication bus or through dedicated input / output ports.
[0156] According to a particular and non-limiting exemplary embodiment, the device 4 comprises a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 comprise one or more of the following interfaces: - RF radio frequency interface, for example Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English “Universal Serial Bus” or “Universal Serial Bus” in French); HDMI interface (from the English “High Definition Multimedia Interface” or “High Definition Multimedia Interface” in French); - LIN interface (from the English “Local Interconnect Network”).
[0157] According to another particular and non-limiting exemplary embodiment, the device 4 comprises a communication interface 43 which makes it possible to establish communication with other devices (such as other computers of the on-board system) via a communication channel 430. The communication interface 43 corresponds for example to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds for example to a wired network of the CAN (Controller Area Network) type, CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by the ISO 17458 standard) or Ethernet (standardized by the ISO / IEC 802-3 standard).
[0158] According to a particular and non-limiting exemplary embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch-sensitive or not, one or more speakers 450 and / or other peripherals 460 (projection system) via the output interfaces 44, respectively. 45, 46. According to a variant, one or other of the external devices is integrated into the device 4.
[0159] Of course, the present invention is not limited to the exemplary embodiments described above but extends to a method for determining a depth by a stereoscopic vision system on board a vehicle, which would include secondary steps without thereby departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0160] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-based motor vehicle, comprising the device 4 of [Fig.5].
Claims
1. Claims Method for determining a depth by a stereoscopic vision system on board a vehicle (10), the stereoscopic vision system comprising a set of cameras of at least three cameras arranged so as to each acquire an image of a three-dimensional scene from a different point of view, said method being characterized in that it comprises the following steps: - reception (21) of data representative of a set of three images comprising a first, a second and a third images acquired respectively by a first (11), a second (12) and a third (13) cameras of said set of cameras at the same acquisition time instant; - determining (22a) first disparities associated with a first set of pixels of the first image from said first set of pixels and a second set of pixels of the second image corresponding to said first set of pixels and determining second disparities associated with a third set of pixels of the second image from said third set of pixels and a fourth set of pixels of the first image corresponding to said third set of pixels, said first and second disparities being predicted by an optical flow method implemented by a first convolutional neural network; - determination (23a) of first and second corrected disparities by correcting the first and second disparities as a function of a disparity induced by a relative position of the first (11) and second (12) cameras; - determining (24a) first and second depths associated respectively with the first and third sets of pixels from respectively the first and second corrected disparities; - generation (25a) of a fourth image from said first image and the first depths and generation of a fifth image from said second image and the second depths; - determination (26a) of a first error by comparing the fourth and second images and determination of a second error by comparing the fifth and first images;
2. - determining (22b) third disparities associated with a fifth set of pixels of the first image from said fifth set of pixels and a sixth set of pixels of the third image corresponding to said fifth set of pixels and determining fourth disparities associated with a seventh set of pixels of the third image from said seventh set of pixels and an eighth set of pixels of the first image corresponding to said seventh set of pixels, said third and fourth disparities being predicted by an optical flow method implemented by said first convolutional neural network; - determination (23b) of third and fourth disparities corrected by correction of the third and fourth disparities as a function of a disparity induced by a relative position of the first (11) and third (13) cameras; - determination (24b) of third and fourth depths associated respectively with the fifth and seventh sets of pixels from respectively the third and fourth corrected disparities; - generation (25b) of a sixth image from said first image and the third depths and generation of a seventh image from said third image and the fourth depths; - determination (26b) of a third error by comparing the sixth and third images and determination of a fourth error by comparing the seventh and first images; - adjusting (27) input parameters of said first convolutional neural network by minimizing a fifth error, said fifth error being determined from an average of said first, second, third and fourth errors. The method of claim 1, wherein said fifth error is determined by the following loss function: Ls = Epavg ( Lx ( p ), L2 ( p ), L3 ( p ), L4(p) ) With: • p the fifth error, • avS an average of arguments, • L / p) the first error for a pixel P of the second image, • L>) the second error for a pixel P of the first image, • L ip) the third error for a pixel P of the third image, and • L^p) the fourth error for a pixel P of the first image.
3. Method according to one of claims 1 to 2, for which an image is generated from an input image and depths associated with the pixels of the input image so as to reproduce an image seen by another camera from the following formula: ) ] )Av“: • Ps a pixel of the generated image, • 77 a function for going from homogeneous coordinates to pixel coordinates by removing one dimension of a vector, • K an intrinsic matrix of the other camera, • F an extrinsic matrix of the stereoscopic vision system formed by the camera having acquired the input image and the other camera, • 0 a projection function in the three-dimensional scene of a pixel as a function of its depth, • K' an intrinsic matrix of the camera having acquired the input image, • Dst^p ) is a depth of the pixel Pt associated with the input image.
4. Method according to one of claims 1 to 3, for which an error is determined by the following loss function: L^p) = Ep[(la) ' \I(p)-I(p)\+a - (l-±SSIM[l(p),î(p)))] With: • L*(p) the error associated with a pixel P, • Kp) a value of the pixel P in an image acquired by a camera, • a value of the pixel P in a generated image' • SSIM a function which takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.
5. Method according to one of claims 1 to 3, for which an error is determined by the following loss function: W. o) = Wfo)|W^P,) With: ,Lsmooth(^ W,o) the error, • DstÇp ) is a depth of a pixel Pt predicted by the stereoscopic vision system; • W is a parameter matrix; • ° is the order of a smoothing gradient; • an L1 norm of second-order depth gradients is computed with W =1 , and ü =2 ; • x and y are the dimensions of an image; • is an environment-dependent hyperparameter; and • Iti.Pi ) is the value of pixel Pt in a generated image.
6. Method according to one of claims 1 to 5, for which a disparity induced by a relative position of two cameras of said set of cameras is determined by a second convolutional neural network.
7. Computer program comprising instructions for implementing the method according to any one of the preceding claims, when these instructions are executed by a processor.
8. A computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to claims 1 to 6.
9. Device (4) for determining a depth by a stereoscopic vision system on board a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for implementing the steps of the method according to any one of claims 1 to 6.
10. Vehicle (10) comprising the device (4) according to claim 9.