Method and device for determining depth by a self-supervised vision system.
The method enhances ADAS systems by using a stereoscopic and monoscopic vision system with convolutional neural networks to improve depth prediction accuracy, addressing data quality issues and enhancing vehicle safety.
Patent Information
- Application Number
- FR2023010304
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-09-28
AI Technical Summary
The quality of data obtained by vehicle cameras affects the performance of Advanced Driver-Assistance Systems (ADAS), leading to potential safety issues due to unreliable depth perception and object detection.
A method and device using a stereoscopic and monoscopic vision system with convolutional neural networks to predict depths by analyzing optical flow and disparity values, and adjusting network parameters through image generation and error minimization to enhance depth prediction accuracy.
Improves the reliability of depth perception and object detection for ADAS systems, providing more precise and reliable data for vehicle safety systems.
Smart Images

Figure 00000030_0000 
Figure 00000031_0000 
Figure 00000031_0001
Abstract
Description
Title of the invention: Method and device for determining a depth by a self-supervised vision system. Technical field
[0001] The present invention relates to methods and devices for determining a depth by a vision system on board a vehicle, for example in a motor vehicle. The present invention also relates to a method and a device for measuring such a depth. The present invention also relates to a method and a device for controlling one or more AD AS systems on board a vehicle from the determined depth. Technological background
[0002] Many modern vehicles are equipped with so-called AD AS (Advanced Driver-Assistance System). Such AD AS systems are passive and active safety systems designed to eliminate the element of human error in the driving of vehicles of all types. AD AS use advanced technologies to assist the driver while driving and thus improve their performance. AD AS use a combination of sensor technologies to perceive the environment around a vehicle, then provide information to the driver or act on certain vehicle systems.
[0003] There are several levels of ADAS, such as rearview cameras and blind spot sensors, lane departure warning systems, adaptive cruise control and automatic parking systems.
[0004] The AD AS embedded in a vehicle are supplied with data obtained from one or more embedded sensors such as, for example, cameras. These cameras make it possible in particular to detect and locate other road users or possible obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle speed; - to act on the braking system in the event of a risk of impact with an object.
[0005] The proper functioning of the driving assistance peripherals using this data therefore depends on the quality of the data emitted by a vision system. Summary of the present invention
[0006] An object of the present invention is to solve at least one of the problems of the technological background described above.
[0007] Another object of the present invention is to improve the quality of the data obtained of these cameras.
[0008] Another object of the present invention is to improve road safety, in particular by improving the operational safety of AD AS systems supplied by data obtained from at least one camera.
[0009] According to a first aspect, the present invention relates to a method for determining a depth by a vision system on board a vehicle, the vision system comprising a set of cameras of at least two cameras arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the method being characterized in that it comprises the following steps: - reception of first and second data respectively representative of a first and second image acquired by respectively a first and second camera of the set of cameras at the same first acquisition time instant; - prediction of first depths associated with a first set of pixels of the first image and second depths associated with a second set of pixels of the second image as a function of optical flow values and first disparity values induced by a relative position of optical axes of the first and second cameras, determined from the first and second sets of pixels, the first and second depths being predicted by a first convolutional neural network associated with a stereoscopic vision system comprising the first and second cameras; - generating a fourth image from the first image and the first depths and a fifth image from the second image and the second depths, the image generation being obtained by applying an image generation function; - training the first convolutional neural network by minimizing a first error, the first error being determined by comparing the first and fifth images and by comparing the fourth and second images; - reception of third data representative of a third image acquired by the first camera at a second acquisition time instant different from the first acquisition time instant of the first image; - determination of data representative of a movement of the first camera between the first time instant of acquisition and the second time instant of acquisition as a function of the first and third images; - prediction of third depths associated with a third set of pixels of the first image and of fourth depths associated with a fourth set of pixels of the third image as a function of optical flow values and second disparity values induced by a relative position of optical axes of the first camera at the first and second acquisition time instants determined from third and fourth sets of pixels, the third and fourth depths being predicted by a second convolutional neural network associated with a monoscopic vision system comprising the first camera; - generating a sixth image from the first image and the third depths and a seventh image from the third image and the fourth depths by applying the image generation function; - training the second convolutional neural network by minimizing a second error, the second error being determined by comparing the sixth and third images and by comparing the seventh and first images; - adjustment of parameters of the second convolutional neural network according to the first depths corresponding to training data.
[0010] Such a method makes it possible to predict more reliable first depths for the stereoscopic vision system, thus making it possible to provide the monoscopic vision system with better quality training data, called “annotated data”. The prediction of second depths by the monoscopic vision system is then more precise and achieves a metric precision available to the stereoscopic vision system.
[0011] A driving assistance system using the first or second depths as input data then has reliable data regarding the distance separating the vision system from the objects present in the field of vision of one of the cameras of this vision system.
[0012] According to a variant of the method, the values of first and second induced disparities are determined by a third convolutional neural network.
[0013] According to another variant of the method, the image generation function is defined by the following function, respectively for the stereoscopic vision system and for the monoscopic vision system: = 7F ( [ W ( p^, D st ( p f ) j ) With : • Ps a pixel of a generated image, • 71 a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • K an intrinsic matrix of the second camera, respectively of the first camera, • M an extrinsic matrix of stereoscopic vision system, respectively a matrix representing the movement of the first camera between the first and second time instants, • 0 a projection function in the three-dimensional scene of a pixel as a function of its depth, • K' an intrinsic matrix of the first camera, and • D^p ) is a depth of the pixel Pt determined by the first convolutional neural network, respectively by the second convolutional neural network.
[0014] According to another variant of the method, the first error is determined by a first loss function defined by: 4 - L p avg(L5X / A LM) Avec : • £0a first error, • avS an average of arguments, • Ls^p) an error for a pixel P of the first image compared to the fifth image, • Lx / p) an error for a pixel P of the second image compared to the fourth image.
[0015] According to yet another variant of the method, the first error is determined by a first loss function defined by: L, = LM)With: • £0a first error, • min a minimum value of an argument, • Ls^p) an error for a pixel P of the first image compared to the fifth image, • LSi(p) an error for a pixel P of the second image compared to the fourth image.
[0016] According to another variant, the second error is determined by a second loss function defined by: Lm = L / àVg(Lmi(p\ L„M)With: • y the second error, • av8 an average of arguments, • Lmfô)) an error for a pixel P of the first image compared to the seventh image, • Lm^p) an error for a pixel P of the third image compared to the sixth image.
[0017] According to an additional method variant, the second error is determined by a second loss function defined by: L m = L fJ min(L mf (p), L ms (p)) Avec : • T the second error, • min a minimum value of an argument, • Lm£p} an error for a pixel P of the first image compared to the seventh image, • Lms(p) an error for a pixel P of the third image compared to the sixth image.
[0018] According to another variant of the method, the third image is replaced by a black image, a black image being an image of which a set of colorimetric values associated with the pixels of the image is zero.
[0019] According to a second aspect, the present invention relates to a device for determining a depth by a vision system on board a vehicle, the device comprising a memory associated with at least one processor configured for implementing the steps of the method according to the first aspect of the present invention.
[0020] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0021] According to a fourth aspect, the present invention relates to a computer program which comprises instructions adapted for executing the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0022] Such a computer program may use any programming language and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0023] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to the first aspect of the present invention.
[0024] On the one hand, the recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM memory, a CD-ROM or a microelectronic circuit type ROM memory, or a magnetic recording means or a hard disk.
[0025] Furthermore, this recording medium may also be a transmissible medium such as an electrical or optical signal, such a signal being able to be conveyed via an electrical or optical cable, by conventional or hertzian radio or by self-directed laser beam or by other means. The computer program according to the present invention may in particular be downloaded from an Internet-type network.
[0026] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to perform or to be used in performing the method in question. Brief description of the figures
[0027] Other characteristics and advantages of the present invention will emerge from the description of the particular and non-limiting exemplary embodiments of the present invention below, with reference to the appended figures 1 to 3, in which:
[0028] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting example of embodiment of the present invention;
[0029] [Fig.2] illustrates a flowchart of the different steps of a method for determining a depth by a vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;
[0030] [Fig.3] schematically illustrates a device configured for determining a depth by a vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention. Description of the exemplary embodiments
[0031] A method and a device for determining a depth by means of a vision system on board a vehicle will now be described in the following with joint reference to FIGS. 1 to 3. The same elements are identified with the same reference signs throughout the description which follows.
[0032] The terms "first(s)", "second(s)" (or "first(s)", "second(s)"), etc. are used in this document by arbitrary convention to enable different elements (such as operations, means, etc.) implemented in the embodiments described below to be identified and distinguished. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0033] According to a particular and non-limiting example of embodiment of the present invention, a method for determining a depth by a vision system on board a vehicle is for example implemented by a computer of the on-board system of the vehicle controlling this vision system.
[0034] The vision system comprises a set of cameras of at least two cameras arranged so as to each acquire an image of a scene from a different point of view.
[0035] For this purpose, the method for determining a depth by a vision system on board a vehicle comprises the steps of receiving first, second and third data respectively representative of first, second and third images.
[0036] The first and second images are acquired respectively by the first and second cameras at the same first acquisition time instant. The first and second cameras then form a stereoscopic vision system.
[0037] The third image is acquired by the first camera at a second acquisition time instant different from the first acquisition time instant. The first camera then forms a monoscopic vision system, the first camera having different viewpoints of the three-dimensional scene depending on a movement of the vehicle 10.
[0038] According to a first particular embodiment, the second acquisition time instant is prior to the first acquisition time instant.
[0039] According to a second particular embodiment, the second acquisition time instant is later than the first acquisition time instant.
[0040] First and second depths are predicted by a first convolutional neural network associated with the stereoscopic vision system and third and fourth depths are predicted by a second convolutional neural network associated with the monoscopic vision system.
[0041] Fourth and fifth images are generated from, respectively, the first and second images and the first and second depths by applying an image generation function.
[0042] Similarly, sixth and seventh images are generated from, respectively, the first and third images and the third and fourth depths by applying the image generation function.
[0043] Each vision system is then learned by minimizing errors determined by comparing the generated images to acquired images and the predictions of second depths by the second convolutional neural network associated with the monoscopic vision system are then improved by adjusting the parameters of the second convolutional neural network as a function of the first depths predicted by the first convolutional neural network, the stereoscopic vision system providing learning data, called “annotated data”.
[0044] [Fig. 1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.
[0045] Such an environment 1 corresponds, for example, to a road environment formed of a network of roads accessible to the vehicle 10.
[0046] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, with an electric motor(s) or even a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as an automobile, a truck, a bus, a motorcycle. Finally, the vehicle 10 corresponds to an autonomous vehicle or not, that is to say a vehicle traveling according to a determined level of autonomy or under the full supervision of the driver.
[0047] The vehicle 10 advantageously comprises several on-board cameras 11, 12, each configured to acquire images of a scene in the environment of the vehicle 10. This set of cameras 11, 12 forms the stereoscopic vision system. Two cameras 11 and 12 are illustrated in [Fig.l]. The present invention is however not limited to a stereoscopic vision system comprising two cameras but extends to any vision system comprising 2 or more cameras, for example 2, 3, 4 or 5 cameras.
[0048] The two cameras 11, 12 have known intrinsic parameters. These parameters consist in particular of: - the focal length fl of the first camera 11; - the focal length f2 of the second camera 12; - distortions which are due to imperfections in the optical system of each camera; - the direction Cl of the optical axis of the first camera 11; - the direction C2 of the optical axis of the second camera 12; and - the respective resolutions of cameras 11, 12.
[0049] The intrinsic parameters characterize the transformation which associates, for an image point, the camera coordinates with the pixel coordinates, in each camera. These parameters do not change if the camera is moved.
[0050] The distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions which generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, that is to say that the distortions are not taken into account or that their correction is processed at the time of image acquisition.
[0051] These two cameras 11, 12 are arranged so as to each acquire an image of a scene from a different point of view, the first point of view is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10, the second point of view is for example located on or in the right rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. In the case where the two cameras are located at the top of the windshield of the vehicle, they are then placed at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, the second camera 12 is located in the right rearview mirror of the vehicle 10.
[0052] A first reference point is associated with the first camera 11: - the direction of the y axis is defined by the position of the second camera 12, so as to place the second camera 12 on the y axis of the first camera 11. The distance B separating the two cameras 11, 12 is called the reference base (in English “baseline”) and the direction separating the two cameras 11, 12 is that of the y axis; - the direction of the x axis is defined orthogonal to that of the y axis and orthogonal to that of the optical axis Cl of the first camera 11; - the direction of the z axis is defined orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal reference frame.
[0053] The extrinsic parameters linked to the position of the cameras 11, 12 are the following parameters: - 3 translations in the x, y and z directions: Tx, Ty and Tz constituting the translation vector T; and - 3 rotations around the x, y and z axes: Rx, Ry and Rz, constituting the rotation matrix R.
[0054] Determining the extrinsic parameters constitutes the problem of calibrating a stereoscopic vision system.
[0055] A main constraint of the stereoscopic vision system used in automobiles is, for example, the large distance between the two cameras. Indeed, to be able to cover a measurement range of 200 meters, the reference base must reach 60cm for the cameras commonly used in this field.
[0056] The two cameras 11, 12 acquire images of a scene located in front of the vehicle 10, the first camera 11 covering only a first acquisition field 13, the second camera 12 covering only a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic vision of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic vision of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic vision of the scene by the stereoscopic vision system composed of the two cameras 11, 12.
[0057] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.
[0058] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system composed of the second camera 12.
[0059] Field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable using the monoscopic vision system composed of the first camera 11.
[0060] Finally, field 19 is not visible from any of the cameras. The part of the scene present in this field 19 is therefore not observable.
[0061] According to an exemplary embodiment, the directions C1, C2 of the optical axes representative of an orientation of the field of vision of each camera are oriented non-parallel so as to obtain the third acquisition field 15 of the environment 1 as wide as possible.
[0062] According to another exemplary embodiment, the directions C1, C2 of the optical axes representative of an orientation of the field of vision of each camera are oriented parallel.
[0063] It is obvious that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with differently placed and oriented cameras.
[0064] The images acquired by the cameras 11, 12 at a given acquisition time instant tl are presented in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colors and brightness of the objects in the observed scene in the form, for example, of RGB (from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness) colorimetric coordinates.
[0065] The images acquired by the cameras 11, 12 represent views of the same scene taken from different viewpoints, the positions of the cameras being distinct. On this scene are found, for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0066] These images are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.
[0067] A process for determining a depth by a vision system on board the vehicle 10 is advantageously implemented by the vehicle 10, that is to say by a computer or a combination of computers of the on-board system of the vehicle 10, for example by the computer(s) in charge of the stereoscopic vision system of the vehicle 10.
[0068] In a first operation, the computer receives first data representative of a first image acquired by a first camera 11 of the set of cameras at a first acquisition time instant.
[0069] In a second operation, the computer receives second data representative of a second image acquired by a second camera 12 of the set of cameras at the same first acquisition time instant.
[0070] The two images received correspond to two views of the same three-dimensional scene taking place around the vehicle 10 at the same first given acquisition time instant.
[0071] In a third operation, first depths associated with a first set of pixels of the first image are predicted based on optical flow values determined from the first set of pixels and a second set of pixels of the second image corresponding to the first set of pixels.
[0072] The fields of vision of the first 11 and the second 12 camera being concurrent, a part of the pixels of the first image and a part of the pixels of the second image are associated with the same objects present in the three-dimensional scene.
[0073] The third operation is carried out by implementing a method called optical flow calculation. Such a method is notably described in “PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume” by Deqing Sun, Xiaodong Yang, Ming-Yu Liu and Jan Kautz of June 25, 2018 and in “RAFT: Recurrent Ail-Pairs Field Transforms for Optical Flow” by Zachary Teed and Jia Deng of August 25, 2020.
[0074] The optical flow calculation method is implemented by a first convolutional neural network (CNN), a tool commonly used in image processing. An intermediate data item of the third operation is an optical flow representative of a displacement between each pixel of the first set of pixels and a second pixel corresponding to each first pixel in the second image, the set of second pixels corresponding to the second set of pixels.
[0075] The rotation and translation to pass from the optical axis of the first camera 11 to the second camera 12 create in particular a first disparity called “induced”. This first induced disparity is, moreover, included in the optical flow determined from the first and second images.
[0076] The rotation is for example due to the fact that the optical axes of the first 11 and second 12 cameras are not parallel. The translation is for example due to the fact that the first 11 and second 12 cameras are not placed at the same height on the vehicle 10.
[0077] According to a particular embodiment, in order not to take into account this first induced disparity likely to distort the prediction of first and second depths, a third convolutional neural network of the FCN type (from the English "Fully Convolutional network"), subsequently called an FCN, is implemented to determine the first induced disparity and filter the optical flow value before predicting the depth of a pixel.
[0078] The input data of such an FCN is a one-dimensional matrix HxWxN with H the height of an image, W the width of an image and N a number of channels including the unnullified elements of the rotation matrix (maximum 9), the unnullified elements of the translation vector (maximum 3) as well as the two optical flows (horizontal and vertical). The maximum number of channels is therefore 14. The number of intermediate layers and the number of channels of the matrix at each layer depends on the deviation of the parallel system.
[0079] The structure of this FCN is constructed by experimentation in such a way as to minimize the size of the FCN while providing an acceptable result.
[0080] The structure of this FCN is constructed by experimentation in such a way as to minimize the size of the FCN while providing an acceptable result.
[0081] Thus, according to this particular exemplary embodiment, in a fourth operation, a first induced disparity is determined by the FCN from the first and second images making it possible to subtract the first induced disparity from the optical flow to obtain a corrected optical flow, the first depths are then predicted as a function of this corrected optical flow, a depth being proportional to a norm of an optical flow from which the first induced disparity is subtracted.
[0082] In the same manner as for the third and fourth operations, in fifth and sixth operations, second depths associated with the second set of pixels are predicted based on optical flow values determined from the second set of pixels of the second image and the first set of pixels of the first image. The second depths are also predicted by the first convolutional neural network.
[0083] Such operations are permitted because the optical flow can be defined in both directions, i.e. from the first image to the second image and vice versa.
[0084] In a seventh operation, a fourth image is generated from the first image and the first predicted depths by applying an image generation function and in an eighth operation a fifth image is generated from the second image and the second depths by applying the same image generation function.
[0085] According to a particular exemplary embodiment, the image generation function is defined by the following function: [Math.l] Ps = tt (K [M¢ (pDst (pf))])
[0086] With: • Ps one pixel of the fourth image, respectively fifth image •77 a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • K an intrinsic matrix of the second camera 12, respectively of the first camera 11, • M an extrinsic matrix of the stereoscopic vision system, • 0 a projection function in the three-dimensional scene of a pixel as a function of its depth, • K' an intrinsic matrix of the first camera 11, respectively of the second camera 12, • D^pj is a first depth of the pixel Pt of the first image, respectively a second depth of a pixel Pt of the second image.
[0087] The principle of image generation consists of projecting a pixel of an image acquired by the first camera 11 (the first image) towards a point in the three-dimensional scene according to its first depth, then reprojecting this point into an image as it would be acquired from the point of view of the second camera 12 (the fourth image). Thus, if the first depth value associated with a pixel of the first image is correct, then the point of the scene is well defined and the second image should have, at the location of this projected point of the scene, a pixel of the same value.
[0088] The same applies in the other direction, that is to say by projecting a pixel of an image acquired by the second camera 12 (the second image) towards a point in the three-dimensional scene as a function of its second depth, then reprojecting this point into an image as it would be acquired from the point of view of the first camera 11 (the fifth image).
[0089] In a ninth operation, the first convolutional neural network associated with the stereoscopic vision system comprising the first 11 and second 12 cameras is learned by minimizing a first error, the first error being determined by comparing the first and fifth images and by comparing the fourth and second images.
[0090] According to a first particular exemplary embodiment, a first reconstruction error for a pixel of a generated image is further determined by comparing the generated image to an image acquired by a function defined as follows: [Math.2] = EJ(la) ■ \I(p)-I(p)\+a-(1-jSSIM(i(p),I(p)))]
[0091] With: • a first reconstruction error for a pixel P, • Kp} a value of the pixel P in an acquired image (the first image, respectively the second image), • 2^ a value of the pixel P in a generated image (the fifth image, respectively the fourth image), • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.
[0092] This function is based on photometric error. Once an image is generated, it is compared to the true image acquired by the stereoscopic vision system.
[0093] According to a second particular exemplary embodiment, a second reconstruction error of a generated image is further determined by the following function: [Math.3] (p,). W,o) = H^P,)IV^Pj )
[0094] With: • LmMOth^O. W, o) a second reconstruction error, • Dt^ p ) is the first depth, respectively second depth Dt of a pixel Pt predicted by the stereoscopic vision system; • VF is a parameter matrix; • 0 is the order of a smoothing gradient; • an L1 norm of the second-order depth gradients is calculated with VF =1, and ° =2; •1 and 3' are the dimensions of the second image, respectively first image; • is an environment-dependent hyperparameter; and • / ( p ) is the value of pixel Pt in the fourth image, respectively fifth image.
[0095] This second loss function is generally used to deal with discontinuity at the edge of objects (in English “edge aware smoothness”).
[0096] The reconstruction error is thus defined from the first and second reconstruction errors previously defined.
[0097] According to a particular exemplary embodiment, the first error is determined by one of the following loss functions.
[0098] A first loss function is defined by: [Math.4] L s = E p avg(£ Ar ( / ?), Mp))
[0099] With: • The first mistake, • avS an average of arguments, • Lj^p) a reconstruction error for a pixel P of the first image compared to the fifth image, • Ls£p) a reconstruction error for a pixel P of the second image compared to the fourth image.
[0100] By calculating the average of the two reconstruction errors, the two predictions are treated equivalently, which improves the prediction reliability when one of the predictions is better in one direction than in the other (from the first image to the second image or vice versa), i.e. when the reliability of the first and second depths is not equivalent.
[0101] A second loss function is defined by: [Math.5] = lLpmm(Lsl{p\
[0102] With: • the first mistake, • min a selection of the minimum value among arguments, • Lst(p) a reconstruction error for a pixel P of the first image compared to the fifth image, • L^(p) a reconstruction error for a pixel P of the second image compared to the fourth image.
[0103] This second loss function is faster and allows the number of iterations in the learning operation to be reduced.
[0104] A third loss function is defined by: [Math.6] Ls=Ep(Lsl(p)+Ljp))
[0105] With: • the first mistake, • Lx£p) a reconstruction error for a pixel P of the first image compared to the fifth image, • Ls£p) a reconstruction error for a pixel P of the second image compared to the fourth image.
[0106] The operations for predicting the first and second depths are thus implemented by the first self-supervised convolutional neural network. Such a convolutional neural network is known to those skilled in the art, as PWCNet (from the English "Pyramidal processing, Warping, and the use of a Cost volume Network").
[0107] This learning also makes it possible to find or refine the correct values of the parameters (called "weights" in English) of a chosen convolutional neural network to thus make the first predicted depths more reliable.
[0108] In a tenth operation, third data representative of a third image acquired by the first camera 11 at a second acquisition time instant different from the first acquisition time instant of the first image are received.
[0109] The third data have, for example, been saved in a memory associated with the computer or in a memory of a device on board the vehicle 10 and accessible to the computer implementing the process.
[0110] If the vehicle 10 is moving, then the third image corresponds to a third view of the scene taken from a third point of view, that of the first camera 11 at its position at the second time instant of acquisition.
[0111] This position of the first camera 11 at the second acquisition time instant is defined by the movement of the vehicle 10 between the first and second acquisition time instants. This movement of the first camera 11 is therefore linked to the speed and direction of movement of the vehicle 10 during the time separating the first and second acquisition time instants.
[0112] The first camera 11, in the two positions defined at the first and second acquisition time instants, forms a monoscopic vision system. The intrinsic parameters of this system remain the same as those defined previously related to the first camera 11 for the stereoscopic vision system. The extrinsic parameters of this monoscopic vision system are the following parameters: - 3 translations in the x, y and z directions: Tx', Ty' and Tz' constituting the translation vector T'; and - 3 rotations around the x, y and z axes: Rx', Ry' and Rz', constituting the rotation matrix R'.
[0113] In an eleventh operation, data representative of a movement of the first camera 11 between the first acquisition time instant and the second acquisition time instant are determined as a function of the first and third images. The extrinsic parameters of the monoscopic vision system are, for example, determined by a computer associated with this same monoscopic vision system. The determination of the extrinsic parameters of the monoscopic vision system is known to those skilled in the art and presented, for example, in the document Unsupervised Leaming of Depth and Ego-Motion from Video by Tinghui Zhou, Matthew Brown, Noah Snavely and David G. Lowe published on August 1, 2017.
[0114] According to a particular exemplary embodiment, in a twelfth operation, the third image is replaced by a black image or by the first image when the third image is identical to the first image or there is no second image, a black image being an image of which a set of colorimetric values associated with the pixels of the image is zero or corresponds to a black-colored object.
[0115] The purpose of the black image here is to be able to predict the depth for the first image of a video according to the first embodiment, i.e. when there is no previous image.
[0116] The objective of the identical image is to be able to predict the depth when the vehicle is stationary.
[0117] Such an operation makes it possible to limit the time of the following operations and thus promotes the learning of the monoscopic vision system.
[0118] In a manner similar to that of the third, in a thirteenth operation, third depths associated with a third set of pixels of the first image are predicted as a function of optical flow and second disparity values induced by a relative position of optical axes of the first camera 11 at the first and second acquisition time instants determined from the third set of pixels and a fourth set of pixels of the third image corresponding to the third set of pixels.
[0119] The thirteenth operation is performed by implementing the method known as optical flow calculation, implemented by a second convolutional neural network associated with the monoscopic vision system comprising the first camera 11. An intermediate data item of the thirteenth operation is an optical flow representative of a displacement between each pixel of the third set of pixels and a fourth pixel corresponding to each third pixel in the second image, the set of fourth pixels corresponding to the fourth set of pixels.
[0120] The rotation and translation to move from the optical axis of the first camera 11 to the first time instant and to the second time instant create in particular a second induced disparity which is, moreover, included in the optical flow determined from the first and third images.
[0121] The third FCN type convolutional neural network is implemented in a fourteenth operation to determine the second induced disparity and filter the optical flow value before predicting the depth of a pixel.
[0122] It should be noted that the FCN training for the monoscopic vision system is longer than for a stereoscopic vision system. Indeed, the parameters ex intrinsic characteristics of the monoscopic vision system constantly change depending on the movements of the vehicle 10.
[0123] Thus, according to this particular exemplary embodiment, in a fourteenth operation, a second induced disparity is determined by the FCN from the first and third images making it possible to subtract the second induced disparity from the optical flow to obtain a corrected optical flow, the third depths are then predicted as a function of this corrected optical flow, a depth being proportional to a norm of an optical flow from which the induced disparity is subtracted.
[0124] In a similar manner to the thirteenth and fourteenth operations, fourth depths associated with the fourth set of pixels of the third image are predicted in a fifteenth operation as a function of optical flow and disparity values induced by a relative position of optical axes of the first camera 11 at the first and second acquisition time instants determined from the fourth set of pixels and the third set of pixels of the first image corresponding to the fourth set of pixels. The induced disparities are predicted by the FCN in a sixteenth operation.
[0125] The fourth depths are thus also predicted by the second convolutional neural network.
[0126] In a manner similar to that of the seventh and eighth operations, in a seventeenth operation a sixth image is generated from the first image and the third depths by applying the image generation function and, in an eighteenth operation, a seventh image is generated from the third image and the fourth depths.
[0127] According to a particular exemplary embodiment, the image generation function is defined by the following function for the monoscopic vision system: [Math.7] Ps = ^(K[M¢iP^ DÀPt] ) ] )
[0128] With: • Ps a pixel of the sixth image, respectively seventh image, • 77 a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • K an intrinsic matrix of the first camera 11 • M a matrix representing the movement of the first camera 11 between the first and second time instants, • 0 a projection function in the three-dimensional scene of a pixel as a function of its depth, * D^p ) is a third depth of the pixel Pt of the first image, respec- tively a fourth depth of the pixel Pt of the third image.
[0129] In a nineteenth operation, the second convolutional neural network associated with the monoscopic vision system comprising the first camera 11 is learned by minimizing a second error, the second error being determined by comparing the sixth and third images and by comparing the seventh and first images.
[0130] According to the first particular exemplary embodiment, a third reconstruction error is determined in a similar manner to the first reconstruction error by the following function: [Math. 8] L.(p) = L,,[ (ia) • | / (p)-r(p) I +«■ Up) )) ]
[0131] With: • LAp) a third reconstruction error for a pixel P, • I(p) a value of the pixel P in an acquired image (the first image, respectively the third image), * / (a value of pixel P in a generated image (the seventh image, respectively the sixth image), • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.
[0132] According to the second particular embodiment, a fourth reconstruction error is determined by the following function: [Math.9] W, o) = W( P,} )
[0133] With: • LP(h(O, W, o) the fourth reconstruction error, • Dt(p ) is the third depth Dt of a pixel Pt; respectively the fourth depth Dt of a pixel Pt, - W is a parameter matrix; - 0 is the order of a smoothing gradient; - an L1 norm of the second-order depth gradients is calculated with W =1, and° =2; - x and P are the dimensions of the first image; - is an environment-dependent hyperparameter; and ■ is the value of the pixel Pt in the sixth image, respectively seventh image.
[0134] According to a particular exemplary embodiment, the second error is determined by one of the following loss functions.
[0135] A fourth loss function defined by: [Math. 10] Lm = E / ?avg(Lm^p\ Lm^p))
[0136] With: • the second error, • avg an average of arguments, • Lmt(p) a reconstruction error for a pixel P of the first image compared to the seventh image, • Lmskp) a reconstruction error for a pixel P of the third image compared to the sixth image.
[0137] A fifth loss function is defined by: [Math. 11] Lm = Lms(p))
[0138] With: • £' the second error, • avS an average of arguments, • a reconstruction error for a pixel P of the first image compared to the seventh image, • LmXp) a reconstruction error for a pixel P of the third image compared to the sixth image.
[0139] A sixth loss function is defined by: [Math. 12] Lm 52^ ( Lm, ( p ) + )
[0140] With: • T the second error, • Lm^p) a reconstruction error for a pixel P of the first image compared to the seventh image, • Lm^p) a reconstruction error for a pixel P of the third image compared to the sixth image.
[0141] In a twentieth operation, the parameters of the second convolutional neural network are adjusted according to the first depths determined by the first convolutional neural network, the first depths then corresponding to the training data. Thus, the parameters of the second convolutional neural network are adjusted according to any method of learning from training data known to those skilled in the art.
[0142] The first convolutional neural network associated with the stereoscopic vision system thus predicts reliable first depths, thus making it possible to provide the second convolutional neural network associated with the monoscopic vision system with good quality training data.
[0143] This training data is used to adjust the parameters of the second convolutional neural network, according to any learning method known to those skilled in the art.
[0144] If the ADAS uses the first or second depths as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another user present on the road, the ADAS is then able to determine this distance precisely. For example, if the ADAS has the function of acting on a braking system of the vehicle 10 in the event of a risk of collision with another road user and the distance separating the vehicle 10 from this same road user decreases significantly, then the ADAS is able to detect this sudden approach and act on the braking system of the vehicle 10 to avoid a possible accident, including if the other user is seen by only one camera.
[0145] [Fig. 2] illustrates a flowchart of the different steps of a method 2 for determining a depth by a vision system on board a vehicle, for example the vehicle of [Fig. 1], according to a particular and non-limiting example of the present invention. The method 2 is for example implemented by a device on board the vehicle 10 or by the device 4 of [Fig. 3].
[0146] In a step 21, first data representative of a first image acquired by a first camera 11 of said set of cameras at a first acquisition time instant are received.
[0147] In a step 22, second data representative of a second image acquired by a second camera 12 of the set of cameras at the same first acquisition time instant are received.
[0148] In a step 23, first depths associated with a first set of pixels of the first image and second depths associated with a second set of pixels of the second image are predicted as a function of optical flow values and first disparity induced by a relative position of optical axes of the first 11 and second 12 cameras.
[0149] The optical flow and induced disparity values are determined from the first and second sets of pixels, the first induced disparity coming from a relative position of optical axes of the first 11 and second 12 cameras.
[0150] The first and second depths are then predicted by a first convolutional neural network associated with a stereoscopic vision system comprising the first 11 and second 12 cameras.
[0151] In a step 24, a fourth image is generated from the first image and the first depths by applying an image generation function and a fifth image is generated from the second image and the second depths, the image generation being obtained by applying the same image generation function.
[0152] In a step 25, the first convolutional neural network is learned by minimizing a first error determined by comparing the first and fifth images and by comparing the fourth and second images.
[0153] In a step 31, third data representative of a third image acquired by the first camera 11 at a second acquisition time instant are received.
[0154] The second acquisition time instant is different from the first acquisition time instant of the first image.
[0155] In a step 32, data representative of a movement of the first camera 11 between the first acquisition time instant and the second acquisition time instant are determined as a function of the first and third images.
[0156] In a step 33, third depths associated with a third set of pixels of the first image and fourth depths associated with a fourth set of pixels of the third image are predicted as a function of optical flow values and second induced disparity.
[0157] The optical flow and second induced disparity values are determined from the third and fourth sets of pixels, the second disparity being induced by a relative position of optical axes of the first camera 11 at the first and second acquisition time instants.
[0158] The third and fourth depths are predicted by a second convolutional neural network associated with a monoscopic vision system comprising the first camera 11.
[0159] In a step 34, a sixth image is generated from the first image and the third depths by applying the image generation function and a seventh image is generated from the third image and the fourth depths by applying the image generation function.
[0160] In a step 35, the second convolutional neural network is learned by minimizing a second error, the second error being determined by comparing the sixth and third images and by comparing the seventh and first images.
[0161] In a step 36, the parameters of the second convolutional neural network are adjusted according to the first depths, the first depths corresponding to training data.
[0162] According to a variant, the variants and examples of the operations described in relation to [Fig.l] apply to steps of method 2 of [Fig.2].
[0163] [Fig. 3] schematically illustrates a device 4 configured for the determination of a depth by a vision system on board a vehicle 10, according to a particular and non-limiting exemplary embodiment of the present invention. The device 4 corresponds for example to a device on board the first vehicle 10, for example a computer.
[0164] The device 4 is for example configured for the implementation of the operations described with regard to [Fig.l] and / or steps described with regard to [Fig.2]. Examples of such a device 4 include, but are not limited to, on-board electronic equipment such as an on-board computer of a vehicle, an electronic calculator such as an ECU (“Electronic Control Unit”), a smartphone, a tablet, a laptop. The elements of the device 4, individually or in combination, can be integrated in a single integrated circuit, in several integrated circuits, and / or in discrete components. The device 4 can be produced in the form of electronic circuits or software (or computer) modules or even a combination of electronic circuits and software modules.
[0165] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the method and / or for executing the instructions of the software(s) embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41 corresponding for example to a volatile and / or non-volatile memory and / or comprises a memory storage device which may comprise volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.
[0166] The computer code of the embedded software(s) comprising the instructions to be loaded and executed by the processor is for example stored in the 4L memory.
[0167] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Telematic Control Unit”), for example via a communication bus or through dedicated input / output ports.
[0168] According to a particular and non-limiting exemplary embodiment, the device 4 comprises a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 comprise one or more of the following interfaces: - RF radio frequency interface, for example Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English “Universal Serial Bus” or “Universal Serial Bus” in French); HD MI interface (from the English “High Definition Multimedia Interface” or “High Definition Multimedia Interface” in French); - LIN interface (from the English “Local Interconnect Network”).
[0169] According to another particular and non-limiting exemplary embodiment, the device 4 comprises a communication interface 43 which makes it possible to establish communication with other devices (such as other computers of the on-board system) via a communication channel 430. The communication interface 43 corresponds for example to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds for example to a wired network of the CAN (Controller Area Network) type, CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by the ISO 17458 standard) or Ethernet (standardized by the ISO / IEC 802-3 standard).
[0170] According to a particular and non-limiting exemplary embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch-sensitive or not, one or more speakers 450 and / or other peripherals 460 (projection system) via the output interfaces 44, 45, 46 respectively. According to a variant, one or other of the external devices is integrated into the device 4.
[0171] Of course, the present invention is not limited to the embodiments described above but extends to a method for determining a depth and measuring such a quantity by a vision system on board a vehicle, which would include secondary steps without thereby departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0172] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-based motor vehicle, comprising the device 4 of [Fig.3].
Claims
Claims
1. Method for determining a depth by a vision system on board a vehicle (10), the vision system comprising a set of cameras of at least two cameras (11, 12) arranged so as to each acquire an image of a three-dimensional scene from a different point of view, said method being characterized in that it comprises the following steps: - reception (21, 22) of first and second data respectively representative of a first and second image acquired by respectively a first and second camera (11, 12) of said set of cameras at the same first acquisition time instant; - prediction (23) of first depths associated with a first set of pixels of the first image and second depths associated with a second set of pixels of the second image as a function of optical flow values and first disparity values induced by a relative position of optical axes of the first (11) and second (12) cameras, said optical flow values and first induced disparity being determined from said first and second sets of pixels, said first and second depths being predicted by a first convolutional neural network associated with a stereoscopic vision system comprising the first (11) and second (12) cameras; - generation (24) of a fourth image from said first image and the first depths and of a fifth image from said second image and the second depths, the image generation being obtained by applying an image generation function; - training (25) said first convolutional neural network by minimizing a first error, the first error being determined by comparing the first and fifth images and by comparing the fourth and second images; - reception (31) of third data representative of a third image acquired by said first camera (11) at a second acquisition time instant different from said first acquisition time instant; - determination (32) of data representative of a movement of the
2.
3. first camera (11) between the first acquisition time instant and the second acquisition time instant as a function of said first and third images; - prediction (33) of third depths associated with a third set of pixels of the first image and of fourth depths associated with a fourth set of pixels of the third image as a function of optical flow values and of second disparity values induced by a relative position of optical axes of the first camera (11) at the first and second acquisition time instants, said optical flow values and second induced disparity being determined from said third and fourth sets of pixels, said third and fourth depths being predicted by a second convolutional neural network associated with a monoscopic vision system comprising the first camera (11); - generation (34) of a sixth image from said first image and the third depths and of a seventh image from said third image and the fourth depths by applying said image generation function; - training (35) said second convolutional neural network by minimizing a second error, the second error being determined by comparing the sixth and third images and by comparing the seventh and first images; - adjustment (36) of parameters of said second convolutional neural network as a function of said first depths corresponding to training data. The method of claim 1, wherein said first and second induced disparity values are determined by a third convolutional neural network. Method according to one of claims 1 to 2, for which said image generation function is defined by the following function, respectively for the stereoscopic vision system and for the monoscopic vision system: Ps = Æ (K [(P^ Dst (Pt))) With : • Ps a pixel of a generated image, • 77 a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • K an intrinsic matrix of the second camera (12), respectively of the first camera (11), • M an extrinsic matrix of the stereoscopic vision system, respectively a matrix representing the movement of the first camera (11) between the first and second time instants, • 0 a projection function in the three-dimensional scene of a pixel as a function of its depth, • K' an intrinsic matrix of the first camera (11), and • D^p ) is a depth of the pixel Pt determined by the first convolutional neural network, respectively by the second convolutional neural network.
4. Method according to one of claims 1 to 3, for which the first error is determined by a first loss function defined by: Lx = E;?avg( L^p), Lss(p) ) With: • L the first error, • avS an average of arguments, • Lslp) an error for a pixel P of the first image compared to the fifth image, • Lxx(p) an error for a pixel P of the second image compared to the fourth image.
5. Method according to one of claims 1 to 3, for which the first error is determined by a first loss function defined by: Lx = Epmin(£yf(p), Lss(p)) With: • L the first error, • min a minimum value of an argument, • Ls^p) an error for a pixel P of the first image compared to the fifth image, • Lxx(p) an error for a pixel P of the second image compared to the fourth image.
6. Method according to one of claims 1 to 5, for which the second error is determined by a second loss function defined by: Lm = Epavg(Lwj77), Lm^p}) With: • T the second error, • with an average of arguments, • an error for a pixel P of the first image compared to the seventh image, • Lms(p) an error for a pixel P of the third image compared to the sixth image.
7. Method according to one of claims 1 to 5, for which the second error is determined by a second loss function defined by: Lm = L,m(p)) With: • l' the second error, • min a minimum value of an argument, • Lm^p) an error for a pixel P of the first image compared to the seventh image, • Lms(p) an error for a pixel P of the third image compared to the sixth image.
8. Method according to one of claims 1 to 7, for which said third image is replaced by a black image, a black image being an image of which a set of colorimetric values associated with the pixels of the image is zero.
9. Device (4) for determining a depth by a vision system on board a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for implementing the steps of the method according to any one of claims 1 to 8.
10. Vehicle (10) comprising the device (4) according to claim 9.