Method and device for learning a depth prediction model of a set of pixels of an image associated with a stereoscopic vision system on board a vehicle.
A convolutional neural network-based depth prediction model for stereoscopic vision systems addresses image distortions and occlusions, enhancing ADAS performance and safety by accurately predicting depth across various environments.
Patent Information
- Application Number
- FR2024003164
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-28
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-03-28
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method and device for learning a depth prediction model of a set of pixels of an image associated with a stereoscopic vision system on board a vehicle. Technical field
[0001] The present invention relates to methods and devices for learning a depth prediction model associated with a stereoscopic vision system on board a vehicle, for example in a motor vehicle. The present invention also relates to a method and a device for determining a depth and / or measuring a distance separating an object from a vehicle carrying a vision system. Technological background
[0002] Many modern vehicles are equipped with so-called AD AS (Advanced Driver-Assistance System). Such AD AS systems are passive and active safety systems designed to eliminate the element of human error in the driving of vehicles of all types. AD AS use advanced technologies to assist the driver while driving and thus improve their performance. AD AS use a combination of sensor technologies to perceive the environment around a vehicle, then provide information to the driver or act on certain vehicle systems.
[0003] There are several levels of ADAS, such as rearview cameras and blind spot sensors, lane departure warning systems, adaptive cruise control and automatic parking systems.
[0004] The AD AS embedded in a vehicle are supplied with data obtained one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or possible obstacles present around a vehicle in order, for example: • to adapt the vehicle's lighting according to the presence of other users; • automatically regulate the vehicle speed; • to act on the braking system in the event of a risk of impact with an object.
[0005] In order to have an extended view of the vehicle's environment, i.e. of a three-dimensional scene taking place around the vehicle, a wide-angle camera, i.e. a camera with a wide field of vision, is recommended. Indeed, Using a wide-angle camera has many advantages over “standard” cameras: • a wider field of vision allowing a larger part of the scene to be captured, which is particularly important in the context of driving, where it is essential to monitor the environment both to the sides and in front of the vehicle, for example, • greater efficiency in perceiving complex environments, such as intersections, tight turns, parking spaces, etc., by minimizing blind spots and providing a more complete view of driving situations, and • increased safety by more easily detecting obstacles, other vehicles, pedestrians and cyclists in areas adjacent to the vehicle.
[0006] Although wide-angle cameras offer many advantages, they can also introduce distortions into the captured images, and these distortions can lead to certain problems such as: • the distortion of straight lines, for example barrel or pincushion, causing curvature of straight lines in an image acquired by the wide-angle camera, making it difficult to estimate the actual distances between objects, particularly towards the edges of the image, • stretching or compressing objects, especially towards the edges of the image, changing the apparent size of objects, which can be problematic when judging the distance or actual size of objects, • changing the proportions of objects, making them larger or smaller than their actual size and more complex to identify, and • the difficulty of rectifying or correcting distortion in post-processing which can be complex and can lead to a loss of information.
[0007] Thus, the processing of an image acquired by a wide-angle camera requires special processing, in particular because of the strong distortion present in the image. Determining a distance separating the vehicle carrying the camera from an object in the scene is then not achievable with the methods commonly used for standard cameras used in certain vision systems.
[0008] Some known methods also have the disadvantage of not being able to accurately predict, or even predict, a depth associated with each pixel of an image. Thus, the “holes”, i.e. the pixels with which no depth is associated, generate inconsistencies or lack of information that can disrupt the operation of an AD AS.
[0009] In addition, images acquired at the same time instant by a stereoscopic vision system, i.e. a vision system comprising several cameras acquiring images of the same three-dimensional scene, sometimes comprise occluded areas, i.e. areas visible in an image acquired by a camera having pixels associated with an object in the three-dimensional scene not visible in an image acquired by another camera in the stereoscopic vision system. Predicting a depth associated with a pixel corresponding to an object in the three-dimensional scene not visible by all the cameras in the stereoscopic vision system is complex and is a source of prediction error. Summary of the present invention
[0010] An object of the present invention is to solve at least one of the problems of the technological background described above.
[0011] Another object of the present invention is to improve the quality of the data resulting from the processing of an image acquired by a vision system, in particular by a depth prediction model implemented by a neural network associated with this stereoscopic vision system.
[0012] Another object of the present invention is to improve road safety, in particular by improving the operational safety of AD AS systems supplied with data obtained from a wide-angle camera.
[0013] According to a first aspect, the present invention relates to a method for learning a depth prediction model implemented by a convolutional neural network associated with a stereoscopic vision system on board a vehicle, the stereoscopic vision system comprising a first camera and a second camera arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception of a first image and a second image acquired by the first camera and the second camera respectively at the same acquisition time instant; - generating a set of first pairs of feature maps by a feature extractor from the first and second images, each first pair of feature maps comprising a first feature map associated with the first image and a second feature map associated with the second image and each first pair of feature maps having a different definition; - prediction, for each first pair of feature maps, of directions associated with pixels of a set of pixels of the first feature map, called first pixels, and of directions associated with pixels of a set of pixels of the second feature map, called second pixels, by a model of direction prediction from the first feature map and the second feature map respectively; - prediction, for each first pair of feature maps, of depths associated with the first pixels and the second pixels by the depth prediction model from respectively the first feature map and the second feature map; - association, with each first pair of characteristic cards, of a second pair of characteristic cards comprising a third characteristic card and a fourth characteristic card, the third feature map being generated from the first feature map, the directions and depths associated with the first pixels and extrinsic parameters of the stereoscopic vision system, and the fourth feature map being generated from the second feature map, the directions and depths associated with the second pixels and the extrinsic parameters of the stereoscopic vision system; - determining, for each first pair of feature maps, a loss error from a first error determined for each first pixel by comparing each first pixel of the first feature map to a pixel of the fourth feature map corresponding to each first pixel and a second error determined for each second pixel by comparing each second pixel of the second feature map to a pixel of the third feature map corresponding to each second pixel; - learning the depth prediction model by minimizing a global error determined from the loss errors determined for the set of first pairs of feature maps.
[0014] According to a variant of the method, the overall error is the sum of the loss errors, determined by the following function: With : • LG the global error, • LR a loss error for a first pair of feature maps R, and •Es( a sum over the set of first pairs of feature maps.
[0015] According to another variant of the method, colorimetric values are associated with each pixel of the first, second, third and fourth characteristic maps.
[0016] According to another variant of the method, the first error, respectively second error, is determined by the following function: L*(p) = EJ ( 1-a) • \l{p)-Hp} | +«• ( \-^SSIM(l{p), ï(p) ) ) ] With : • Lp* ( p ) the first error noted L ( ( p ), respectively the second error noted ^2(P), p being a pixel defined by its coordinates in a feature map, • l(p) a colorimetric value of the pixel p in the first feature map, respectively second feature map, * l(p) a colorimetric value of pixel p in the fourth feature map, respectively third feature map, • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment in which the vehicle operates.
[0017] According to an additional variant of the method, the loss error is further determined from a construction error determined by the following function: L^^p) ^W(p) KD^p)|*<«) With: • ^smoothip} the construction error Ls3( p) for a pixel p of the third feature map, respectively the construction error L^p) for a pixel p of the fourth feature map, • D{p) is a depth of a pixel p obtained from the third depth map, respectively obtained from the fourth depth map; • VF is a parameter matrix; • 0 is the order of a smoothing gradient; • an L1 norm of second-order depth gradients is calculated with VF =1, et0 =2; • x and are the dimensions of the feature maps; • P is a hyperparameter dependent on the environment in which the vehicle operates; and • A(Pf) is a colorimetric value of pixel p in the third feature map, respectively fourth feature map.
[0018] According to yet another variant of the method, the loss error is determined by the following function: LR = ^p(min{LYp), L^pY} + LSY p) + LSY P)) With : • LR the loss error, * L{(p) 'a first error for a pixelp, p being a pixel defined by its coordinates in a feature map, • L2 the second error for pixel p, • Ls3 the construction error for a pixel of the third feature map, and • the construction error for a pixelp of the fourth feature map.
[0019] According to yet another variant of the method, the third and fourth characteristic maps are generated using the following function: p, With : • Ps a pixel of a generated feature map corresponding to the third feature map, respectively the fourth feature map, •77 a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • K a direction prediction model associated with the second camera, respectively with the first camera, • K' a direction prediction model associated with the first camera, respectively with the second camera, • T an extrinsic matrix of the stereoscopic vision system, • 0 a projection function in the three-dimensional scene of a pixel as a function of its depth, and * D^p ) is a depth associated with a pixel pt of the first feature map, respectively of the second feature map.
[0020] According to a second aspect, the present invention relates to a device configured to learn a depth prediction model by a vision system on board a vehicle, the device comprising a memory associated with at least one processor configured to implement the steps of the method according to the first aspect of the present invention.
[0021] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0022] According to a fourth aspect, the present invention relates to a computer program which comprises instructions adapted for executing the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0023] Such a computer program may use any programming language and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0024] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which a program is recorded. computer system comprising instructions for carrying out the steps of the method according to the first aspect of the present invention.
[0025] On the one hand, the recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM memory, a CD-ROM or a microelectronic circuit type ROM memory, or a magnetic recording means or a hard disk.
[0026] Furthermore, this recording medium may also be a transmissible medium such as an electrical or optical signal, such a signal being able to be conveyed via an electrical or optical cable, by conventional or hertzian radio or by self-directed laser beam or by other means. The computer program according to the present invention may in particular be downloaded from an Internet-type network.
[0027] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to perform or to be used in performing the method in question. Brief description of the figures
[0028] Other characteristics and advantages of the present invention will emerge from the description of the particular and non-limiting exemplary embodiments of the present invention below, with reference to the appended figures 1 to 4, in which:
[0029] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting example of embodiment of the present invention;
[0030] [Fig.2] illustrates a flowchart of the different steps of a method for determining a depth of a pixel of an image by a depth prediction model associated with a vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;
[0031] [Fig.3] illustrates a flowchart of the different steps of a method for learning the depth prediction model used in the method of [Fig.2], according to a particular and non-limiting exemplary embodiment of the present invention; and
[0032] [Fig.4] schematically illustrates a device configured to learn a depth prediction model by a vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention. Description of examples of implementation
[0033] A method and a device for learning a depth prediction model implemented by a convolutional neural network associated with a stereoscopic vision system on board a vehicle will now be described in what follows with joint reference to Figures 1 to 4. The same elements are identified with the same reference signs throughout the description which follows.
[0034] The terms "first(s)", "second(s)" (or "first(s)", "second(s)"), etc. are used in this document by arbitrary convention to enable different elements (such as operations, means, etc.) implemented in the embodiments described below to be identified and distinguished. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0035] For the entire description, reception of an image is understood to mean the reception of data representative of an image. Similarly, for the generation of a feature map, it is understood to mean the generation of data representative of a feature map. These shortcuts are only intended to simplify the description; however, since the methods are implemented by one or more processors, it is obvious that the input and output data of the different steps of a method are computer data.
[0036] According to a particular and non-limiting example of embodiment of the present invention, a depth prediction model associated with a stereoscopic vision system comprising two cameras is learned in a learning phase.
[0037] Indeed, the method comprises the generation of a set of first pairs of characteristic maps from images acquired by the stereoscopic vision system, each first pair of characteristic maps comprising first and second characteristic maps associated respectively with each acquired image, each first pair of characteristic maps having a different definition.
[0038] A second pair of feature maps is associated with each first pair of features and comprises two other feature maps generated from, in particular, predicted depths and directions for the pixels of the first and second feature maps.
[0039] The depth prediction model is learned by minimizing an error determined by comparing the feature maps for the different resolutions.
[0040] [Fig. 1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.
[0041] Such an environment 1 corresponds, for example, to a road environment formed of a network of roads accessible to the vehicle 10.
[0042] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, with an electric motor(s) or even a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as an automobile, a truck, a bus, a motorcycle. Finally, the vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say a vehicle circulating according to a determined level of autonomy or under the total supervision of the driver.
[0043] The vehicle 10 advantageously comprises at least two on-board cameras, a first camera 11 and a second camera 12, configured to acquire images of a three-dimensional scene taking place in the environment of the vehicle 10 from separate observation positions. The first camera 11 and the second camera 12 form a stereoscopic vision system when used together as illustrated in [Fig.l]. The first camera 11 forms a monoscopic vision system when used alone, likewise the second camera 12 forms another monoscopic vision system when used alone. The present invention, however, extends to any vision system comprising at least two cameras, for example 2, 3 or 5 cameras.
[0044] The intrinsic parameters of the first camera 11 characterize the transformation which associates, for an image point, subsequently called “point”, its three-dimensional coordinates in the reference frame of the first camera 11 with the pixel coordinates in an image acquired by the first camera 11. These parameters do not change if the first camera 11 is moved. The intrinsic parameters of the first camera 11 include in particular a first focal distance fl associated with the first camera 11.
[0045] The intrinsic parameters of the second camera 12 characterize, for their part, the transformation which associates, for an image point, its three-dimensional coordinates in the reference frame of the second camera 12 with the pixel coordinates in an image acquired by the second camera 12. These parameters do not change if the second camera 12 is moved. The intrinsic parameters of the second camera 12 include in particular a second focal length f2 associated with the second camera 12.
[0046] The distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions which generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, that is to say that the distortions are not taken into account, that their correction is processed at the time of image acquisition or at the time of calibration.
[0047] These two cameras 11, 12 are arranged so as to each acquire an image of a scene from a different point of view, the first point of view being for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10, the second point of view is for example located on or in the right rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. In the case where the two cameras are located at the top of the windshield of the vehicle, they are then placed at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, the second camera 12 is located in the right rearview mirror of the vehicle 10.
[0048] A first reference point is associated with the first camera 11: - the direction of the x axis is defined horizontal and normal to the optical axis Cl of the first camera 11. The distance B separating the optical center of the first camera 11 from the projection of the optical center of the second camera 12 on the horizontal plane passing through the optical center of the first camera 11 is called the reference base (in English “baseline”); - the direction of the y axis is defined vertical and normal to the optical axis Cl of the first camera 11; - the direction of the z axis is defined orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal reference frame.
[0049] The optical axis C1 of the first camera 11 and the optical axis C2 of the second camera 12 are not necessarily parallel or even included in the same plane.
[0050] According to a variant, optical axes of the first and second cameras are not coplanar.
[0051] The extrinsic parameters linked to the position of the cameras 11, 12 are the following parameters: - three translations in the x, y and z directions: Tx, Ty and Tz constituting the translation vector T; and - three rotations in the x, y and z directions: 0x, 0y and 0z.
[0052] An extrinsic matrix of the vision system then includes the previously defined extrinsic parameters.
[0053] The extrinsic parameters are determined, for example, during a calibration phase of the stereoscopic vision system comprising the first camera 11 and the second camera 12.
[0054] A main constraint of the stereoscopic vision system used in automobiles is, for example, the large distance between the two cameras. Indeed, to be able to cover a measurement range of 200 meters, the reference base must reach 60cm for the cameras commonly used in this field.
[0055] The two cameras 11, 12 acquire images of a scene located in front of the vehicle 10, the first camera 11 covering only a first acquisition field 13, the second camera 12 covering only a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic vision of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic vision of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic vision of the scene by the stereoscopic vision system composed of the two cameras 11, 12.
[0056] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.
[0057] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system comprising the second camera 12.
[0058] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable using the monoscopic vision system comprising the first camera 11.
[0059] Finally, field 19 is not visible to any of the cameras. The part of the scene present in this field 19 is therefore not observable.
[0060] According to a particular exemplary embodiment, the field of vision of the second camera 12 covers at least half of the field of vision of the first camera 11.
[0061] It is obvious that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with differently placed and oriented cameras.
[0062] The images acquired by the cameras 11, 12 at an acquisition time instant are presented in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colors and brightness of objects in the observed scene in the form, for example, of RGB colorimetric coordinates (from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness).
[0063] Each pixel of the acquired image is representative of an object in the three-dimensional scene present in the camera's field of vision. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a luminous point resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, captured by a photosensitive sensor of the camera after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates a image faithfully representing the physical object observed by the camera. An image point previously presented is thus a point on a surface of an object in the three-dimensional scene.
[0064] The images acquired by the cameras 11, 12 represent views of the same scene taken from different viewpoints, the positions of the cameras being distinct. On this scene are for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0065] According to a particular embodiment, the first camera 11 and / or the second camera 12 is of the “wide-angle” type, a wide-angle camera being for example equipped with a lens designed to acquire an image representative of a three-dimensional scene perceived according to a wider field of vision than that of a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens makes it possible to capture a larger portion of the three-dimensional scene taking place in front of or around the wide-angle camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by this camera. The angle a of the field of vision of the wide-angle camera is for example equal to 120°, 145°, 180° or 360°, whereas a standard camera offers, for example, an open field of vision following an angle of 45° or less.Such a wide-angle camera is, for example, a camera equipped with mirrors or a "fisheye" camera. Wide-angle lenses have a shorter focal length than standard lenses, which makes them suitable for capturing images of landscapes, architecture, road intersections or any other subject requiring a wide perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with an extended depth of field.
[0066] According to a particular embodiment, an image acquired by the first camera 11 and / or an image acquired by the second camera 12 comprises a distortion equal to 0.5%, 0.8% or greater than 1%. The measurement of such a distortion corresponds to the determination of a ratio between: - the maximum spacing of a pixel of the image from a straight line of the first three-dimensional scene whose image is a line touching the longest edge of the first image, either at the center of the edge of the image, or at the corners of the edge of the image, and - the length of this edge.
[0067] Commonly, distortion is considered, in the world of photography, as: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • bothersome if it is greater than or equal to 1% or more.
[0068] A barrel distortion is characterized by a positive percentage, while a crescent distortion is characterized by a negative percentage.
[0069] According to a particular exemplary embodiment, a field of view of the first camera 11 covers at least half of a field of view of the second camera 12 and a field of view of the second camera 12 covers at least half of a field of view of the first camera 11. In other words, more than half of the pixels of an image acquired by the first camera 11 correspond to an object of the three-dimensional scene seen by the second camera 12, pixels of an image acquired by the second camera 12 also corresponding to this object of the three-dimensional scene. Similarly, more than half of the pixels of an image acquired by the second camera 12 correspond to an object of the three-dimensional scene seen by the first camera 11, pixels of an image acquired by the first camera 11 also corresponding to this object of the three-dimensional scene.
[0070] The images acquired by the first camera 11 and by the second camera 12 are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.
[0071] A method for determining a depth by a vision system on board the vehicle 10 is advantageously implemented by the vehicle 10, that is to say by a processor, a computer or a combination of computers of the on-board system of the vehicle 10, for example by the computer(s) in charge of the vision system of the vehicle 10.
[0072] [Fig. 2] illustrates a flowchart of the different steps of a method 2 for determining a depth of a pixel of an image by depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, for example in the vehicle 10 of [Fig. 1], according to a particular and non-limiting exemplary embodiment of the present invention. The method 2 is for example implemented by a device of the vision system embedded in the vehicle 10 or by the device 4 of [Fig. 4].
[0073] In a step 21, data representative of an image acquired by the first camera 11 and of an image acquired by the second camera 12 are received.
[0074] In a step 22, depths associated with a set of pixels of one of the received images are determined by the depth prediction model from the two received images.
[0075] Each determined depth then corresponds to a distance separating the vehicle 10 or a part of the vehicle 10 from an object of the three-dimensional scene with which a pixel is associated, the determination of a depth of a pixel then corresponding to a measurement of a distance separating an object from the vehicle carrying the vision system.
[0076] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another user present on the road, the ADAS is then able to determine this distance precisely. For example, if the ADAS has the function of acting on a braking system of the vehicle 10 in the event of a risk of collision with another road user and the distance separating the vehicle 10 from this same road user decreases significantly, then the ADAS is able to detect this sudden approach and act on the braking system of the vehicle 10 to avoid a possible accident.
[0077] [Fig. 3] illustrates a flowchart of the different steps of a method for learning the depth prediction model used in a method for determining a depth of a pixel of an image, for example in method 2 of [Fig. 2], according to a particular and non-limiting exemplary embodiment of the present invention.
[0078] The learning method 3 is for example implemented by the device on board the vehicle 10 implementing the method for determining a depth by the vision system on board a vehicle or by the device 4 of [Fig.4].
[0079] In a step 31, a first image and a second image are received, the first image being acquired by the first camera 11 at an acquisition time instant and the second image being acquired by the second camera at the same acquisition time instant.
[0080] According to a particular exemplary embodiment, the first image and second image have the same definition, that is to say they comprise the same number of pixels, have the same number of pixels according to their height and the same number of pixels according to their width.
[0081] According to another particular embodiment, the first image and second image are not of the same definition. An additional step then consists of them resize or crop them to obtain a first image and a second image of the same definition.
[0082] In a step 32, a set of first pairs of feature maps is generated by a feature extractor from the first and second images, each first pair of feature maps comprising a first feature map associated with the first image and a second feature map associated with the second image and each first pair of feature maps having a different definition.
[0083] The dotted box in [Fig.3] entitled 'R' groups together the different steps in which each first pair of characteristic maps of the set of first pairs of characteristics are processed, that is to say that these same steps are applied as many times as there are first pairs of characteristic maps.
[0084] According to a particular embodiment, among these first pairs of feature maps, a single pair of features has the same definition as the first and second images. The other first pairs of features comprise first and second feature maps of lower definition, for example whose definition is a sub-multiple of the definition of the first and second images.
[0085] It should be noted that each first feature map of a first pair of feature maps has the same definition as the second feature map of this first pair of feature maps.
[0086] The number of first pairs of characteristics is greater than or equal to two, for example this number is equal to 3, 5, 20 or 64.
[0087] Such a feature extractor is known to those skilled in the art and is for example presented in the document “Unifying Flow, Stereo and Depth Estimation” written by Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao and Andréas Geiger, published in July 2023 or in the document “PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume” written by Deqing Sun, Xiaodong Yang, Ming-Yu Liu and Jan Kautz, published in June 2018. The last document cited presents an advantageous feature extractor which offers better coverage of the range of features by proposing fewer candidates for each level of the feature map but several feature maps at different resolutions.
[0088] Such characteristics are for example representative of information relating to a shape of an object in an image and / or a texture of a set of pixels of an image and / or a color of a set of pixels of an image.
[0089] In a step 33, for each first pair of feature maps, directions associated with pixels of a set of pixels of the first feature map, called first pixels, and directions associated with pixels of a set of pixels of the second feature map, called second pixels, are predicted by a direction prediction model from the first feature map and the second feature map respectively.
[0090] Such a direction prediction model is known to those skilled in the art; it is notably presented in the document “Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion” written by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich and Adrien Gaidon, published in August 2020. This direction prediction model implements, for example, a convolutional neural network different from that implemented by the depth prediction model.
[0091] The predicted direction is representative of the direction in which the point of the three-dimensional scene corresponding to the pixel or to a set of pixels of an acquired image and of which the pixel of the characteristic map is the image is located, the direction being expressed in the frame of reference of the camera having acquired the image at the origin of the characteristic map. This direction prediction model is in particular capable of predicting a direction associated with a pixel even when an image has a large distortion as is the case for images acquired by a wide-angle camera or by a “fish-eye” camera. In addition, this direction prediction model can also be used for an uncalibrated camera, and replaces, for example, an intrinsic matrix corresponding to a pinhole camera.
[0092] The predicted direction associated with a pixel of an acquired image represents the direction in which an incident light ray encounters the camera acquiring an image, the incident light ray corresponding to this same pixel.
[0093] In a step 34, for each first pair of feature maps, depths associated with the first pixels and the second pixels are predicted by the depth prediction model from the first feature map and the second feature map respectively.
[0094] Such a depth prediction model, implemented by a convolutional neural network, is known to those skilled in the art and is for example presented in the document “PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume”. The predicted depths are then associated with the pixels of the first and second feature maps, for example in an additional channel, or are recorded in a first and a second depth map, the first depth map being associated with the first feature map and the second depth map being associated with the second feature map.
[0095] It should be noted that the document “Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion” also presents a depth prediction model usable in this invention.
[0096] In a step 35, a second pair of characteristic maps comprising a third characteristic map and a fourth characteristic map is associated with each first pair of characteristic maps.
[0097] The third feature map is generated from: • of the first characteristics card, • directions and depths associated with the first pixels, and • extrinsic parameters of the stereoscopic vision system.
[0098] The fourth feature map is generated from: • of the second characteristics card, • directions and depths associated with the second pixels, and • extrinsic parameters of the stereoscopic vision system.
[0099] The generation of a third or fourth feature map from a first or second feature map consists of reprojecting a pixel of the first or second feature map into the three-dimensional scene in the form of a point whose coordinates are expressed in the frame of reference of a camera, then projecting this point into the image plane of another camera of the vision system, so as to obtain a feature map corresponding to a view of the three-dimensional scene from the point of view of the other camera. The image plane of a camera corresponds to a plane defined in the frame of reference of the camera, normal to the optical axis of the camera and located at the focal length of the camera. Thus, the third feature map generated from the first feature map is comparable to the second feature map.Similarly, the fourth feature map generated from the second feature map is comparable to the first feature map. Since the cameras' fields of view are not confused and objects can mask other objects in the scene, the generated feature maps are not identical to acquired feature maps. In addition, the prediction of depths and directions, as well as the models used to generate the feature maps, are not error-free. Thus, comparing a generated feature map to an acquired feature map allows us to evaluate the relevance of the different models used.
[0100] According to a particular exemplary embodiment, the third and fourth characteristic maps are generated using the following function:
[0101] [Math.l] p=7ï(K[T$(p}^\ D(pt))])
[0102] With: • P^ a pixel of a generated feature map corresponding to the third feature map, respectively the fourth feature map, • Æ a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • K a direction prediction model associated with the second camera 12, respectively with the first camera 11, • K' a direction prediction model associated with the first camera 11, respectively with the second camera 12, • T an extrinsic matrix of the stereoscopic vision system, • 0 a projection function in the three-dimensional scene of a pixel as a function of its depth, and * I^p ) is a depth associated with a pixel Pt of the first feature map, respectively of the second feature map.
[0103] Note that the extrinsic matrix is not the same for the generation of the two images; in fact, a first extrinsic matrix makes it possible to move from a reference frame associated with the first camera 11 to a reference frame associated with the second camera 12 during the generation of the third characteristics map, while a second extrinsic matrix makes it possible to move from a reference frame associated with the second camera 12 to a reference frame associated with the first camera 11 during the generation of the fourth characteristics map.
[0104] The projections and reprojections are inverse functions obtained from the direction prediction model and are a function of a depth, such a projection model is notably presented in the document “Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion”.
[0105] In a step 36, for each first pair of feature maps, a loss error is determined from a first error determined for each first pixel by comparing each first pixel of the first feature map to a pixel of the fourth feature map corresponding to each first pixel, i.e. having the same coordinates, and from a second error determined for each second pixel by comparing each second pixel of the second feature map to a pixel of the third feature map corresponding to each second pixel, i.e. having the same coordinates.
[0106] According to a particular exemplary embodiment, colorimetric values are associated with each pixel of the first, second, third and fourth characteristic maps. Note that in the case where a characteristic map has the same definition as the image to which it corresponds, also called the source image, then the colorimetric values of a pixel of the characteristic map are identical to those of the pixel corresponding to it in the image. In the case where the definition of the characteristic map is lower than the definition of the source image, then the colorimetric value of a pixel of the characteristic map is determined as a function of colorimetric values of several pixels corresponding to it in the source image.
[0107] According to a first variant embodiment, the first and second errors are photometric errors (in English “photometric error”) as presented in the document “Digging Into Self-Supervised Monocular Depth Estimation” by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow published in August 2019 and are determined respectively by the following function:
[0108] [Math.2]
[0109] With: • Lp* ( p ) the first error noted L ( ( p ), respectively the second error noted ^2(P), p being a pixel defined by its coordinates in a feature map, • l(p) a colorimetric value of the pixel p in the first feature map, respectively second feature map, * 'Kp) a colorimetric value of pixel p in the fourth feature map, respectively third feature map, • SSIM a function which takes into account a local structure, and • has a weighting factor depending in particular on the type of environment in which the vehicle 10 operates.
[0110] According to a second variant embodiment, the loss error is further determined from a construction error determined by the following function: [Math.3] ^smooth(P) ~ (W(p)\W>*(p)\rMWi) [YES] With: • ^smoah^P) the construction error L^p) for a pixel p of the third feature map, respectively the construction error Lsit{p} for a pixel p of the fourth feature map, • D(p) is a depth of one pixel obtained from the third depth map, respectively obtained from the fourth depth map; • VF is a parameter matrix; • 0 is the order of a smoothing gradient; • an L1 norm of second-order depth gradients is calculated with W =1, and0 =2; • x and are the dimensions of the feature maps; • P is a hyperparameter dependent on the environment in which the vehicle 10 operates; and • lt(P^) is a colorimetric value of pixel p in the third feature map, respectively fourth feature map.
[0112] This second function is generally used to deal with discontinuity at the edge of objects (in English “edge aware smoothness”).
[0113] The loss error is thus defined, for example, from the photometric errors and reconstruction errors previously defined.
[0114] According to a particular exemplary embodiment, the loss error is determined by the following function: [Math.4] LR^ / imntL^p), L2(p) )+Ls3(p) + Ls4(p))
[0115] With: • Lr the loss error for a first pair of feature maps, * Lp}(p) 'a first error for a pixel p, p being a pixel defined by its coordinates in a feature map, * the second error for pixel p, • / the construction error for a pixel p of the third map of characteristics, and * J- / p) the construction error for a pixel p of the fourth feature map.
[0116] In a step 37, the depth prediction model is learned by minimizing a global error determined from the loss errors determined for the set of first pairs of feature maps, i.e. for the set of definitions applied to the feature maps.
[0117] According to a particular exemplary embodiment, the overall error is the sum of the loss errors, determined by the following function: [Math.5] lg = ( lr )
[0118] With: • Lg the overall error, • Lr a loss error for a first pair of feature maps R, and * ) a sum over the set of first pairs of feature maps.
[0119] Training the depth prediction model consists of adjusting input parameters of the convolutional neural network in order to minimize the previously calculated loss error.
[0120] An occluded or non-visible object in the field of vision of one camera and visible in the field of vision of the other camera does not impact the loss error thanks to the minimization function making the loss error, and therefore the overall error, insensitive to occlusion. Thus, the depth prediction model used for the depth prediction of a pixel of an image acquired by one of the cameras of the stereoscopic vision system is made reliable thanks to this learning method.
[0121] The advantage of using several feature map resolutions is that it avoids "holes", i.e. having pixels in the feature map with a resolution equal to that of the acquired images without predicted depth. Indeed, thanks to the different definitions, it is always possible to predict a depth for a pixel of a feature map. This architecture makes it possible to find a good compromise between the computational cost and the range of features covered by this model. By using feature maps at different resolutions, the model is better able to cover features of different sizes.
[0122] This learning is carried out from data acquired by the on-board vision system and therefore does not require data annotated by another on-board system or storage of a library of learning images. In addition, the learning data is representative of the data received when the system is in operation or in production, in fact the learning data is representative of real environments in which the vehicle carrying the stereoscopic vision system evolves or moves, this learning data is therefore particularly relevant.
[0123] [Fig. 4] schematically illustrates a device 4 configured to learn a depth prediction model by a vision system embedded in a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention. The device 4 corresponds for example to a device embedded in the first vehicle 10, for example a computer associated with the stereoscopic vision system.
[0124] The device 4 is for example configured for the implementation of the operations described with regard to figures 1 and 4 and / or steps described with regard to figures 2 and 3. Examples of such a device 4 include, but are not limited to, on-board electronic equipment such as a vehicle on-board computer, an electronic calculator such as an ECU (“Electronic Control Unit”), a smartphone, a tablet, a laptop. The elements of the device 4, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. The device 4 may be implemented in the form of electronic circuits or software (or computer) modules or even a combination of electronic circuits and software modules.
[0125] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the method and / or for executing the instructions of the software(s) embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41 corresponding for example to a volatile and / or non-volatile memory and / or comprises a memory storage device which may comprise volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.
[0126] The computer code of the embedded software(s) comprising the instructions to be loaded and executed by the processor is for example stored in the 4L memory.
[0127] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Telematic Control Unit”), for example via a communication bus or through dedicated input / output ports.
[0128] According to a particular and non-limiting exemplary embodiment, the device 4 comprises a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 comprise one or more of the following interfaces: - RF radio frequency interface, for example Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English “Universal Serial Bus” or “Universal Serial Bus” in French); HD MI interface (from the English “High Definition Multimedia Interface” or “High Definition Multimedia Interface” in French); - LIN interface (from the English “Local Interconnect Network”).
[0129] According to another particular and non-limiting exemplary embodiment, the device 4 comprises a communication interface 43 which makes it possible to establish communication with other devices (such as other computers of the on-board system) via a communication channel 430. The communication interface 43 corresponds for example to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds for example to a wired network of the CAN (Controller Area Network) type, CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by the ISO 17458 standard) or Ethernet (standardized by the ISO / IEC 802-3 standard).
[0130] According to a particular and non-limiting exemplary embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch-sensitive or not, one or more speakers 450 and / or other peripherals 460 via the output interfaces 44, 45, 46 respectively. According to a variant, one or other of the external devices is integrated into the device 4.
[0131] Of course, the present invention is not limited to the exemplary embodiments described above but extends to a method for determining the depth of a pixel of an image acquired by a vision system, and / or for measuring a distance separating an object from a vehicle carrying a vision system, the depth and / or the distance being predicted and / or measured via a depth prediction model learned according to the learning method described above, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0132] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-based motor vehicle, comprising the device 4 of [Fig.4].
Claims
1. Claims Method for learning a depth prediction model implemented by a convolutional neural network associated with a stereoscopic vision system embedded in a vehicle (10), the stereoscopic vision system comprising a first camera (11) and a second camera (12) arranged so as to each acquire an image of a three-dimensional scene from a different point of view, said method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception (31) of a first image and a second image acquired by the first camera (11) and the second camera (12) respectively at the same acquisition time instant; - generation (32) of a set of first pairs of feature maps by a feature extractor from the first and second images, each first pair of feature maps comprising a first feature map associated with the first image and a second feature map associated with the second image and each first pair of feature maps having a different definition; - prediction (33), for each first pair of feature maps, of directions associated with pixels of a set of pixels of the first feature map, called first pixels, and of directions associated with pixels of a set of pixels of the second feature map, called second pixels, by a direction prediction model from respectively the first feature map and the second feature map; - prediction (34), for each first pair of feature maps, of depths associated with the first pixels and the second pixels by said depth prediction model from respectively the first feature map and the second feature map; - association (35), with each first pair of characteristic cards, of a second pair of characteristic cards comprising a third characteristic card and a fourth characteristic card, the third feature map being generated from the first feature map, the directions and depths associated with the first pixels and extrinsic parameters of the stereoscopic vision system, and the fourth feature map being generated from the second feature map, the directions and depths associated with the second pixels and extrinsic parameters of the stereoscopic vision system;- determining (36), for each first pair of feature maps, a loss error from a first error determined for each first pixel by comparing said each first pixel of the first feature map to a pixel of the fourth feature map corresponding to said each first pixel and a second error determined for each second pixel by comparing said each second pixel of the second feature map to a pixel of the third feature map corresponding to said each second pixel; - learning (37) the depth prediction model by minimizing a global error determined from the loss errors determined for the set of first pairs of feature maps.;
2. Method according to claim 1, for which the global error is the sum of the loss errors, determined by the following function: Lg = 1Lr(Lr) With: • La the global error, • Lr a loss error for a first pair of feature maps R, and * R( ) a sum over the set of first pairs of feature maps.
3. The method of claim 1 or 2, wherein colorimetric values are associated with each pixel of the first, second, third, and fourth feature maps.
4. Method according to claim 3, for which the first error, respectively second error, is determined by the following function: l,(P) = LJ (i-« » ■ \i(p)-i(p ) Ji4ss / w( / (p), ï(p) ) ) ] With: • Lp*(p) the first error noted L^p), respectively the second error noted L2(p), p being a pixel defined by its coordinates in a characteristic map, • l(p) a colorimetric value of the pixel p in the first characteristic map, respectively second characteristic map, * l(p} a colorimetric value of the pixelp in the fourth characteristic map, respectively third characteristic map, • SSIM a function which takes into account a local structure, and • “ a weighting factor depending in particular on the type of environment in which the vehicle operates (10).
5. Method according to one of claims 3 to 4, for which the loss error is further determined from a construction error determined by the following function: L^p) W(p) MMp) With: • ^smootiÀP) the construction error Ls3(p) for a pixel p of the third feature map, respectively the construction error L^p) for a pixel p of the fourth feature map, • D(p) is a depth of a pixel p obtained from the third depth map, respectively obtained from the fourth depth map; • VF is a parameter matrix; • 0 is the order of a smoothing gradient; • an L1 norm of the second-order depth gradients is calculated with VF =1, and0 =2; • x and are the dimensions of the feature maps; • fi is a hyperparameter dependent on the environment in which the vehicle (10) is moving;and • It[Pt) is a colorimetric value of the pixelp in the third feature map, respectively fourth feature map.;
6. A method according to claim 5, wherein the loss error is determined by the following function: LR = Ep{min(Ll(p^ L2(p) ) + Ls3(p) +Ls4(p)) With: • Lr the loss error, * Lp^p) 'a first error for a pixelp, p being a pixel defined by its coordinates in a feature map, * L^p) 'a second error for the pixelp, • / the construction error for a pixelp of the third feature map, and * L, / p) the construction error for a pixelp of the fourth feature map.
7. Method according to one of claims 1 to 3, for which the third and fourth feature maps are generated using the following function: p )] ) With: • Ps a pixel of a generated feature map corresponding to the third feature map, respectively the fourth feature map, • n a function for going from homogeneous coordinates to pixel coordinates by removing one dimension of a vector, • a direction prediction model associated with the second camera (12), respectively with the first camera (11), • K' a direction prediction model associated with the first camera (11), respectively with the second camera (12), • T an extrinsic matrix of the stereoscopic vision system, • 0 a projection function in the three-dimensional scene of a pixel as a function of its depth,and * ) is a depth associated with a pixel pt of the first feature map, respectively of the second feature map.,
8. Computer program comprising instructions for implementing the method according to any one of the preceding claims, when these instructions are executed by a processor.
9. Device (4) configured to learn a depth prediction model by a vision system on board a vehicle (10), said device (4) comprising a memory (41) associated with at least one of the following: 28 minus one processor (40) configured to implement the steps of the method according to any one of claims 1 to 7.
10. Vehicle (10) comprising the device (4) according to claim 9.
Citation Information
Patent Citations
Stereo depth estimation using deep neural networks
US20190295282A1