Method and device for determining the depth of a pixel of an image by a neural network associated with a vision system on board a vehicle

A neural network corrects image distortions in wide-angle cameras by predicting depth and direction, improving ADAS system precision and reliability.

FR3158381A1Pending Publication Date: 2025-07-18STELLANTIS AUTO SAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
FR2024000450
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Wide-angle cameras in vehicles introduce distortions in captured images, making it difficult to accurately estimate distances and object sizes, which affects the precision of ADAS systems relying on these images.

Method used

A neural network is trained to predict depth and direction of pixels using a learning phase that involves receiving images from different observation positions, determining displacement, and minimizing errors through an analytical projection model to correct distortions.

Benefits of technology

The method accurately predicts the distance between the vehicle and objects, enhancing the precision and reliability of ADAS systems by correcting image distortions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method or device for determining a depth of a pixel of an image by a neural network associated with a vision system embedded in a vehicle comprising a camera. Indeed, the neural network is learned in a learning phase comprising the reception (31, 32) of a first and a second images representative of a three-dimensional scene observed from two observation positions by minimizing (37) an error in generating a third image, the error being obtained by comparing the third image to the second image, the third image being generated (36) from pixels of the first image projected successively into the three-dimensional scene (35) then into the third image using an analytical projection model associated with the camera, from predicted depths and directions (34) for the pixels of the first image and a displacement (33) between the two observation positions. Figure for the abstract: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for determining the depth of a pixel of an image by a neural network associated with a vision system on board a vehicle Technical field

[0001] The present invention relates to methods and devices for determining a depth by a vision system on board a vehicle, for example in a motor vehicle. The present invention also relates to a method and a device for measuring a distance separating an object from a vehicle carrying a vision system. Technological background

[0002] Many modern vehicles are equipped with so-called AD AS (Advanced Driver-Assistance System). Such AD AS systems are passive and active safety systems designed to eliminate the element of human error in driving vehicles of all types. AD AS use advanced technologies to assist the driver while driving and thus improve their performance. AD AS use a combination of sensor technologies to perceive the environment around a vehicle, then provide information to the driver or act on certain vehicle systems.

[0003] There are several levels of ADAS, such as rearview cameras and blind spot sensors, lane departure warning systems, adaptive cruise control and automatic parking systems.

[0004] The AD AS embedded in a vehicle are supplied with data obtained one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or possible obstacles present around a vehicle in order, for example: • to adapt the vehicle lighting according to the presence of other users; • automatically regulate the vehicle speed; • to act on the braking system in the event of a risk of impact with an object.

[0005] In order to have an extended view of the vehicle's environment, i.e. a three-dimensional scene taking place around the vehicle, a wide-angle camera, i.e. a camera with a wide field of vision, is recommended. Indeed, the use of a wide-angle camera has many advantages over "standard" cameras: • a wider field of vision allowing you to capture a larger part of the scene, which is particularly important in the context of driving, where it is essential to monitor the environment both to the sides and in front of the vehicle, for example, • greater efficiency in perceiving complex environments, such as intersections, tight turns, parking spaces, etc., by minimizing blind spots and providing a more complete view of driving situations, and • increased safety by more easily detecting obstacles, other vehicles, pedestrians and cyclists in areas adjacent to the vehicle.

[0006] Although wide-angle cameras offer many advantages, they can also introduce distortions into the captured images, and these distortions can lead to certain problems such as: • distortion of straight lines, for example barrel or pincushion, causing straight lines in an image acquired by the wide-angle camera to bend, making it difficult to estimate the actual distances between objects, especially towards the edges of the image, • stretching or compressing objects, especially towards the edges of the image, changing the apparent size of objects, which can be problematic when judging the distance or actual size of objects, • changing the proportions of objects, making them larger or smaller than their actual size and more complex to identify, and • the difficulty of rectifying or correcting distortion in post-processing which can be complex and can lead to a loss of information.

[0007] Thus, the processing of an image acquired by a wide-angle camera requires special processing, in particular because of the strong distortion present in the image. Determining a distance separating the vehicle carrying the camera from an object in the scene is then not achievable with the methods commonly used for standard cameras used in certain so-called monoscopic or stereoscopic vision systems.

[0008] Furthermore, the quality of the data emitted by a vision system, i.e. the accuracy of the positioning of the objects in the three-dimensional scene, has a direct impact on the proper functioning of the driving assistance systems using this data. It is then important to control the accuracy of a model for predicting the distance and position of an object in the scene. Summary of the present invention

[0009] An object of the present invention is to solve at least one of the problems of the technological background described above.

[0010] Another object of the present invention is to improve the quality of the data resulting from the processing of an image acquired by a vision system, in particular by a network of neurons associated with this vision system.

[0011] Another object of the present invention is to improve road safety, in particular by improving the operational safety of AD AS systems supplied by data obtained from a wide-angle camera.

[0012] According to a first aspect, the present invention relates to a method for determining a depth of a pixel of an image by a neural network associated with a vision system on board a vehicle, characterized in that the neural network is learned in a learning phase comprising the following steps: - reception of a set of pixels of a first image representative of a set of points of a three-dimensional scene observed from a first observation position; - reception of a set of pixels of a second image representative of a set of points of the three-dimensional scene observed from a second observation position; - determining a displacement between the first observation position and the second observation position by a displacement prediction model from the sets of pixels of the first and second images; - determining a depth and a direction associated with each pixel of the set of pixels of the second image by a depth and direction prediction model comprising the neural network; - determining a set of point positions of the three-dimensional scene, each point position of the set of point positions being determined from a depth and a direction associated with a pixel of the set of pixels of the first image; - generation of a set of pixels of a third image, each pixel of the set of pixels of the third image corresponding to a projection of a position of a point of the set of positions of points towards the second observation position from an analytical projection model associated with a camera of the vision system, the determined displacement and colorimetric information of the set of pixels of the first image; and - learning the neural network by minimizing an error determined by comparing all the pixels of the second image to all the pixels of the third image.

[0013] Such a neural network thus makes it possible to accurately predict the depth of a pixel of an image acquired by the camera, i.e. the distance separating the vehicle carrying the camera from a physical object of the three-dimensional scene associated with this pixel.

[0014] According to a variant of the method, the first and second images are acquired by the camera at two distinct time instants.

[0015] According to yet another variant of the method, the movement prediction model comprises another neural network learned by minimizing the error.

[0016] According to an additional variant of the method, another phase of learning the neural network comprises the following steps: - receiving a set of projected pixels and annotated data associated with each pixel of the other set of pixels, each projected pixel being an image of a virtual point of a virtual three-dimensional scene projected towards a third observation position from the analytical projection model associated with the camera, and the annotated data associated with the projected pixel being representative of coordinates of the virtual point in the virtual three-dimensional scene of which the projected pixel is the image; - determination of a depth and direction associated with each pixel projected by the depth and direction prediction model; - determination, for each projected pixel, of a position of a reprojected point of the virtual three-dimensional scene associated with the projected pixel from a depth and a direction associated with the projected pixel; - learning the neural network by minimizing another error determined by comparing, for each projected pixel, the position of a reprojected point associated with the projected pixel to the annotated data associated with the projected pixel.

[0017] According to another variant of the method, the other error is determined from Euclidean distances separating, for each projected pixel, the position of the reprojected point associated with the projected pixel from the position of a virtual point of which the projected pixel is the image.

[0018] According to yet another variant of the method, the analytical projection model associated with the camera is determined by calibrating the camera using a test pattern.

[0019] According to a further variant of the method, the camera is a wide-angle camera.

[0020] According to a second aspect, the present invention relates to a device for determining Depth determination by a vision system on board a vehicle, the device comprising a memory associated with at least one processor configured for implementing the steps of the method according to the first aspect of the present invention.

[0021] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0022] According to a fourth aspect, the present invention relates to a computer program which comprises instructions adapted for the execution of the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0023] Such a computer program may use any programming language and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0024] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to the first aspect of the present invention.

[0025] On the one hand, the recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM memory, a CD-ROM or a microelectronic circuit type ROM memory, or even a magnetic recording means or a hard disk.

[0026] Furthermore, this recording medium may also be a transmissible medium such as an electrical or optical signal, such a signal being able to be conveyed via an electrical or optical cable, by conventional or hertzian radio or by self-directed laser beam or by other means. The computer program according to the present invention may in particular be downloaded from an Internet-type network.

[0027] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to perform or to be used in performing the method in question. Brief description of the figures

[0028] Other characteristics and advantages of the present invention will emerge from the description of the particular and non-limiting exemplary embodiments of the present invention below, with reference to the appended figures 1 to 5, in which:

[0029] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting example of embodiment of the present invention;

[0030] [Fig.2] illustrates a flowchart of the different steps of a method for determining a depth of a pixel of an image by a neural network associated with a vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;

[0031] [Fig.3] illustrates a flowchart of the different steps of a first method of learning the neural network used in the method of [Fig.2], according to a particular and non-limiting exemplary embodiment of the present invention;

[0032] [Fig.4] illustrates a flowchart of the different stages of a second process of learning the neural network used in the method of [Fig.2], according to a particular and non-limiting exemplary embodiment of the present invention;

[0033] [Fig.5] schematically illustrates a device configured to determine a depth of a pixel of an image by a neural network associated with a vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention. Description of examples of implementation

[0034] A method and a device for determining a depth of a pixel of an image by a neural network associated with a vision system on board a vehicle will now be described in what follows with joint reference to Figures 1 to 5. The same elements are identified with the same reference signs throughout the description which follows.

[0035] The terms "first(s)", "second(s)" (or "first(s)", "second(s)"), etc. are used in this document by arbitrary convention to enable different elements (such as operations, means, etc.) implemented in the embodiments described below to be identified and distinguished. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0036] According to a particular and non-limiting example of embodiment of the present invention, a method for determining a depth of a pixel of an image by a neural network associated with a vision system on board a vehicle is for example implemented by a computer of the on-board system of the vehicle controlling this vision system.

[0037] The vision system comprises at least one camera arranged so as to acquire an image of a three-dimensional scene from a determined point of view.

[0038] For this purpose, the method for determining a depth of a pixel of an image comprises a depth and direction prediction model comprising a neural network associated with the vision system on board the vehicle. This neural network is learned in a learning phase comprising the reception of data representative of sets of pixels of a first and a second image representative of a three-dimensional scene observed from two observation positions, the learning of the neural network being obtained by minimizing an error in generating a third image, the error being obtained by comparing pixels of the third image to pixels of the second image, the third image being generated from pixels of the first image reprojected into the three-dimensional scene and then projected into the third image.

[0039] The reprojection into the three-dimensional scene is performed from predicted depths and directions for pixels of the first image and the projection into the third image is performed from a determined displacement between the two observation positions and from an analytical projection model associated with the vision system camera.

[0040] The neural network for determining a depth of a pixel in an image is then learned, the reliability and precision of the depths predicted by the depth and direction prediction model implementing this neural network are then improved. A driving assistance system or AD AS receiving data representative of previously determined depths then gains in precision and reliability.

[0041] [Fig. 1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.

[0042] Such an environment 1 corresponds, for example, to a road environment formed of a network of roads accessible to the vehicle 10.

[0043] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, with an electric motor(s) or even a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as an automobile, a truck, a bus, a motorcycle. Finally, the vehicle 10 corresponds to an autonomous vehicle or not, that is to say a vehicle traveling according to a determined level of autonomy or under the total supervision of the driver.

[0044] The vehicle 10 advantageously comprises at least one on-board camera 11, configured to acquire images of a three-dimensional scene taking place in the environment of the vehicle 10 from a current observation position. The camera 11 forms a monoscopic vision system if it is used alone as illustrated in [Fig.l]. The present invention is however not limited to a monoscopic vision system comprising a single camera but extends to any vision system comprising at least one camera, for example 1, 2, 3 or 5 cameras.

[0045] The camera 11 has intrinsic parameters, in particular: - a focal length, - distortions which are due to imperfections in the optical system of the camera 11, - a direction of the optical axis of the camera 11, and - a resolution.

[0046] The intrinsic parameters characterize the transformation which associates, for an image point, subsequently called “point”, its three-dimensional coordinates in the frame of reference of the camera 11 with the pixel coordinates in an image acquired by the camera 11. These parameters do not change if the camera is moved.

[0047] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, that is to say that the distortions are not taken into account, that their correction is processed at the time of image acquisition or at the time of calibration.

[0048] The camera 11 is arranged so as to acquire an image of a three-dimensional scene according to a defined point of view, the point of view is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10 as illustrated in [Fig. 1].

[0049] The camera 11 for example acquires images of a three-dimensional scene located in front of the vehicle 10, the first camera 11 covering an acquisition field 12. An object 13 is placed in the acquisition field 12 of the camera 11, defining an occlusion field 14 for the vision system, an object present in the occlusion field 14 not being observable by the camera 11 from its current observation position.

[0050] It is obvious that it is possible to use such a vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently, the invention not being limited to the observation of a three-dimensional scene taking place in front of the vehicle 10 carrying the camera 11.

[0051] According to a particular embodiment, the camera 11 is of the “wide-angle” type, a wide-angle camera being for example equipped with a lens designed to acquire an image representative of a three-dimensional scene perceived according to a wider field of vision than that of a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens makes it possible to capture a larger portion of the three-dimensional scene taking place in front of or around the camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by this camera. The angle a of the field of vision of the camera 11 is for example equal to 120°, 145°, 180° or 360°, whereas a standard camera offers, for example, an open field of vision following an angle of 45° or less.Such a camera 11 corresponds, for example, to a camera equipped with mirrors or even to a "fisheye" camera (in French "fish eye"). Wide-angle lenses have a shorter focal length compared to standard lenses, which makes them suitable for acquiring images of landscapes, architecture, road intersections or any other subject requiring an extended perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with a depth of field. extent.

[0052] An image acquired by the camera 11 at an acquisition time instant is presented in the form of data representing pixels characterized by: - coordinates in the image; and - data relating to the colors and brightness of objects in the observed scene in the form, for example, of RGB colorimetric coordinates (from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness).

[0053] Each pixel of the acquired image is representative of an object of the three-dimensional scene present in the field of vision of the camera 11. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a luminous point resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When the light strikes this object, photons are emitted or reflected, captured by a photosensitive sensor of the camera 11 after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image faithfully representing the physical object observed by the camera 11. An image point previously presented is thus a point on a surface of an object of the three-dimensional scene.

[0054] According to a particular embodiment, the image acquired by the camera 11 comprises a distortion equal to 0.5%, 0.8% or greater than 1%. The measurement of such a distortion corresponds to the determination of a ratio between: - the maximum spacing of a pixel of the image from a straight line of the first three-dimensional scene whose image is a line touching the longest edge of the first image, either at the center of the edge of the image, or at the corners of the edge of the image, and - the length of this edge.

[0055] Commonly, distortion is considered, in the world of photography, as: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • bothersome if it is greater than or equal to 1% or more.

[0056] A barrel distortion is characterized by a positive percentage, while a crescent distortion is characterized by a negative percentage.

[0057] When the vehicle 10 is in motion, then two images acquired by the camera 11 at two distinct time instants represent views of the same three-dimensional scene. mentional taken from different viewpoints or observation positions, the observation positions of the camera 11 being distinct. On this three-dimensional scene are for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.

[0058] Each image is for example sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.

[0059] A method for determining a depth by a vision system on board the vehicle 10 is advantageously implemented by the vehicle 10, that is to say by a processor, a computer or a combination of computers of the on-board system of the vehicle 10, for example by the computer(s) in charge of the vision system of the vehicle 10.

[0060] [Fig. 2] illustrates a flowchart of the different steps of a method 2 for determining a depth of a pixel of an image by a neural network associated with a vision system on board a vehicle, for example in the vehicle 10 of [Fig. 1], according to a particular and non-limiting exemplary embodiment of the present invention. The method 2 is for example implemented by a device of the vision system on board the vehicle 10 or by the device 5 of [Fig. 5].

[0061] In a step 21, data representative of an image acquired by the first camera 11 are received.

[0062] In a step 22, depths associated with a set of pixels of the image are determined by a depth and direction prediction model, the depth and direction prediction model comprising a neural network learned in a training phase.

[0063] Each determined depth then corresponds to a distance separating the vehicle 10 or a part of the vehicle 10 from an object of the three-dimensional scene with which a pixel is associated, the determination of a depth of a pixel then corresponding to a measurement of a distance separating an object from the vehicle carrying the vision system.

[0064] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another user present on the road, the ADAS is then able to determine this distance precisely. For example, if the ADAS has the function of acting on a braking system of the vehicle 10 in the event of a risk of collision with another user of the road and the distance separating vehicle 10 from this same road user decreases significantly, then the ADAS is able to detect this sudden approach and act on the braking system of vehicle 10 to avoid a possible accident.

[0065] [Fig. 3] illustrates a flowchart of the different steps of a first method of learning the neural network used in a method of determining a depth of a pixel of an image, for example in method 2 of [Fig. 2], according to a particular and non-limiting exemplary embodiment of the present invention.

[0066] The first learning method 3 is for example implemented by the device on board the vehicle 10 implementing the method for determining a depth by a vision system on board a vehicle or by the device 5 of [Fig.5].

[0067] In a step 31, a set of pixels of a first image is received, the set of pixels of the first image being representative of a set of points of a three-dimensional scene observed from a first observation position.

[0068] In a step 32, a set of pixels of a second image is received, the set of pixels of the second image being representative of a set of points of the same three-dimensional scene observed from a second observation position.

[0069] Thus, the first and second images correspond to two images of the same three-dimensional scene seen or observed from two observation positions.

[0070] According to a particular exemplary embodiment, the first and second images are acquired by the camera 11 at two distinct time instants. The three-dimensional scene is then a real scene, taking place around the vehicle 10. If the vehicle 10 is moving, then these two observation positions are distinct.

[0071] It should be noted that the first and second images may be associated with annotated data, for example associated with data from another depth measurement mode such as a LIDAR. Images alone or associated with annotated data may be combined, the training being carried out by iterations of the first learning method with pairs of first and second images acquired by the camera 11 and sometimes associated with annotated data.

[0072] Each training or each iteration of the process must ensure the convergence of the solution, namely the prediction accuracy on the annotated data. For example, 20% of the training data is annotated data allowing the convergence of the learned neural network to be validated.

[0073] In a step 33, a displacement between the first observation position and the second observation position is determined by a displacement prediction model from the sets of pixels of the first and second images.

[0074] According to a particular exemplary embodiment, the displacement prediction model ("pose estimation model" in English) includes another neural network. Such a displacement prediction model is known to those skilled in the art, and is for example presented in the document "Digging Into Self-Supervised Monocular Depth Estimation", written by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow published in August 2019.

[0075] In a step 34, a depth and a direction associated with each pixel of the set of pixels of the second image are determined by the depth and direction prediction model comprising the neural network.

[0076] Such a depth and direction prediction model is known to those skilled in the art; it is notably presented in the document “Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion” written by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich and Adrien Gaidon published in August 2020.

[0077] It should be noted that this document also presents a displacement prediction model.

[0078] In a step 35, a set of point positions of the three-dimensional scene is determined, each point position of the set of point positions being determined from a depth and a direction associated with a pixel of the set of pixels of the first image. Each point thus corresponds to the reprojection of a pixel of the first image in the three-dimensional scene. In the case of perfect reprojection, this point corresponds for example to an image point of the scene with which the pixel is associated during the acquisition or generation of the first image.

[0079] In a step 36, a set of pixels of a third image is generated, each pixel of the set of pixels of the third image corresponding to a projection of a position of a point of the set of positions of points towards the second observation position from the analytical projection model associated with the camera 11 of the vision system, the previously determined displacement and colorimetric information of the set of pixels of the first image.

[0080] Unlike the method presented in the document “Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion”, the projection of a position of a point from the set of point positions to the second observation position is not performed by a neural network but by the analytical projection model associated with the camera 11 of the vision system, this step is then controlled and does not require additional learning, a calibration of the camera 11 being sufficient to obtain this analytical model.

[0081] According to a particular exemplary embodiment, the analytical projection model is determined by calibrating the camera 11 using a test pattern. Such a method calibration is known to those skilled in the art, and is for example presented in the document “Single View Point Omnidirectional Camera Calibration from Planar Grids”, written by Christopher Mei and Patrick Rives published in April 2007. Thus, the intrinsic parameters of the camera 11 are determined and can be used to analytically determine the coordinates of a pixel of an image acquired by the camera 11 from the relative position of a point in the three-dimensional scene with respect to the observation position of the camera 11, the pixel being associated with this point.

[0082] In a step 37, the neural network is learned by minimizing an error determined by comparing all the pixels of the second image to all the pixels of the third image. Such an error corresponds for example to a photometric error as presented in the document “Digging Into Self-Supervised Monocular Depth Estimation” and corresponds to a loss function.

[0083] According to a particular exemplary embodiment, the other neural network used by the displacement prediction model is also learned by minimizing this error. Thus, the depth and direction prediction and displacement prediction models are learned simultaneously.

[0084] Thus, the depth and direction prediction model used for the depth prediction of a pixel of an image acquired by the camera 11 is made reliable thanks to this first learning method.

[0085] [Fig.4] illustrates a flowchart of the different steps of a second method of learning the neural network used in a method of determining a depth of a pixel of an image, for example in the method of [Fig.2], according to a particular and non-limiting exemplary embodiment of the present invention.

[0086] In a step 41, a set of projected pixels and annotated data are received. Each projected pixel is an image of a virtual point of a virtual three-dimensional scene projected to a third observation position from the analytical projection model associated with the camera 11. The annotated data associated with the projected pixel are representative of the coordinates of the virtual point in the virtual three-dimensional scene of which the projected pixel is the image.

[0087] Thus, each pixel of the set of projected pixels corresponds to a virtual point of the virtual scene whose position is known and received via the annotated data.

[0088] To ensure coverage of the points of the volume measurable by the camera 11, for example for a wide-angle camera, the angular field of vision can be 120° horizontally, 60° vertically and the measurable depth 200 meters, the points arranged in the virtual three-dimensional scene must then cover a conical space uniformly. The density of the points in this space is for example defined according to a maturity level of a neural network learned via this second learning process, starting for example with one point per degree of the angular field of view and one point per meter of depth. The density of the points is then increased, for example with a rate of 200% for each iteration of this process. The iterations of the second learning process are then terminated when an increase in point density no longer has an effect on the depth prediction. Such a definition of the learning data, i.e. of the set of virtual pixels, makes it possible to limit the quantity of training data and thus limits the learning time of a trained neural network.

[0089] In a step 42, a depth and direction associated with each projected pixel are determined by the depth and direction prediction model.

[0090] In a step 43, a position of a reprojected point of the virtual three-dimensional scene associated with the projected pixel is determined for each projected pixel from a depth and a direction associated with the projected pixel.

[0091] According to a particular exemplary embodiment, a position of each point of the virtual scene is defined by coordinates along three axes, a first coordinate x along a first axis, a second coordinate y along a second axis and a third coordinate z along a third axis, the three axes defining an orthonormal reference frame. Thus, for each projected pixel, the virtual point of the virtual scene of which it is the image has coordinates (x; y; z) and the associated reprojected point has coordinates (x'; y'; z'). In other words, the annotated data associated with the projected point includes the coordinates (x; y; z) of the virtual point of the virtual scene of which it is the image.

[0092] In a step 44, the neural network is learned by minimizing another error determined by comparing, for each projected pixel, the position of the reprojected point associated with the projected pixel to the annotated data associated with the projected pixel.

[0093] According to the previous particular embodiment, the other error is determined from Euclidean distances separating, for each projected pixel, the position of the reprojected point associated with the projected pixel from the position of the virtual point of which the projected pixel is the image. This Euclidean distance is then obtained by the following formula:

[0094] [Math.l] D = ^(x - x')2 + (y - y')2 + (z - z')2 ]

[0095] With: • D the Euclidean distance determined for a projected pixel, • the “square root” function, • (x;y;z) the coordinates of the virtual point of the virtual scene of which the projected pixel is the image, and • (x' ;y' ;z') the coordinates of the reprojected point associated with the projected pixel.

[0096] According to a particular exemplary embodiment, the other error is determined from the Euclidean distances determined for the set of projected pixels, for example the other error is equal to the sum, the average or the median of these Euclidean distances.

[0097] [Fig. 5] schematically illustrates a device 5 configured for the determination of a depth by a vision system on board a vehicle 10, according to a particular and non-limiting exemplary embodiment of the present invention. The device 5 corresponds for example to a device on board the first vehicle 10, for example a computer.

[0098] The device 5 is for example configured for the implementation of the operations described with regard to [Fig.l] and / or steps described with regard to [Fig.2]. Examples of such a device 5 include, but are not limited to, on-board electronic equipment such as an on-board computer of a vehicle, an electronic calculator such as an ECU (“Electronic Control Unit”), a smartphone, a tablet, a laptop. The elements of the device 5, individually or in combination, can be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. The device 5 can be produced in the form of electronic circuits or software (or computer) modules or even a combination of electronic circuits and software modules.

[0099] The device 5 comprises one (or more) processor(s) 50 configured to execute instructions for carrying out the steps of the method and / or for executing the instructions of the software(s) embedded in the device 5. The processor 50 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 5 further comprises at least one memory 51 corresponding for example to a volatile and / or non-volatile memory and / or comprises a memory storage device which may comprise volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.

[0100] The computer code of the embedded software(s) comprising the instructions to be loaded and executed by the processor is for example stored in the memory 51.

[0101] According to various particular and non-limiting embodiments, the device 5 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Telematic Control Unit”), for example via a communication bus or through dedicated input / output ports.

[0102] According to a particular and non-limiting exemplary embodiment, the device 5 comprises a block 52 of interface elements for communicating with external devices. The interface elements of the block 52 comprise one or more of the following interfaces: - RF radio frequency interface, for example Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English “Universal Serial Bus” or “Universal Serial Bus” in French); HDMI interface (from the English “High Definition Multimedia Interface” or “High Definition Multimedia Interface” in French); - LIN interface (from the English “Local Interconnect Network”).

[0103] According to another particular and non-limiting exemplary embodiment, the device 5 comprises a communication interface 53 which makes it possible to establish communication with other devices (such as other computers of the on-board system) via a communication channel 530. The communication interface 53 corresponds for example to a transmitter configured to transmit and receive information and / or data via the communication channel 530. The communication interface 53 corresponds for example to a wired network of the CAN (Controller Area Network) type, CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by the ISO 17458 standard) or Ethernet (standardized by the ISO / IEC 802-3 standard).

[0104] According to a particular and non-limiting exemplary embodiment, the device 5 can provide output signals to one or more external devices, such as a display screen 540, touch-sensitive or not, one or more speakers 550 and / or other peripherals 560 via the output interfaces 54, 55, 56 respectively. According to a variant, one or other of the external devices is integrated into the device 5.

[0105] Of course, the present invention is not limited to the exemplary embodiments described above but extends to a method for determining a depth by a vision system on board a vehicle, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0106] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-based motor vehicle, comprising the device 5 of [Fig.5].

Claims

Claims

1. Method for determining a depth of a pixel of an image by a neural network associated with a vision system on board a vehicle (10), the method being implemented by a processor and characterized in that the neural network is learned in a learning phase comprising the following steps: - reception (31) of a set of pixels of a first image representative of a set of points of a three-dimensional scene observed from a first observation position; - reception (32) of a set of pixels of a second image representative of a set of points of the three-dimensional scene observed from a second observation position; - determination (33) of a displacement between the first observation position and the second observation position by a displacement prediction model from the sets of pixels of the first and second images; - determining (34) a depth and a direction associated with each pixel of the set of pixels of the second image by a depth and direction prediction model comprising the neural network; - determining (35) a set of point positions of the three-dimensional scene, each point position of the set of point positions being determined from a depth and a direction associated with a pixel of the set of pixels of the first image; - generation (36) of a set of pixels of a third image, each pixel of the set of pixels of the third image corresponding to a projection of a position of a point of the set of positions of points towards the second observation position from an analytical projection model associated with a camera (11) of the vision system, the determined displacement (33) and colorimetric information of the set of pixels of the first image; and - training (37) of the neural network by minimizing an error determined by comparing the set of pixels of the second image to the set of pixels of the third image.

2. Method according to claim 1, for which the first and second images are acquired by the camera (11) at two distinct time instants.

3. A method according to claim 1 or 2, wherein the movement prediction model comprises another neural network learned by minimizing the error.

4. Method according to one of claims 1 to 3, which comprises a further phase of training the neural network comprising the following steps: - receiving (41) a set of projected pixels and annotated data associated with each pixel of the other set of pixels, each projected pixel being an image of a virtual point of a virtual three-dimensional scene projected towards a third observation position from the analytical projection model associated with the camera (11), and the annotated data associated with the projected pixel being representative of coordinates of the virtual point in the virtual three-dimensional scene of which the projected pixel is the image; - determining (42) a depth and a direction associated with each projected pixel by the depth and direction prediction model;- determining (43), for each projected pixel, a position of a reprojected point of the virtual three-dimensional scene associated with the projected pixel from a depth and a direction associated with the projected pixel; - learning (44) the neural network by minimizing another error determined by comparing, for each projected pixel, the position of a reprojected point associated with the projected pixel to the annotated data associated with the projected pixel.;

5. Method according to claim 4, for which the other error is determined from Euclidean distances separating, for each projected pixel, the position of the reprojected point associated with the projected pixel from the position of a virtual point of which the projected pixel is the image.

6. Method according to one of claims 1 to 5, for which the analytical projection model associated with the camera (11) is determined by calibrating the camera (11) using a test pattern.

7. Method according to one of the preceding claims, for which the camera (11) is a wide-angle camera.

8. Computer program comprising instructions for implementing the method according to any one of the preceding claims, when these instructions are executed by a processor.

9. Device (5) for determining a depth by a vision system on board a vehicle (10), said device (5) comprising a memory (51) associated with at least one processor (50) configured for implementing the steps of the method according to any one of claims 1 to 7.

10. Vehicle (10) comprising the device (5) according to claim 9.

Citation Information

Patent Citations

  • Computer-implemented method of self-supervised learning in neural network for robust and unified estimation of monocular camera ego-motion and intrinsics

    EP4216107A1