Method and device for learning a depth prediction model insensitive to the presence of dynamic objects in training images.

The method enhances depth prediction models by partitioning images into blocks, predicting depths, and minimizing photometric errors to exclude dynamic objects, addressing inaccuracies caused by moving objects and improving ADAS system safety.

FR3160798A1Active Publication Date: 2025-10-03STELLANTIS AUTO SAS
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
FR2024003394
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-02
Publication Date
2025-10-03
Estimated Expiration
2044-04-02

AI Technical Summary

Technical Problem

Existing depth prediction models in monocular vision systems for vehicles are compromised by the presence of dynamic objects, leading to degraded learning quality and inaccurate depth predictions.

Method used

A method for learning a depth prediction model using a convolutional neural network that partitions images into blocks, predicts depths, generates corresponding blocks based on camera movement, and minimizes loss error based on photometric errors to exclude dynamic objects during the learning phase.

Benefits of technology

Enables accurate depth prediction even in the presence of dynamic objects, improving the learning phase and operational safety of ADAS systems by ensuring precise depth measurements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Method or device for learning a depth prediction model associated with a monocular vision system. Indeed, the method comprises the partitioning (32) of a first and a second image into a set of first and second blocks to form first pairs of blocks each comprising a first and a second block. For each first pair of blocks, a second pair of blocks is associated and comprises a third and a fourth block generated (36) from the first and second blocks and from predicted depths (33) for the pixels of the first and second blocks. The model is learned (39) by minimizing a loss error determined from at least one determined photometric error (37), for at least a first selected pair of blocks (38), by comparing the first and fourth blocks and by comparing the second and third blocks. Figure for abstract: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for learning a depth prediction model insensitive to the presence of dynamic objects in training images. Technical field

[0001] The present invention relates to methods and devices for learning a depth prediction model associated with a stereoscopic vision system on board a vehicle, for example in a motor vehicle. The present invention also relates to a method and a device for determining a depth and / or measuring a distance separating an object from a vehicle carrying a vision system. Technological background

[0002] Many modern vehicles are equipped with so-called AD AS (Advanced Driver-Assistance System). Such AD AS systems are passive and active safety systems designed to eliminate the element of human error in the driving of vehicles of all types. AD AS use advanced technologies to assist the driver while driving and thus improve their performance. AD AS use a combination of sensor technologies to perceive the environment around a vehicle, then provide information to the driver or act on certain vehicle systems.

[0003] There are several levels of ADAS, such as rearview cameras and blind spot sensors, lane departure warning systems, adaptive cruise control and automatic parking systems.

[0004] The AD AS embedded in a vehicle are supplied with data obtained one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or possible obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle speed; - to act on the braking system in the event of a risk of impact with an object.

[0005] A position of another user or of an obstacle is for example determined by a vision system comprising a model for predicting a depth associated with a pixel or a distance separating the vision system from an object in a three-dimensional scene. Such a model is for example learned using images, these images being obtained from a universal database, for example Kitti® or Sceneflow®. Kitti®, for example, provides images of a city-center road environment, but such a database does not include all the road environments in which a vehicle can operate. The training data is therefore unsuitable for training the prediction model for a vehicle traveling in other road environments.

[0006] The quality of the training of the depth or distance prediction model is however very important, in fact, the depths or distances predicted by the prediction model represent for example distances at which the other users or obstacles present in the road environment of the vehicle carrying the vision system and the AD AS are located. The proper functioning of the driving assistance peripherals using this data therefore depends on the quality of the data emitted by the vision system.

[0007] A vision system comprising a single camera, hereinafter called a monocular vision system, uses several consecutive images to predict depths associated with pixels of one of the images. This monocular vision system has the disadvantage of not being able to accurately predict depths associated with pixels of an image when objects of the observed three-dimensional scene are moving in this same scene, these objects being hereinafter called dynamic objects. Indeed, the depth is obtained by comparing positions of pixels corresponding to the same object of the three-dimensional scene in different images.However, if the observed object is in motion, then the pixels associated with this object are displaced between two images acquired at different time instants and the displacement of these pixels is due both to the displacement of the monocular vision system on board the vehicle and to the specific displacement of the dynamic object in the three-dimensional scene. A depth prediction model is then not able to precisely determine the position of this dynamic object. The presence of a dynamic object in a learning phase then disrupts the learning of the depth prediction model, that is to say that the quality of the learning of the depth prediction model is degraded when pixels corresponding to dynamic objects are taken into consideration during the learning phase. Summary of the present invention.

[0008] An object of the present invention is to solve at least one of the problems of the technological background described above.

[0009] Another object of the present invention is to improve the learning phase of a depth prediction model from images acquired by a vision system and representing objects of a three-dimensional scene in motion in this same scene, the depth prediction model being in particular implemented by a neural network and associated with a vision system embedded in a vehicle.

[0010] Another object of the present invention is to improve road safety, in particular by improving the operational safety of AD AS systems supplied by data obtained from a camera of a vision system.

[0011] According to a first aspect, the present invention relates to a method for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, the vision system comprising a camera arranged so as to acquire an image of a three-dimensional scene taking place in the surroundings of the vehicle, the method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception of a first image and a second image acquired by the camera at two distinct acquisition time instants; - partitioning the first image into a set of first blocks and the second image into a set of second blocks, the first and second blocks having the same dimensions, each second block being associated with a first block to form a first pair of blocks, the first and second blocks of a first pair of blocks having the same position in the first image and the second image respectively; - for each first pair of blocks comprising the first block and the second block: • prediction of depths associated with the pixels included in the first block and in the second block by the depth prediction model from the first and second blocks; • generation of a second pair of blocks corresponding to each first pair of blocks and comprising a third and a fourth block, the third block being generated from the first block, predicted depths for the pixels of the first block and extrinsic parameters corresponding to a movement of the camera between the first and second time instants, and the fourth block being generated from the second block, predicted depths for the pixels of the second block and extrinsic parameters; • determination of a photometric error associated with each first pair of blocks from a first photometric error determined by comparison of the first and fourth blocks and a second photometric error determined by comparison of the second and third blocks; - selection of at least a first pair of blocks according to the photometric error associated with each first pair of blocks to obtain at least a first selected pair of blocks; - learning the depth prediction model by minimizing a loss error determined from at least one photometric error associated with said at least one first pair of selected blocks.

[0012] The method advantageously makes it possible to determine the loss error only on the blocks having, for example, the smallest photometric error. The presence of a dynamic object in an image acquired by the vision system then does not impact the learning of the depth prediction model, a part of the image comprising pixels corresponding to this dynamic object being excluded during the selection of blocks on which to minimize the loss error. It is then possible to learn a depth prediction model associated with the vision system even when images acquired by this vision system comprise pixels representative of a dynamic object present in the observed three-dimensional scene.

[0013] According to a variant of the method, the photometric error associated with each first pair of blocks is determined by the following function: = W / ?))] With : • Lp corresponding to the photometric error associated with the first pair of blocks F, • LpiÇp) the first photometric error determined for a pixel P in the first block of the first pair of blocks, the pixel P being defined by its coordinates in two dimensions, and • Lp2(p) the second photometric error determined for pixel P in a second block of the first pair of blocks.

[0014] The fact of retaining only the minimum value among the first and the second error makes it possible in particular not to impact the learning when an object of the three-dimensional scene is occluded in one of the blocks, that is to say when a block of a first pair of blocks comprises pixels corresponding to this object and the other block does not comprise pixels corresponding to this same object.

[0015] According to another variant of the method, the loss error is equal to an average of the at least one photometric error.

[0016] According to a further variant of the method, the first, respectively the second, photometric error is determined by the following function: Lp*M = C1-®) ■ U(p)-l(p)\+a- With : • Lp^( p) the first photometric error noted Lpi( p), respectively the second photometric error noted Lp2(p), P being a pixel defined by its coordinates in two dimensions, • I(p) a value of the pixel P in the first block, respectively in the second block, • a value of the pixel P in the fourth block, respectively in the third block, • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment in which the vehicle operates.

[0017] Such a function is known to those skilled in the art and is commonly used to determine a photometric error by comparing two images or two blocks of images.

[0018] According to another variant of the method, the third block, respectively fourth block, is generated by the following function: Ps = [T^P^ D(Pt) ) ] ) With : • Ps the coordinates of a pixel of the third block, respectively of the fourth block, • 77 a function to go from homogeneous coordinates to pixel coordinates by removing a dimension of a vector, • K a direction prediction model associated with the camera, • T a matrix comprising the extrinsic parameters, and • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates in the first image, respectively second image, and the depth [^pj associated with it.

[0019] According to yet another variant of the method, the first image and the second image each comprise the same number of blocks horizontally and vertically.

[0020] According to an additional variant of the method, the same number of blocks is greater than or equal to three.

[0021] With the same number of blocks, a dynamic object cannot alone cover each block of an image, at least one block of the image then does not include pixels corresponding to a dynamic object.

[0022] According to a second aspect, the present invention relates to a device configured to learn a depth prediction model by a vision system on board a vehicle, the device comprising a memory associated with at least one processor configured to implement the steps of the method according to the first aspect of the present invention.

[0023] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0024] According to a fourth aspect, the present invention relates to a computer program which comprises instructions adapted for executing the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0025] Such a computer program may use any programming language and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0026] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to the first aspect of the present invention.

[0027] On the one hand, the recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM memory, a CD-ROM or a microelectronic circuit type ROM memory, or a magnetic recording means or a hard disk.

[0028] Furthermore, this recording medium may also be a transmissible medium such as an electrical or optical signal, such a signal being able to be conveyed via an electrical or optical cable, by conventional or hertzian radio or by self-directed laser beam or by other means. The computer program according to the present invention may in particular be downloaded from an Internet-type network.

[0029] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to perform or to be used in performing the method in question. Brief description of the figures

[0030] Other characteristics and advantages of the present invention will emerge from the description of the particular and non-limiting exemplary embodiments of the present invention below, with reference to the appended figures 1 to 5, in which:

[0031] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting example of embodiment of the present invention;

[0032] [Fig.2] illustrates a flowchart of the different steps of a method for determining a depth of a pixel of an image by a depth prediction model associated with a vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;

[0033] [Fig.3] illustrates a flowchart of the different steps of a method for learning the depth prediction model used in the method of [Fig.2], according to a particular and non-limiting example of embodiment of the present invention; and

[0034] [Fig.4] schematically illustrates a device configured to learn a model depth prediction by a vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention; and

[0035] [Fig.5] schematically illustrates a first pair of images acquired by the camera on board the vehicle of [Fig.1], according to a particular and non-limiting exemplary embodiment of the present invention. Description of examples of implementation

[0036] A method and a device for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle will now be described in what follows with joint reference to Figures 1 to 5. The same elements are identified with the same reference signs throughout the description which follows.

[0037] The terms "first(s)", "second(s)" (or "first(s)", "second(s)"), etc. are used in this document by arbitrary convention to enable different elements (such as operations, means, etc.) implemented in the embodiments described below to be identified and distinguished. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0038] For the entire description, reception of an image is understood to mean the reception of data representative of an image. Similarly, for obtaining and generating a block, we mean the obtaining and generation of data representative of an image block and for determining a depth or an error, we mean the determination of data representative of a depth or an error. These shortcuts are only intended to simplify the description, however, since the methods described below are implemented by one or more processors, it is obvious that the input and output data of the different steps of a method are computer data.

[0039] According to a particular and non-limiting example of embodiment of the present invention, the depth prediction model is learned in a learning phase comprising the partitioning of a first image into a set of first blocks and the partitioning of a second image into a set of second blocks to form first pairs of blocks each comprising a first block and a second block associated with the first block. The images are in particular two images acquired, by a camera of the vision system on board the vehicle, at two distinct time instants.

[0040] For each first pair of blocks, a second pair of blocks is associated and comprises a third and a fourth block generated from the first and second blocks and predicted depths for the pixels of the first and second blocks.

[0041] The model is learned by minimizing a loss error determined from at least one photometric error, the photometric error being determined, for at least a first pair of selected blocks, by comparing the first and fourth blocks and by comparing the second and third blocks.

[0042] [Fig. 1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.

[0043] The vehicle 10 is located in an environment 1 corresponding, for example, to a road environment formed of a network of roads accessible to the vehicle 10.

[0044] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, with an electric motor(s) or even a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as an automobile, a truck, a bus, a motorcycle. Finally, the vehicle 10 corresponds to an autonomous vehicle or not, that is to say a vehicle traveling according to a determined level of autonomy or under the total supervision of the driver.

[0045] The vehicle 10 advantageously comprises at least one on-board camera 11, configured to acquire images of a three-dimensional scene taking place in the environment of the vehicle 10 from a current observation position. The camera 11 forms a monocular vision system if it is used alone as illustrated in [Fig.l]. The present invention is however not limited to a monocular vision system comprising a single camera but extends to any vision system comprising at least one camera, for example 1, 2, 3 or 5 cameras.

[0046] The camera 11 has intrinsic parameters, in particular: - a focal length, - distortions which are due to imperfections in the optical system of the camera 11, - a direction of the optical axis of the camera 11, and - a resolution.

[0047] The intrinsic parameters characterize the transformation which associates, for an image point, subsequently called “point”, its three-dimensional coordinates in the frame of reference of the camera 11 with the pixel coordinates in an image acquired by the camera 11. These parameters do not change if the camera is moved.

[0048] The distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions which generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are not taken into account, their correction is processed at the time of image acquisition or at the time of calibration.

[0049] The camera 11 is arranged so as to acquire an image of a three-dimensional scene according to a defined point of view, the point of view is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10 as illustrated in [Fig.l].

[0050] The camera 11 for example acquires images of a three-dimensional scene located in front of the vehicle 10, the first camera 11 covering an acquisition field 12. An object 13 is placed in the acquisition field 12 of the camera 11, defining an occlusion field 14 for the vision system, an object present in the occlusion field 14 not being observable by the camera 11 from its current observation position.

[0051] It is obvious that it is possible to use such a vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently, the invention not being limited to the observation of a three-dimensional scene taking place in front of the vehicle 10 carrying the camera 11.

[0052] According to a particular embodiment, the camera 11 is of the “wide-angle” type, a wide-angle camera being for example equipped with a lens designed to acquire an image representative of a three-dimensional scene perceived according to a wider field of vision than that of a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens makes it possible to capture a larger portion of the three-dimensional scene taking place in front of or around the camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by this camera. The angle a of the field of vision of the camera 11 is for example equal to 120°, 145°, 180° or 360°, whereas a standard camera offers, for example, an open field of vision following an angle of 45° or less.Such a camera 11 corresponds, for example, to a camera equipped with mirrors or even to a "fisheye" camera (in French "fish eye"). Wide-angle lenses have a shorter focal length compared to standard lenses, which makes them suitable for acquiring images of landscapes, architecture, road intersections or any other subject requiring an extended perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with an extended depth of field.

[0053] An image acquired by the camera 11 at an acquisition time instant is presented in the form of data representing pixels characterized by: - coordinates in the image; and - data relating to the colors and brightness of objects in the observed scene in the form, for example, of RGB colorimetric values ​​(from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness).

[0054] Each pixel of the acquired image is representative of an object in the three-dimensional scene present in the field of vision of the camera 11. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a luminous point resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When the light strikes this object, photons are emitted or reflected, captured by a photosensitive sensor of the camera 11 after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image faithfully representing the physical object observed by the camera 11. An image point previously presented is thus a point on a surface of an object in the three-dimensional scene.

[0055] When the vehicle 10 is in motion, then two images acquired by the camera 11 at two distinct time instants represent views of the same three-dimensional scene taken from different viewpoints or observation positions, the observation positions of the camera 11 being distinct. On this three-dimensional scene are for example: - buildings; - road infrastructure; - other users or stationary objects, for example a parked vehicle; and / or - other users or moving objects, for example another vehicle, a cyclist or a moving pedestrian.

[0056] According to a particular embodiment, an image acquired by the camera 11 comprises a distortion equal to 0.5%, 0.8% or greater than 1%. The measurement of such a distortion corresponds to the determination of a ratio between: - the maximum spacing of a pixel of the image from a straight line of the first three-dimensional scene whose image is a line touching the longest edge of the first image, either at the center of the edge of the image, or at the corners of the edge of the image, and - the length of this edge.

[0057] Commonly, distortion is considered, in the world of photography, as: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • bothersome if it is greater than or equal to 1% or more.

[0058] A barrel distortion is characterized by a positive percentage, while a crescent distortion is characterized by a negative percentage.

[0059] Each image acquired by the camera 11 is for example sent to a processor, for example a computer of a device equipping the vehicle 10, or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10. It is then used during the implementation of a method for determining a depth of one of its pixels and / or during the implementation of a method for learning a depth prediction model associated with this camera of the vision system.

[0060] [Fig. 2] illustrates a flowchart of the different steps of a method 2 for determining a depth of a pixel of an image by depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, for example in the vehicle 10 of [Fig. 1], according to a particular and non-limiting exemplary embodiment of the present invention. The method 2 is for example implemented by a device of the vision system embedded in the vehicle 10 or by the device 4 of [Fig. 4].

[0061] In a step 21, at least one image acquired by the camera 11 is received.

[0062] In a step 22, depths associated with a set of pixels of an image received are determined by the depth prediction model from the at least one received image.

[0063] Each determined depth then corresponds to a distance separating the vehicle 10 or a part of the vehicle 10 from an object of the three-dimensional scene with which a pixel is associated, the determination of a depth of a pixel then corresponding to a measurement of a distance separating an object from the vehicle carrying the vision system.

[0064] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another user present on the road, the ADAS is then able to determine this distance precisely. For example, if the ADAS has the function of acting on a braking system of the vehicle 10 in the event of a risk of collision with another road user and the distance separating the vehicle 10 from this same road user decreases significantly, then the ADAS is able to detect this sudden approach and act on the braking system of the vehicle 10 to avoid a possible accident.

[0065] [Fig.3] illustrates a flowchart of the different stages of a method for learning the depth prediction model used in a decoding method. termination of a depth of a pixel of an image, for example in method 2 of [Fig.2], according to a particular and non-limiting exemplary embodiment of the present invention.

[0066] The learning method 3 is for example implemented by the device on board the vehicle 10 implementing the method for determining a depth for the vision system on board a vehicle 10 or by the device 4 of [Fig.4].

[0067] In a step 31, data representative of a first image and a second image are received, the first image being acquired by the camera 11 at a first acquisition time instant and the second image being acquired by the camera 11 at a second acquisition time instant distinct from the first acquisition time instant.

[0068] According to a particular exemplary embodiment, the first acquisition time instant is prior to the second acquisition time instant.

[0069] According to another particular exemplary embodiment, the second acquisition time instant is prior to the first acquisition time instant.

[0070] According to a particular exemplary embodiment, the first image and second image have the same definition, that is to say they comprise the same number of pixels, have the same number of pixels according to their height and the same number of pixels according to their width.

[0071] According to another particular exemplary embodiment, the first image and second image are not of the same definition. An additional step then consists of resizing or cropping them to obtain a first image and a second image of the same definition.

[0072] In a step 32, the first image is partitioned into a set of first blocks and the second image is partitioned into a set of second blocks.

[0073] According to a particular exemplary embodiment, the first image and the second image each comprise the same number of blocks horizontally and vertically. The number of first blocks or second blocks is then the square of a number, for example 4, 9 or 16 blocks.

[0074] The aim being to obtain blocks not comprising pixels corresponding to a dynamic object, that is to say to a moving object in the three-dimensional scene such as a pedestrian, a bicycle or a car, the same number of blocks is for example greater than or equal to three. Indeed, by having so many blocks, it is unlikely that a dynamic object covers a field of vision of the camera 11 so as to have pixels corresponding to this dynamic object in all the blocks. Indeed, if the dynamic object is of acceptable size, pixels corresponding to it are present at most in two or four blocks depending on whether these pixels are located at the junction of one or multiple blocks in an image.

[0075] [Fig. 5] schematically illustrates a first pair of images acquired by a camera, for example by the camera 11 on board the vehicle 10 of [Fig. 1], according to a particular and non-limiting exemplary embodiment of the present invention.

[0076] An image 5 corresponding to the first image is partitioned into a set of nine blocks 51, 52, 53, 54, 55, 56, 57, 58 and 59 and comprises the same number of blocks horizontally and vertically, this same number of blocks being equal to three.

[0077] Similarly, an image 6 corresponding to the second image is partitioned into a set of nine blocks 61, 62, 63, 64, 65, 66, 67, 68 and 69 and comprises the same number of blocks horizontally and vertically, this same number of blocks also being equal to three.

[0078] A set of pixels 500 of the image 5 corresponds to a dynamic object present in the three-dimensional scene observed at the first time instant of acquisition. Still according to this example, pixels of this set of pixels 500 are located simultaneously in a block 52, in a block 54, in a block 55 and in a block 56, the set of pixels 500 being located at the junction of these four blocks 52, 53, 55 and 56.

[0079] This same dynamic object is also in the three-dimensional scene observed at the second time instant of acquisition. Thus, a set of pixels 600 of the image 6 corresponds to this same dynamic object, pixels of this set of pixels 600 are located at the same time in a block 62, in a block 64, in a block 65 and in a block 66, the set of pixels 600 being located at the junction of these four blocks 62, 63, 65 and 66.

[0080] The partition of the first and second images makes it possible to obtain first and second blocks of the same dimensions and each second block is associated with a first block to form a first pair of blocks, the first and second blocks of a first pair of blocks having the same position in the first image and the second image respectively.

[0081] According to the particular embodiment illustrated in [Fig.5], nine first pairs of blocks are constituted and comprise respectively: • the first block 51 and the second block 61, noted {51;61}, • the first block 54 and the second block 64, noted {54;64], • the first block 57 and the second block 67, noted {57;67], • the first block 58 and the second block 68, noted {58;68], and • the first block 59 and the second block 69, noted {59;69], these first five pairs of blocks not including pixels corresponding to a dynamic object in the three-dimensional scene, and • the first block 52 and the second block 62, noted {52;62], • the first block 53 and the second block 63, noted {53;63], • the first block 55 and the second block 65, noted {55;65], and • the first block 56 and the second block 66, noted {56;66], these first four pairs of blocks comprising pixels corresponding to a dynamic object in the three-dimensional scene.

[0082] For each first pair of blocks comprising the first block and the second block, in a step 33, depths associated with the pixels included in the first block and in the second block are predicted by the depth prediction model from the first and second blocks.

[0083] Determining a depth of a pixel of an image acquired by a moving camera from a depth prediction model implemented by a convolutional neural network is known to those skilled in the art, for example with monodepth2® described in the document “Digging Into Self-Supervised Monocular Depth Estimation” written by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow and published in August 2019 or with an algorithm called NRS described in the document “Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion” written by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich and Adrien Gaidon, published in August 2020.

[0084] It should be noted that the depth prediction by the depth prediction model associated with the monocular vision system, i.e. comprising only the camera 11, is precise for a pixel corresponding to a static object, but is much less precise for a pixel corresponding to a dynamic object.

[0085] Still for each first pair of blocks comprising the first block and the second block, in a step 36, a second pair of blocks corresponding to each first pair of blocks is generated and comprises a third and a fourth block. The third block is generated from the first block, predicted depths for the pixels of the first block and extrinsic parameters corresponding to a movement of the camera 11 between the first and second time instants. The fourth block is generated from the second block, predicted depths for the pixels of the second block and extrinsic parameters.

[0086] The generation of a third block from a first block acquired by the camera 11 consists of reprojecting a pixel of the first block acquired at the first time instant into the three-dimensional scene in the form of a point, then projecting this point into the image plane of the camera 11 at its position at the second time instant of acquisition, so as to obtain an image corresponding to a view of the three-dimensional scene from the point of view of the camera 11 at the second time instant of acquisition. The image plane of the camera 11 corresponds to a plane defined in the frame of reference of the camera 11, normal to the optical axis of the camera 11 and located at the focal length of the camera 11. Thus, the third block generated from the first block is comparable to the second block. Similarly, the fourth block generated from of the second block is comparable to the first block.

[0087] Since the positions of the camera 11 at the first and second acquisition time instants are not the same when the vehicle 10 is moving and since objects may mask other objects in the scene or even move between the two acquisition time instants, the generated blocks are not identical to the blocks obtained from an image acquired by the camera 11. In addition, the prediction of the depths, just like the models used to generate the blocks, are not error-free. Thus, comparing a generated block to a block obtained from an acquired image makes it possible to evaluate the relevance of the different models used.

[0088] According to a particular exemplary embodiment, the third and fourth blocks are respectively generated by the following function:

[0089] [Math.l] D(pt) ) ] )

[0090] With: • Ps the coordinates of a pixel of the third block, respectively of the fourth block, •77 a function to go from homogeneous coordinates to pixel coordinates by removing a dimension of a vector, • K a direction prediction model associated with camera 11, • T a matrix comprising the extrinsic parameters, and • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates in the first image, respectively second image, and the depth ) associated with it.

[0091] According to a first variant, the direction prediction model corresponds to an intrinsic matrix of the camera 11, or even to an intrinsic matrix of the camera 11 adapted to the processed block. For example, the intrinsic matrix of the camera 11 is defined by: [Math.2] '4 cx K = fy cy 1.

[0092] With: • K the intrinsic matrix of camera 11, • / a first component of the focal length of the camera 11, • fy a second component of the focal length of the camera 11, • a first component of the coordinates of the optical center of the camera 11, and • c>' a second component of the coordinates of the optical center of the camera 11.

[0093] If the intrinsic matrix of the camera 11 is adapted to the different blocks, for example in the particular embodiment illustrated in [Fig.5], then the coordinates of the optical center of the camera 11 are modified as follows for the calculations relating to the pixels of each of the blocks, the coordinates of the pixels included in each of the blocks then having their coordinates expressed in a reference frame relating to the block comprising them and having as origin the upper left corner of the block comprising them: [Math.3] [H = ~ 2 2 -w' c1 - -— 2 cj = 2 / 4 - 2£ — 2 &x = VP 2 - w' t A 2 -h cl- 12 cl = M' 2 - h' ( 4 = hh ~ 2

[0094] With: • 4 the abscissa of the optical center for the first block 51 and the second block 61, • cj. the ordinate of the optical center for the first block 51 and the second block 61, • 4 the abscissa of the optical center for the first block 52 and the second block 62, • 4 the ordinate of the optical center for the first block 52 and the second block 62, • 4 the abscissa of the optical center for the first block 53 and the second block 63, • 4 the ordinate of the optical center for the first block 53 and the second block 63, • 4 the abscissa of the optical center for the first block 54 and the second block 64, • cj the ordinate of the optical center for the first block 54 and the second block 64, • 4 the abscissa of the optical center for the first block 55 and the second block 65, • 4 the ordinate of the optical center for the first block 55 and the second block 65, • 4 the abscissa of the optical center for the first block 56 and the second block 66, • 4 the ordinate of the optical center for the first block 56 and the second block 66, • 4 the abscissa of the optical center for the first block 57 and the second block 67, • 4 the ordinate of the optical center for the first block 57 and the second block 67, • 4 the abscissa of the optical center for the first block 58 and the second block 68, • 4 the ordinate of the optical center for the first block 58 and the second block 68, • 4 the abscissa of the optical center for the first block 59 and the second block 69, • 4 the ordinate of the optical center for the first block 59 and the second block 69, • w the width of the first and second images, • w' the width of the first and second blocks, • h the height of the first and second images, and • the height of the first and second blocks.

[0095] This first variant is particularly applicable to pinhole and calibrated cameras.

[0096] According to a second variant, the direction prediction model corresponds to a learned model implementing a second neural network. Such a direction prediction model is known to those skilled in the art; it is notably presented in the document “Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion”.

[0097] This second variant is suitable for wide-angle or “fish-eye” type cameras or even for uncalibrated cameras.

[0098] According to this second variant, a step 34 of direction prediction associated with the pixels of the first block and the pixels of the second block is necessary.

[0099] The matrix comprising the extrinsic parameters of the moving camera 11 can be obtained in several ways. For example, from data emitted by a driving assistance system of the vehicle 10, called AD AS system, or from sensors embedded in the vehicle 10. According to another example, the extrinsic parameters corresponding to the movement of the camera 11 between the first and second time instants are determined in a step 35 by a motion prediction model, for example by that presented in the document “Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion”.

[0100] Steps 34 and 35 are therefore optional and their implementation depends on the type of camera and the data available to be implemented.

[0101] Still for each first pair of blocks comprising the first block and the second block, in a step 37, a photometric error associated with each first pair of blocks is determined from a first photometric error determined by comparison of the first and fourth blocks and a second photometric error determined by comparison of the second and third blocks.

[0102] If the depth prediction model is accurate, then a first or second photometric error is small. Conversely, if the depth prediction model lacks accuracy, for example because its training is not successful, then a first and / or second photometric error is not negligible.

[0103] It should also be noted that if a block includes pixels associated with a dynamic object present in the observed three-dimensional scene, then the photometric error associated with the corresponding pixels is significant.

[0104] According to a first particular embodiment, the first and second photometric errors are respectively determined by the following function: [Math.4] hAp)^ ​​(l-«) ■ U(p)A(p)\+a-

[0105] With : • Lp«{p) the first photometric error noted Lpi(p), respectively the second photometric error noted Lpz(p), P being a pixel defined by its coordinates in two dimensions, • I(p) a value of the pixel P in the first block, respectively in the second block, a value of the pixel P in the fourth block, respectively in the third block, • SSIM a function that takes into account a local structure, and • ® a weighting factor depending in particular on the type of environment in which the vehicle 10 operates.

[0106] According to a second particular embodiment, the first and second photometric errors include a determination of a reconstruction error of a generated block determined, in addition, by the following function: [Math.5] p) = wW l )

[0107] With : • Lsnworth( p ) the reconstruction error l, : Çp^ for a pixel p of the third block, respectively the construction error L^( p) for a pixel p of the fourth block, • D(p) is a depth associated with pixel P in the third block, respectively in the fourth block; • W is a parameter matrix; • 0 is the order of a smoothing gradient; • an L1 norm of second-order depth gradients is calculated with W =1, and0 =2 ; • x and y are the third and fourth block dimensions; • is a hyperparameter dependent on the environment in which the vehicle 10 operates; and • l\ p ) is a colorimetric value of pixel P in the third characteristic map, respectively fourth characteristic map.

[0108] This second function is generally used to deal with discontinuity at the edge of objects (in English “edge aware smoothness”).

[0109] The photometric error is thus defined, for example, from the photometric errors and reconstruction errors previously defined.

[0110] According to the first particular embodiment, the photometric error associated with each first pair of blocks is determined by the following function: [Math.6] L^F) = [ min(L / ?1 ( p ), Lp2 ( P ) ) ] [YES] With: • Lp corresponding to the photometric error associated with the first pair of blocks F, • Lp] ( p) the first photometric error determined for a pixel P in the first block of the first pair of blocks, the pixel P being defined by its coordinates in two dimensions, and • Lp2{p ) the second photometric error determined for pixel P in a second block of the first pair of blocks.

[0112] According to the second particular embodiment, the photometric error associated with each first pair of blocks is determined by the following function: [Math.7] £^F) = Lp2(p'))+Ls3(p) + 44(p)]

[0113] With: • Lp corresponding to the photometric error associated with the first pair of blocks F • Lpi(p) the first photometric error determined for a pixel P in the first block of the first pair of blocks, the pixel P being defined by its coordinates in two dimensions, and • Lp2(p) the second photometric error determined for pixel P in a second block of the first pair of blocks.

[0114] In a step 38, at least a first pair of blocks is selected based on the photometric error associated with each first pair of blocks to obtain at least a first selected pair of blocks.

[0115] For example, the first pairs of blocks whose photometric error is lowest are selected. A dynamic object detector is for example integrated into the method and makes it possible to identify a number of dynamic objects present in the observed three-dimensional scene, the number of first pairs of blocks to be selected being defined in relation to the number of dynamic objects present. For example, when two dynamic objects are present, the first three pairs of blocks having the highest photometric errors are eliminated, the first pairs of blocks selected being the first remaining pairs of blocks.

[0116] According to another example, the photometric error of each first pair of blocks is compared to a threshold value, only the first pairs of blocks having a photometric error lower than this threshold value are selected.

[0117] According to the particular embodiment presented in [Fig.5], the first pairs of blocks comprising the blocks {52;62], {53;63], {55;65] and {56;66] have significant photometric errors due to the presence of pixels of the sets of pixels 500, 600 corresponding to a dynamic object in these blocks. These first pairs of blocks are then not retained, only the first pairs of blocks comprising the blocks {51;61}, {54;64}, {57;67}, {58;68} and {59;69} are selected.

[0118] In a step 39, the depth prediction model is learned by minimizing a loss error determined from at least one photometric error associated with at least a first pair of selected blocks.

[0119] According to a particular exemplary embodiment, the loss error is equal to an average of the at least one photometric error.

[0120] According to the particular embodiment illustrated in [Fig.5], the loss error is then determined by the following function: [Math. 8] L = mean(Lp ( {51 ; 61} ) ; Lp ( {54 ; 64} ) ; Lp ( {57 ; 67} ) ; Lp ( {58 ; 68} ) ; Lp ( {59 ; 69} ) )

[0121] With: • L corresponding to the loss error, • Lp( { 51 ; 61} ) the photometric error determined for the first pair of blocks {51 ;61], • Lp( { 54 ; 64} ) the photometric error determined for the first pair of blocks {54 ;64], • Lp( { 57 ; 67} ) the photometric error determined for the first pair of blocks {57 ;67], • Lp( {58 ; 68} ) the photometric error determined for the first pair of blocks {58 ; 68}, and • Lp( {59 ; 69} ) the photometric error determined for the first pair of blocks {59 ; 69}.

[0122] Training the depth prediction model consists of adjusting input parameters of the convolutional neural network in order to minimize the previously calculated loss error.

[0123] An object visible in the field of vision of the camera 11 at the first time instant of acquisition and masked in the field of vision of the camera 11 at the second time instant of acquisition or vice versa does not impact the loss error thanks to the min function, this method is therefore insensitive to occlusions.

[0124] The learning is only performed on blocks not comprising pixels associated with a dynamic object or in which the first or second photometric error associated with pixels corresponding to a dynamic object is negligible. Such a learning method thus makes it possible to learn a depth prediction model from images comprising pixels corresponding to dynamic objects. namics without losing quality compared to learning from images free of pixels corresponding to a dynamic object.

[0125] Thus, the depth prediction model used for the depth prediction of a pixel of an image acquired by the camera 11 is made reliable thanks to this learning method.

[0126] This learning is carried out from data acquired by the on-board vision system and therefore does not require data annotated by another on-board system or storage of a library of learning images. In addition, the learning data is representative of the data received when the system is in operation or in production, in fact the learning data is representative of real environments in which the vehicle carrying the vision system evolves or moves, this learning data is therefore particularly relevant.

[0127] [Fig. 4] schematically illustrates a device 4 configured to learn a depth prediction model by a vision system embedded in a vehicle and / or to predict a depth associated with a pixel of an image acquired by a camera, according to a particular and non-limiting exemplary embodiment of the present invention. The device 4 corresponds for example to a device embedded in the first vehicle 10, for example a computer associated with the stereoscopic vision system.

[0128] The device 4 is for example configured for the implementation of the operations described with regard to figures 1 and 4 and / or steps described with regard to figures 2 and 3. Examples of such a device 4 include, but are not limited to, on-board electronic equipment such as an on-board computer of a vehicle, an electronic calculator such as an ECU (“Electronic Control Unit”), a smartphone, a tablet, a laptop. The elements of the device 4, individually or in combination, can be integrated in a single integrated circuit, in several integrated circuits, and / or in discrete components. The device 4 can be produced in the form of electronic circuits or software (or computer) modules or even a combination of electronic circuits and software modules.

[0129] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the method and / or for executing the instructions of the software(s) embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41 corresponding for example to a volatile and / or non-volatile memory and / or comprises a memory storage device which may comprise volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.

[0130] The computer code of the embedded software(s) comprising the instructions to be loaded and executed by the processor is for example stored in the memory 41.

[0131] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Telematic Control Unit”), for example via a communication bus or through dedicated input / output ports.

[0132] According to a particular and non-limiting exemplary embodiment, the device 4 comprises a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 comprise one or more of the following interfaces: - RF radio frequency interface, for example Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English “Universal Serial Bus” or “Universal Serial Bus” in French); HDMI interface (from the English “High Definition Multimedia Interface” or “High Definition Multimedia Interface” in French); - LIN interface (from the English “Local Interconnect Network”).

[0133] According to another particular and non-limiting exemplary embodiment, the device 4 comprises a communication interface 43 which makes it possible to establish communication with other devices (such as other computers of the on-board system) via a communication channel 430. The communication interface 43 corresponds for example to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds for example to a wired network of the CAN (Controller Area Network) type, CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by the ISO 17458 standard) or Ethernet (standardized by the ISO / IEC 802-3 standard).

[0134] According to a particular and non-limiting exemplary embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch-sensitive or not, one or more speakers 450 and / or other peripherals 460 via the output interfaces 44, 45, 46 respectively. According to a variant, one or other of the external devices is integrated into the device 4.

[0135] Of course, the present invention is not limited to the exemplary embodiments described above but extends to a method for determining the depth of a pixel of an image acquired by a vision system, and / or for measuring a distance separating an object from a vehicle carrying a vision system, the depth and / or the distance being predicted and / or measured via a depth prediction model learned according to the learning method described above, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0136] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-based motor vehicle, comprising the device 4 of [Fig.4].

Claims

Claims

1. Method for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle (10), the vision system comprising a camera (11) arranged to acquire an image of a three-dimensional scene taking place in the vicinity of the vehicle (10), said method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception (31) of a first image and a second image acquired by the camera (11) at two distinct acquisition time instants; - partitioning (32) the first image into a set of first blocks and the second image into a set of second blocks, the first and second blocks having the same dimensions, each second block being associated with a first block to form a first pair of blocks, the first and second blocks of a first pair of blocks having the same position in the first image and the second image respectively; - for each first pair of blocks comprising the first block and the second block: • prediction (33) of depths associated with the pixels included in said first block and in said second block by the depth prediction model from said first and second blocks; • generation (36) of a second pair of blocks corresponding to said each first pair of blocks and comprising a third and a fourth block, the third block being generated from said first block, predicted depths for the pixels of the first block and extrinsic parameters corresponding to a movement of the camera (11) between the first and second time instants, and the fourth block being generated from said second block, predicted depths for the pixels of the second block and extrinsic parameters; • determination (37) of a photometric error associated with said each first pair of blocks from a first photometric error determined by comparison of the first and fourth blocks and a second photometric error determined by comparison of the second and third blocks; - selection (38) of at least a first pair of blocks as a function of said photometric error associated with each first pair of blocks to obtain at least a first selected pair of blocks; - learning (39) of the depth prediction model by minimizing a loss error determined from at least one photometric error associated with said at least a first selected pair of blocks.

2. Method according to claim 1, for which the photometric error associated with each first pair of blocks is determined by the following function: L^F) = ^p[min(Lp।(p), Lpl(p))] With: • Lp corresponding to the photometric error associated with the first pair of blocks F, • Lp] (p) the first photometric error determined for a pixel P in the first block of the first pair of blocks, the pixel P being defined by its coordinates in two dimensions, and • Lp2^p} the second photometric error determined for the pixel P in a second block of the first pair of blocks.

3. Method according to one of claims 1 to 2, for which the loss error is equal to an average of said at least one photometric error.

4. Method according to claim 3, for which the first, respectively the second, photometric error is determined by the following function: Lp^P) = (M' \Kp)-Hp)\+a - (l-^SSIM(l(p),ï(p)Y) With: • Lp*(p} the first photometric error noted Lppp}, respectively the second photometric error noted Lp2(p), ? being a pixel defined by its coordinates in two dimensions, • I(p) a value of the pixel P in the first block, respectively in the second block, • a value of the pixel P in the fourth block, respectively in the third block, • SSIM a function which takes into account a local structure, and • has a weighting factor depending in particular on the type of environment in which the vehicle operates (10).

5. Method according to claim 4, for which the third block, respectively fourth block, is generated by the following function: Ps = T0(pliK ^ D(pt) ) ] ) With: • Ps the coordinates of a pixel of the third block, respectively of the fourth block, • 77 a function for going from homogeneous coordinates to pixel coordinates by removing one dimension of a vector, • K a direction prediction model associated with the camera (11), • T a matrix comprising the extrinsic parameters, and • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates in the first image, respectively second image, and of the depth D^p^ which is associated with it.

6. Method according to one of claims 1 to 5, for which the first image and the second image each comprise the same number of blocks horizontally and vertically.

7. The method of claim 6, wherein said same number of blocks is greater than or equal to three.

8. Computer program comprising instructions for implementing the method according to any one of the preceding claims, when these instructions are executed by a processor.

9. Device (4) configured to learn a depth prediction model by a vision system on board a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured to implement the steps of the method according to any one of claims 1 to 7.

10. Vehicle (10) comprising the device (4) according to claim 9.

Citation Information

Patent Citations

  • Self-supervised training of a depth estimation system

    US20190356905A1

  • Depth estimation using a neural network

    US20210183088A1

  • Self-supervised training of a depth estimation model using depth hints

    US20210218950A1

  • Method for training depth estimation model, training apparatus, and electronic device applying the method

    US20230401737A1