Method and device for learning a depth prediction model insensitive to the presence of dynamic objects in training images.

The method for learning a depth prediction model using block partitioning and photometric error minimization addresses the challenge of dynamic objects in monocular vision systems, enabling accurate training and enhancing ADAS reliability.

FR3160798B1Active Publication Date: 2026-02-13STELLANTIS AUTO SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2024003394
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-04-02
Publication Date
2026-02-13
Estimated Expiration
2044-04-02

AI Technical Summary

Technical Problem

Existing depth prediction models in monocular vision systems for vehicles are unable to accurately determine the position of dynamic objects due to their movement, which disrupts the learning process and degrades the quality of the model's training.

Method used

A method for learning a depth prediction model using a convolutional neural network that partitions images into blocks, predicts depths for each block, generates corresponding blocks based on camera movement, determines photometric errors, and selects blocks with minimal errors to minimize loss, excluding dynamic objects from the training process.

Benefits of technology

The method enables accurate training of a depth prediction model even when dynamic objects are present, improving the reliability of ADAS systems by ensuring precise depth estimation and enhancing road safety.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A method or device for learning a depth prediction model associated with a monocular vision system. The method comprises partitioning (32) a first and second image into a set of first and second blocks to form first pairs of blocks, each comprising a first and a second block. For each first pair of blocks, a second pair of blocks is associated, comprising a third and a fourth block generated (36) from the first and second blocks and from predicted depths (33) for the pixels of the first and second blocks. The model is learned (39) by minimizing a loss error determined from at least one photometric error determined (37), for at least one selected first pair of blocks (38), by comparing the first and fourth blocks and by comparing the second and third blocks. Figure for the abstract: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for learning a depth prediction model insensitive to the presence of dynamic objects in training images. technical field

[0001] The present invention relates to methods and devices for learning a depth prediction model associated with a stereoscopic vision system embedded in a vehicle, for example, in a motor vehicle. The present invention also relates to a method and device for determining depth and / or measuring the distance separating an object from a vehicle equipped with a vision system. Technological background

[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS systems use advanced technologies to assist the driver while driving and thus improve performance. ADAS systems use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.

[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.

[0004] The AD AS systems embedded in a vehicle are powered by data obtained one or more onboard sensors, such as cameras. These cameras make it possible to detect and locate other road users or potential obstacles around a vehicle in order to, for example: - to adapt the vehicle's lighting according to the presence of other road users; - to automatically regulate the vehicle's speed; - to act on the braking system in case of risk of impact with an object.

[0005] The position of another user or an obstacle is, for example, determined by a vision system comprising a model for predicting the depth associated with a pixel or the distance separating the vision system from an object in a three-dimensional scene. Such a model is, for example, learned using images, these images being obtained from a universal database, for example Kitti® or Sceneflow®. Kitti®, for example, provides images of a city center road environment, but such a database does not encompass all the road environments a vehicle might encounter. The training data is therefore unsuitable for training the predictive model for a vehicle operating in other road environments.

[0006] The quality of the training of the depth or distance prediction model is, however, very important. Indeed, the depths or distances predicted by the model represent, for example, the distances to other road users or obstacles present in the road environment of the vehicle equipped with the vision system and ADAS. The proper functioning of the driver assistance devices using this data therefore depends on the quality of the data emitted by the vision system.

[0007] A vision system comprising a single camera, hereafter referred to as a monocular vision system, uses several consecutive images to predict depths associated with pixels in one of the images. This monocular vision system has the drawback of not being able to accurately predict depths associated with pixels in an image when objects in the observed three-dimensional scene are moving within that same scene; these objects are hereafter referred to as dynamic objects. Indeed, depth is obtained by comparing the positions of pixels corresponding to the same object in the three-dimensional scene in different images.However, if the observed object is in motion, then the pixels associated with that object are displaced between two images acquired at different times. This displacement is due both to the movement of the monocular vision system mounted in the vehicle and to the inherent movement of the dynamic object within the three-dimensional scene. A depth prediction model is therefore unable to accurately determine the position of this dynamic object. The presence of a dynamic object during a training phase disrupts the learning of the depth prediction model; that is, the quality of the model's training is degraded when pixels corresponding to dynamic objects are considered during the training phase. Summary of the present invention.

[0008] One object of the present invention is to solve at least one of the problems of the technological background described above.

[0009] Another object of the present invention is to improve the learning phase of a depth prediction model from images acquired by a vision system and representing objects in a three-dimensional scene moving within that same scene, the depth prediction model being implemented in particular by a neural network and associated with a vision system embedded in a vehicle.

[0010] Another object of the present invention is to improve road safety, in particular by improving the reliability of AD AS systems powered by data obtained from a camera of a vision system.

[0011] According to a first aspect, the present invention relates to a method for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, the vision system comprising a camera arranged to acquire an image of a three-dimensional scene taking place around the vehicle, the process being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception of a first image and a second image acquired by the camera at two distinct acquisition times; - partitioning of the first image into a set of first blocks and of the second image into a set of second blocks, the first and second blocks having the same dimensions, each second block being associated with a first block to form a first pair of blocks, the first and second blocks of a first pair of blocks having the same position in respectively the first image and the second image; - for each first pair of blocks comprising the first block and the second block: • prediction of depths associated with pixels included in the first block and in the second block by the depth prediction model from the first and second blocks; • generation of a second pair of blocks corresponding to each first pair of blocks and including a third and a fourth block, the third block being generated from the first block, from predicted depths for the pixels of the first block and from extrinsic parameters corresponding to a camera movement between the first and second time instants, and the fourth block being generated from the second block, predicted depths for the pixels of the second block and extrinsic parameters; • determination of a photometric error associated with each first pair of blocks from a first photometric error determined by comparison of the first and fourth blocks and a second photometric error determined by comparison of the second and third blocks; - selection of at least one first pair of blocks based on the photometric error associated with each first pair of blocks to obtain at least one first pair of blocks selected; - learning the depth prediction model by minimizing a loss error determined from at least one photometric error associated with at least one first pair of selected blocks.

[0012] The method advantageously allows the loss error to be determined only on blocks exhibiting, for example, the smallest photometric error. The presence of a dynamic object in an image acquired by the vision system does not then impact the training of the depth prediction model, as a portion of the image containing pixels corresponding to this dynamic object is excluded when selecting blocks on which to minimize the loss error. It is therefore possible to train a depth prediction model associated with the vision system even when images acquired by this vision system include pixels representing a dynamic object present in the observed three-dimensional scene.

[0013] According to a variant of the method, the photometric error associated with each first pair of blocks is determined by the following function: = W / ?))] With : • Lp corresponding to the photometric error associated with the first pair of blocks F, • Lpi(p) the first photometric error determined for a pixel P in the first block of the first pair of blocks, the pixel P being defined by its two-dimensional coordinates, and • Lp2(p) the second photometric error determined for pixel P in a second block of the first pair of blocks.

[0014] The fact of retaining only the minimum value between the first and second error makes it possible in particular not to impact the learning when an object of the three-dimensional scene is occluded in one of the blocks, that is to say when one block of a first pair of blocks includes pixels corresponding to this object and the other block does not include pixels corresponding to this same object.

[0015] According to another variant of the method, the loss error is equal to an average of at least one photometric error.

[0016] According to a further variant of the method, the first, respectively the second, photometric error is determined by the following function: Lp*M = C1-®) ■ U(p)-l(p)\+a- With : • Lp^( p) the first photometric error denoted Lpi( p), respectively the second photometric error denoted Lp2(p), P being a pixel defined by its two-dimensional coordinates, • I(p) is a value of pixel P in the first block, respectively in the second block, • a value of pixel P in the fourth block, respectively in the third block, • SSIM is a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment in which the vehicle operates.

[0017] Such a function is known to a person skilled in the art and is commonly used to determine a photometric error by comparing two images or two blocks of images.

[0018] According to yet another variant of the process, the third block, respectively fourth block, is generated by the following function: Ps = [T^P^ D(Pt) ) ] ) With : • Ps the coordinates of a pixel in the third block, respectively in the fourth block, • 77 a function to go from homogeneous coordinates to pixel coordinates by removing one dimension from a vector, • K is a camera-associated direction prediction model, • Let T be a matrix comprising the extrinsic parameters, and • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates in the first image, respectively second image, and of the depth [^pj] associated with it.

[0019] According to yet another variant of the method, the first image and the second image each comprise the same number of blocks horizontally and vertically.

[0020] According to a further variant of the method, the same number of blocks is greater than or equal to three.

[0021] With the same number of blocks, a dynamic object cannot by itself cover every block of an image, at least one block of the image then does not include pixels corresponding to a dynamic object.

[0022] According to a second aspect, the present invention relates to a device configured to learn a depth prediction model by a vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.

[0023] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0024] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0025] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0026] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.

[0027] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.

[0028] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.

[0029] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures

[0030] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 5, in which:

[0031] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a a particular and non-limiting example of a realization of the present invention;

[0032] [Fig.2] illustrates a flowchart of the different steps of a process for determining the depth of a pixel of an image by a depth prediction model associated with a vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention;

[0033] [Fig. 3] illustrates a flowchart of the different stages of a method for learning the depth prediction model used in the method of [Fig. 2], according to a a particular and non-limiting example of an embodiment of the present invention; and

[0034] [Fig.4] schematically illustrates a device configured to learn a model depth prediction by a vision system embedded in the vehicle of [Fig. 1], according to a particular and non-limiting embodiment of the present invention; and

[0035] [Fig.5] schematically illustrates a first pair of images acquired by the camera mounted in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements

[0036] A method and device for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle will now be described in what follows with joint reference to Figures 1 to 5. The same elements are identified with the same reference signs throughout the description that follows.

[0037] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0038] For the purposes of this description, image reception means receiving data representative of an image. Similarly, obtaining and generating a block means obtaining and generating data representative of an image block, and determining a depth or an error means determining data representative of a depth or an error. These abbreviations are intended solely to simplify the description; however, since the processes described below are implemented by one or more processors, it is clear that the input and output data of the various steps of a process are computer data.

[0039] According to a particular and non-limiting embodiment of the present invention, the depth prediction model is learned in a learning phase comprising the partitioning of a first image into a set of first blocks and the partitioning of a second image into a set of second blocks to form first pairs of blocks, each pair comprising a first block and a second block associated with the first block. The images are, in particular, two images acquired by a camera of the vehicle's onboard vision system at two distinct time points.

[0040] For each first pair of blocks, a second pair of blocks is associated and comprises a third and a fourth block generated from the first and second blocks and predicted depths for the pixels of the first and second blocks.

[0041] The model is learned by minimizing a loss error determined from at least one photometric error, the photometric error being determined, for at least a first pair of selected blocks, by comparing the first and fourth blocks and by comparing the second and third blocks.

[0042] Fig. 1 schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.

[0043] The vehicle 10 is located in an environment 1 corresponding, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.

[0044] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.

[0045] The vehicle 10 advantageously comprises at least one onboard camera 11, configured to acquire images of a three-dimensional scene unfolding in the environment of the vehicle 10 from a common viewing position. The camera 11 forms a monocular vision system if used alone, as illustrated in [Fig. 1]. However, the present invention is not limited to a monocular vision system comprising a single camera but extends to any vision system comprising at least one camera, for example, 1, 2, 3, or 5 cameras.

[0046] The camera 11 has intrinsic parameters, including: - a focal length, - distortions that are due to imperfections in the optical system of camera 11, - a direction of the optical axis of camera 11, and - a resolution.

[0047] The intrinsic parameters characterize the transformation which associates, for an image point, hereafter called "point", its three-dimensional coordinates in the camera 11 reference frame with the pixel coordinates in an image acquired by the camera 11. These parameters do not change if the camera is moved.

[0048] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in curvature, lens parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are not taken into account, and that their correction is addressed at the time of image acquisition or at the time of calibration.

[0049] The camera 11 is arranged to acquire an image of a three-dimensional scene from a defined viewpoint, the viewpoint is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10 as illustrated in [Fig.1].

[0050] The camera 11, for example, acquires images of a three-dimensional scene located in front of the vehicle 10, the first camera 11 covering an acquisition field 12. An object 13 is placed in the acquisition field 12 of the camera 11, defining an occlusion field 14 for the vision system, an object present in the occlusion field 14 not being observable by the camera 11 from its current observation position.

[0051] It is evident that it is possible to use such a vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently, the invention not being limited to the observation of a three-dimensional scene taking place in front of the vehicle 10 carrying the camera 11.

[0052] According to one particular embodiment, the camera 11 is a wide-angle camera, a wide-angle camera being, for example, equipped with a lens designed to acquire a representative image of a three-dimensional scene seen over a wider field of view than a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens makes it possible to capture a larger portion of the three-dimensional scene unfolding in front of or around the camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by this camera. The angle α of the field of view of the camera 11 is, for example, equal to 120°, 145°, 180°, or 360°, whereas a standard camera offers, for example, a field of view open at an angle of 45° or less.Such a camera 11 corresponds, for example, to a camera equipped with mirrors or a "fisheye" camera. Wide-angle lenses have a shorter focal length compared to standard lenses, making them suitable for capturing images of landscapes, architecture, road intersections, or any other subject requiring a wide perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with an extended depth of field.

[0053] An image acquired by the camera 11 at a given acquisition time is in the form of data representing pixels characterized by: - coordinates in the image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric values ​​(from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).

[0054] Each pixel of the acquired image represents an object in the three-dimensional scene present in the field of view of the camera 11. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a point of light resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, which are captured by a photosensitive sensor of the camera 11 after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image that faithfully represents the physical object observed by the camera 11. An image point described above is therefore a point on the surface of an object in the three-dimensional scene.

[0055] When the vehicle 10 is in motion, then two images acquired by the camera 11 at two distinct time instants represent views of the same three-dimensional scene taken from different viewpoints or observation positions, the observation positions of the camera 11 being distinct. For example, this three-dimensional scene may contain: - buildings; - road infrastructure; - other users or stationary objects, for example a parked vehicle; and / or - other users or moving objects, for example another vehicle, a cyclist or a moving pedestrian.

[0056] According to a particular embodiment, an image acquired by the camera 11 includes a distortion equal to 0.5%, 0.8%, or greater than 1%. The measurement of such distortion corresponds to determining a ratio between: - the maximum spacing of a pixel in the image of a straight line in the first three-dimensional scene whose image is a line touching the longest edge of the first image, either at the center of the image edge or at the corners of the image edge, and - the length of this edge.

[0057] In the world of photography, distortion is commonly considered to be: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • bothersome if it is greater than or equal to 1% or more.

[0058] Barrel distortion is characterized by a positive percentage, while crescent distortion is characterized by a negative percentage.

[0059] Each image acquired by the camera 11 is, for example, sent to a processor, for example a computer of a device equipping the vehicle 10, or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10. It is then used during the implementation of a method for determining the depth of one of its pixels and / or during the implementation of a method for learning a depth prediction model associated with this camera of the vision system.

[0060] Figure 2 illustrates a flowchart of the different steps of a method 2 for determining the depth of a pixel in an image by means of a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, for example in vehicle 10 of Figure 1, according to a particular and non-limiting embodiment of the present invention. Method 2 is implemented, for example, by a device of the vision system embedded in vehicle 10 or by device 4 of Figure 4.

[0061] In a step 21, at least one image acquired by the camera 11 is received.

[0062] In a step 22, depths associated with a set of pixels of an image received are determined by the depth prediction model from at least one received image.

[0063] Each determined depth then corresponds to a distance separating the vehicle 10 or a part of the vehicle 10 from an object in the three-dimensional scene to which a pixel is associated, the determination of a depth of a pixel then corresponding to a measurement of a distance separating an object from the vehicle carrying the vision system.

[0064] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another road user, the ADAS is then able to determine this distance precisely. For example, if the ADAS's function is to activate the braking system of the vehicle 10 in the event of a risk of collision with another road user, and the distance between the vehicle 10 and that same road user decreases sharply, then the ADAS is able to detect this sudden approach and activate the braking system of the vehicle 10 to avoid a possible accident.

[0065] Figure 3 illustrates a flowchart of the different stages of a method for learning the depth prediction model used in a decoding process. termination of a depth of one pixel of an image, for example in method 2 of [Fig.2], according to a particular and non-limiting embodiment of the present invention.

[0066] The learning process 3 is for example implemented by the device on board the vehicle 10 implementing the depth determination process for the vision system on board a vehicle 10 or by the device 4 of the [Fig.4].

[0067] In a step 31, representative data of a first image and a second image are received, the first image being acquired by the camera 11 at a first time instant of acquisition and the second image being acquired by the camera 11 at a second time instant of acquisition distinct from the first time instant of acquisition.

[0068] According to a particular embodiment, the first time instant of acquisition is prior to the second time instant of acquisition.

[0069] According to another particular embodiment, the second time instant of acquisition is prior to the first time instant of acquisition.

[0070] According to a particular embodiment, the first image and second image are of the same definition, that is to say they have the same number of pixels, have the same number of pixels according to their height and the same number of pixels according to their width.

[0071] According to another particular embodiment, the first and second images are not of the same resolution. An additional step then consists of resizing or cropping them to obtain a first and second image of the same resolution.

[0072] In a step 32, the first image is partitioned into a set of first blocks and the second image is partitioned into a set of second blocks.

[0073] According to a particular embodiment, the first image and the second image each comprise the same number of blocks horizontally and vertically. The number of first blocks or second blocks is then the square of a number, for example 4, 9 or 16 blocks.

[0074] The aim being to obtain blocks that do not contain pixels corresponding to a dynamic object, that is, a moving object in the three-dimensional scene such as a pedestrian, a bicycle, or a car, the same number of blocks is, for example, greater than or equal to three. Indeed, with so many blocks, it is unlikely that a dynamic object will cover a field of view of the camera 11 in such a way as to have pixels corresponding to this dynamic object in all the blocks. In fact, if the dynamic object is of acceptable size, pixels corresponding to it are present at most in two or four blocks, depending on whether these pixels are located at the junction of one or several blocks in an image.

[0075] [Fig.5] schematically illustrates a first pair of images acquired by a camera, for example by the camera 11 mounted in the vehicle 10 of [Fig.1], according to a particular and non-limiting embodiment of the present invention.

[0076] An image 5 corresponding to the first image is partitioned into a set of nine blocks 51, 52, 53, 54, 55, 56, 57, 58 and 59 and comprises the same number of blocks horizontally and vertically, this same number of blocks being equal to three.

[0077] Similarly, an image 6 corresponding to the second image is partitioned into a set of nine blocks 61, 62, 63, 64, 65, 66, 67, 68 and 69 and comprises the same number of blocks horizontally and vertically, this same number of blocks also being equal to three.

[0078] A set of 500 pixels in the image 5 corresponds to a dynamic object present in the three-dimensional scene observed at the first time instant of acquisition. According to this example, pixels of this set of 500 pixels are located simultaneously in a block 52, in a block 54, in a block 55 and in a block 56, the set of 500 pixels being located at the junction of these four blocks 52, 53, 55 and 56.

[0079] This same dynamic object is also in the three-dimensional scene observed at the second time instant of acquisition. Thus, a set of pixels 600 of the image 6 corresponds to this same dynamic object, pixels of this set of pixels 600 are located simultaneously in a block 62, in a block 64, in a block 65 and in a block 66, the set of pixels 600 being located at the junction of these four blocks 62, 63, 65 and 66.

[0080] The partitioning of the first and second images allows us to obtain first and second blocks of the same dimensions and each second block is associated with a first block to form a first pair of blocks, the first and second blocks of a first pair of blocks having the same position in respectively the first image and the second image.

[0081] Following the particular embodiment illustrated in [Fig. 5], the first nine pairs of blocks are formed and comprise respectively: • the first block 51 and the second block 61, denoted {51 ;61}, • the first block 54 and the second block 64, noted {54 ;64], • the first block 57 and the second block 67, noted {57 ;67], • the first block 58 and the second block 68, noted {58;68], and • the first block 59 and the second block 69, noted {59 ;69], these first five pairs of blocks do not include pixels corresponding to a dynamic object in the three-dimensional scene, and • the first block 52 and the second block 62, noted {52 ;62], • the first block 53 and the second block 63, noted {53 ;63], • the first block 55 and the second block 65, noted {55;65], and • the first block 56 and the second block 66, noted {56;66], these first four pairs of blocks comprising pixels corresponding to a dynamic object in the three-dimensional scene.

[0082] For each first pair of blocks comprising the first block and the second block, in a step 33, depths associated with the pixels included in the first block and in the second block are predicted by the depth prediction model from the first and second blocks.

[0083] The determination of the depth of a pixel of an image acquired by a moving camera from a depth prediction model implemented by a convolutional neural network is known to those skilled in the art, for example with monodepth2® described in the document "Digging Into Self-Supervised Monocular Depth Estimation" written by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow and published in August 2019 or with an algorithm called NRS described in the document "Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion" written by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich and Adrien Gaidon, published in August 2020.

[0084] It should be noted that the depth prediction by the depth prediction model associated with the monocular vision system, i.e. including only the camera 11, is accurate for a pixel corresponding to a static object, but is much less accurate for a pixel corresponding to a dynamic object.

[0085] Again, for each first pair of blocks comprising the first block and the second block, in a step 36, a second pair of blocks corresponding to each first pair of blocks is generated and comprising a third and a fourth block. The third block is generated from the first block, predicted depths for the pixels of the first block, and extrinsic parameters corresponding to a movement of the camera 11 between the first and second time instants. The fourth block is generated from the second block, predicted depths for the pixels of the second block, and the extrinsic parameters.

[0086] Generating a third block from a first block acquired by camera 11 consists of reprojecting a pixel from the first block acquired at the first time instant into the three-dimensional scene as a point, and then projecting this point into the image plane of camera 11 at its position at the second time instant of acquisition, so as to obtain an image corresponding to a view of the three-dimensional scene from the viewpoint of camera 11 at the second time instant of acquisition. The image plane of camera 11 corresponds to a plane defined in the camera 11's frame of reference, normal to the optical axis of camera 11 and located at the focal length of camera 11. Thus, the third block generated from the first block is comparable to the second block. Similarly, the fourth block generated from the second block is comparable to the first block.

[0087] Since the positions of the camera 11 at the first and second acquisition times are not identical when the vehicle 10 is moving, and since objects may obscure other objects in the scene or even move between the two acquisition times, the generated blocks are not identical to the blocks obtained from an image acquired by the camera 11. Furthermore, the depth prediction, as well as the models used to generate the blocks, are not error-free. Thus, comparing a generated block to a block obtained from an acquired image allows for an evaluation of the accuracy of the different models used.

[0088] According to a particular embodiment, the third and fourth blocks are respectively generated by the following function:

[0089] [Math.l] D(pt) ) ] )

[0090] With: • Ps the coordinates of a pixel in the third block, respectively in the fourth block, •77 a function to go from homogeneous coordinates to pixel coordinates by removing one dimension of a vector, • K is a direction prediction model associated with camera 11, • Let T be a matrix comprising the extrinsic parameters, and • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates in the first image, respectively second image, and the depth ) associated with it.

[0091] According to a first variant, the direction prediction model corresponds to an intrinsic matrix of camera 11, or even to an intrinsic matrix of camera 11 adapted to the processed block. For example, the intrinsic matrix of camera 11 is defined by: [Math.2] '4 cx K = fy cy 1.

[0092] With: • K, the intrinsic matrix of camera 11, • / a first component of the focal length of camera 11, • fy a second component of the focal length of camera 11, • a first component of the coordinates of the optical center of camera 11, and • c>' a second component of the coordinates of the optical center of camera 11.

[0093] If the intrinsic matrix of the camera 11 is adapted to the different blocks, by In the particular embodiment illustrated in [Fig. 5], the coordinates of the optical center of camera 11 are modified as follows for calculations relating to the pixels of each block, the coordinates of the pixels included in each block then having their coordinates expressed in a reference frame relative to the block containing them and having as its origin the upper left corner of the block containing them: [Math.3] [H = ~ 2 2 -w' c1 - -— 2 cj = 2 / 4 - 2£ — 2 &x = VP 2 - w' t A 2 -h cl- 12 cl = M' 2 - h' ( 4 = hh ~ 2

[0094] With: • 4 the abscissa of the optical center for the first block 51 and the second block 61, • cj. the ordinate of the optical center for the first block 51 and the second block 61, • 4 the abscissa of the optical center for the first block 52 and the second block 62, • 4 the ordinate of the optical center for the first block 52 and the second block 62, • 4 the abscissa of the optical center for the first block 53 and the second block 63, • 4 the ordinate of the optical center for the first block 53 and the second block 63, • 4 the abscissa of the optical center for the first block 54 and the second block 64, • cj the ordinate of the optical center for the first block 54 and the second block 64, • 4 the abscissa of the optical center for the first block 55 and the second block 65, • 4 the ordinate of the optical center for the first block 55 and the second block 65, • 4 the abscissa of the optical center for the first block 56 and the second block 66, • 4 the ordinate of the optical center for the first block 56 and the second block 66, • 4 the abscissa of the optical center for the first block 57 and the second block 67, • 4 the ordinate of the optical center for the first block 57 and the second block 67, • 4 the abscissa of the optical center for the first block 58 and the second block 68, • 4 the ordinate of the optical center for the first block 58 and the second block 68, • 4 the abscissa of the optical center for the first block 59 and the second block 69, • 4 the ordinate of the optical center for the first block 59 and the second block 69, • w the width of the first and second images, • w' the width of the first and second blocks, • h the height of the first and second images, and • the height of the first and second blocks.

[0095] This first variant is particularly applicable to pinhole and calibrated cameras.

[0096] According to a second variant, the direction prediction model corresponds to a The learned model implements a second neural network. Such a direction prediction model is known to those skilled in the art; it is notably presented in the document "Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion".

[0097] This second variant is suitable for wide-angle or "fish-eye" type cameras or even for uncalibrated cameras.

[0098] According to this second variant, a direction prediction step 34 associated with the pixels of the first block and the pixels of the second block is necessary.

[0099] The matrix comprising the extrinsic parameters of the moving camera 11 can be obtained in several ways. For example, from data emitted by a vehicle driver assistance system 10, called the AD AS system, or from sensors on board the vehicle 10. According to another example, the extrinsic parameters corresponding to the movement of the camera 11 between the first and second time instants are determined in a step 35 by a motion prediction model, for example by the one presented in the document "Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion".

[0100] Steps 34 and 35 are therefore optional and their implementation depends on the type of camera and the data available to be implemented.

[0101] Again for each first pair of blocks comprising the first block and the second block, in a step 37, a photometric error associated with each first pair of blocks is determined from a first photometric error determined by comparison of the first and fourth blocks and a second photometric error determined by comparison of the second and third blocks.

[0102] If the depth prediction model is accurate, then a first or second photometric error is small. Conversely, if the depth prediction model lacks accuracy, for example because its training is incomplete, then a first and / or second photometric error is not negligible.

[0103] It should also be noted that if a block includes pixels associated with a dynamic object present in the observed three-dimensional scene, then the photometric error associated with the corresponding pixels is significant.

[0104] According to a first particular embodiment, the first and second photometric errors are respectively determined by the following function: [Math.4] hAp)^ ​​(l-«) ■ U(p)A(p)\+a-

[0105] With : • Lp«{p) the first photometric error denoted Lpi( p ), respectively the second photometric error denoted Lpz(p), P being a pixel defined by its two-dimensional coordinates, • I(p) is a value of pixel P in the first block, respectively in the second block, a value of pixel P in the fourth block, respectively in the third block, • SSIM is a function that takes into account a local structure, and • ® a weighting factor depending in particular on the type of environment in which the vehicle operates 10.

[0106] According to a second particular embodiment, the first and second Photometric errors include determining a reconstruction error of a generated block determined, moreover, by the following function: [Math.5] p) = wW l )

[0107] With : • Lsnwoth(p) is the reconstruction error l, : Çp^ for a pixel p of the third block, respectively the construction error L^(p) for a pixel p of the fourth block, • D(p) is a depth associated with pixel P in the third block, respectively in the fourth block; • W is a parameter matrix; • 0 is the order of a smoothing gradient; • an L1 norm of the second-order depth gradients is calculated with W = 1, et0 =2 ; • x and y are the dimensions of the third and fourth blocks; • is a hyperparameter dependent on the environment in which the vehicle 10 is operating; and • l\ p ) is a colorimetric value of pixel P in the third feature map, respectively fourth feature map.

[0108] This second function is generally used to deal with the discontinuity at the edge of objects (in English "edge aware smoothness").

[0109] The photometric error is thus defined, for example, from the photometric errors and reconstruction errors previously defined.

[0110] According to the first particular embodiment, the photometric error associated with each first pair of blocks is determined by the following function: [Math.6] L^F) = [ min(L / ?1 ( p ), Lp2 ( P ) ) ] [YES] With: • Lp corresponding to the photometric error associated with the first pair of F blocks, • Lp] ( p) the first photometric error determined for a pixel P in the first block of the first pair of blocks, the pixel P being defined by its two-dimensional coordinates, and • Lp2{p ) the second photometric error determined for pixel P in a second block of the first pair of blocks.

[0112] According to the second particular embodiment, the photometric error associated with each first pair of blocks is determined by the following function: [Math.7] £^F) = Lp2(p'))+Ls3(p) + 44(p)]

[0113] With: • Lp corresponding to the photometric error associated with the first pair of F blocks • Lpi(p) the first photometric error determined for a pixel P in the first block of the first pair of blocks, the pixel P being defined by its two-dimensional coordinates, and • Lp2(p) the second photometric error determined for pixel P in a second block of the first pair of blocks.

[0114] In a step 38, at least a first pair of blocks is selected according to the photometric error associated with each first pair of blocks to obtain at least a first pair of selected blocks.

[0115] For example, the first pairs of blocks with the lowest photometric error are selected. A dynamic object detector is integrated into the method, for instance, to identify a number of dynamic objects present in the observed three-dimensional scene. The number of first pairs of blocks to be selected is defined relative to the number of dynamic objects present. For example, when two dynamic objects are present, the first three pairs of blocks with the highest photometric errors are eliminated, and the first pairs of blocks selected are the first remaining pairs.

[0116] According to another example, the photometric error of each first pair of blocks is compared to a threshold value, only the first pairs of blocks having a photometric error less than this threshold value are selected.

[0117] According to the particular embodiment shown in [Fig. 5], the first pairs of blocks comprising blocks {52;62], {53;63], {55;65], and {56;66] exhibit significant photometric errors due to the presence of pixels from the Sets of pixels 500, 600 corresponding to a dynamic object in these blocks. These first pairs of blocks are then not retained, only the first pairs of blocks including blocks {51;61}, {54;64}, {57;67}, {58;68} and {59;69} are selected.

[0118] In a step 39, the depth prediction model is learned by minimizing a loss error determined from at least one photometric error associated with at least one first pair of selected blocks.

[0119] According to a particular embodiment, the loss error is equal to an average of at least one photometric error.

[0120] According to the particular embodiment illustrated in [Fig.5], the loss error is then determined by the following function: [Math. 8] L = average(Lp({51; 61}); Lp({54; 64}); Lp({57; 67}); Lp({58; 68}); Lp({59; 69}))

[0121] With: • L corresponding to the loss error, • Lp( { 51 ; 61} ) the photometric error determined for the first pair of blocks {51 ;61], • Lp( { 54 ; 64} ) the photometric error determined for the first pair of blocks {54 ;64], • Lp( { 57 ; 67} ) the photometric error determined for the first pair of blocks {57 ;67], • Lp( {58 ; 68} ) the photometric error determined for the first pair of blocks {58 ;68}, and • Lp( {59 ; 69} ) the photometric error determined for the first pair of blocks {59 ;69}.

[0122] The training of the depth prediction model consists of adjusting the input parameters of the convolutional neural network in order to minimize the previously calculated loss error.

[0123] An object visible in the field of view of the camera 11 at the first time instant of acquisition and masked in the field of view of the camera 11 at the second time instant of acquisition or vice versa does not impact the loss error thanks to the min function, this process is therefore insensitive to occlusions.

[0124] The training is performed only on blocks that do not contain pixels associated with a dynamic object or in which the first or second photometric error associated with pixels corresponding to a dynamic object is negligible. Such a learning process thus makes it possible to learn a depth prediction model from images containing pixels corresponding to dynamic objects. namics without losing quality compared to learning from pixel-free images corresponding to a dynamic object.

[0125] Thus, the depth prediction model used for the depth prediction of a pixel of an image acquired by the camera 11 is made more reliable thanks to this learning process.

[0126] This learning process is performed using data acquired by the on-board vision system and therefore does not require data annotated by another on-board system or the storage of a training image library. Furthermore, the training data is representative of the data received when the system is in operation or in production; indeed, the training data is representative of real-world environments in which the vehicle carrying the vision system operates or moves, making this training data particularly relevant.

[0127] Figure 4 schematically illustrates a device 4 configured to learn a depth prediction model from a vehicle-mounted vision system and / or to predict the depth associated with a pixel in an image acquired by a camera, according to a particular, non-limiting embodiment of the present invention. The device 4 corresponds, for example, to a device mounted in the first vehicle 10, for example, a computer associated with the stereoscopic vision system.

[0128] Device 4 is, for example, configured to carry out the operations described opposite Figures 1 and 4 and / or the steps described opposite Figures 2 and 3. Examples of such a device 4 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 4, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 4 may be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.

[0129] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41, corresponding, for example, to volatile and / or non-volatile memory, and / or includes a memory storage device which may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disc.

[0130] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 41.

[0131] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.

[0132] According to a particular and non-limiting embodiment, the device 4 includes a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 include one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HDMI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").

[0133] According to another particular and non-limiting embodiment, the device 4 includes a communication interface 43 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 430. The communication interface 43 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds, for example, to a wired network of the CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3) type.

[0134] According to a particular and non-limiting embodiment, device 4 can provide output signals to one or more external devices, such as a display screen 440, touch or not, one or more speakers 450 and / or other peripherals 460 via output interfaces 44, 45, 46 respectively. In one variant, one or more of the external devices is integrated into device 4.

[0135] Of course, the present invention is not limited to the embodiments described above but extends to a method for determining the depth of a pixel in an image acquired by a vision system, and / or for measuring the distance between an object and a vehicle equipped with a vision system, the depth and / or distance being predicted and / or measured via a depth prediction model learned according to the learning method described above, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0136] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 4 of [Fig.4].

Claims

Demands

1. A method for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle (10), the vision system comprising a camera (11) arranged to acquire an image of a three-dimensional scene taking place around the vehicle (10), said method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception (31) of a first image and a second image acquired by the camera (11) at two distinct time moments of acquisition; - partitioning (32) of the first image into a set of first blocks and of the second image into a set of second blocks, the first and second blocks having the same dimensions, each second block being associated with a first block to form a first pair of blocks, the first and second blocks of a first pair of blocks having the same position in respectively the first image and the second image; - for each first pair of blocks comprising the first block and the second block: • prediction (33) of depths associated with the pixels included in said first block and in said second block by the depth prediction model from said first and second blocks; • generation (36) of a second pair of blocks corresponding to each first pair of blocks and comprising a third and a fourth block, the third block being generated from said first block, with predicted depths for the pixels of the first block and extrinsic parameters corresponding to a camera movement (11) between the first and second time instants, and the fourth block being generated from said second block, with predicted depths for the pixels of the second block and extrinsic parameters; • determination (37) of a photometric error associated with each first pair of blocks from a first photometric error determined by comparison of the first and fourth blocks and a second photometric error determined by comparison of the second and third blocks; - selection (38) of at least one first pair of blocks as a function of said photometric error associated with each first pair of blocks to obtain at least one first pair of blocks selected; - learning (39) of the depth prediction model by minimizing a loss error determined from at least one photometric error associated with said at least one first pair of blocks selected.

2. A method according to claim 1, wherein the photometric error associated with each first pair of blocks is determined by the following function: L^F) = ^p[min(Lp2^p(p), Lpl(p))] With: • Lp corresponding to the photometric error associated with the first pair of blocks F, • Lp] ( p) the first photometric error determined for a pixel P in the first block of the first pair of blocks, the pixel P being defined by its coordinates in two dimensions, and • Lp2^p} the second photometric error determined for the pixel P in a second block of the first pair of blocks.

3. A method according to any one of claims 1 to 2, wherein the loss error is equal to an average of said at least one photometric error.

4. A method according to claim 3, wherein the first, and respectively the second, photometric error is determined by the following function: Lp^P) = (M' \Kp)-Hp)\+a - (l-^SSIM(l(p),ï(p)Y) With: • Lp*(p) the first photometric error, denoted Lppp, and respectively the second photometric error, denoted Lp2(p), where ? is a pixel defined by its two-dimensional coordinates, • I(p) a value of pixel P in the first block, and respectively in the second block, • a value of pixel P in the fourth block, and respectively in the third block, • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment in which the vehicle operates (10).

5. A method according to claim 4, wherein the third block, respectively fourth block, is generated by the following function: Ps = T0(pliK ^ D(pt) ) ] ) With: • Ps the coordinates of a pixel of the third block, respectively of the fourth block, • 77 a function for going from homogeneous coordinates to pixel coordinates by removing one dimension of a vector, • K a direction prediction model associated with the camera (11), • T a matrix comprising the extrinsic parameters, and • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates in the first image, respectively second image, and of the depth D^p^ associated with it.

6. A method according to any one of claims 1 to 5, wherein the first image and the second image each comprise the same number of blocks horizontally and vertically.

7. Method according to claim 6, wherein said same number of blocks is greater than or equal to three.

8. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by a processor.

9. Device (4) configured to learn a depth prediction model by a vision system embedded in a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured to implement the steps of the method according to any one of claims 1 to 7.

10. Vehicle (10) comprising the device (4) according to claim 9.