Method and device for learning a prediction model capable of predicting a depth associated with different objects.

The method improves depth prediction for dynamic objects in monocular vision systems by learning with synthetic and real images, enhancing ADAS system reliability through precise distance measurements.

FR3168049A1Pending Publication Date: 2026-05-01STELLANTIS AUTO SAS
2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
STELLANTIS AUTO SAS
Filing Date
2024-10-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Monocular vision systems in vehicles struggle to accurately predict the depth of dynamic objects due to displacement of pixels caused by both the movement of the vehicle and the objects, leading to degraded learning and unreliable depth prediction models for ADAS systems.

Method used

A method using a convolutional neural network to learn a depth prediction model with synthetic images representing dynamic objects, involving initial learning with computer-generated images and optional refinement with real images, to improve accuracy and reliability.

Benefits of technology

Enhances the accuracy of depth prediction for dynamic objects, thereby improving the reliability of ADAS systems by providing precise distance measurements to road users and obstacles.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A method or device for learning a depth prediction model associated with a monocular vision system. Specifically, the method comprises determining (21) types of training objects and types of training environments, and receiving (22) a synthetic image from a set of synthetic images representing at least one object observed at a defined distance. The observed object is of a type of training object and integrated into a training environment, and target depths are associated with pixels representing the object. The synthetic image and target depths are generated by a synthetic image generation model. Depths associated with pixels of the synthetic image are then predicted (23) by the depth prediction model, which is learned (24) by minimizing a loss error determined by comparing the predicted depths to the target depths. Figure for the abstract: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for learning a prediction model capable of predicting a depth associated with different objects. technical field

[0001] The present invention relates to methods and devices for learning a depth prediction model associated with a vision system embedded in a vehicle, for example, in a motor vehicle. The present invention also relates to a method and device for determining depth and / or measuring the distance separating a static or dynamic object from a vehicle equipped with a vision system. Technological background

[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS systems use advanced technologies to assist the driver while driving and thus improve performance. ADAS systems use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.

[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.

[0004] The ADAS systems installed in a vehicle are powered by data obtained from one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or any obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle's speed; - to act on the braking system in case of risk of impact with an object.

[0005] The position of another user or an obstacle in a three-dimensional scene is, for example, determined by a vision system comprising a prediction model for a depth associated with a pixel or a distance separating the vision system from an object in a three-dimensional scene. However, the accuracy of a predicted depth or distance for a dynamic object, that is, an object in movement in the three-dimensional scene is generally not satisfactory for a vision system comprising only a single camera, also called a monocular vision system.

[0006] Indeed, a monocular vision system uses several consecutive images to predict depths associated with pixels in one of the images. This monocular vision system has the drawback of not being able to accurately predict depths associated with pixels in an image when objects in the observed three-dimensional scene are moving within that same scene, because the depth is obtained by comparing the positions of pixels corresponding to the same object in the three-dimensional scene in different images. However, if the observed object is moving, then the pixels associated with that object are displaced between two images acquired at different times, and this displacement is due both to the movement of the monocular vision system mounted in the vehicle and to the inherent movement of the dynamic object within the three-dimensional scene.A depth prediction model is therefore unable to accurately determine the position of this dynamic object.

[0007] Moreover, the presence of a dynamic object in a learning phase then disrupts the learning of the depth prediction model, that is to say that the quality of the learning of the depth prediction model is degraded when pixels corresponding to dynamic objects are taken into consideration during the learning phase.

[0008] The quality of the training of the depth or distance prediction model is, however, very important. Indeed, the depths or distances predicted by the model represent, for example, the distances to other road users or obstacles present in the road environment of the vehicle equipped with the vision system, whose output data feeds ADAS (Advanced Driver Assistance Systems). The proper functioning of the driver assistance devices using this data therefore depends on the quality of the data emitted by the vision system. Summary of the present invention

[0009] One object of the present invention is to solve at least one of the problems of the technological background described above.

[0010] Another object of the present invention is to improve the learning phase of a depth prediction model from synthetic images representing dynamic objects, the depth prediction model being implemented in particular by a neural network and associated with a vision system embedded in a vehicle.

[0011] Another object of the present invention is to improve road safety, in particular by improving the reliability of AD AS systems powered by data obtained from a camera of a vision system.

[0012] According to a first aspect, the present invention relates to a method for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, the vision system comprising a camera arranged to acquire an image of a three-dimensional scene taking place near the vehicle, the method being implemented by at least one processor, and comprising the following steps: - Determining the types of learning objects and the type of learning environment, - receiving a synthetic image from a set of synthetic images, each synthetic image from the set of synthetic images representing at least one first object observed at a defined distance and corresponding to a first type of learning object among the types of learning objects integrated into an environment among the types of learning environments, target depths being associated with pixels representing at least one first object, each synthetic image and the target depths associated with the pixels of each synthetic image being generated by a synthetic image generation model; - prediction of initial depths associated with pixels of the synthetic image by the depth prediction model; - first learning of the depth prediction model by minimizing a first loss error determined by comparing the first depths to the target depths associated with the pixels of the synthetic image.

[0013] Such a method makes it possible to learn a depth prediction model associated with a vision system comprising a single camera for different types of training objects corresponding to dynamic objects, which are moving within an observed three-dimensional scene. A depth determined by the learned depth prediction model is then more accurate for a dynamic object.

[0014] According to a variant of the method, a distribution of distances defined for objects of each type of learning object among the types of learning objects in the set of synthetic images is uniform over a range of distances, the range of distances corresponding to a measurement range associated with the vision system.

[0015] According to another variant of the method, a distribution of positions of objects of each type of learning object in the set of synthetic images is uniform along the axes of the abscissa and ordinate.

[0016] According to yet another variant of the process, the synthetic image generation model generates images based on intrinsic parameters of the vision system camera.

[0017] According to a further variant of the method, synthetic images from the set of synthetic images represent several objects.

[0018] According to yet another variant, the method further comprises the following steps: - receiving a real image representing at least a second object observed at a defined distance and corresponding to a second type of learning object among the types of learning objects, target depths being associated with pixels representing at least a second object, - prediction of second depths associated with pixels of the real image by the depth prediction model; - second learning of the depth prediction model by minimizing a second loss error determined by comparing the second depths to the target depths associated with the pixels of the real image.

[0019] According to one variant of the method, a type of learning object belongs to a set of learning objects comprising: - a car, - a truck, - a coach, - a motorcycle, and - a bicycle, and the type of learning environment belongs to a set of learning environments comprising: - a city center, - a residential area, - an industrial area, - a parking area, - a highway, and - a rural area.

[0020] According to a second aspect, the present invention relates to a device configured to learn a depth prediction model by a vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.

[0021] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0022] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0023] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0024] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.

[0025] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.

[0026] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.

[0027] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures

[0028] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 3, in which:

[0029] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a a particular and non-limiting example of a realization of the present invention;

[0030] [Fig.2] illustrates a flowchart of the different stages of a method for learning a depth prediction model associated with the vision system of [Fig.1], according to a particular and non-limiting embodiment of the present invention; and

[0031] [Fig.3] schematically illustrates a device configured to learn a depth prediction model associated with the vehicle's onboard vision system of [Fig.1], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements

[0032] A method and device for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle will now be described in what follows with joint reference to Figures 1 to 3. The same elements are identified with the same reference signs throughout the description that follows.

[0033] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0034] Fig. 1 schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.

[0035] The vehicle 10 is located in an environment 1 corresponding, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.

[0036] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.

[0037] The vehicle 10 advantageously comprises at least one onboard camera 11, configured to acquire images of a three-dimensional scene unfolding in the environment of the vehicle 10 from a common viewing position. The camera 11 forms a monocular vision system if used alone, as illustrated in [Fig. 1]. However, the present invention is not limited to a monocular vision system comprising a single camera but extends to any vision system comprising at least one camera, for example, 1, 2, 3, or 5 cameras.

[0038] The camera 11 has intrinsic parameters, including: - a focal length, - distortions which are due to imperfections in the optical system of camera 11, - a direction of the optical axis of camera 11, and - a resolution.

[0039] The intrinsic parameters characterize the transformation which associates, for an image point, hereafter called a "point", its three-dimensional coordinates in the camera 11 reference frame at pixel coordinates in an image acquired by camera 11. These parameters do not change if the camera is moved.

[0040] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in lens curvature, parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are not taken into account, and their correction is addressed during image acquisition or calibration.

[0041] The camera 11 is arranged to acquire an image of a three-dimensional scene from a defined viewpoint, the viewpoint is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10 as illustrated in [Fig.1].

[0042] The camera 11, for example, acquires images of a three-dimensional scene located in front of the vehicle 10, the first camera 11 covering an acquisition field 12. An object 13 is placed in the acquisition field 12 of the camera 11, defining an occlusion field 14 for the vision system, an object present in the occlusion field 14 not being observable by the camera 11 from its current observation position.

[0043] It is evident that it is possible to use such a vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently, the invention not being limited to the observation of a three-dimensional scene taking place in front of the vehicle 10 carrying the camera 11.

[0044] According to one particular embodiment, the camera 11 is of the "wide-angle" type, a wide-angle camera being, for example, equipped with a lens designed to acquire a representative image of a three-dimensional scene seen over a wider field of view than a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens makes it possible to capture a larger portion of the three-dimensional scene unfolding in front of or around the camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by this camera. The angle α of the field of view of the camera 11 is, for example, equal to 120°, 145°, 180°, or 360°, whereas a standard camera offers, for example, a field of view open at an angle of 45° or less. Such a camera 11 corresponds, for example, to a camera equipped with mirrors or a fisheye camera. Wide-angle lenses have a shorter focal length compared to standard lenses, making them suitable for capturing images of landscapes, architecture, road intersections, or any other subject requiring a wide perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with an extended depth of field.

[0045] An image acquired by the camera 11 at a given acquisition time is in the form of data representing pixels characterized by: - ​​coordinates in the image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric values ​​(from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).

[0046] Each pixel of the acquired image represents an object in the three-dimensional scene present in the field of view of the camera 11. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a point of light resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, which are captured by a photosensitive sensor of the camera 11 after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image that faithfully represents the physical object observed by the camera 11. An image point described above is therefore a point on the surface of an object in the three-dimensional scene.

[0047] When the vehicle 10 is in motion, then two images acquired by the camera 11 at two distinct time instants represent views of the same three-dimensional scene taken from different viewpoints or observation positions, the observation positions of the camera 11 being distinct. On this three-dimensional scene there are, for example: - buildings; - road infrastructure; - other users or stationary objects, for example a parked vehicle; and / or - other users or moving objects, for example another vehicle, a cyclist or a moving pedestrian.

[0048] According to a particular embodiment, an image acquired by the camera 11 includes a distortion equal to 0.5%, 0.8%, or greater than 1%. The measurement of such distortion corresponds to determining a ratio between: - the maximum spacing of a pixel in the image of a straight line in the first three-dimensional scene whose image is a line touching the longest edge of the first image, either at the center of the image edge or at the corners of the image edge, and - the length of this edge.

[0049] In the world of photography, distortion is commonly considered to be: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • problematic if it is greater than or equal to 1% or more.

[0050] Barrel distortion is characterized by a positive percentage, while crescent distortion is characterized by a negative percentage.

[0051] Each image acquired by the camera 11 is for example sent to a processor, for example a computer of a device equipping the vehicle 10, or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10. It is then used during the implementation of a method for determining the depth of one of its pixels and / or during the implementation of a method for learning a depth prediction model associated with this camera of the vision system.

[0052] Figure 2 illustrates a flowchart of the different stages of a method for learning a depth prediction model, such a depth prediction model being for example associated with the vision system on board the vehicle 10 and comprising the first camera 11 and also being used in a method for determining the depth of a pixel of an image, according to a particular and non-limiting embodiment of the present invention.

[0053] The learning process 2 is for example implemented by a device on board the vehicle 10 implementing the depth determination process for the vision system on board a vehicle 10 or by the device 3 of [Fig.3].

[0054] Such a learning process 2 comprises a first learning phase described below and optionally a second learning phase also described below.

[0055] The first learning phase is based on the processing of synthetic images, while the optional second learning phase is based on the processing of annotated real images.

[0056] The learning process 2 first includes steps associated with the first phase of learning.

[0057] In a first step 21, types of learning objects and types of learning environments are determined, for example through a human-machine interface by a user or from the analysis of data such as a set of images acquired by one or more vision systems mounted on vehicles traveling in different road environments. Indeed, these images are acquired in real-world situations, and objects present in these images are representative of real objects observable by the vision system mounted on the vehicle 10 and are therefore entirely relevant.

[0058] According to a particular embodiment, a type of learning object belongs to a set of learning objects comprising: - a car, - a truck, - a coach, - a motorcycle, and - a bicycle.

[0059] Such training objects are dynamic objects, meaning they can be in motion within a three-dimensional scene, and depths associated with pixels representing such objects are generally unreliable when predicted by a depth prediction model associated with a monocular vision model, i.e., one comprising only a single camera. It is therefore necessary to improve the accuracy of the depth prediction model for these dynamic objects, which is why this learning method is implemented.

[0060] According to yet another particular embodiment, a type of learning environment belongs to a set of learning environments comprising: - a city centre, - a residential area, - an industrial area, - a parking area, - a highway, and - a rural environment.

[0061] Such environments are road environments commonly traveled by a vehicle and are therefore relevant for learning a depth prediction model associated with the vision system embedded in the vehicle 10.

[0062] In a second step, a synthetic image is received, this synthetic image belonging to a set of synthetic images in which each image of The synthesis represents at least one initial object observed at a defined distance. This initial object corresponds to a first type of training object among the previously determined types of training objects and is integrated or immersed in an environment among the previously determined types of training environments. Thus, some pixels of the synthesis image represent the initial object, while other pixels surrounding it represent the environment.

[0063] It should be noted that a single computer-generated image comprises, according to a particular embodiment, several objects. These objects are then of the same type of learning object or of different types of learning object.

[0064] The computer-generated images in the computer-generated image set are generated by a computer-generated image generation model. During the generation of a computer-generated image, target depths are associated with the pixels representing objects, i.e., the first objects. Such images are generated by a computer-generated image generation model known to those skilled in the art, for example, using the computer-generated image generation model of the aiSim® simulation tool developed by aiMotive®.

[0065] Thus, for each pixel of the synthetic image generated by the synthetic image generation model, a target depth is associated. This target depth is then annotated data, which is reliable and usable for training the depth prediction model. Indeed, the rendering of each object is performed according to the object's orientation and the distance separating it from the observation point associated with the generated synthetic image, the depths being then determined based on this distance. The synthetic image and the target depths are then saved, either in the same file or in separate but linked files, for example in memory accessible to the depth prediction model.Receiving the synthetic image and target depths then consists of receiving data representing the pixels and target depths, via a wireless connection with a server and / or a data bus embedded in a vehicle, for example.

[0066] In order to use a set of synthetic images representative of numerous situations, this set then comprises, for example, several thousand synthetic images. According to this example, a distribution of distances defined for objects of each type of learning object among the types of learning objects determined in the set of synthetic images is uniform over a range of distances, this range of distances corresponding to a measurement range associated with the vision system embedded in the vehicle 10. In other words, according to this example, learning objects of each type of learning object are represented in synthetic images from the set of synthetic images under Several distinct representations or forms correspond to their appearance when observed from different distances. Note that different renderings can be used for the same type of object; thus, an object of the automotive type may appear as a city car in some computer-generated images and as a family minivan in other computer-generated images, in different colors for example.

[0067] In the case where one thousand (1000) synthetic images represent objects of the same type of learning object, these one thousand examples are determined for distances uniformly distributed within the measurement range of the vision system onboard the vehicle 10, for example from 10 (ten) to 200 m (two hundred meters) with an interval of 1 m (one meter). Optionally, it is possible to choose a series of non-homogeneous distance intervals, with the intervals being narrower for shorter distances in order to facilitate the learning of the depth prediction model for short distances.

[0068] Similarly, the distribution of positions of objects of each type of training object in the set of computer graphics is, for example, uniform along the x- and y-axes. Thus, objects of the same type of training object are represented at different positions in computer graphics of the set of computer graphics, and for each type of training object and for each area of ​​computer graphics, there is at least one image in the set of computer graphics representing an object of that type of training object. Such a set of computer graphics is a series of images representing "floating" objects.

[0069] To optimize the size of the database containing all the computer-generated images, it is preferable to have a maximum number of objects in a single computer-generated image to reduce the number of images. For example, an image can contain many motorcycles in different areas of the image at different viewing distances.

[0070] According to a particular embodiment, and in order to make the synthetic images as relevant as possible, the synthetic image generation model generates images based on intrinsic parameters of the camera 11 of the vision system. Thus, the synthetic images represent the objects as they would be perceived by the camera 11 itself. Therefore, if the camera 11 acquires images with a certain distortion, the synthetic images also exhibit this same distortion. It should be noted that in the case where the camera 11 has a pinhole lens, a generic synthetic image generation model may suffice; that is, it is not necessary to use the camera's intrinsic parameters to generate the synthetic images. allowing the same set of synthetic images to be used for several camera models.

[0071]

[0072] In a third step 23, initial depths associated with pixels of the synthetic image are predicted by the depth prediction model. Such a depth prediction model is known to those skilled in the art, and is described, for example, in the following documents: - “Digging Into Self-Supervised Monocular Depth Estimation” by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow published in August 2019, and - “HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation” by Xiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong, Lina Liu, Yong Liu, Xinxin Chen and Yi Yuan published in December 2020.

[0073] In a fourth step 24 called first learning, the depth prediction model is learned by minimizing a first loss error determined by comparing the first depths to the target depths associated with the pixels of the synthetic image.

[0074] The first error is for example equal to the sum of the errors determined for pixels of the synthetic image, each error determined for a pixel of the synthetic image being equal to the absolute value of the difference between the predicted depth and the target depth associated with that pixel.

[0075] The first learning of the depth prediction model consists of adjusting the input parameters of the convolutional neural network implementing the depth prediction model in order to minimize the first loss error previously calculated.

[0076] According to one particular embodiment, the first learning phase is repeated for more computer-generated images from the set of computer-generated images, or even for all the computer-generated images from the set of computer-generated images.

[0077] Optionally, a second learning is then implemented in order to improve the reliability of the depth prediction model trained during the first learning phase.

[0078] In a fifth step 25, a real image is received. This real image represents at least a second object observed at a defined distance and corresponding to a second type of training object among the types of training objects determined in the first step 21. Target depths are in particular associated with pixels representing at least one second object.

[0079] As before, a real image represents, for example, a plurality of different objects. This real image is, for example, acquired by camera 11 or by a similar camera mounted on a vehicle and thus represents scenes real three-dimensional images are observed by this camera. Target depths are obtained, for example, from a LiDAR® system onboard the vehicle carrying the camera that acquires real images, or are determined by another depth prediction model.

[0080] In a sixth step 26, similar to the third step, second depths associated with pixels of the real image are predicted by the depth prediction model.

[0081] In a seventh step 27 called second learning, similar to the fourth step, the depth prediction model is learned by minimizing a second loss error determined by comparing the second depths to the target depths associated with the pixels of the real image.

[0082] As with the first learning, the second learning of the depth prediction model consists of adjusting the input parameters of the convolutional neural network implementing the depth prediction model in order to minimize the second loss error previously calculated.

[0083] Once the training of the depth model is complete, for example when a first and / or a second error is less than a threshold value defining an acceptable level of accuracy, it is then possible to use the depth prediction model to predict depths associated with pixels of images acquired by the camera 11.

[0084] Each determined depth then corresponds to a distance separating the vehicle 10 or a part of the vehicle 10 from an object in the three-dimensional scene to which a pixel is associated, the determination of a depth of a pixel then corresponding to a measurement of a distance separating an object from the vehicle carrying the vision system.

[0085] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another road user, the ADAS is then able to determine this distance precisely. For example, if the ADAS's function is to activate the braking system of the vehicle 10 in the event of a risk of collision with another road user, and the distance between the vehicle 10 and that same road user decreases sharply, then the ADAS is able to detect this sudden approach and activate the braking system of the vehicle 10 to avoid a possible accident.

[0086] Figure 3 schematically illustrates a device configured to learn a depth prediction model associated with a vehicle-mounted vision system and / or to predict a depth associated with a pixel of an image acquired by a camera of the vehicle-mounted vision system, according to an example of a particular and non-limiting embodiment of the present invention. Device 3 corresponds, for example, to a device embedded in the vehicle 10, for example a computer associated with the monocular vision system.

[0087] Device 3 is, for example, configured to carry out the steps described opposite Figures 1 and 2. Examples of such a device 3 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 3, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 3 may be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.

[0088] The device 3 comprises one (or more) processor(s) 30 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 3. The processor 30 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 3 further comprises at least one memory 31, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.

[0089] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 31.

[0090] According to various particular and non-limiting embodiments, the device 3 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English "Telematic Control Unit" or in French "Unité de Contrôle Télématique"), for example via a communication bus or through dedicated input / output ports.

[0091] According to a particular and non-limiting embodiment, the device 3 includes a block 32 of interface elements for communicating with external devices. The interface elements of the block 32 include one or more of the following interfaces: - Radio frequency (RF) interface, for example, Wi-Fi® (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox using UBN (Ultra Narrow Band) radio technology narrow), or LoRa in the 868 MHz frequency band, LTE (from the English "Long-Term Evolution" or in French "Evolution à long terme"), LTE-Advanced (or in French LTE-avancé); - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HD MI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").

[0092] According to another particular and non-limiting embodiment, the device 3 includes a communication interface 33 which allows communication to be established with other devices (such as other computers in the embedded system) via a communication channel 330. The communication interface 33 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 330. The communication interface 33 corresponds, for example, to a wired network of the CAN (Controller Area Network) type, CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3).

[0093] According to a particular and non-limiting embodiment, the device 3 can provide output signals to one or more external devices, such as a display screen 340, touch or not, one or more speakers 350 and / or other peripherals 360 via the output interfaces 34, 35, 36 respectively. According to a variant, one or more of the external devices is integrated into the device 3.

[0094] Of course, the present invention is not limited to the embodiments described above but extends to a method for determining the depth of a pixel in an image acquired by a vision system, and / or for measuring the distance between an object and a vehicle equipped with a vision system, the depth and / or distance being predicted and / or measured via a depth prediction model learned according to the learning method described above, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0095] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 3 of [Fig.3].

Claims

Demands

1. A method for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle (10), the vision system comprising a camera (11) arranged to acquire an image of a three-dimensional scene taking place near the vehicle (10), said method being implemented by at least one processor, and being characterized in that it comprises the following steps: - determination (21) of types of training objects and types of training environments, - reception (22) of a synthetic image from a set of synthetic images, each synthetic image from said set of synthetic images representing at least a first object observed at a defined distance and corresponding to a first type of training object from among said types of training objects integrated into an environment from among said types of training environments,target depths being associated with pixels representing said at least one first object, said each synthetic image and said target depths associated with the pixels of said each synthetic image being generated by a synthetic image generation model; - prediction (23) of first depths associated with pixels of the synthetic image by the depth prediction model; - first learning (24) of the depth prediction model by minimizing a first loss error determined by comparing the first depths to the target depths associated with said pixels of the synthetic image.

2. A method according to claim 1, wherein a distribution of distances defined for objects of each type of learning object among said types of learning objects in said set of synthetic images is uniform over a range of distances, said range of distances corresponding to a measurement range associated with the vision system.

3. A method according to claim 2, wherein a distribution of positions of the objects of said each type of learning object in said set of computer-generated images is uniform along the x-axis and y-axis.

4. A method according to any one of claims 1 to 3, wherein said synthetic image generation model generates images based on intrinsic parameters of the camera (11) of the vision system.

5. A method according to any one of claims 1 to 4, wherein computer-generated images of said set of computer-generated images represent several objects.

6. A method according to any one of claims 1 to 5, further comprising the following steps: - receiving (25) a real image representing at least a second object observed at a defined distance and corresponding to a second type of training object among said types of training objects, target depths being associated with pixels representing said at least a second object, - predicting (26) second depths associated with pixels of the real image by the depth prediction model; - second training (27) of the depth prediction model by minimizing a second loss error determined by comparing the second depths to the target depths associated with said pixels of the real image.

7. A method according to any one of claims 1 to 6, wherein a type of learning object belongs to a set of learning objects comprising: - a car, - a truck, - a coach, - a motorcycle, and - a bicycle, and said type of learning environment belongs to a set of learning environments comprising: - a city center, - a residential area, - an industrial area, - a parking area, - a motorway, and - a rural area.

8. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by a processor.

9. Device (3) configured to learn a depth prediction model by a vision system mounted in a vehicle (10), said device (3) comprising a memory (31) associated with at least one processor (30) configured to implement the steps of the method according to any one of claims 1 to 7.

10. Vehicle (10) comprising the device (3) according to claim 9.

Citation Information

Patent Citations

  • Monocular depth estimation

    GB2605621A

  • Mixed-batch training of a multi-task network

    US20220156525A1