Method and device for acquiring training data for a vision system.

The process and device for acquiring a learning database for depth prediction models in vehicle vision systems address the challenge of diverse road environments by collecting and recording real scene data, resulting in improved accuracy and safety for ADAS systems.

FR3155341A1Pending Publication Date: 2025-05-16STELLANTIS AUTO SAS
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
FR2023012183
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing depth prediction models for vision systems in vehicles are not adequately trained on diverse road environments, leading to reduced accuracy and safety in ADAS systems.

Method used

A process and device for acquiring a learning database for a depth prediction model, utilizing a vision system with multiple cameras to collect and selectively record data representative of real scenes, including object types, depths, and road environments, to create a tailored database for each vehicle's environment.

Benefits of technology

The proposed solution enhances the relevance and accuracy of learning data for depth prediction models, improving the operating security of ADAS systems by providing more precise depth predictions tailored to specific road environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method or device for acquiring a training database for a depth prediction model for a vehicle-mounted vision system, the vision system comprising a first camera (11) and a second camera (12) arranged to each acquire an image of a scene from a different viewpoint. Specifically, the method comprises receiving representative data from a set of four images, acquired by the first and second cameras at two distinct time points, determining a depth associated with the current object using the depth prediction model based on bounding boxes, and storing in the training database a current element comprising representative data of the current object and its depth, as well as the four images. Figure 1 (for the abstract)
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for acquiring learning data from a vision system. Technical field

[0001] The present invention relates to methods and devices for acquiring training data for a depth prediction model for a vision system embedded in a vehicle, for example in a motor vehicle. The present invention also relates to a method for training a depth prediction model for a vision system embedded in a vehicle from training data acquired by the onboard vision system. Technological background

[0002] Many modern vehicles are equipped with so-called AD AS (Advanced Driver-Assistance System). Such AD AS systems are passive and active safety systems designed to eliminate the element of human error in the driving of vehicles of all types. AD AS use advanced technologies to assist the driver while driving and thus improve their performance. AD AS use a combination of sensor technologies to perceive the environment around a vehicle, then provide information to the driver or act on certain vehicle systems.

[0003] There are several levels of ADAS, such as rearview cameras and blind spot sensors, lane departure warning systems, adaptive cruise control and automatic parking systems.

[0004] The AD AS embedded in a vehicle are supplied with data obtained from one or more embedded sensors such as, for example, cameras. These cameras make it possible in particular to detect and locate other road users or possible obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle speed; - to act on the braking system in the event of a risk of impact with an object.

[0005] A position of another user or of an obstacle is for example determined by a vision system comprising a model for predicting a depth or a distance. Such a model is for example learned using images, these images being obtained from a universal database, for example Kitti® or Sceneflow®. Kitti® for example provides images of a road environment in a city center, but such a database does not, however, include the entirety of the surroundings. road conditions in which a vehicle can operate. The training data is then unsuitable for training the prediction model for a vehicle traveling in other road environments.

[0006] The quality of the training of the prediction model is however very important, in fact, the depths or distances predicted by the prediction model represent for example distances at which the other users or obstacles present in the road environment of the vehicle equipped with the vision system and the AD AS are located. The proper functioning of the driving aid peripherals using this data therefore depends on the quality of the data emitted by the vision system. Summary of the present invention

[0007] An object of the present invention is to solve at least one of the problems of the technological background described above.

[0008] Another object of the present invention is to improve the relevance of the training data of a depth prediction model for an on-board vision system in a vehicle.

[0009] Another object of the present invention is to improve road safety, in particular by improving the operational safety of AD AS systems supplied by data obtained from a vision system on board the vehicle.

[0010] According to a first aspect, the present invention relates to a method for acquiring a training database of a depth prediction model for a vision system embedded in a vehicle, the vision system comprising a set of cameras comprising a first camera and a second camera arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the training database comprising a first set of elements, each element of the first set of elements comprising first data representative of: • a type of object, • a first depth, and • a first set of four images, the method being characterized in that it comprises the following steps: - determining a first depth distribution associated with a target object type from the first depths of a second set of elements belonging to the first set of elements of the learning database, each object type of an element of the second set of elements corresponding to the target object type; - receiving second data representative of a second set of four images, the second set of four images comprising: • a first and a second image acquired respectively by the first camera and the second camera at a first time instant, • a third and a fourth images acquired respectively by the first camera and the second camera at a second time instant different from the first time instant; - detection of a current object, in each image of the second set of four images, a type of the current object corresponding to the type of target object; - determination, in each image of the second set of four images, of a bounding box comprising pixels associated with the current object; - determination of a second depth associated with the current object by the depth prediction model from the bounding boxes; - adding the second depth to the first depth distribution to generate a second depth distribution; and - recording in the learning database a current element comprising third data representative of the current object and the second depth and the second data, the recording being a function of the second depth distribution.

[0011] Such a method thus makes it possible to obtain a database adapted to the road environment encountered by the vehicle. Indeed, the types of object and the images associated with these objects are acquired by the vision system of a vehicle and are therefore representative of real scenes taking place around a vehicle. Since the recording is a function of the second depth distribution, the database is fed selectively in order to limit its size while retaining relevant elements.

[0012] According to a variant of the method, each element of the learning database further comprising data representative of a type of road environment, the method further comprises a step of determining a current type of road environment associated with the current element from the second data, data representative of the current type of environment being recorded in the learning database.

[0013] The different elements of the database can thus be arranged by type of road environment.

[0014] According to another variant of the method, the depth prediction model is trained from the second data when the current object type corresponds to the target object type.

[0015] The depth prediction model is thus trained to improve the depth prediction accuracy for this type of target object.

[0016] According to an additional variant of the method, the second depth is equal to an average depth of a set of pixels of the bounding box of the first image.

[0017] According to a further variant of the method, the current element is recorded when the current object is entirely visible in each of the images of the second set of four images.

[0018] According to another variant of the method, the current element is recorded if a total number of elements of the second set of elements is less than a total threshold.

[0019] According to yet another variant of the method, the current element is recorded if a number of elements of a part of the second set of elements comprising a first depth belonging to the same depth class as the second depth is between a minimum threshold and a maximum threshold, the minimum and maximum thresholds being defined as a function of a total number of elements of the second set of elements and a total number of depth classes.

[0020] According to an additional variant of the method, the maximum threshold is determined by the following function: S„ = 4xNt / (3xIt) with : • Smax the maximum threshold, • Nt the total number of elements in the second set of elements, and • IT the number of intervals.

[0021] According to a second aspect, the present invention relates to a device for acquiring a training database of a depth prediction model for a vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for implementing the steps of the method according to the first aspect of the present invention.

[0022] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0023] According to a fourth aspect, the present invention relates to a computer program which comprises instructions adapted for executing the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0024] Such a computer program may use any programming language and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0025] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to the first aspect of the present invention.

[0026] On the one hand, the recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM memory, a CD-ROM or a microelectronic circuit type ROM memory, or a magnetic recording means or a hard disk.

[0027] Furthermore, this recording medium may also be a transmissible medium such as an electrical or optical signal, such a signal being able to be conveyed via an electrical or optical cable, by conventional or hertzian radio or by self-directed laser beam or by other means. The computer program according to the present invention may in particular be downloaded from an Internet-type network.

[0028] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to perform or to be used in performing the method in question. Brief description of the figures

[0029] Other characteristics and advantages of the present invention will emerge from the description of the particular and non-limiting exemplary embodiments of the present invention below, with reference to the appended figures 1 to 4, in which:

[0030] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting example of embodiment of the present invention;

[0031] [Fig.2] illustrates a flowchart of the different steps of a method for acquiring a training database of a vision model embedded in the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;

[0032] [Fig.3] illustrates a diagram representing the statistical distribution of a number of elements of the learning database by depth class, according to a particular and non-limiting exemplary embodiment of the present invention;

[0033] [Fig.4] schematically illustrates a device configured for the acquisition of a training database of a vision model embedded in the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention. Description of examples of implementation

[0034] A method and a device for acquiring a training database of a vision model embedded in a vehicle will now be described in what follows with joint reference to FIGS. 1 to 4. The same elements are identified with the same reference signs throughout the description which follows.

[0035] The terms “first(s)”, “second(s)” (or “first(s)”, “second(s)”), etc. are used in this document by arbitrary convention to identify and distinguish different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0036] According to a particular and non-limiting example of embodiment of the present invention, a method for acquiring a training database of a depth prediction model for a vision system on board a vehicle is for example implemented by a computer of the on-board system of the vehicle controlling this vision system.

[0037] The vision system comprises a set of cameras of at least two cameras comprising a first camera and a second camera arranged so as to each acquire an image of a three-dimensional scene from a different point of view.

[0038] The learning database comprises a first set of elements, each element of the first set of elements comprising first data representative of: • a type of object, • a first depth, and • a first set of four images.

[0039] Indeed, the method comprises the determination of a first distribution of depths associated with a target object type from the first depths of a second set of elements belonging to the first set of elements of the learning database, each object type of an element of the second set of elements corresponding to the target object type.

[0040] The method also comprises receiving second data representative of a second set of four images, acquired by the first and second cameras at two distinct time instants and detecting a current object in each image of the second set of four images, a type of the current object corresponding to the type of target object.

[0041] A bounding box comprising pixels associated with the current object is then determined in each image of the second set of four images and a second depth associated with the current object is determined by the depth prediction model from the bounding boxes.

[0042] The second depth is added to the first depth distribution to generate a second depth distribution and a current element comprising third data representative of the current object and the second depth and the second data is recorded in the training database, the recording being a function of the second depth distribution.

[0043] Such a method thus makes it possible to obtain a database adapted to the road environment and to the objects encountered by the vehicle. Indeed, the types of object and the images associated with these objects are acquired by the vision system of a vehicle and are therefore representative of real scenes taking place around a vehicle and as perceived by such an on-board vision system.

[0044] Since the recording is a function of the second depth distribution, the database is fed selectively in order to limit its size while keeping relevant elements, the relevant elements making it possible to have more training data representative of a set of three-dimensional scenes.

[0045] [Fig. 1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.

[0046] An environment 1 corresponds, for example, to a road environment formed of a network of roads accessible to the vehicle 10.

[0047] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, with an electric motor(s) or even a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as an automobile, a truck, a bus, a motorcycle. Finally, the vehicle 10 corresponds to an autonomous vehicle or not, that is to say a vehicle traveling according to a determined level of autonomy or under the total supervision of the driver.

[0048] The vehicle 10 advantageously comprises a set of cameras comprising a first camera 11 and a second camera 12 on board, each configured to acquire images of a three-dimensional scene in the environment 1 of the vehicle 10. This set of cameras forms the vision system. Two cameras are illustrated in [Fig.l]. The present invention is however not limited to a vision system comprising two cameras but extends to any vision system comprising 2 or more cameras, for example 2, 3, 4 or 5 cameras.

[0049] The first 11 and second 12 cameras have known intrinsic parameters. These parameters consist in particular of: - the focal length fl of the first camera 11; - the focal length f2 of the second camera 12; - distortions which are due to imperfections in the optical system of each camera; - the direction Cl of the optical axis of the first camera 11; - the direction C2 of the optical axis of the second camera 12; and - the respective resolutions of cameras 11, 12.

[0050] The intrinsic parameters characterize the transformation which associates, for an image point, the camera coordinates with the pixel coordinates, in each camera. These parameters do not change if the camera is moved.

[0051] The distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions which generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, that is to say that the distortions are not taken into account or that their correction is processed at the time of image acquisition.

[0052] These first 11 and second 12 cameras are arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the first point of view is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10, the second point of view is for example located on or in the right rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. In the case where the two cameras are located at the top of the windshield of the vehicle, they are then placed at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, the second camera 12 is located in the right rearview mirror of the vehicle 10.

[0053] A first marker is associated with the first camera 11: - the direction of the y axis is defined by the position of the second camera 12, so as to place the second camera 12 on the y axis of the first camera 11. The distance B separating the two cameras 11, 12 is called the reference base (in English “baseline”) and the direction separating the two cameras 11, 12 is that of the y axis; - the direction of the x axis is defined orthogonal to that of the y axis and orthogonal to that of the optical axis Cl of the first camera 11; - the direction of the z axis is defined orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal reference frame.

[0054] The extrinsic parameters linked to the position of the cameras 11, 12 are the following parameters: - 3 translations in the x, y and z directions: Tx, Ty and Tz constituting the translation vector T; and - 3 rotations around the x, y and z axes: Rx, Ry and Rz, constituting the rotation matrix R.

[0055] The determination of the extrinsic parameters is carried out for example during the calibration of the vision system.

[0056] A stereoscopic vision system is a vision system comprising a plurality of cameras, for example the first camera 11 and the second camera 12. A main constraint of a stereoscopic vision system used in The automotive industry, for example, is characterized by the large distance between the two cameras. In fact, to be able to cover a measuring range of 200 meters, the reference base must reach 60 cm for the cameras commonly used in this field.

[0057] The first 11 and second 12 cameras acquire images of a three-dimensional scene located in front of the vehicle 10, the first camera 11 covering only a first acquisition field 13, the second camera 12 covering only a second acquisition field 14 and the first 11 and second 12 cameras both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic vision of the three-dimensional scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic vision of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic vision of the scene by the stereoscopic vision system composed of the first 11 and second 12 cameras.

[0058] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.

[0059] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system composed of the second camera 12.

[0060] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable using the monoscopic vision system composed of the first camera 11.

[0061] Finally, field 19 is not visible from any of the cameras. The part of the scene present in this field 19 is therefore not observable.

[0062] When the vehicle 10 is in motion, each of the first 11 and second 12 cameras forms a monoscopic vision system.

[0063] According to an exemplary embodiment, the directions C1, C2 of the optical axes representative of an orientation of the field of vision of each camera are oriented non-parallel so as to obtain the third acquisition field 15 of the environment 1 as wide as possible.

[0064] According to another exemplary embodiment, the directions C1, C2 of the optical axes representative of an orientation of the field of vision of each camera are oriented parallel.

[0065] It is obvious that it is possible to use such a vision system, stereoscopic or monoscopic, to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently.

[0066] The images acquired by the first 11 and second 12 cameras at a time given acquisition time are presented in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colors and brightness of objects in the observed scene in the form, for example, of RGB colorimetric coordinates (from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness).

[0067] The images acquired by the first 11 and second 12 cameras at the same time instant represent views of the same scene taken from different viewpoints, the positions of the cameras being distinct. On this scene there are for example objects, each object belonging to a type of object such as: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.

[0068] These images are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.

[0069] A process for acquiring a learning database of a depth prediction model for the vision system on board the vehicle 10 is advantageously implemented by the vehicle 10, i.e. by a computer or a combination of computers of the on-board system of the vehicle 10, for example by the computer(s) in charge of the vision system of the vehicle 10.

[0070] The training database comprises a first set of elements, each element of the first set of elements comprising first data representative of: • a type of object, • a first depth, and • a first set of four images.

[0071] It should be noted that the method works even if the learning database does not yet include any elements, in fact the aim of this method is to complete an existing learning database or to constitute a new learning database.

[0072] A first depth is for example a distance, expressed in meters, separating the vision system having acquired a first set of images of an object.

[0073] A first set of four images comprises, for example, two pairs of images acquired by two separate cameras at two separate time instants. For example, example by the first camera 11 and by the second camera 12 of the vision system on board the vehicle 10, or by another pair of cameras belonging to a similar vision system. Indeed, according to a first particular embodiment, the database is entirely made up of elements comprising data acquired by the vision system of the vehicle 10, while according to a second particular embodiment, the database is community-based or comes from a remote server and is made up of elements comprising data acquired by different similar vision systems on board several vehicles. The similar vision systems on board a plurality of vehicles are, for example, vision systems comprising cameras of the same technical specifications and arranged at the same locations in the plurality of vehicles.

[0074] According to a particular exemplary embodiment, each element of the learning database further comprises data representative of a type of road environment. A type of road environment belongs for example to a set of types of road environments comprising: - a parking lot, - a city center in an agglomeration, - an agglomeration (outside the city center), - a departmental or national road in a rural area, - a road with regulated access, and - a highway.

[0075] In a first operation, a first depth distribution associated with a target object type is determined from the first depths of a second set of elements belonging to the first set of elements of the training database, each object type of an element of the second set of elements corresponding to the target object type.

[0076] In other words, the first operation consists of selecting from the learning database the elements of the first set of elements for which the object type corresponds to a target object type. Such a target object type is for example determined by a user or by a program determining a target object type according to a need, for example in order to acquire additional elements corresponding to this target object type. Once these elements have been selected, they form the second set of elements. From this second set of elements are extracted the first depths associated with each of these elements, and a first depth distribution is determined.

[0077] According to an exemplary embodiment, the first depth distribution is determined by depth classes. The depth classes are for example of equal width or different width, progressive as a function of the depth. average of each.

[0078] [Fig. 3] illustrates an example of a first distribution of depths by class in the form of a histogram 3. Each bar 31, 32, 33, 34, 35 of the histogram 3 represents the ratio between the number of elements of the second set of elements whose first length is included in a class and the number of elements of the second set of elements. For example: - the first class represented by a first bar 31 comprises all the elements of the second set of elements whose first depth is between 0 and Dl, - the second class represented by a second bar 32 comprises all the elements of the second set of elements whose first depth is between D1 and D2, - the third class represented by a third bar 33 comprises all the elements of the second set of elements whose first depth is between D2 and D3, - the fourth class represented by a fourth bar 34 comprises all the elements of the second set of elements whose first depth is between D3 and D4, and - the fifth class represented by a fifth bar 35 includes all the elements of the second set of elements whose first depth is greater than D4.

[0079] The abscissa axis represents the depth D (in English "Depth"), for example expressed in meters (m) while the ordinate axis represents the ratio or the probability P that each class represents. The ratio is for example expressed as a percentage (%).

[0080] According to a particular exemplary embodiment, each class has a ratio between a minimum threshold Smin and a maximum threshold Smax. Such limits thus make it possible to ensure that each class is represented in a homogeneous manner. Thus, for each depth range, a similar number of elements corresponding to the target object type is recorded in the learning database.

[0081] The minimum and maximum thresholds are for example defined as a function of a total number of elements of the second set of elements and a total number of depth classes, for example the minimum threshold is determined by the following function: Smin = 2 x NT / (3 x IT) With: • Smin the minimum threshold, • Nt the total number of elements in the second set of elements, and • IT the number of intervals, and the maximum threshold Smax is determined by the following function: S - ? x S ^max ^min*

[0082] In a second operation, second data representative of a second set of four images are received. The second set of four images comprises: • a first and a second image acquired respectively by the first camera 11 and the second camera 12 at a first time instant, • a third and a fourth images acquired respectively by the first camera 11 and the second camera 12 at a second time instant different from the first time instant.

[0083] The images acquired respectively by the first camera 11 and the second camera 12 at a first time instant make it possible, for example, to determine a depth of an object present in these images by a stereoscopic vision system comprising the first 11 and second 12 cameras.

[0084] The images acquired by the first camera 11 at the two distinct time instants make it possible, for example, to determine the depth of an object present in these images by a monoscopic vision system comprising the first camera 11.

[0085] The images acquired by the second camera 12 at the two distinct time instants make it possible, for example, to determine a depth of an object present in these images by a monoscopic vision system comprising the second camera 12.

[0086] In a third operation, a current object is detected in each image of the second set of four images, a type of the current object corresponding to the target object type and in a fourth operation, a bounding box comprising pixels associated with the current object is determined in each image of the second set of four images.

[0087] The detection of an object in an image and the determination of a bounding box are known to those skilled in the art, these operations being carried out for example using algorithms such as Yolo7®, centemet® or vfnet®.

[0088] According to a particular exemplary embodiment, the following operations are only implemented if the current object is entirely visible in each of the images of the second set of four images. Here, “entirely visible” means that no other object partially masks the current object, i.e. that the current object is not occluded by another object, or that no part of the object is located outside the field of vision of each of the cameras at the two time instants of acquisition of the images of the second set of four images. This makes it possible to subsequently save only current elements comprising the second set of four images in which the current object is entirely visible.

[0089] In a fifth operation, a second depth associated with the current object is determined by the depth prediction model from the bounding boxes. The depth prediction model is for example associated with the stereoscopic vision system comprising the first 11 and second 12 cameras and known to those skilled in the art, for example using the CoEx® model, or with the monoscopic vision system comprising the first moving camera 11 or the second moving camera 12 and known to those skilled in the art, for example using the HR-depth® model. This operation consists firstly of determining the depths of the pixels included in the bounding boxes.

[0090] According to a particular exemplary embodiment, the depths are predicted for pixels included in reduced bounding boxes, the reduced bounding boxes being included in the previously determined bounding boxes from which peripheral pixels are subtracted, the reduced bounding boxes being for example of a width and height reduced by 25% compared to the previously determined bounding boxes and centered on these same bounding boxes. This reduction of the bounding boxes thus makes it possible not to take into account background pixels when determining depths and which would distort the determination of the depth of the current object.

[0091] Second, the second depth is determined to be equal to an average depth of a set of pixels of the bounding box of the first image.

[0092] In a sixth operation, the second depth is added to the first depth distribution to generate a second depth distribution.

[0093] According to the particular embodiment example previously presented in which each element of the learning database further comprises data representative of a type of road environment, a current type of road environment associated with the current element is determined in a seventh operation from the second data.

[0094] In an eighth operation, a current element comprising third data representative of the current object and the second depth and the second data is recorded in the learning database according to the second depth distribution.

[0095] According to the particular embodiment previously presented in which each element of the learning database further comprises data representative of a type of road environment, data representative of the current type of road environment are also recorded in the learning database.

[0096] According to a particular exemplary embodiment, the current element is recorded if a total number of elements in the second set of elements is less than a total threshold. Such a condition thus makes it possible to avoid overloading or saturating the database and avoids using too much memory capacity to store the training database.

[0097] According to yet another particular exemplary embodiment, the current element is recorded if a number of elements of a part of the second set of elements comprising a first depth belonging to the same depth class as the second depth is between the minimum threshold and the maximum threshold previously defined. This condition makes it possible to record in the database only elements whose second depth is included in a depth class not being saturated in number of elements. Thus, the current element is saved only if it corresponds to a depth class not comprising the maximum number of elements, corresponding to the maximum threshold Smax.

[0098] It should be noted that each first set of four images and the second set of four images comprise four images. However, the invention is not limited to a vision system comprising two cameras and extends to any vision system comprising a plurality of cameras. The first and second sets of four images can then be sets of six, eight or more images, the number of images being for example twice the number of cameras included in the vision system on board the vehicle. Each element of the database then comprises first data representative of a set of a plurality of images and the different operations of the process are then implemented on this plurality of images.

[0099] According to a particular exemplary embodiment, the depth prediction model is trained from the second data when the current object type corresponds to the target object type.

[0100] According to yet another particular exemplary embodiment, the depth prediction model is trained from the acquired database and comprising the current element. The training of such a depth prediction model is known in particular to those skilled in the art. Indeed, many depth prediction models are implemented by convolutional neural networks. The convolutional neural network is then trained, that is to say its input parameters are adjusted, for example by minimizing loss functions, for example determined from reconstruction errors.Training such a depth prediction model from the training database is also called self-supervision and is known to those skilled in the art, for example through the use of algorithms such as UnOs® for a prediction model associated with a stereoscopic vision system or monodepth2® for a prediction model associated with a . monoscopic vision system.

[0101] This learning data is thus used to adjust the parameters of the depth prediction model associated with the vision system on board the vehicle, making the predicted depths more reliable and increasing the robustness of the depth prediction model.

[0102] If an ADAS uses depths predicted by this prediction model as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another user present on the road, the ADAS is then able to determine this distance precisely. For example, if the ADAS has the function of acting on a braking system of the vehicle 10 in the event of a risk of collision with another road user and the distance separating the vehicle 10 from this same road user decreases sharply, then the ADAS is able to detect this sudden approach and act on the braking system of the vehicle 10 to avoid a possible accident, including if the other user is seen by only one camera. The safety of the users is thus improved.

[0103] This database can further be shared, for example by saving it on a remote server accessible to the plurality of vehicles carrying similar vision systems. The latter then benefit from a relevant learning database making it possible to train the depth prediction model they carry.

[0104] [Fig.2] illustrates a flowchart of the different steps of a method 2 for acquiring a training database of a depth prediction model for a vision system embedded in a vehicle comprising a first camera 11 and a second camera 12, for example in the vehicle 10 of [Fig.1], according to a particular and non-limiting exemplary embodiment of the present invention. The method 2 is for example implemented by a device embedded in the vehicle 10 or by the device 4 of [Fig.4].

[0105] The learning database notably comprises a first set of elements, each element of the first set of elements comprising first data representative of: • a type of object, • a first depth, and • a first set of four images.

[0106] In a step 21, a first depth distribution associated with a target object type is determined from the first depths of a second set of elements belonging to the first set of elements of the learning database, each object type of an element of the second set of elements corresponding to the target object type.

[0107] In a step 22, second data representative of a second set of four images are received, the second set of four images comprising: • a first and a second image acquired respectively by the first camera 11 and the second camera 12 at a first time instant, • a third and a fourth images acquired respectively by the first camera 11 and the second camera 12 at a second time instant different from the first time instant.

[0108] In a step 23, a current object is detected in each image of the second set of four images, a type of the current object corresponding to the type of target object.

[0109] In a step 24, a bounding box comprising pixels associated with the current object is determined in each image of the second set of four images.

[0110] In a step 25, a second depth associated with the current object is determined by the depth prediction model from the bounding boxes.

[0111] In a step 26, the second depth is added to the first depth distribution to generate a second depth distribution.

[0112] In a step 27, a current element comprising third data representative of the current object and of the second depth and the second data is recorded in the learning database, the recording being a function of the second depth distribution

[0113] According to a variant, the variants and examples of the operations described in relation to figures 1 and 3 apply to the steps of method 2 of [Fig.2].

[0114] [Fig. 4] schematically illustrates a device 4 configured for the acquisition of a training database of a depth prediction model for a vision system embedded in a vehicle, for example in the vehicle 10 of the first figure, according to a particular and non-limiting exemplary embodiment of the present invention. The device 4 corresponds for example to a device embedded in the first vehicle 10, for example a computer.

[0115] The device 4 is for example configured for the implementation of the operations described with regard to figures 1 and 3 and / or steps described with regard to [Fig.2]. Examples of such a device 4 include, but are not limited to, on-board electronic equipment such as an on-board computer of a vehicle, an electronic calculator such as an ECU (“Electronic Control Unit”), a smartphone, a tablet, a laptop. The elements of the device 4, individually or in combination, can be integrated in a single integrated circuit, in several integrated circuits, and / or in discrete components. The device 4 can be produced in the form of electronic circuits or software (or computer) modules or even a combination of electronic circuits and modules software.

[0116] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the method and / or for executing the instructions of the software(s) embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41 corresponding for example to a volatile and / or non-volatile memory and / or comprises a memory storage device which may comprise volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.

[0117] The computer code of the embedded software(s) comprising the instructions to be loaded and executed by the processor is for example stored in the 4L memory.

[0118] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Telematic Control Unit”), for example via a communication bus or through dedicated input / output ports.

[0119] According to a particular and non-limiting exemplary embodiment, the device 4 comprises a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 comprise one or more of the following interfaces: - RF radio frequency interface, for example Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English “Universal Serial Bus” or “Universal Serial Bus” in French); HDMI interface (from the English “High Definition Multimedia Interface” or “High Definition Multimedia Interface” in French); - LIN interface (from the English “Local Interconnect Network”).

[0120] According to another particular and non-limiting exemplary embodiment, the device 4 comprises a communication interface 43 which makes it possible to establish communication with other devices (such as other computers of the on-board system) via a communication channel 430. The communication interface 43 corresponds for example to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds for example to a wired network of the CAN (Controller Area Network) type, CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by the ISO 17458 standard) or Ethernet (standardized by the ISO / IEC 802-3 standard).

[0121] According to a particular and non-limiting exemplary embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch-sensitive or not, one or more speakers 450 and / or other peripherals 460 (projection system) via the output interfaces 44, 45, 46 respectively. According to a variant, one or other of the external devices is integrated into the device 4.

[0122] Of course, the present invention is not limited to the exemplary embodiments described above but extends to a method for learning a depth prediction model for a vision system embedded in a vehicle from learning data acquired by the onboard vision system, which would include secondary steps without thereby departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0123] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-based motor vehicle, comprising the device 4 of [Fig.4].

Claims

1. Claims Method for acquiring a training database of a depth prediction model for a vision system embedded in a vehicle, the vision system comprising a set of cameras comprising a first camera (11) and a second camera (12) arranged so as to each acquire an image of a three-dimensional scene from a different point of view, said training database comprising a first set of elements, each element of said first set of elements comprising first data representative of: • a type of object, • a first depth, and • a first set of four images, said method being characterized in that it comprises the following steps: - determining (21) a first depth distribution associated with a target object type from the first depths of a second set of elements belonging to said first set of elements of the learning database, each object type of an element of said second set of elements corresponding to said target object type; - reception (22) of second data representative of a second set of four images, said second set of four images comprising: • a first and a second image acquired respectively by the first camera (11) and the second camera (12) at a first time instant, • a third and a fourth images acquired respectively by the first camera (11) and the second camera (12) at a second time instant different from the first time instant; - detection (23) of a current object, in each image of said second set of four images, a type of the current object corresponding to said type of target object; - determination (24), in each image of said second set of four images, of a bounding box comprising pixels associated with said current object; - determination (25) of a second depth associated with said object current by said depth prediction model from the bounding boxes; - adding (26) said second depth to said first depth distribution to generate a second depth distribution; and - recording (27) in said training database a current element comprising third data representative of the current object and the second depth and the second data, said recording being a function of said second depth distribution.

2. Method according to claim 1, each element of said learning database further comprising data representative of a type of road environment, further comprising a step of determining a current type of road environment associated with said current element from said second data, data representative of said current type of environment being recorded in said learning database.

3. Method according to one of claims 1 to 2, for which said depth prediction model is trained from the second data when said current object type corresponds to said target object type.

4. Method according to one of claims 1 to 3, for which said second depth is equal to an average depth of a set of pixels of the bounding box of the first image.

5. Method according to one of claims 1 to 4, for which said current element is recorded when said current object is entirely visible in each of the images of the second set of four images.

6. Method according to one of claims 1 to 5, for which said current element is recorded if a total number of elements of the second set of elements is less than a total threshold.

7. Method according to one of claims 1 to 5, for which said current element is recorded if a number of elements of a part of the second set of elements comprising a first depth belonging to the same depth class as the second depth is between a minimum threshold and a maximum threshold, said minimum and maximum thresholds being defined as a function of a total number of elements of the second set of elements and a total number of depth classes.

8. A method according to claim 7, wherein said maximum threshold is determined by the following function: S_ = 4 x NT / (3 x IT) with: • Smax the minimum threshold, • Nt the total number of elements of the second set of elements, and • IT the number of intervals.

9. Device (4) for acquiring a learning database of a vision system on board a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for implementing the steps of the method according to any one of claims 1 to 8.

10. Vehicle (10) comprising the device (4) according to claim 9.

Citation Information

Patent Citations

  • Generating stereo image data from monocular images

    US11805236B2

  • Method and system for object centric stereo in autonomous driving vehicles

    US20180348780A1

  • Network architecture for monocular depth estimation and object detection

    US20220301202A1

  • Systems and methods for self-supervised depth estimation

    US20230037731A1

  • Monocular depth estimation method, apparatus and device

    WO2022165722A1