Method and device for learning a depth prediction model to predict a depth for any distance traveled between the acquisition of two images.
The method improves ADAS system reliability by training a depth prediction model in two stages, first at low speeds and then at high speeds, ensuring accurate depth predictions across varying vehicle speeds, thus enhancing safety.
Patent Information
- Application Number
- FR2024003628
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-04-09
- Publication Date
- 2026-02-20
- Estimated Expiration
- 2044-04-09
AI Technical Summary
Existing depth prediction models for vehicle vision systems are trained on data that do not accurately represent diverse driving environments, particularly at varying speeds, leading to unreliable depth predictions and compromised safety in ADAS systems.
A method for learning a depth prediction model using image sequences acquired by a vehicle camera, classified by speed, with two-stage training: first at low speeds and then at high speeds, adjusting the model to accurately predict depths across different vehicle speeds.
The method enables the depth prediction model to reliably predict depths at any speed, using real-world driving conditions for training, thereby enhancing the safety and reliability of ADAS systems.
Smart Images

Figure 00000030_0000 
Figure 00000031_0000 
Figure 00000031_0001
Abstract
Description
Title of the invention: Method and device for learning a depth prediction model to predict a depth for any distance traveled between the acquisition of two images. technical field
[0001] The present invention relates to methods and devices for learning a depth prediction model associated with a vision system embedded in a vehicle, for example, in a motor vehicle. The present invention also relates to a method and device for determining depth and / or measuring the distance separating an object from a vehicle equipped with a vision system. Technological background
[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS systems use advanced technologies to assist the driver while driving and thus improve performance. ADAS systems use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.
[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.
[0004] The AD AS systems embedded in a vehicle are powered by data obtained one or more onboard sensors, such as cameras. These cameras make it possible to detect and locate other road users or potential obstacles around a vehicle in order to, for example: - to adapt the vehicle's lighting according to the presence of other road users; - to automatically regulate the vehicle's speed; - to act on the braking system in case of risk of impact with an object.
[0005] The position of another user or an obstacle is, for example, determined by a vision system comprising a model for predicting the depth associated with a pixel or the distance separating the vision system from an object in a three-dimensional scene. Such a model is, for example, implemented by a convolutional neural network and learned using images, these images being obtained from a database of Universal data, for example Kitti® or Sceneflow®. Kitti®, for instance, provides images of a city center road environment, but such a database does not encompass all the road environments a vehicle might encounter. Crucially, the image sequences acquired by a single moving camera—a sequence of images constituting a video—are obtained when that camera makes small movements, typically on the order of a few centimeters. Indeed, the image sequences available in Kitti®, for example, are representative of images acquired by a moving camera, each image being acquired from a viewpoint approximately 15 cm (fifteen centimeters) away from the viewpoint of the previously acquired image, which corresponds to a vehicle moving at a speed close to 18 km / h (eighteen kilometers per hour).The vehicle carrying the camera acquires the images of the image sequence at a frequency of 30 Hz (thirty hertz). Such images are useful for training a predictive model for a vehicle traveling at low speed in an urban environment, but are not suitable for a vehicle traveling in another type of environment such as a highway or motorway, where the vehicle's speed is significantly higher, for example, on the order of 110 or 130 km / h, and the vehicle's environment is also different. The training data is then unsuitable for training the predictive model for a vehicle traveling in this other type of road environment.
[0006] However, the quality of the training of the depth or distance prediction model is very important, as the depths or distances predicted by the depth prediction model represent the distances to other road users or obstacles present in the road environment of the vehicle equipped with the vision system and ADAS. Indeed, the proper functioning of the driver assistance devices using this data depends on the quality of the data emitted by the vision system.
[0007] A vision system comprising a single camera, hereinafter referred to as a monocular vision system, uses several consecutive images to predict depths associated with pixels in one of the images acquired by the camera using the depth prediction model. It is therefore important to train this depth prediction model with images or sequences of images acquired under conditions similar to those encountered when the vehicle is moving, particularly on the road network and at different speeds of the vehicle carrying the vision system. Summary of the present invention
[0008] An object of the present invention is to solve at least one of the problems of the technological background described previously.
[0009] Another object of the present invention is to improve the learning phase of a depth prediction model from image sequences acquired by a camera from viewpoints located at different distances from each other.
[0010] Another object of the present invention is to improve road safety, in particular by improving the reliability of AD AS systems powered by data obtained from a camera of a vision system.
[0011] According to a first aspect, the present invention relates to a method for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, the vision system comprising a camera arranged to acquire an image of a three-dimensional scene of an environment external to the vehicle, the process being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception of image sequences, each image sequence comprising a set of images acquired by the camera at different consecutive acquisition times separated by the same time interval, each image sequence being classified according to a vehicle speed at the time of image acquisition in a first class when the vehicle speed is less than a threshold speed and in a second class when the vehicle speed is greater than the threshold speed; - for each image sequence included in the first class, first training of the depth prediction model from first pairs of images, each first pair of images comprising a first image and a second image selected from each image sequence included in the first class; - for each image sequence included in the second class, second training of the depth prediction model from second pairs of images, each second pair of images comprising a third image and a fourth image selected from each image sequence included in the second class.
[0012] The method advantageously allows the depth prediction model to be learned in two stages. In the first stage, the depth prediction model is learned from image sequences acquired from a camera mounted in a vehicle moving at a speed below the threshold speed, i.e., at a reduced speed. In the second stage, the depth prediction model is learned from image sequences acquired from a camera mounted in a vehicle moving at a speed of above the threshold speed, i.e., at a high speed. Thus, the prediction model is learned from image sequences representing different vehicle speeds, enabling it to accurately predict depths or distances when the vehicle is moving at any speed, whether low or high. Furthermore, the image sequences used during training are, for example, acquired by the vehicle itself and are therefore perfectly representative of real-world driving conditions, and thus of three-dimensional scenes actually observed by the vehicle's onboard camera.
[0013] According to a variant of the method, the most recent image in each image sequence is assigned in each first pair of images as the first image and each image acquired prior to the first image is assigned in a first pair of images as the second image, each second image being different in each first pair of images.
[0014] The first pairs of images thus obtained comprise images acquired by a camera positioned at viewpoints located at a different distance for each first pair of images obtained from the same image sequence. The prediction model is then learned from pairs of images acquired from viewpoints positioned at a variable distance.
[0015] According to another variant of the method, the first learning comprises several iterations, each iteration using a different first pair of images corresponding to the first pair of images for which a time interval separating the acquisition of the first image and the second image is the smallest among the first pairs not used in a previous iteration, the depth prediction model being learned from the first pair of images used.
[0016] At each iteration, the distance separating the image acquisition viewpoints of a first pair of images is increasing, the prediction model is thus first learned for a small distance separating the image acquisition viewpoints and then for increasingly larger distances, until it is able to predict reliable depths for a large distance, this large distance being on the order, for example, of a multiple of the small distance, the multiple being equal to the number of iterations.
[0017] According to a further variant of the method, the first pairs are generated by selecting the first image as corresponding to the most recent image in each image sequence and the second image as corresponding to an image prior to the first image in each image sequence, the second image being selected according to a number of time intervals separating the acquisition of the first image from the acquisition of the second image, the number of intervals being determined according to a determined probability law.
[0018] The probability law thus makes it possible to control the frequency of use of the first pairs of images, particularly as a function of the number of time intervals between the acquisition times of the images in the first pair of images. It is then possible to favor or not the first pairs of images comprising images acquired at few time intervals or, conversely, at many time intervals. The sensitivity of the depth prediction model to the number of time intervals, and therefore to the distance separating the image acquisition points of the first pairs of images, is thus adjusted.
[0019] According to yet another variant of the method, the probability law is obtained by the following function: f(x) = r with : • / (x) a probability associated with an interval denoted x, • ° a standard deviation.
[0020] Such a probability distribution is representative of a normal distribution with a mean of zero. This probability distribution favors the use of initial image pairs comprising images acquired at short time intervals. Varying the standard deviation also allows for adjustment. Indeed, increasing the standard deviation reduces the occurrence of initial image pairs comprising images acquired at short time intervals and, conversely, increases the occurrence of initial image pairs comprising images acquired at long time intervals.
[0021] According to yet another embodiment of the method, the threshold speed is determined as a function of a target distance and an image acquisition frequency by the camera. This target distance makes it possible, for example, to distinguish between image sequences comprising images acquired when the vehicle is moving at low speed, for example in an urban environment, and image sequences comprising images acquired when the vehicle is moving at higher speed, for example in a rural environment, on a restricted access road, and on a motorway.
[0022] According to a further variant of the method, the depth prediction model is learned by minimizing a loss error determined from each pair of images of a plurality of pairs of images by comparing the images of each pair of images to images reconstructed from the images of each pair of images and depths predicted by the depth prediction model and associated with pixels of the images of said each pair of images, the plurality of pairs of images corresponding to the first pairs of images during the first learning and to the second pairs of images during the second learning.
[0023] Determining this loss error enables self-supervised training, that is, training solely from images acquired by the onboard camera and without requiring the use of annotated data from, for example, another onboard system such as a LiDAR®. The reconstruction of the two images is called bilateral reconstruction. It allows training both the depth prediction model and, if necessary, a model for predicting the camera's movement in the three-dimensional scene between the two time points of image acquisition for the first pair of images.
[0024] According to a second aspect, the present invention relates to a device configured to learn a depth prediction model by a vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.
[0025] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0026] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0027] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0028] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.
[0029] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.
[0030] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.
[0031] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures
[0032] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 7, in which:
[0033] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a a particular and non-limiting example of a realization of the present invention;
[0034] [Fig.2] illustrates a flowchart of the different stages of a determination process mination of a depth of one pixel of an image by a depth prediction model associated with a vision system embedded in the vehicle of the [Fig.1], according to a particular and non-limiting embodiment of the present invention;
[0035] [Fig.3] illustrates a flowchart of the different stages of a learning process of the depth prediction model used in the process of [Fig. 2], according to a particular and non-limiting embodiment of the present invention; and
[0036] [Fig.4] schematically illustrates a device configured to learn a model depth prediction by a vision system embedded in the vehicle of the [Fig.1], according to a particular and non-limiting embodiment of the present invention;
[0037] [Fig.5] schematically illustrates the generation of first pairs of images from of a sequence of images acquired by the camera mounted in the vehicle of the [Fig.1], according to a particular and non-limiting embodiment of the present invention;
[0038] [Fig.6] illustrates a flowchart of different stages of a learning process of the depth prediction model from a pair of images used in the training process of the depth prediction model of [Fig. 3], according to a particular and non-limiting embodiment of the present invention; and
[0039] [Fig.7] illustrates a graph representing a distance travelled by the vehicle of the [Fig.1] depending on its speed and a number of time intervals between the acquisition of two images, according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements
[0040] A method and device for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle will now be described in the following, with joint reference to Figures 1 to 7. The same elements are identified with the same reference signs throughout the description that follows.
[0041] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0042] For the purposes of this description, receiving an image or a sequence of images means receiving data representative of an image or, respectively, a sequence of images. Similarly, determining a depth or an error means determining data representative of a depth or an error. These abbreviations are intended solely to simplify the description; however, since the processes described below are implemented by one or more processors, it is clear that the input and output data of the various stages of a process are computer data.
[0043] According to a particular and non-limiting embodiment of the present invention, the depth prediction model is learned in a learning phase comprising a first and a second learning.
[0044] The first training of the depth prediction model is carried out from first pairs of images, each first pair of images being formed from images of the same sequence of images acquired by a camera mounted in a vehicle and moving at a speed less than a threshold speed in a three-dimensional scene.
[0045] The second learning of the depth prediction model is carried out from second pairs of images, each second pair of images being formed from images of the same sequence of images acquired by the camera on board the vehicle and moving at a speed greater than the threshold speed in the three-dimensional scene.
[0046] Fig. 1 schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.
[0047] The vehicle 10 is located in an environment 1 corresponding, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.
[0048] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.
[0049] The vehicle 10 advantageously comprises at least one on-board camera 11, configured to acquire images of a three-dimensional scene unfolding in the environment of vehicle 10 from a current viewing position. The camera 11 forms a monocular vision system when used alone as illustrated in [Fig. 1]. However, the present invention is not limited to a monocular vision system comprising a single camera but extends to any vision system comprising at least one camera, for example, 1, 2, 3, or 5 cameras.
[0050] The camera 11 has intrinsic parameters, including: - a focal length, - distortions that are due to imperfections in the optical system of camera 11, - a direction of the optical axis of camera 11, and - a resolution.
[0051] The intrinsic parameters characterize the transformation which associates, for an image point, hereafter called "point", its three-dimensional coordinates in the camera 11 reference frame with the pixel coordinates in an image acquired by the camera 11. These parameters do not change if the camera is moved.
[0052] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in lens curvature, parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are not taken into account, and their correction is addressed during image acquisition or calibration.
[0053] The camera 11 is arranged to acquire an image of a three-dimensional scene from a defined viewpoint, the viewpoint being, for example, located on or in the left-hand rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10 as illustrated in [Fig. 1]. The viewpoint of the camera 11 is then fixed in the frame of reference of the vehicle 10, and when the vehicle 10 moves, the viewpoint of the camera 11 also moves in the frame of reference of the observed three-dimensional scene which is located in the environment of the vehicle 10.
[0054] The camera 11, for example, acquires images of a three-dimensional scene located in front of the vehicle 10, the first camera 11 covering an acquisition field 12. An object 13 is placed in the acquisition field 12 of the camera 11, defining an occlusion field 14 for the vision system, an object present in the occlusion field 14 not being observable by the camera 11 from its current observation position.
[0055] It is evident that it is possible to use such a vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently, the invention not being limited to the observation of a three-dimensional scene taking place in front of the vehicle 10 carrying the camera 11.
[0056] According to one particular embodiment, the camera 11 is a wide-angle camera, a wide-angle camera being, for example, equipped with a lens designed to acquire a representative image of a three-dimensional scene seen over a wider field of view than a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens makes it possible to capture a larger portion of the three-dimensional scene unfolding in front of or around the camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by this camera. The angle α of the field of view of the camera 11 is, for example, equal to 120°, 145°, 180°, or 360°, whereas a standard camera offers, for example, a field of view open at an angle of 45° or less.Such a camera 11 corresponds, for example, to a camera equipped with mirrors or a "fisheye" camera. Wide-angle lenses have a shorter focal length compared to standard lenses, making them suitable for capturing images of landscapes, architecture, road intersections, or any other subject requiring a wide perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with an extended depth of field.
[0057] An image acquired by the camera 11 at a given acquisition time is in the form of data representing pixels characterized by: - coordinates in the image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric values (from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).
[0058] Each pixel of the acquired image represents an object in the three-dimensional scene present in the field of view of the camera 11. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a point of light resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, which are captured by a photosensitive sensor of the camera 11 after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates a image faithfully representing the physical object observed by the camera 11. A previously presented image point is thus a point on the surface of an object in the three-dimensional scene.
[0059] When the vehicle 10 is in motion, then two images acquired by the camera 11 at two distinct time instants represent views of the same three-dimensional scene taken from different viewpoints or observation positions, the observation positions of the camera 11 being distinct. For example, this three-dimensional scene may contain: - buildings; - road infrastructure; - other users or stationary objects, for example a parked vehicle; and / or - other users or moving objects, for example another vehicle, a cyclist or a moving pedestrian.
[0060] The distance between the different acquisition viewpoints depends in particular on the speed of the vehicle 10 and the image acquisition frequency of the camera 11. It should be noted that if the vehicle moves at a constant speed and the camera 11 acquires images regularly, i.e., at a constant frequency, then the distance between viewpoints associated with consecutively acquired images is constant. If the vehicle's speed varies, this distance also varies proportionally. Thus, when the vehicle 10 moves at low speed, the distance between viewpoints when two consecutive images are acquired by the camera 11 is small, and conversely, when the vehicle 10 moves at high speed, the distance between viewpoints when two consecutive images are acquired by the camera 11 is large.
[0061] According to a particular embodiment, an image acquired by the camera 11 includes a distortion equal to 0.5%, 0.8%, or greater than 1%. The measurement of such distortion corresponds to determining a ratio between: - the maximum spacing of a pixel in the image of a straight line in the first three-dimensional scene whose image is a line touching the longest edge of the first image, either at the center of the image edge or at the corners of the image edge, and - the length of this edge.
[0062] In the world of photography, distortion is commonly considered to be: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • problematic if it is greater than or equal to 1% or more.
[0063] Barrel distortion is characterized by a positive percentage, while crescent distortion is characterized by a negative percentage.
[0064] Each image acquired by the camera 11 is for example sent to a processor, for example a computer of a device equipping the vehicle 10, or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10. It is then used during the implementation of a method for determining the depth of one of its pixels and / or during the implementation of a method for learning a depth prediction model associated with this camera of the vision system.
[0065] Figure 2 illustrates a flowchart of the different steps of a method 2 for determining the depth of a pixel in an image by means of a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, for example in vehicle 10 of Figure 1, according to a particular and non-limiting embodiment of the present invention. Method 2 is implemented, for example, by a device of the vision system embedded in vehicle 10 or by device 4 of Figure 4.
[0066] In a step 21, two images acquired by the camera 11 are received.
[0067] In a step 22, depths associated with a set of pixels of one of the two received images are determined by the depth prediction model from the two received images.
[0068] Each determined depth then corresponds to a distance separating the vehicle 10 or a part of the vehicle 10 from an object in the three-dimensional scene to which a pixel is associated, the determination of a depth of a pixel then corresponding to a measurement of a distance separating an object from the vehicle carrying the vision system.
[0069] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another road user, the ADAS is then able to determine this distance precisely. For example, if the ADAS's function is to activate the braking system of the vehicle 10 in the event of a risk of collision with another road user, and the distance between the vehicle 10 and that same road user decreases sharply, then the ADAS is able to detect this sudden closeness and activate the braking system of the vehicle 10 to avoid a possible accident.
[0070] Figure 3 illustrates a flowchart of the different steps in a method for learning the depth prediction model used in a method for determining the depth of a pixel in an image, for example in method 2 of Figure 2, according to a particular and non-limiting embodiment of the present invention.
[0071] The learning process 3 is for example implemented by the device on board the vehicle 10 implementing the depth determination process for the vision system on board a vehicle 10 or by the device 4 of the [Fig.4].
[0072] In a step 31, image sequences are received. Each image sequence comprises a set of images acquired by the camera 11 at different consecutive acquisition times separated by the same time interval. For example, each received image sequence comprises six images acquired by the camera 11.
[0073] According to a particular embodiment, the acquired images are pre-filtered so that no image contains a pixel corresponding to a dynamic object, that is, an object moving within the three-dimensional scene observed by the camera at the time of image acquisition. Indeed, the presence of pixels corresponding to a dynamic object in the three-dimensional scene impairs learning, as a monocular vision system has difficulty estimating depth for such a pixel because a disparity or a determined optical flux between pixels of two images associated with the same object then comprises two components: a first component resulting from a movement of the camera 11 within the three-dimensional scene, that is, of the vehicle 10, and a second component resulting from an inherent movement of the dynamic object within the observed three-dimensional scene.
[0074] According to another example, the camera is always in motion between two consecutive acquired images. Indeed, to predict depths, the monocular vision system needs to be in motion so as to observe the three-dimensional scene from different viewpoints.
[0075] It should be noted that in the case where a sequence includes images containing pixels linked to dynamic objects or images acquired from the same viewpoint, it is possible to filter them using a filtering process or with the help of a technician. Indeed, sorting such a sequence does not require an analysis of each image in the image sequence but an analysis of the image sequence itself, an image sequence corresponding to a video.
[0076] For the remainder of this description, the acquired images are considered to be of the same resolution. Indeed, since they are acquired by the same camera 11, their resolution should not differ. However, if the resolution of images in the same sequence were different, it is possible to add an additional step consisting of resizing or cropping certain images to obtain a sequence of images comprising only images of the same resolution, that is to say comprising the same number of pixels according to their height and their width.
[0077] Each image sequence is classified according to the vehicle speed 10 during the acquisition of the images in the sequence: • in first class when the said speed of the vehicle 10 is below a threshold speed, and • in a second class when said vehicle speed 10 is greater than the threshold speed.
[0078] The speed of the vehicle 10 is for example associated with each image or each sequence of images, the speed being obtained for example from systems on board the vehicle 10, for example from an odometer or a geolocation system.
[0079] According to a particular embodiment, the threshold speed is determined as a function of a target distance and an image acquisition frequency by the camera 11. The threshold speed is then determined, for example, by the following function: Vscibie / acquisition threshold? aVCC . • vSeuii the threshold speed, • dCibie the target distance, and • ficquisition the image acquisition frequency by the camera 11.
[0080] Image sequences classified in the first class thus correspond to image sequences acquired at low speed, while image sequences classified in the second class correspond to image sequences acquired at high speed.
[0081] The aim of this learning process 3 is then to learn the depth prediction model from sequences of images acquired at low speed, that is, images acquired from closely spaced viewpoints, and then progressively to learn the depth prediction model from sequences of images acquired at higher speed, that is, images acquired from more distant viewpoints. The learning process 3 is thus divided into a first learning phase, followed by a second learning phase.
[0082] For each image sequence included in the first class, in a step 32, the depth prediction model is learned in a first learning from first pairs of images, each first pair of images comprising a first image and a second image selected from each image sequence.
[0083] Learning a depth prediction model from images is known to those skilled in the art. An example of a learning method 6 for the depth prediction model from a pair of images is described in particular below with reference to [Fig.6].
[0084] According to a first variant of the learning process 3, the most recent image in each image sequence is assigned to each first pair of images as being the first image and each image acquired prior to the first image is assigned in a first pair of images as being the second image, each second image being different in each first pair of images.
[0085] According to a particular embodiment of the first variant, the first learning comprises several iterations, each iteration using a different first pair of images corresponding to the first pair of images for which a time interval separating the acquisition of the first image and the second image is the smallest among the first pairs not used in a previous iteration, the depth prediction model being learned from the first pair of images used.
[0086] Fig. 5 schematically illustrates the generation of first pairs of images from a sequence of images acquired by the camera 11 mounted in the vehicle 10, according to a particular and non-limiting embodiment of the present invention.
[0087] A sequence of images 500 comprises, according to this particular embodiment, seven images acquired temporally in a precise order, the camera 11 having acquired each of the images of the sequence of images 500 in the following order: image 51, image 52, image 53, image 54, image 55, image 56 and lastly image 57. Image 57 is therefore the most recent image of the sequence of images 500, that is to say the last image acquired temporally in the sequence of images 500, while image 51 is the oldest image acquired in the sequence of images 500, that is to say the first image acquired temporally in the sequence of images 500.
[0088] During the different iterations, the prediction model is learned from initial pairs of images. • During the first iteration, the first pair of images 510 includes as the first image the most recent image corresponding to image 57 and the second image for which only one time interval separates its acquisition from the acquisition of the first image, here image 57, the second image then corresponding to image 56. • During the second iteration, the first pair of images 520 includes as the first image the most recent image, always corresponding to image 57, and the second image for which the smallest or minimum time interval separates its acquisition from the acquisition of the first image among the images of the sequence of images 500 not previously used, the second image then corresponding to image 55. Indeed, two time intervals separate the instant of acquisition of image 57 from the instant of acquisition of image 55 and image 56 was used during the previous iteration. • During the third iteration, a first pair of images 530 is formed and includes image 57 as the first image and image 54 as the second image. • During the fourth iteration, a first pair of images 540 is formed, comprising image 57 as the first image and image 53 as the second image. • During the fifth iteration, a first pair of images 550 is formed, comprising image 57 as the first image and image 52 as the second image. • During the sixth and final iteration, a first pair of images 560 is formed, comprising image 57 as the first image and image 51 as the second image.
[0089] The vehicle 10, traveling at a certain speed at the time of image acquisition for a sequence of images, and therefore the camera 11 mounted on the vehicle 10, has traveled a determined distance between each image acquisition point of the same image sequence.
[0090] Fig. 7 illustrates a graph representing a distance travelled 'd' by vehicle 10 as a function of its speed V10 and a number 'i' of time intervals between the acquisition of two images.
[0091] According to the illustrated example, the camera 11 acquires images at a frequency of 30 Hz (thirty Hertz), or thirty images per second. When the vehicle 10 is traveling at a speed of 15 km / h (fifteen kilometers per hour), or a speed of 4.2 m / s (four point two meters per second), it travels 14 cm (fourteen centimeters) between the viewpoint for acquiring one image and the viewpoint for acquiring the next consecutive image. Similarly, it travels 14 cm between the viewpoints for acquiring two other images. The vehicle has therefore traveled 28 cm (twenty-eight centimeters) between the viewpoints for acquiring three consecutive images, i.e., during two time intervals.Learning the depth prediction model from images separated by two time intervals is equivalent to learning the depth prediction model from images acquired consecutively, i.e., at a time interval, while vehicle 10 is traveling twice as fast.
[0092] Using similar reasoning, the vehicle travels 42 cm (forty-two centimeters) between image acquisition points at three time intervals. This distance of 42 cm is close to the distance of 46 cm (forty-six centimeters) traveled between two consecutive image acquisition points at a speed of 50 km / h (fifty kilometers per hour). Learning the depth prediction model with these images separated by three time intervals is equivalent to learning the depth prediction model for a vehicle traveling three times faster, i.e., a simulated 50 km / h versus an actual 15 km / h.
[0093] Similarly, the distance traveled by vehicle 10 between the acquisition of images acquired at two time intervals at a speed of 50 km / h is close to the distance traveled by vehicle 10 between the acquisition of two consecutive images at a speed of 100 km / h. A speed of 100 km / h is then simulated at based on images acquired at a speed of 50km / h.
[0094] Thus, during each iteration, the depth prediction model is learned with an increasing distance between the image acquisition viewpoints of the first pair of images. The depth prediction model is thus learned for a small distance between the image acquisition viewpoints, and then, when the depth prediction model is sufficiently accurate, a new iteration is implemented, and the previously learned depth prediction model continues its learning with a larger distance between the image acquisition viewpoints. Increasing this distance is equivalent to simulating a higher speed of the vehicle 10. Indeed, the distance traveled between the image acquisition viewpoints is doubled between the first and second iterations, then tripled between the first and third iterations, and so on.It is then possible to learn the depth prediction model by successive steps of distances and thus to achieve progressive learning with an increasing simulated speed. In this way, the learning of the depth prediction model is controlled and consistent over wide ranges of vehicle speeds 10.
[0095] According to a second variant of the learning process 3, the first pairs are generated by selecting the first image as corresponding to the most recent image in each image sequence and the second image as corresponding to an image prior to the first image in each image sequence, the second image being selected according to a number of time intervals separating the acquisition of the first image from the acquisition of the second image, the number of intervals being determined according to a determined probability law.
[0096] The probability law is, for example, obtained by the following function: [Math.l] 7 crdlû
[0097] With: * f(x) a probability associated with an interval denoted x, • J is a standard deviation.
[0098] The standard deviation a varies, for example, at each iteration of the first training, so as to define <72 as a function of iteration i, for example with <72 = i. If the number of images in a sequence of images is seven, at the first iteration <?2 = 1 tandis qu’à la dernière itération, <72 = 6.
[0099] A function generating a random result according to the probability law is for example defined in Python® language by the function numpy.random.normal, the Python® language being particularly used in the field of image processing.
[0100] Depending on a result of this function, a first pair of images is selected according to the number of time intervals separating the acquisition times of the first and second images of the first pair of images, this selection being determined by the following function: [Math.2] 2.1 < |s| <2 3.3 < 15'| S 4 4.4 < |.s| <5 5.5 < H < 6 6.151 >6
[0101] With : • x the number of time intervals calculated, and • $ the result of a function generating a random result according to the probability law.
[0102] The function presented above is obviously adapted to a sequence comprising seven images, and therefore up to six time intervals.
[0103] Such a probability distribution corresponds to a normal distribution centered on zero, which results in a higher probability of using a first pair of images when the first and second images of the first pair are acquired at a small number of time intervals, and a lower probability of using a first pair of images when the first and second images of the first pair are acquired at a large number of time intervals. Thus, compared to the pairs illustrated in [Fig. 5], the first pair of images 510 is more likely to be used for training the depth prediction model than the first pair 520, which itself is more likely to be used than the first pair 530. The first pair 560 is therefore the one with the lowest probability of being used.
[0104] Unlike the first variant, in this second variant not all possible pairs of images are exploited and the increase in the distance separating the image acquisition viewpoints of a second pair of images is not progressive.
[0105] The first learning phase ends when the depth prediction model is sufficiently reliable. This step 32 ends when the depth prediction model achieves, for example, the same performance as in the publication "Digging Into Self-Supervised Monocular Depth Estimation" by Clément Godard, Oisin Mac Aodha, Michael Firman, and Gabriel Brostow, published in August 2019. The performance is to be assessed qualitatively, such as by comparing maps. of depths generated for the same image, a relative error being for example less than 25%.
[0106] For each image sequence included in the second class, in a step 33, the depth prediction model is learned in a second learning from second pairs of images, each second pair of images comprising a third image and a fourth image selected from each image sequence.
[0107] The second training is thus similar to the first training, but with second pairs of images as input data, corresponding to images acquired at higher speeds. The high speed is therefore no longer simulated as before; the depth prediction model continues its training with real and representative data, for example, of extra-urban environments where the vehicle 10 actually travels at high speeds. Indeed, during the first training, the images are, for example, representative of a city center, but the simulated speed is greater than 100 km / h, which does not correspond to a truly observable driving situation.
[0108] The first and second learning stages correspond to the determination of input parameters for the convolutional neural network implementing the depth prediction model. Figure 6 illustrates a flowchart of different steps in such a learning process 6 of the depth prediction model from a pair of images used in the learning process 3 of the depth prediction model, according to a particular and non-limiting embodiment of the present invention.
[0109] According to the particular embodiment illustrated in [Fig.6], the depth prediction model is learned by minimizing a loss error determined from each pair of images of a plurality of pairs of images by comparing the images of each pair of images to images reconstructed from the images of each pair of images and depths predicted by the depth prediction model and associated with pixels of the images of each pair of images, the plurality of pairs of images corresponding to the first pairs of images during said first learning and to the second pairs of images during said second learning.
[0110] In a step 61, a pair of images comprising a fifth and a sixth image is received. The fifth image then corresponds to a first image during a first learning or to a third image during a second learning, the sixth image corresponds to a second image during a first learning or to a fourth image during a second learning.
[0111] In a step 62, depths are predicted by the pro prediction model founders for pixels of the fifth and sixth images from the fifth and sixth images.
[0112] The determination of the depth of a pixel of an image acquired by a moving camera from a depth prediction model implemented by a convolutional neural network is known to those skilled in the art, for example with monodepth2® described in the paper “Digging Into Self-Supervised Monocular Depth Estimation” written by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow and published in August 2019 or with an algorithm called NRS described in the paper “Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion” written by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich and Adrien Gaidon, published in August 2020.
[0113] In step 63, seventh and eighth images are generated. The seventh image is generated from the fifth image, the depths associated with the pixels of the fifth image, and extrinsic parameters of the moving monocular vision system, while the eighth image is generated from the sixth image, the depths associated with the pixels of the sixth image, and extrinsic parameters of the moving monocular vision system. This step 63 is commonly referred to as bilateral image reconstruction.
[0114] Generating a seventh image from a fifth image acquired by camera 11 consists of reprojecting a pixel from the fifth image acquired from a first viewpoint at the time of acquisition of the fifth image as a point, and then projecting this point onto the image plane of camera 11 at its position corresponding to a second viewpoint at the time of acquisition of the sixth image, so as to obtain an image corresponding to a view of the three-dimensional scene from the second viewpoint of camera 11. The image plane of camera 11 corresponds to a plane defined in the camera 11's frame of reference, normal to the optical axis of camera 11 and located at the focal length of camera 11. Thus, the seventh image generated from the fifth image is comparable to the sixth image. Similarly, the eighth image generated from the sixth image is comparable to the fifth image.
[0115] Since the first and second positions of the camera 11 at the acquisition times of the fifth and sixth images respectively are not coincident when the vehicle 10 is in motion, and since objects may obscure other objects in the scene, the generated images are not identical to the images acquired by the camera 11. Furthermore, the prediction of depth and / or the prediction of movement of the camera 11 between the first and second viewpoints is not error-free. Thus, comparing a generated image to an acquired image allows for the evaluation of the accuracy of the different models used: the prediction model of depth and / or camera motion prediction model 11.
[0116] According to a particular embodiment, the seventh and eighth images are respectively generated by the following function:
[0117] [Math.3] Ps = D(pt) ) ] )
[0118] With: • Ps the coordinates of a pixel in the seventh image, respectively in the eighth image, • 77 a function to convert from homogeneous coordinates to pixel coordinates by removing one dimension from a vector, • K is a direction prediction model associated with camera 11, • Let T be a matrix comprising the extrinsic parameters, and • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates in the fifth image, respectively sixth image, and of the depth j associated with it.
[0119] According to one variant, the direction prediction model corresponds to an intrinsic matrix of the camera 11. This variant is particularly applicable to pinhole and calibrated cameras.
[0120] According to another variant, the direction prediction model is a learned model implementing a second neural network. Such a direction prediction model is known to those skilled in the art; it is notably presented in the document "Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion." This other variant is suitable for wide-angle or fisheye cameras, or even for uncalibrated cameras. According to this variant, a direction prediction step associated with the pixels of the fifth image and the pixels of the sixth image is necessary.
[0121] The matrix comprising the extrinsic parameters of the moving camera 11 can be obtained in several ways. For example, from data emitted by a vehicle driver assistance system 10, called an ADAS system, or from sensors on board the vehicle 10. According to another example, the extrinsic parameters corresponding to the movement of the camera 11 between the time instants of acquisition of the fifth and sixth images are determined in an additional step by a motion prediction model, for example by the one presented in the document "Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion".
[0122] In a step 64, the depth prediction model is learned by minimizing a loss error.
[0123] The loss error includes, for example, a first photometric error determined by comparison of the fifth and eighth images and a second photometric error determined by comparison of the sixth and seventh images.
[0124] If the depth prediction model is accurate, then a first or second photometric error is small. Conversely, if the depth prediction model lacks accuracy, for example because its training is incomplete, then a first and / or second photometric error is not negligible.
[0125] According to a first particular embodiment, the first and second photometric errors are respectively determined by the following function: [Math.4] ■ U(p) ~Kp) I + a ■
[0126] With: • p) the first photometric error denoted Lp\ (p), respectively the second photometric error denoted Lp2(p), P being a pixel defined by its two-dimensional coordinates, • / (p) a value of pixel P in the fifth image, respectively in the sixth image, * a value of pixel P in the eighth image, respectively in the seventh picture, • SSIM is a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment in which the vehicle operates 10.
[0127] According to a second particular embodiment, the first and second photometric errors include a determination of a reconstruction error of a generated image determined by the following function: [Math.5]
[0128] With: • Lsmooth(p) the reconstruction error L^( p) for a pixel p of the seventh image, respectively the construction error L^(p) for a pixel p of the eighth image; • D(p) is a depth associated with pixel P in the seventh image, respectively in the eighth image; • W is a parameter matrix; • 0 is the order of a smoothing gradient; • an L1 norm of the second order depth gradients is calculated with W = 1, et0 = 2; • x and y are the dimensions of the seventh and eighth; • is a hyperparameter dependent on the environment in which the vehicle 10 is operating; and • It(p) is a colorimetric value of pixel P in the seventh image, respectively eighth image.
[0129] This second function is generally used to deal with the discontinuity at the edge of objects (in English "edge aware smoothness").
[0130] The photometric error is thus defined, for example, from the photometric errors and reconstruction errors previously defined.
[0131] According to the first particular embodiment, the photometric error associated with each first pair of blocks is determined by the following function: [Math.6] L = [ min(Lpl ( p ), Lp2 ( P ) ) ]
[0132] With: • Lp corresponding to the loss error, • Lpi(p) the first photometric error determined for a pixel P in the fifth image, the pixel P being defined by its two-dimensional coordinates, and • Lp2{p ) the second photometric error determined for pixel P in the sixth image.
[0133] According to the second particular embodiment, the photometric error associated with each first pair of blocks is determined by the following function: [Math.7] L = Ejmin(L pl ]>)' LX?) ) +L sl (p) ]
[0134] With: • L corresponding to the loss error, • Lp} ( p) the first photometric error determined for a pixel P in the fifth image, the pixel P being defined by its two-dimensional coordinates, and • Lp2 ( p ) the second photometric error determined for pixel P in the eighth image, • the reconstruction error for a pixel p of the seventh image, and • L^(p) the reconstruction error for a pixel p of the eighth image.
[0135] The training of the depth prediction model consists of adjusting the input parameters of the convolutional neural network in order to minimize the previously calculated loss error.
[0136] An object visible in the field of view of the camera 11 at the time of acquisition of the fifth image and masked in the field of view of the camera 11 at the time of acquisition of the sixth image or vice versa does not impact the loss error thanks to the min function, this process is therefore insensitive to occlusions.
[0137] Thus, the depth prediction model used for the depth prediction of a pixel of an image acquired by the camera 11 is made more reliable thanks to this learning process.
[0138] This learning process is performed using data acquired by the on-board vision system and therefore does not require data annotated by another on-board system or the storage of a training image library. Furthermore, the training data is representative of the data received when the system is in operation or in production; indeed, the training data is representative of real-world environments in which the vehicle carrying the monocular vision system operates or moves, making this training data particularly relevant.
[0139] Figure 4 schematically illustrates a device 4 configured to learn a depth prediction model associated with a vehicle-mounted vision system and / or to predict the depth associated with a pixel of an image acquired by a camera from the depth prediction model associated with it, according to a particular and non-limiting embodiment of the present invention. The device 4 corresponds, for example, to a device mounted in the first vehicle 10, for example, a computer associated with the stereoscopic vision system.
[0140] Device 4 is, for example, configured to carry out the operations described opposite Figures 1 and 4 and / or the steps described opposite Figures 2 and 3. Examples of such a device 4 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 4, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 4 may be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.
[0141] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41 corresponding for example to volatile and / or non-volatile memory and / or includes a memory storage device which may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.
[0142] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 41.
[0143] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.
[0144] According to a particular and non-limiting embodiment, the device 4 includes a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 include one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HDMI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").
[0145] According to another particular and non-limiting embodiment, the device 4 includes a communication interface 43 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 430. The communication interface 43 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds, for example, to a wired CAN (Controller Area Network) or CAN FD (Controller Area Network Flexible Data-Rate) type network. flexible data rate controllers”), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3).
[0146] According to a particular and non-limiting embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch or not, one or more speakers 450 and / or other peripherals 460 via the output interfaces 44, 45, 46 respectively. According to a variant, one or more of the external devices is integrated into the device 4.
[0147] Of course, the present invention is not limited to the embodiments described above but extends to a method for determining the depth of a pixel in an image acquired by a vision system, and / or for measuring the distance between an object and a vehicle equipped with a vision system, the depth and / or distance being predicted and / or measured via a depth prediction model learned according to the learning method described above, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0148] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 4 of [Fig.4].
Claims
Demands
1. A method for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle (10), the vision system comprising a camera (11) arranged to acquire an image of a three-dimensional scene of an environment external to the vehicle (10), said method being implemented by at least one processor, and being characterized in that it comprises the following steps: - receiving (31) image sequences, each image sequence comprising a set of images acquired by the camera (11) at different consecutive acquisition times separated by the same time interval,Each image sequence is classified according to the vehicle's speed (10) during image acquisition, into a first class when said vehicle speed (10) is less than a threshold speed and into a second class when said vehicle speed (10) is greater than the threshold speed; - for each image sequence in the first class, first training (32) of the depth prediction model from first pairs of images, each first pair of images comprising a first image and a second image selected from said image sequence; - for each image sequence in the second class, second training (33) of the depth prediction model from second pairs of images, each second pair of images comprising a third image and a fourth image selected from said image sequence.
2. A method according to claim 1, wherein the most recent image in said each image sequence is assigned in each first pair of images as being the first image and each image acquired prior to the first image is assigned in a first pair of images as being the second image, each second image being different in each first pair of images.
3. A method according to claim 2, wherein the first learning comprises several iterations, each iteration using a different first pair of images corresponding to the first pair of images for which a time interval separating the acquisition from the first image and the second image is the weakest among the first pairs not used in a previous iteration, the depth prediction model being learned from said first pair of images used.
4. A method according to claim 1, wherein said first pairs are generated by selecting said first image as corresponding to the most recent image in said each image sequence and said second image as corresponding to an image prior to the first image in said each image sequence, the second image being selected according to a number of time intervals separating the acquisition of the first image from the acquisition of the second image, said number of intervals being determined according to a determined probability law.
5. Method according to claim 4, wherein the probability law is obtained by the following function: with: • a probability associated with an interval denoted x, • has a standard deviation.
6. A method according to any one of claims 1 to 5, wherein the threshold speed is determined as a function of a target distance and an image acquisition frequency by the camera (11).
7. A method according to any one of claims 1 to 5, wherein the depth prediction model is learned by minimizing a loss error determined from each pair of images of a plurality of pairs of images by comparing the images of said each pair of images to images reconstructed from the images of said each pair of images and depths predicted by the depth prediction model and associated with pixels of the images of said each pair of images, said plurality of pairs of images corresponding to said first pairs of images during said first learning and to said second pairs of images during said second learning.
8. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by a processor.
9. Device (4) configured to learn a depth prediction model from a vision system embedded in a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for carrying out the steps of the process according to any one of claims 1 to 7.
10. Vehicle (10) comprising the device (4) according to claim 9.