Method and device for learning a depth prediction model from sorted image sequences.
The method improves depth prediction model training for monocular vision systems in vehicles by filtering images based on photometric differences and errors, enhancing accuracy and safety in ADAS systems.
Patent Information
- Application Number
- FR2024003857
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-15
- Publication Date
- 2025-10-17
AI Technical Summary
Existing monocular vision systems in vehicles struggle to accurately predict depths from images acquired when stationary, leading to unsatisfactory accuracy in ADAS systems due to the lack of effective training data from diverse road environments.
A method for learning a depth prediction model using a convolutional neural network that filters and selects images based on photometric differences and errors, ensuring images are acquired from different viewpoints, thereby improving the training process for moving vehicles.
Enhances the accuracy of depth prediction models by minimizing loss errors, making them reliable for dynamic environments and improving the operational safety of ADAS systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method and device for learning a depth prediction model from sorted image sequences. Technical field
[0001] The present invention relates to methods and devices for learning a depth prediction model associated with a stereoscopic vision system on board a vehicle, for example in a motor vehicle. The present invention also relates to a method and a device for determining a depth and / or measuring a distance separating an object from a vehicle carrying a vision system. Technological background
[0002] Many modern vehicles are equipped with so-called AD AS (Advanced Driver-Assistance System). Such AD AS systems are passive and active safety systems designed to eliminate the element of human error in the driving of vehicles of all types. AD AS use advanced technologies to assist the driver while driving and thus improve their performance. AD AS use a combination of sensor technologies to perceive the environment around a vehicle, then provide information to the driver or act on certain vehicle systems.
[0003] There are several levels of ADAS, such as rearview cameras and blind spot sensors, lane departure warning systems, adaptive cruise control and automatic parking systems.
[0004] The AD AS embedded in a vehicle are supplied with data obtained from one or more embedded sensors such as, for example, cameras. These cameras make it possible in particular to detect and locate other road users or possible obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle speed; - to act on the braking system in the event of a risk of impact with an object.
[0005] A position of another user or of an obstacle is for example determined by a vision system comprising a model for predicting a depth associated with a pixel or a distance separating the vision system from an object in a three-dimensional scene. Such a model is for example learned using images, these images being obtained from a universal database, for example Kitti® or Sceneflow®. Kitti®, for example, provides images of a city-center road environment, but such a database does not include all the road environments in which a vehicle can operate. The training data is therefore unsuitable for training the prediction model for a vehicle traveling in other road environments.
[0006] The quality of the training of the depth or distance prediction model is however very important, in fact, the depths or distances predicted by the prediction model represent for example distances at which the other users or obstacles present in the road environment of the vehicle carrying the vision system and the AD AS are located. The proper functioning of the driving assistance peripherals using this data therefore depends on the quality of the data emitted by the vision system.
[0007] A vision system comprising a single camera, hereinafter called a monocular vision system, uses several consecutive images to predict depths associated with pixels of one of the images. This monocular vision system has the defect of not being able to accurately predict depths associated with pixels of an image from images acquired from the same point of view, that is to say acquired when the monocular vision system is not in motion. Indeed, when this monocular vision system is on board a vehicle, the accuracy of the depths predicted by this onboard monocular vision system is not satisfactory if the vehicle is stationary. It is therefore appropriate to learn the depth prediction model associated with this monocular vision system from sequences of images acquired when the vehicle carrying it is in motion. Summary of the present invention
[0008] An object of the present invention is to solve at least one of the problems of the technological background described above.
[0009] Another object of the present invention is to improve the learning phase of a depth prediction model from images acquired by a vision system moving in a three-dimensional scene, the depth prediction model being in particular implemented by a neural network and associated with a vision system embedded in a vehicle.
[0010] Another object of the present invention is to improve road safety, in particular by improving the operational safety of AD AS systems supplied by data obtained from a camera of a vision system.
[0011] According to a first aspect, the present invention relates to a method for learning a depth prediction model implemented by a convolutional neural network associated with a vision system on board a vehicle, the system vision comprising a camera arranged to acquire an image of a three-dimensional scene of an environment outside the vehicle, the method being implemented by at least one processor, and being characterized in that it comprises the following steps: - receiving a batch of images acquired by the camera comprising a set of images comprising at least one sequence of images, each sequence of images of the set of images comprising a source image and at least one reference image, each reference image being acquired at an acquisition time instant determined relative to an acquisition time instant of the source image, - for each reference image of each image sequence: • determination of a photometric difference associated with pixels of the reference image by comparing the reference image to the source image of the image sequence; • selecting a set of pixels in the reference image, each pixel in the set of pixels having an associated photometric difference greater than a first threshold value; • deletion of the reference image in the batch of images if the set of pixels includes a number of pixels lower than a second threshold value; - for each undeleted reference image of each image sequence, determination of a photometric error associated with the undeleted reference image from: • a first photometric error determined by comparing pixels of a first image to pixels of the undeleted reference image, the first image being generated from the source image and first depths associated with pixels of the source image, and • a second photometric error determined by comparing pixels of a second image to pixels of the source image, the second image being generated from the undeleted reference image and second depths associated with pixels of the undeleted reference image, the first and second depths being predicted by the depth prediction model from the unsuppressed source and reference images; - learning the depth prediction model by minimizing a loss error determined from: • photometric errors associated with undeleted reference images in the image batch, and • a number of undeleted reference images included in the batch of images.
[0012] This learning method thus makes it possible to delete reference images if related to a source image, that is to say acquired from the same point of view in the three-dimensional scene. Indeed, the training of the depth prediction model is not effective from images acquired by a stationary or static camera in the three-dimensional scene. The first threshold value makes it possible in particular to estimate a photometric equivalence between pixels of two images and the second threshold value makes it possible to qualify two images as sufficiently similar to be acquired from the same point of view.
[0013] According to a variant of the learning method, the photometric error associated with the reference image is determined by the following function: [ min(Lp ( ( p ), Lp2 ( p ) ) ] With : • Lp corresponding to the photometric error associated with the undeleted reference image noted Ref -, • Lp\( p) the first photometric error determined for a pixel P in the reference image, the pixel P being defined by its two-dimensional coordinates, and • Lp2( p) the second photometric error determined for pixel P in the source image of the image sequence including the undeleted reference image.
[0014] The fact of retaining only the minimum value among the first and second photometric errors makes it possible in particular not to impact the learning when an object of the three-dimensional scene is occluded in one of the images, that is to say when one image comprises pixels corresponding to this object and the other image does not comprise pixels corresponding to this same object.
[0015] According to another variant of the learning method, the loss error is equal to an average of the photometric error determined for each undeleted reference image.
[0016] According to yet another variant of the learning method, the photometric difference associated with the reference image is defined by the following function: D(Ref\ p)=\I Ref (p)- ISource 0)1 With : • D corresponding to the photometric difference associated with a pixel P of the undeleted reference image, the reference image being noted Rej, the pixel P being defined by its coordinates in two dimensions, • IRef(p) a colorimetric value determined for a pixel P in the reference image, and * ISource(p) a colorimetric value determined for pixel P in the source image of the image sequence comprising the reference image.
[0017] According to a further variant of the learning method, the first, respectively the second, photometric error is determined by the following function: Lp*= Lp(la) ■ \I(p)-l(p)\+a-(l-^SSIM(l(p)J(p))) With : • Lp* the first photometric error noted Lpi ( p ), respectively the second photometric error noted Lp2^p^ ■> • I(p) a colorimetric value of pixel P in the reference image, respectively in the source image, P being a pixel defined by its coordinates in two dimensions, • a value of pixel P in the first image, respectively in the second picture, • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment in which the vehicle operates.
[0018] According to an additional variant of the learning method, the first image, respectively second image, is generated by the following function: P s= 71 ( K [( pX1' D ( Pt) ) ] ) With : • P s the coordinates of a pixel of the first image, respectively of the second image, • 77 a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • K a direction prediction model associated with the camera, • T a matrix comprising the extrinsic parameters, and • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates in the source image, respectively in the reference image, and of the depth associated with it.
[0019] According to another variant of the learning method, the first threshold is between 1 and 2 when a colorimetric value of a pixel is coded on eight bits.
[0020] According to yet another variant of the learning method, the second threshold is between 60% and 90%.
[0021] According to a second aspect, the present invention relates to a device configured to learn a depth prediction model by a vision system on board a vehicle, the device comprising a memory associated with at least one processor configured to implement the steps of the method according to the first aspect of the present invention.
[0022] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0023] According to a fourth aspect, the present invention relates to a computer program which comprises instructions adapted for executing the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0024] Such a computer program may use any programming language and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0025] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to the first aspect of the present invention.
[0026] On the one hand, the recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM memory, a CD-ROM or a microelectronic circuit type ROM memory, or a magnetic recording means or a hard disk.
[0027] Furthermore, this recording medium may also be a transmissible medium such as an electrical or optical signal, such a signal being able to be conveyed via an electrical or optical cable, by conventional or hertzian radio or by self-directed laser beam or by other means. The computer program according to the present invention may in particular be downloaded from an Internet-type network.
[0028] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to perform or to be used in performing the method in question. Brief description of the figures
[0029] Other characteristics and advantages of the present invention will emerge from the description of the particular and non-limiting exemplary embodiments of the present invention below, with reference to the appended figures 1 to 4, in which:
[0030] [Fig-1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting example of embodiment of the present invention;
[0031] [Fig.2] illustrates a flowchart of the different steps of a method for determining a depth of a pixel of an image by a depth prediction model associated with a vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;
[0032] [Fig.3] illustrates a flowchart of the different steps of a method for learning the depth prediction model used in the method of [Fig.2], according to a particular and non-limiting example of the present invention; and
[0033] [Fig.4] schematically illustrates a device configured to learn a depth prediction model by a vision system on board the vehicle of [Fig.l], according to a particular and non-limiting example of the present invention. Description of examples of implementation
[0034] A method and a device for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle will now be described in what follows with joint reference to Figures 1 to 4. The same elements are identified with the same reference signs throughout the description which follows.
[0035] The terms "first(s)", "second(s)" (or "first(s)", "second(s)"), etc. are used in this document by arbitrary convention to enable different elements (such as operations, means, etc.) implemented in the embodiments described below to be identified and distinguished. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0036] For the entire description, reception of an image or a batch of images is understood to mean the reception of data representative of an image or a batch of images. Generally speaking, any reception, determination or generation of an object or a value amounts to the reception, determination or generation of data representative of this object or value. These shortcuts are only intended to simplify the description, however, the methods described below being implemented by one or more processors, it is obvious that the input and output data of the different steps of a method are computer data.
[0037] According to a particular and non-limiting example of embodiment of the present invention, the depth prediction model is learned in a learning phase from a batch of images acquired by a camera of the vision system associated with the depth prediction model. The batch of images comprises at least one sequence of images, each sequence of images comprising a source image and at least one reference image acquired at an acquisition time instant determined relative to an acquisition time instant of the source image.
[0038] The reference images are first filtered. For each reference image of an image sequence, a photometric difference associated with pixels of the reference image is determined by comparing the reference image to the source image of the image sequence and a set of pixels in the reference image is selected, each pixel in the set of pixels having an associated photometric difference greater than a first threshold value. The image is then deleted if the set of pixels comprises a number of pixels less than a second value. threshold.
[0039] For each unsuppressed reference image of the image sequence, a photometric error associated with the unsuppressed reference image is determined from a first and a second photometric error determined by comparing pixels of the unsuppressed and source reference images to a first and a second image, respectively, the first and second images being generated from the source image, respectively the unsuppressed reference image, and depths predicted for these unsuppressed and source reference images by the depth prediction model.
[0040] The depth prediction model is then learned by minimizing a loss error determined from the photometric errors associated with the undeleted reference images in said batch of images and a number of undeleted reference images included in the batch of images.
[0041] [Fig. 1] schematically illustrates a vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.
[0042] The vehicle 10 is located in an environment 1 corresponding, for example, to a road environment formed of a network of roads accessible to the vehicle 10.
[0043] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, with an electric motor(s) or even a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as an automobile, a truck, a bus, a motorcycle. Finally, the vehicle 10 corresponds to an autonomous vehicle or not, that is to say a vehicle traveling according to a determined level of autonomy or under the total supervision of the driver.
[0044] The vehicle 10 advantageously comprises at least one on-board camera 11, configured to acquire images of a three-dimensional scene taking place in the environment of the vehicle 10 from a current observation position. The camera 11 forms a monocular vision system if it is used alone as illustrated in [Fig.l]. The present invention is however not limited to a monocular vision system comprising a single camera but extends to any vision system comprising at least one camera, for example 1, 2, 3 or 5 cameras.
[0045] The camera 11 has intrinsic parameters, in particular: - a focal length, - distortions which are due to imperfections in the optical system of the camera 11, - a direction of the optical axis of the camera 11, and - a resolution.
[0046] The intrinsic parameters characterize the transformation which associates, for an image point, subsequently called “point”, its three-dimensional coordinates in the frame of reference of the camera 11 with the pixel coordinates in an image acquired by the camera 11. These settings do not change if you move the camera.
[0047] The distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions which generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, that is to say that the distortions are not taken into account, that their correction is processed at the time of image acquisition or at the time of calibration.
[0048] The camera 11 is arranged so as to acquire an image of a three-dimensional scene according to a defined point of view, the point of view is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10 as illustrated in [Fig.l].
[0049] The camera 11 for example acquires images of a three-dimensional scene located in front of the vehicle 10, the first camera 11 covering an acquisition field 12. An object 13 is placed in the acquisition field 12 of the camera 11, defining an occlusion field 14 for the vision system, an object present in the occlusion field 14 not being observable by the camera 11 from its current observation position.
[0050] It is obvious that it is possible to use such a vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently, the invention not being limited to the observation of a three-dimensional scene taking place in front of the vehicle 10 carrying the camera 11.
[0051] According to a particular embodiment, the camera 11 is of the “wide-angle” type, a wide-angle camera being for example equipped with a lens designed to acquire an image representative of a three-dimensional scene perceived according to a wider field of vision than that of a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens makes it possible to capture a larger portion of the three-dimensional scene taking place in front of or around the camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by this camera. The angle a of the field of vision of the camera 11 is for example equal to 120°, 145°, 180° or 360°, whereas a standard camera offers, for example, an open field of vision following an angle of 45° or less.Such a camera 11 corresponds for example to a camera equipped with mirrors or even to a “fisheye” camera (in French “oeil . fisheye lens"). Wide-angle lenses have a shorter focal length than standard lenses, making them suitable for capturing images of landscapes, architecture, road intersections, or any other subject requiring an extended perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with an extended depth of field.
[0052] An image acquired by the camera 11 at an acquisition time instant is presented in the form of data representing pixels characterized by: - coordinates in the image; and - data relating to the colors and brightness of objects in the observed scene in the form, for example, of RGB colorimetric values (from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness).
[0053] Each pixel of the acquired image is representative of an object in the three-dimensional scene present in the field of vision of the camera 11. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a luminous point resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When the light strikes this object, photons are emitted or reflected, captured by a photosensitive sensor of the camera 11 after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image faithfully representing the physical object observed by the camera 11. An image point previously presented is thus a point on a surface of an object in the three-dimensional scene.
[0054] When the vehicle 10 is in motion, then two images acquired by the camera 11 at two distinct time instants represent views of the same three-dimensional scene taken from different viewpoints or observation positions, the observation positions of the camera 11 being distinct. On this three-dimensional scene are for example: - buildings; - road infrastructure; - other users or stationary objects, for example a parked vehicle; and / or - other users or moving objects, for example another vehicle, a cyclist or a moving pedestrian.
[0055] According to a particular exemplary embodiment, an image acquired by the camera 11 comprises a distortion equal to 0.5%, 0.8% or greater than 1%. The measurement of such a distortion corresponds to the determination of a ratio between: - the maximum distance of a pixel from the image to a straight line in the first scene three-dimensional whose image is a line touching the longest edge of the first image, either at the center of the image edge or at the corners of the image edge, and - the length of this edge.
[0056] Commonly, distortion is considered, in the world of photography, as: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • annoying if it is greater than or equal to 1% or more.
[0057] A barrel distortion is characterized by a positive percentage, while a crescent distortion is characterized by a negative percentage.
[0058] Each image acquired by the camera 11 is for example sent to a processor, for example a computer of a device equipping the vehicle 10, or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10. It is then used during the implementation of a method for determining a depth of one of its pixels and / or during the implementation of a method for learning a depth prediction model associated with this camera of the vision system.
[0059] [Fig. 2] illustrates a flowchart of the different steps of a method 2 for determining a depth of a pixel of an image by depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle, for example in the vehicle 10 of [Fig. 1], according to a particular and non-limiting exemplary embodiment of the present invention. The method 2 is for example implemented by a device of the vision system embedded in the vehicle 10 or by the device 4 of [Fig. 4].
[0060] In a step 21, at least one image acquired by the camera 11 is received.
[0061] In a step 22, depths associated with a set of pixels of an image received are determined by the depth prediction model from the at least one received image.
[0062] Each determined depth then corresponds to a distance separating the vehicle 10 or a part of the vehicle 10 from an object of the three-dimensional scene with which a pixel is associated, the determination of a depth of a pixel then corresponding to a measurement of a distance separating an object from the vehicle carrying the vision system.
[0063] If the ADAS uses these depths or distances as input data to determine the distance between a part of the vehicle 10, for example the front bumper, and another user present on the road, the ADAS is then able to precisely determine this distance. For example, if the ADAS has the function of acting on a braking system of the vehicle 10 in the event of a risk of collision with another road user and the distance separating the vehicle 10 from this same road user decreases significantly, then the ADAS is able to detect this sudden approach and act on the braking system of the vehicle 10 to avoid a possible accident.
[0064] [Fig. 3] illustrates a flowchart of the different steps of a method for learning the depth prediction model used in a method for determining a depth of a pixel of an image, for example in method 2 of [Fig. 2], according to a particular and non-limiting exemplary embodiment of the present invention.
[0065] The learning method 3 is for example implemented by the device on board the vehicle 10 implementing the method for determining a depth for the vision system on board a vehicle 10 or by the device 4 of [Fig.4].
[0066] In a step 31, a batch of images acquired by the camera 11 is received. The batch of images comprises a set of images comprising at least one sequence of images, each sequence of images of this set of images comprising a source image and at least one reference image, each reference image being acquired at an acquisition time instant determined relative to an acquisition time instant of the source image included in the same sequence as the reference image.
[0067] According to a particular exemplary embodiment, each sequence of images is homogeneous and comprises a source image and a determined number of reference images, the images of the sequence forming an ordered sequence of images acquired successively by the camera 11 and whose time interval separating the acquisition of two successive images is constant, the frequency of acquisition of images by the camera 11 being fixed. The position of the source image in the sequence of images is in particular predefined. Thus, a sequence of images comprises, for example, five images acquired successively by the camera 11 and acquired in the following manner: • a source image acquired at a time instant t, • a reference image acquired at two time intervals preceding the time instant t, i.e. at t-2i, i corresponding to the time interval previously defined, • a reference image acquired at a time interval preceding the time instant t, i.e. at ti, • a reference image acquired at a time interval following the time instant t, i.e. at t+i, and • a reference image acquired at two time intervals following the instant temporal t, or at t+2i i being a previously defined time interval. Obviously the invention is not limited to this type of sequence but to any sequence comprising a source image and at least one reference image, each sequence not necessarily being homogeneous.
[0068] According to a particular exemplary embodiment, the acquired images are filtered beforehand so that no image comprises a pixel corresponding to a dynamic object, that is to say to an object in motion in the three-dimensional scene observed by the camera 11 at the time instant of acquisition of the image. Indeed, the presence of pixels corresponding to a dynamic object in the three-dimensional scene deteriorates the learning, a monocular vision system having difficulties in estimating a depth for such a pixel because a disparity or a determined optical flow between pixels of two images associated with this same object then comprises two components, a first component resulting from a movement of the camera 11 in the three-dimensional scene, that is to say of the vehicle 10, and a second component resulting from a proper movement of the dynamic object in the observed three-dimensional scene.
[0069] Note that in the case where a sequence includes images comprising pixels linked to dynamic objects or images acquired from the same point of view, it is possible to filter them with a filtering method or with the help of a technician. Indeed, the sorting of such a sequence does not require an analysis of each image of the sequence of images but an analysis of the sequence of images itself, a sequence of images corresponding to a video.
[0070] For the remainder of the description, the acquired images are considered to be of the same definition. Indeed, being acquired by the same camera 11, their definition does not have to be different. However, in the case where the definition of images of the same sequence would be different, it is possible to add an additional step consisting of resizing or cropping certain images to obtain a sequence of images comprising only images of the same definition, that is to say comprising the same number of pixels according to their height and according to their width.
[0071] For each reference image of each image sequence, in a step 32, a photometric difference associated with pixels of the reference image is determined by comparing the reference image to the source image of the image sequence comprising the reference image.
[0072] The photometric difference associated with a reference image is, for example, defined by the following function: [Math.l] D(Ref \ P) = j IRef ( P ) - ISource ( P ) I
[0073] With: • D corresponding to the photometric difference associated with a pixel P of the undeleted reference image, the reference image being noted Ref, the pixel P being defined by its coordinates in two dimensions, * lRef(p ) a colorimetric value determined for a pixel P in the reference image, and * ISource(P) a colorimetric value determined for pixel P in the source image of the image sequence comprising the reference image.
[0074] Still for each reference image of each sequence of images, in a step 33, a set of pixels is selected in the reference image, each pixel of the set of pixels having an associated photometric difference greater than a first threshold value.
[0075] This step of selecting a set of pixels amounts to determining a mask of pixels varying from the source image to the reference image, the pixels being defined by their position or coordinates in two dimensions and which are identical in the source image and in the reference image.
[0076] The aim is to determine whether the reference image is acquired from the same viewpoint as the source image. An image acquired from the same viewpoint as another image and in which pixels correspond to fixed objects in the observed three-dimensional scene then comprises pixels having the same colorimetric value as the pixels in the other image. If a large number of pixels in the image are similar to pixels in the other image, then it is highly probable that the image was acquired from the same viewpoint as the other image. However, to ensure proper learning of the depth prediction model, it is necessary to only consider images acquired from different viewpoints.
[0077] According to a particular embodiment, the first threshold is between 1 and 2 when a colorimetric value of a pixel is coded on eight bits. An eight-bit coding makes it possible to obtain a colorimetric value between 0 and 255, a first threshold close to 1.5 is then particularly suitable for comparing two images acquired from the same point of view with a duration separating the temporal instants of acquisition of the two images that is very short, for example with a duration of 1 / 30th of a second corresponding to two images acquired consecutively by a camera acquiring images at a frequency of thirty Hertz (30Hz). Indeed, with such an image acquisition frequency, a variation in lighting of the three-dimensional scene is not perceptible.
[0078] In order not to consider the reference images of a sequence of images acquired from the same point of view as the source image included in this same sequence of images, in a step 34, the reference image is deleted in the batch of images if the pixel set includes a number of pixels less than a second threshold value, in other words if it is highly likely that this reference image was acquired from the same viewpoint as the reference image.
[0079] According to a particular exemplary embodiment, the second threshold is between 60% and 90%. The second threshold is for example between 60% and 70% for an image having a large width compared to its height, which is for example the case of a camera observing a wide scene, which is suitable for use in an urban environment. According to another example, the second threshold is between 80% and 90% if the images have a similar width and height, which is for example the case of a camera observing a distant scene in front of the vehicle 10, which is suitable for use on a fast lane or motorway.
[0080] The rest of the learning process therefore only considers the reference images of the batch of images which were not deleted during step 34.
[0081] For each undeleted reference image of an image sequence, in a step 35, a photometric error associated with the undeleted reference image is determined from: • a first photometric error determined by comparing pixels of a first image to pixels of the undeleted reference image, the first image being generated from the source image of the sequence of images comprising the reference image and first depths associated with pixels of the source image, and • a second photometric error determined by comparing pixels of a second image to pixels of the source image, the second image being generated from the undeleted reference image and second depths associated with pixels of the undeleted reference image, the first and second depths being predicted by the depth prediction model from the unsuppressed source and reference images.
[0082] Determining a depth of a pixel of an image acquired by a moving camera from a depth prediction model implemented by a convolutional neural network is known to those skilled in the art, for example with monodepth2® described in the document “Digging Into Self-Supervised Monocular Depth Estimation” written by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow and published in August 2019 or with an algorithm called NRS described in the document “Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion” written by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich and Adrien Gaidon, published in August 2020.
[0083] It should be noted that the depth prediction by the depth prediction model associated with the monocular vision system, i.e. comprising only the camera 11, is accurate for a pixel corresponding to a static object, but is much less precise for a pixel corresponding to a dynamic object.
[0084] The generation of a first image from a source image acquired by the camera 11 consists of reprojecting a pixel of the source image acquired at a first time instant into the three-dimensional scene in the form of a point, then projecting this point into the image plane of the camera 11 at its position at a second acquisition time instant corresponding to the acquisition time instant of the reference image, so as to obtain an image corresponding to a view of the three-dimensional scene from the viewpoint of the camera 11 at the second acquisition time instant. The image plane of the camera 11 corresponds to a plane defined in the frame of reference of the camera 11, normal to the optical axis of the camera 11 and located at the focal length of the camera 11. Thus, the first image generated from the source image is comparable to the reference image. Similarly, the second image generated from the reference image is comparable to the source image. 。
[0085] Since the positions of the camera 11 at the first and second acquisition time instants are not the same when the vehicle 10 is moving and since objects may mask other objects in the scene or even move between the two acquisition time instants, the generated images are not identical to the images acquired by the camera 11. Furthermore, the prediction of the depths including the prediction of a movement of the camera 11 between the first and second time instants, just like the models used to generate the images, are not error-free. Thus, the comparison of a generated image with an acquired image makes it possible to evaluate the relevance of the different models used and the predicted depth values.
[0086] According to a particular exemplary embodiment, the first image, respectively second image, is generated by the following function:
[0087] [Math.2] Ps = Æ(^[ T^PiK^ DM ) ] )
[0088] With: • Ps the coordinates of a pixel of the first image, respectively of the second image, • 71 a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector, • K a direction prediction model associated with camera 11, • T a matrix comprising the extrinsic parameters, and • 0 a projection function in the three-dimensional scene of a pixel Pi as a function of its coordinates in the source image, respectively in the reference image, and of the depth [){n ) associated with it.
[0089] According to a first variant, the direction prediction model corresponds to an intrinsic matrix of the camera 11.
[0090] This first variant is particularly applicable to pinhole and calibrated cameras.
[0091] According to a second variant, the direction prediction model corresponds to a learned model implementing a second neural network. Such a direction prediction model is known to those skilled in the art; it is notably presented in the document “Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion”.
[0092] This second variant is suitable for wide-angle or “fish-eye” type cameras or even for uncalibrated cameras.
[0093] According to this second variant, a direction prediction step associated with the pixels of the first block and the pixels of the second block is necessary.
[0094] The matrix comprising the extrinsic parameters of the moving camera 11 can be obtained in several ways. For example, from data emitted by a driving assistance system of the vehicle 10, called AD AS system, or from sensors embedded in the vehicle 10. According to another example, the extrinsic parameters corresponding to the movement of the camera 11 between the first and second time instants are determined in an additional step by a motion prediction model, for example by that presented in the document “Neural Ray Surfaces for Self-Supervised Learning of Depth and Ego-motion”.
[0095] A first and / or a second photometric error is all the smaller the more precise the depth prediction model is, that is to say the more effective its training is. Conversely, if the depth prediction model lacks precision, for example because its training is not successful, then a first and / or a second photometric error is not negligible.
[0096] According to a first particular embodiment, the first and second photometric errors are respectively determined by the following function: [Math.3] L p 4p) = (l-«) ' V(p)-Up)\+a- I(p)))
[0097] With: • Lp^p} the first photometric error noted Lp\( p), respectively the second photometric error noted Lp2(p), P being a pixel defined by its coordinates in two dimensions, • Kp) a value of the pixel P in the undeleted reference image, respectively in the source image, • a value of the pixel P in the first image, respectively in the second image, • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment in which the vehicle 10 operates.
[0098] According to a second particular embodiment, the first and second photometric errors comprise a determination of a reconstruction error of a generated image, the reconstruction error being determined by the following function: [Math.4] IP))
[0099] With: • Lw p) the reconstruction error, noted for a pixel P of the first image, respectively the construction error for a pixel P of the second picture , • D(p) is a depth associated with pixel P in the first image, respectively in the second image; • H7 is a parameter matrix; • 0 is the order of a smoothing gradient; • an L1 norm of second-order depth gradients is calculated with W =1, and ° =2; • x and - ' are the dimensions of the first and second images; • / 3 is a hyperparameter dependent on the environment in which the vehicle 10 operates; and • It(j) ) is a colorimetric value of the pixel P in the first image, respectively second image.
[0100] This second function is generally used to deal with discontinuity at the edge of objects (in English “edge aware smoothness”).
[0101] The photometric error is thus defined, for example, from the photometric errors and reconstruction errors previously defined.
[0102] According to the first particular embodiment, the photometric error associated with each undeleted reference image is determined by the following function: [Math.5] L^Ref) = [min(Lp^p), Lp2( / ?))]
[0103] With: • Lp corresponding to the photometric error associated with the undeleted reference image noted Ref, • Lpi(p} the first photometric error determined for a pixel P in the image of reference, the pixel P being defined by its coordinates in two dimensions, and • Lp2(p) the second photometric error determined for the pixel P in the source image of the image sequence comprising the undeleted reference image.
[0104] According to the second particular embodiment, the photometric error associated with each undeleted reference image is determined by the following function: [Math.6] L^Ref ) = ^p [ min ( LpX ( p ), Lp2 ( p ))+Lsl(p) + L,2 ( p ) ]
[0105] With: • Lp corresponding to the photometric error associated with the undeleted reference image noted Ref, • Lpfp) the first photometric error determined for a pixel P in the reference image, the pixel P being defined by its two-dimensional coordinates, • Lp2(p) the second photometric error determined for pixel P in the source image of the image sequence including the undeleted reference image, • Ls{(p) the reconstruction error associated with the first image, and • L^ ( p) the reconstruction error associated with the second image.
[0106] In a step 36, the depth prediction model is learned by minimizing a loss error determined from: • photometric errors associated with undeleted reference images in the image batch, and • a number of undeleted reference images included in the batch of images.
[0107] The loss error is, for example, equal to an average of the photometric error determined for each undeleted reference image. It is for example determined by the following function: [Math.7] ,.....RW) “ * Ref
[0108] With: • L corresponding to the loss error, • Lp corresponding to the photometric error associated with the undeleted reference image noted Ref, and • Ref corresponding to the number of undeleted reference images in the batch of images.
[0109] Training the depth prediction model consists of adjusting input parameters of the convolutional neural network implementing the depth prediction model in order to minimize the previously calculated loss error.
[0110] According to a particular exemplary embodiment, if the loss error determined for the batch of images is not zero, a gradient descent algorithm is used. Such an algorithm is known to those skilled in the art and is, for example, implemented using the loss.backward() function in Pytorch®.
[0111] An object visible in the field of vision of the camera 11 at the first time instant of acquisition and masked in the field of vision of the camera 11 at the second time instant of acquisition or vice versa does not impact the loss error thanks to the min function, this method is therefore insensitive to occlusions.
[0112] The learning is only carried out from comparing images acquired from different viewpoints and for which the first or second photometric error associated with pixels corresponding to a dynamic object is negligible. Such a learning method thus makes it possible to learn a depth prediction model from sequences of images acquired by a camera of the vision system, the camera being in motion.
[0113] Thus, the depth prediction model used for the depth prediction of a pixel of an image acquired by a camera of the vision system is made reliable thanks to this learning method.
[0114] This learning is carried out from data acquired by the on-board vision system and therefore does not require data annotated by another on-board system or storage of a library of learning images. In addition, the learning data are representative of the data received when the system is in operation or in production, in fact the learning data are representative of real environments in which the vehicle carrying the vision system evolves, or moves, these learning data are therefore particularly relevant because they are representative of real conditions for implementing a depth prediction method.
[0115] [Fig.4] schematically illustrates a device 4 configured to learn a depth prediction model by a vision system embedded in a vehicle and / or for predicting a depth associated with a pixel of an image acquired by a camera, according to a particular and non-limiting exemplary embodiment of the present invention. The device 4 corresponds for example to a device embedded in the first vehicle 10, for example a computer associated with the stereoscopic vision system.
[0116] The device 4 is for example configured for the implementation of the operations described with regard to figures 1 and 4 and / or steps described with regard to figures 2 and 3. Examples of such a device 4 include, but are not limited to, on-board electronic equipment such as an on-board computer of a vehicle, an electronic calculator such as an ECU (“Electronic Control Unit”), a smartphone, a tablet, a laptop. The elements of the device 4, individual dually or in combination, can be integrated in a single integrated circuit, in several integrated circuits, and / or in discrete components. The device 4 can be produced in the form of electronic circuits or software (or computer) modules or even a combination of electronic circuits and software modules.
[0117] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the method and / or for executing the instructions of the software(s) embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41 corresponding for example to a volatile and / or non-volatile memory and / or comprises a memory storage device which may comprise volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.
[0118] The computer code of the embedded software(s) comprising the instructions to be loaded and executed by the processor is for example stored in the 4L memory.
[0119] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Telematic Control Unit”), for example via a communication bus or through dedicated input / output ports.
[0120] According to a particular and non-limiting exemplary embodiment, the device 4 comprises a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 comprise one or more of the following interfaces: - RF radio frequency interface, for example Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English “Universal Serial Bus” or “Universal Serial Bus” in French); HDMI interface (from the English “High Definition Multimedia Interface” or “High Definition Multimedia Interface” in French); - LIN interface (from the English “Local Interconnect Network”, or in French “Réseau local interconnected”).
[0121] According to another particular and non-limiting exemplary embodiment, the device 4 comprises a communication interface 43 which makes it possible to establish communication with other devices (such as other computers of the on-board system) via a communication channel 430. The communication interface 43 corresponds for example to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds for example to a wired network of the CAN (Controller Area Network) type, CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by the ISO 17458 standard) or Ethernet (standardized by the ISO / IEC 802-3 standard).
[0122] According to a particular and non-limiting exemplary embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch-sensitive or not, one or more speakers 450 and / or other peripherals 460 via the output interfaces 44, 45, 46 respectively. According to a variant, one or other of the external devices is integrated into the device 4.
[0123] Of course, the present invention is not limited to the exemplary embodiments described above but extends to a method for determining the depth of a pixel of an image acquired by a vision system, and / or for measuring a distance separating an object from a vehicle carrying a vision system, the depth and / or the distance being predicted and / or measured via a depth prediction model learned according to the learning method described above, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0124] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-based motor vehicle, comprising the device 4 of [Fig.4].
Claims
Claims
1. Method for learning a depth prediction model implemented by a convolutional neural network associated with a vision system embedded in a vehicle (10), the vision system comprising a camera (11) arranged to acquire an image of a three-dimensional scene of an environment outside the vehicle (10), said method being implemented by at least one processor, and being characterized in that it comprises the following steps: - reception (31) of a batch of images acquired by the camera (11) comprising a set of images comprising at least one sequence of images, each sequence of images of said set of images comprising a source image and at least one reference image, each reference image being acquired at a time instant of acquisition determined relative to a time instant of acquisition of said source image, - for each reference image of said each sequence of images: • determination (32) of a photometric difference associated with pixels of said reference image by comparison of said reference image with the source image of said sequence of images; • selection (33) of a set of pixels in said reference image, each pixel of said set of pixels having an associated photometric difference greater than a first threshold value; • deletion (34) of said reference image in said batch of images if said set of pixels comprises a number of pixels less than a second threshold value; - for each undeleted reference image of said each sequence of images, determination (35) of a photometric error associated with said undeleted reference image from: • a first photometric error determined by comparing pixels of a first image to pixels of said undeleted reference image, said first image being generated from the source image and first depths associated with pixels of the source image, and • a second photometric error determined by comparing pixels of a second image to pixels of said source image, said second image being generated from the undeleted reference image and second depths associated with pixels of
2.
3.
4. the unsuppressed reference image, said first and second depths being predicted by the depth prediction model from the unsuppressed source and reference images; - learning (36) of the depth prediction model by minimizing a loss error determined from: • photometric errors associated with the undeleted reference images in said batch of images, and • a number of undeleted reference images included in said batch of images. Method according to claim 1, for which the photometric error associated with the reference image is determined by the following function: = E4nûn(Lpl(p)' Lp2(p))] With : • Lp corresponding to the photometric error associated with the undeleted reference image noted Ref, • Lp\ ( p) the first photometric error determined for a pixel P in the reference image, the pixel P being defined by its two-dimensional coordinates, and • p) the second photometric error determined for pixel P in the source image of the image sequence including the undeleted reference image. Method according to one of claims 1 to 2, for which the loss error is equal to an average of said photometric error determined for each undeleted reference image. Method according to one of claims 1 to 3, for which the photometric difference associated with said reference image is defined by the following function: D(Ref ; p) = 11Ref ( p ) -1Som.ce ( p) | With : • D corresponding to the photometric difference associated with a pixel P of the undeleted reference image, the reference image being noted Ref, the pixel P being defined by its coordinates in two dimensions, * a colorimetric value determined for a pixel P in the reference image, and *hmrce(p) a colorimetric value determined for pixel P in the source image of the image sequence comprising the image of reference.
5. Method according to one of claims 1 to 4, for which the first, respectively the second, photometric error is determined by the following function: Lp*= £^(1-«) ' U(p)-Up)\ +«■ ( )) With: • Lp* the first photometric error noted Lp] (p), respectively the second photometric error noted Lp2{p\ • I(p) a colorimetric value of the pixel P in the undeleted reference image, respectively in the source image, P being a pixel defined by its coordinates in two dimensions, • a value of the pixel P in the first image, respectively in the second image, • SSIM a function which takes into account a local structure, and • has a weighting factor depending in particular on the type of environment in which the vehicle (10) is moving.
6. Method according to one of claims 1 to 5, for which the first image, respectively second image, is generated by the following function: With: • Ps the coordinates of a pixel of the first image, respectively of the second image, • 77 a function for going from homogeneous coordinates to pixel coordinates by removing one dimension of a vector, • K a direction prediction model associated with the camera (11), • T a matrix comprising the extrinsic parameters, and • 0 a projection function in the three-dimensional scene of a pixel Pt as a function of its coordinates in the source image, respectively in the unremoved reference image, and of the depth D^p^ which is associated with it.
7. Method according to one of claims 1 to 6, for which the first threshold is between 1 and 2 when a colorimetric value of a pixel is coded on eight bits.
8. Method according to one of claims 1 to 7, for which the second threshold is between 60% and 90%.
9. Device (4) configured to learn a depth prediction model by a vision system on board a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured to implement the steps of the method according to any one of claims 1 to 8.
10. Vehicle (10) comprising the device (4) according to claim 9.
Citation Information
Patent Citations
Camera agnostic depth network
US11257231B2