Method and device for determining a calibration fault in a vehicle-mounted vision system
The method and device for detecting calibration faults in vehicle-mounted vision systems improve ADAS reliability and road safety by analyzing depth prediction model errors and updating non-conformity rates to maintain accurate depth predictions.
Patent Information
- Application Number
- FR2024000564
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-01-19
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-01-19
AI Technical Summary
Existing vehicle-mounted vision systems face challenges in maintaining the accuracy of depth prediction due to potential calibration faults, which can affect the reliability of Advanced Driver-Assistance Systems (ADAS) and overall road safety.
A method and device for determining calibration faults in vehicle-mounted vision systems by analyzing depth prediction model errors using a loss function, comparing image data against thresholds, and updating non-conformity rates to detect recurring errors, which can include meteorological data and graphical alerts.
Ensures the quality of data from the vision system by identifying and addressing calibration faults, improving the reliability of ADAS systems, and enhancing road safety by ensuring accurate depth predictions.
Smart Images

Figure 00000022_0000 
Figure 00000022_0001 
Figure 00000023_0000
Abstract
Description
Title of the invention: Method and device for determining a calibration fault in a vehicle-mounted vision system technical field
[0001] The present invention relates to methods and devices for determining a calibration fault in a vehicle-mounted vision system, for example, in a motor vehicle. The present invention also relates to a method and device for alerting a driver to a calibration fault in a vehicle-mounted vision system. Technological background
[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS systems use advanced technologies to assist the driver while driving and thus improve performance. ADAS systems use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.
[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.
[0004] Vehicle-mounted ADAS systems are powered by data obtained from one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or any obstacles present around a vehicle in order, for example: • to adapt the vehicle's lighting according to the presence of other road users; • to automatically regulate the vehicle's speed; • to act on the braking system in case of risk of impact with an object.
[0005] Furthermore, the quality of the data emitted by a vision system, that is, the accuracy of the positioning of objects in the three-dimensional scene, has a direct impact on the proper functioning of driver assistance systems using this data. It is therefore important to control the accuracy of a model predicting the distance and position of an object in the scene, both after calibration of the vision system and throughout the entire lifespan of the vision system. Indeed, a A defect in the vision system can occur during its lifetime, whether it is temporary or permanent. Summary of the present invention
[0006] One object of the present invention is to solve at least one of the problems of the technological background described above.
[0007] Another object of the present invention is to ensure the quality of data from the processing of an image acquired by a vision system embedded in a vehicle.
[0008] Another object of the present invention is to improve road safety, in particular by improving the reliability of AD AS systems powered by data obtained from a wide-angle camera.
[0009] According to a first aspect, the present invention relates to a method for determining a calibration fault in a vision system embedded in a vehicle, the vision system comprising an assembly of at least one camera, each camera in the assembly being arranged so as to acquire an image of a three-dimensional scene, the process being implemented by at least one processor, and being characterized in that it comprises the following steps: - receiving data representative of a non-conformity rate and a first error of a depth prediction model associated with the vision system, the first error being determined by a loss function during a learning phase of the depth prediction model; - receiving image data representative of a set of images acquired by the vision system; - determination of a second error by applying the loss function to the image data; - determination of a conformity indicator based on the result of a comparison of the second error to a threshold error, the threshold error being determined based on the first error; - updating the non-conformity rate based on the conformity indicator; - determination of a vision system calibration fault based on the result of a comparison of the non-conformity rate to a threshold value.
[0010] Such a method for determining a calibration fault in the vision system thus makes it possible to test the quality or accuracy of depths predicted by the depth prediction system at a current time in comparison with the quality or accuracy of depths predicted by the depth prediction system at the end of the learning phase. In the event of several recurring non-conformities, a calibration fault in the vision system is then determined or detected.
[0011] According to one variant of the method, the threshold error is proportional to the first error, a ratio between the threshold error and the first error being strictly greater than 1.
[0012] According to another variant, the method further includes a step of receiving meteorological data representative of meteorological conditions in a vehicle environment, a test frequency of the vision system being determined based on the meteorological data.
[0013] According to yet another variant, the method further includes a step of checking the display of graphic content on a screen embedded in the vehicle, the graphic content including information representative of the calibration fault.
[0014] According to a further variant of the method, the vision system is a stereoscopic vision system comprising at least two cameras.
[0015] According to another variant of the method, the vision system is a monoscopic vision system comprising a single camera.
[0016] According to yet another variant of the method, a camera of the vision system is a wide-angle camera.
[0017] According to a second aspect, the present invention relates to a device for determining a calibration fault of a vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.
[0018] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0019] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0020] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0021] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.
[0022] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM, or a microelectronic circuit-type ROM, or a magnetic recording means, or a hard drive.
[0023] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.
[0024] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures
[0025] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 4, in which:
[0026] [Fig-1] schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting example of the present invention;
[0027] [Fig.2] schematically illustrates a monoscopic vision system equipping a vehicle, according to a particular and non-limiting example of the present invention;
[0028] [Fig.3] illustrates a flowchart of the different stages of a determination process of a calibration defect of a vision system embedded in the vehicle of [Fig.1] or [Fig.2], according to a particular and non-limiting embodiment of the present invention;
[0029] [Fig.4] schematically illustrates a device configured to determine deter mination of a calibration defect of a vision system embedded in the vehicle of [Fig.1] or of [Fig.2], according to a particular and non-limiting embodiment example of the present invention. Description of examples of achievements
[0030] A method and device for determining a calibration fault of a vision system on board a vehicle will now be described in what follows with joint reference to Figures 1 to 4. The same elements are identified with the same reference signs throughout the description that follows.
[0031] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0032] According to a particular and non-limiting example of an embodiment of the present invention, a method for determining a calibration fault of a vision system embedded in a vehicle is for example implemented by a computer of the vehicle's embedded system controlling this vision system.
[0033] Indeed, the method includes receiving data representative of a first error of a depth prediction model associated with the vision system determined by a loss function during a learning phase of the depth prediction model, receiving image data representative of a set of images acquired by the vision system and determining a second error by applying the loss function to the image data.
[0034] A conformity indicator is determined based on the result of a comparison of the second error to a threshold error determined based on the first error, and a non-conformity rate is updated based on the conformity indicator.
[0035] The calibration defect of the vision system is then determined according to the rate of non-conformities.
[0036] Fig. 1 schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.
[0037] The first vehicle 10 is in a first environment 1 corresponding, for example, to a road environment formed of a network of roads accessible to the first vehicle 10.
[0038] In this example, the first vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. The first vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, the first vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.
[0039] The first vehicle 10 advantageously comprises at least two onboard cameras, a first camera 11 and a second camera 12, configured to acquire images of a three-dimensional scene unfolding in the environment of the first vehicle 10 from distinct observation positions. The first camera 11 and the second camera 12 form a stereoscopic vision system when used together, as illustrated in [Fig. 1]. The first camera 11 forms a monoscopic vision system when used alone, and similarly, the second camera 12 forms another monoscopic vision system when used alone. The present invention, however, extends to any vision system including at least two cameras, for example 2, 3 or 5 cameras.
[0040] The intrinsic parameters of the first camera 11 characterize the transformation which associates, for an image point, hereafter called a "point", its three-dimensional coordinates in the frame of reference of the first camera 11 with the pixel coordinates in an image acquired by the first camera 11. These parameters do not change if the first camera 11 is moved. The intrinsic parameters of the first camera 11 include in particular a first focal length fl associated with the first camera 11.
[0041] The intrinsic parameters of the second camera 12 characterize, for their part, the transformation which associates, for an image point, its three-dimensional coordinates in the reference frame of the second camera 12 with the pixel coordinates in an image acquired by the second camera 12. These parameters do not change if the second camera 12 is moved. The intrinsic parameters of the second camera 12 include in particular a second focal length f2 associated with the second camera 12.
[0042] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in lens curvature, parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are not taken into account, and their correction is addressed during image acquisition or calibration.
[0043] These two cameras 11, 12 are arranged so that each acquires an image of a scene from a different viewpoint. The first viewpoint is, for example, located on or in the left-hand rearview mirror of the first vehicle 10 or at the top of the windshield of the first vehicle 10. The second viewpoint is, for example, located on or in the right-hand rearview mirror of the first vehicle 10 or at the top of the windshield of the first vehicle 10. If both cameras are located at the top of the vehicle's windshield, they are positioned at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the first vehicle 10, and the second camera 12 is located in the right-hand rearview mirror of the first vehicle 10.
[0044] A first marker is associated with the first camera 11: - The direction of the x-axis is defined as horizontal and normal to the optical axis of the first camera 11. The distance B separating the optical center of the first camera 11 from the projection of the optical center of the second camera 12 onto the horizontal plane passing through the optical center of the first camera 11 is called the reference base (in English "baseline"); - the direction of the y-axis is defined as vertical and normal to the optical axis of the first camera 11; - The direction of the z-axis is defined as orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal coordinate system.
[0045] The extrinsic parameters related to the position of cameras 11, 12 are the following parameters: - three translations in the x, y, and z directions: Tx, Ty, and Tz, constituting the translation vector T; and - three rotations in the x, y and z directions: 0x, 0y and 0z.
[0046] The extrinsic parameters are determined, for example, during a calibration phase of the stereoscopic vision system comprising the first camera 11 and the second camera 12.
[0047] A key constraint of stereoscopic vision systems used in automobiles is, for example, the large distance between the two cameras. Indeed, to cover a measurement range of 200 meters, the reference base must be 60 cm for cameras commonly used in this field.
[0048] The first and second cameras 11, 12 acquire images of a scene located in front of the first vehicle 10, the first camera 11 alone covering a first acquisition field 13, the second camera 12 alone covering a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic view of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic view of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic view of the scene by the stereoscopic vision system composed of the two cameras 11, 12.
[0049] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.
[0050] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system comprising the second camera 12.
[0051] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable using the monoscopic vision system comprising the first camera 11.
[0052] Finally, field 19 is not visible to any of the cameras. The part of the scene present in this field 19 is therefore not observable.
[0053] According to one particular embodiment, the field of view of the second camera 12 covers at least half of the field of view of the first camera 11.
[0054] It is evident that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the first vehicle 10 by equipping it with cameras placed and oriented differently.
[0055] The images acquired by the first and second cameras 11, 12 at a given acquisition time are in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colors and brightness of objects in the observed scene, for example in the form of RGB (Red Green Blue) or HSL (Hint, Saturation, Luminosity) colorimetric coordinates.
[0056] Each pixel of an acquired image represents an object in the three-dimensional scene present in the field of view of the first or second camera 11, 12. Indeed, a pixel of an acquired image is the smallest visible unit and corresponds to a point of light resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, captured by a photosensitive sensor of the first or second camera 11, 12 after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels.Each pixel records light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image that faithfully represents the physical object observed by the first or second camera 11, 12. An image point previously presented is therefore a point on the surface of an object in the three-dimensional scene observed by the stereoscopic vision system comprising the first and second cameras 11, 12.
[0057] The images acquired by the first and second cameras 11, 12 represent views of the same scene taken from different viewpoints, the camera positions being distinct. For example, this scene includes: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0058] According to a particular embodiment, the first camera 11 is of the "wide-angle" type, a wide-angle camera being, for example, equipped with a lens designed To acquire a representative image of a three-dimensional scene viewed from a wider field of view than a standard camera, also sometimes called a panoramic lens. In other words, a wide-angle lens captures a larger portion of the three-dimensional scene unfolding in front of or around the camera, which is particularly useful in situations where it is necessary to include more elements in the frame of the image acquired by that camera. The angle of view of the first camera 11 is, for example, 120°, 145°, 180°, or 360°, whereas a standard camera offers, for example, a field of view of 45° or less. Such a first camera 11 corresponds, for example, to a camera equipped with mirrors or a fisheye camera.Wide-angle lenses have a shorter focal length compared to standard lenses, making them suitable for capturing images of landscapes, architecture, road intersections, or any other subject requiring a wide perspective. Wide-angle cameras are, for example, used to capture immersive and dynamic images with an extended depth of field.
[0059] According to a particular embodiment, the image acquired by the first camera 11 includes a distortion equal to 0.5%, 0.8%, or greater than 1%. The measurement of such distortion corresponds to determining a ratio between: - the maximum spacing of a pixel in the image of a straight line in the first three-dimensional scene whose image is a line touching the longest edge of the first image, either at the center of the image edge or at the corners of the image edge, and - the length of this edge.
[0060] In the world of photography, distortion is commonly considered to be: • negligible if it is less than 0.3%, • not very sensitive if it is between 0.3% or 0.4%, • sensitive if it is between 0.5% and 0.6%, • very sensitive if it is between 0.7% and 0.9%, and • problematic if it is greater than or equal to 1% or more.
[0061] Barrel distortion is characterized by a positive percentage, while crescent distortion is characterized by a negative percentage.
[0062] The images acquired by the first and second cameras 11,12 are, for example, sent to a computer of a device equipping the first vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the first vehicle 10.
[0063] Depths associated with pixels of images acquired by the vision system Stereoscopic depths are, for example, determined by a first depth prediction model, such a depth prediction model being, for example, implemented by a neural network. To be reliable and accurate, the first depth prediction model is learned in an initial learning phase by minimizing an error determined by a first loss function. This initial learning phase is known to those skilled in the art. Various methods are known and applicable depending on the types of the first and second cameras11,12 and the extrinsic parameters of the stereoscopic vision system, for example, the method described in the paper: “UnOS: Unified Unsupervised Optical-flow and Stereo-depth Estimation by Watching Videos” by Yang Wang, Peng Wang, Zhenheng Yang, Chenxu Luo, Yi Yang, and Wei Xu, published in June 2019.This method implements several operations, including one to determine a photometric error by comparing images generated from images acquired by the stereoscopic vision system and predicted depths for pixels in these images to images acquired by the same stereoscopic vision system. The function used to determine this photometric error is called the loss function. Another operation minimizes this photometric error by adjusting parameters of the depth prediction model. At the end of the learning process, a residual error deemed acceptable remains. This error then corresponds to the first error received during a process of determining a calibration fault in the vision system embedded in the first vehicle.
[0064] A method for determining a calibration fault of the stereoscopic vision system on board the first vehicle 10 is advantageously implemented by the first vehicle 10, i.e. by a processor, a computer or a combination of computers of the on-board system of the first vehicle 10, for example by the computer or computers in charge of the stereoscopic vision system of the first vehicle 10 or by the device 4 of [Fig.4].
[0065] Figure 2 schematically illustrates a monoscopic vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.
[0066] The second vehicle 20 is in a second environment 2 corresponding, for example, to a road environment consisting of a network of roads accessible to the vehicle 20.
[0067] In this example, the second vehicle 20 corresponds to a vehicle with an internal combustion engine, with electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. The second vehicle 20 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, the second vehicle 20 corresponds to an autonomous or non-autonomous vehicle. that is to say, a vehicle operating according to a determined level of autonomy or under the total supervision of the driver.
[0068] The second vehicle 20 advantageously comprises at least one third on-board camera 21 configured to acquire images of a three-dimensional scene taking place in the environment of the second vehicle 20 from a predetermined observation position. The third camera 21 forms a monoscopic vision system.
[0069] The intrinsic parameters of the third camera 21 characterize the transformation which associates, for an image point, hereafter called a "point", its three-dimensional coordinates in the reference frame of the third camera 21 with the pixel coordinates in an image acquired by the third camera 21. These parameters do not change if the third camera 21 is moved. The intrinsic parameters of the third camera 21 include in particular a third focal length f3 associated with the third camera 21.
[0070] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in the curvature and parallelism of the lenses and the coaxiality of the optical axes. In this example, the third camera 21 is assumed to be perfect, that is to say, that distortions are not taken into account, and that their correction is addressed during image acquisition or calibration.
[0071] This third camera 21 is arranged to acquire an image of a three-dimensional scene from a determined point of view, for example located on or in the left rearview mirror of the second vehicle 20 or at the top of the windshield of the second vehicle 20. In this example, the third camera 21 is located at the top of the windshield of the second vehicle 20.
[0072] The third camera 21 acquires images of a scene located in front of the second vehicle 20, the third camera 21 covering an acquisition field 22 in which, for example, an object 23 is positioned. The presence of the object 23 defines an occlusion field 24 for the monoscopic vision system comprising the third camera 21.
[0073] It is evident that it is possible to use such a monoscopic vision system to take images of scenes located on the sides or behind the second vehicle 20 by equipping it with cameras placed and oriented differently.
[0074] The images acquired by the third camera 21 at a given acquisition time are presented in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colours and brightness of objects in the observed scene in the form of, for example, RGB colourimetric coordinates (from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).
[0075] Each pixel of the image acquired by the third camera 21 represents an object in the three-dimensional scene present within the field of view of the third camera 21. Indeed, a pixel of the acquired image is the smallest visible unit and corresponds to a point of light resulting from the emission or reflection of light by a physical object present in the three-dimensional scene. When light strikes this object, photons are emitted or reflected, which are captured by a photosensitive sensor of the third camera 21 after passing through its lens. This sensor divides the three-dimensional scene into a grid of pixels. Each pixel records the light intensity at a specific location, thus capturing visual details. The combination of millions of pixels creates an image that faithfully represents the physical object observed by the third camera 21.An image point is thus a point on the surface of an object in the three-dimensional scene observed by the third camera 21.
[0076] The images acquired by the third camera 21 represent views of the same scene, for example, acquired at different times during acquisition. When the second vehicle 20 is in motion, these images are acquired by the third camera 21 from different observation positions or viewpoints. This three-dimensional scene includes, for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0077] According to a particular embodiment, the third camera 21 is of the "wide-angle" type. The angle [3] of the field of view of the third camera 21 is, for example, equal to 120°, 145°, 180° or 360°. Such a third camera 21 corresponds, for example, to a camera equipped with mirrors or to a "fisheye" camera.
[0078] According to one particular embodiment, an image acquired by the third camera 21 includes a distortion equal to 0.5%, 0.8% or greater than 1%.
[0079] The images acquired by the third camera 21 are, for example, sent to a computer of a device equipping the second vehicle 20 or stored in a memory of a device accessible to a computer of a device equipping the second vehicle 20.
[0080] A method for determining a calibration fault of the stereoscopic vision system on board the second vehicle 20 is advantageously implemented by the second vehicle 20, i.e. by a processor, a computer or a combination of computers of the on-board system of the second vehicle 20, for example by the computer(s) in charge of the monoscopic vision system of the second vehicle 20 or by the device 4 of [Fig.4].
[0081] Depths associated with pixels in images acquired by the monoscopic vision system are, for example, determined by a second depth prediction model, such a depth prediction model being, for example, implemented by a neural network. To be reliable and accurate, the second depth prediction model is learned, in particular, in a second learning phase by minimizing an error determined by a second loss function. Such a second learning phase is known to those skilled in the art.Different methods are known and applicable depending on the type of third camera 21, for example the method described in the document: “Digging Into Self-Supervised Monocular Depth Estimation” by Clément Godard, Oisin Mac Aodha, Michael Firman and Gabriel Brostow published in August 2019, or, for wide-angle cameras, “Neural Ray Surfaces for Self-Supervised Leaming of Depth and Ego-motion” by Igor Vasiljevic, Vitor Guizilini, Rares Ambrus, Sudeep Pillai, Wolfram Burgard, Greg Shakhnarovich and Adrien Gaidon published in August 2020.
[0082] Generally, any self-supervised or self-learning vision system—that is, one whose depth prediction model is learned from images acquired by the vision system itself and without using annotated training data—includes a loss function for evaluating the performance of the depth prediction model and improving its accuracy through an iterative method implemented during a learning phase. This method aims to minimize an error determined by the loss function. At the end of the learning phase, a residual error remains, corresponding to a minimum error obtained during the learning phase. This residual error then corresponds to the first error received during a process for determining a calibration fault in the vision system embedded in the first vehicle 10.
[0083] Optionally, the learning phases are implemented after a calibration phase of the vision system cameras. The calibration of a vision system camera is also known to those skilled in the art, and is described, for example, in the following document: "A Flexible New Technique for Camera Calibration" by Zhengyou Zhang, published in December 1998, or in this document: "Single View Point Omnidirectional Camera Calibration from Planar Grids" by Christopher Mei and Patrick Rives, published in April 2007.
[0084] Figure [Fig. 3] illustrates a flowchart of the different steps of a method for determining a calibration fault of a vision system on board a vehicle, for example in the first vehicle 10 of [Fig. 1] or in the second vehicle 20 of [Fig. 2], according to a particular and non-limiting embodiment of the present invention.
[0085] In the description associated with [Fig. 3], the term "vehicle 10, 20" refers interchangeably to the first vehicle 10 of [Fig. 1] or to the second vehicle 20 of [Fig. 2], the method 3 for determining a calibration fault in a vision system being applicable to both a stereoscopic vision system and a monoscopic vision system. Similarly, the term "learning phase" refers to the first learning phase when "the vehicle" refers to the first vehicle 10, while it refers to the second learning phase when "the vehicle" refers to the second vehicle 20.
[0086] The method 3 is for example implemented by the device on board the vehicle 10, 20 implementing the method of determining a depth by on-board vision system or by the device 4 of the [Fig.4].
[0087] In step 31, representative data of a non-conformity rate and a first error of a depth prediction model associated with the vision system are received. The first error is determined by a loss function during a training phase of the depth prediction model as described with reference to Figures 1 or 2.
[0088] According to a particular embodiment, the nonconformity rate is, for example, equal to 0 at the beginning of the process. When the subsequent steps, i.e., from 33 to 37, are repeated or iterative, then the nonconformity rate is the rate previously determined, for example, from the last iterations. Further details on this calculation are provided below in the description of step 36.
[0089] In a step 33, image data representative of a set of images acquired by the vision system are received. The number of images depends in particular on the type of vision system.
[0090] In step 34, a second error is determined by applying the loss function to the image data. The loss function is the same loss function used during the training phase of the depth prediction model. Thus, the second error corresponds to a typical error, determined under the normal operating conditions of the vision system.
[0091] In step 35, a conformity indicator is determined based on the result of a comparison of the second error to a threshold error, the threshold error being determined based on the first error. Thus, the first error corresponds to the minimum error obtained at the end of the learning phase of the prediction model. depth. This is then ideal and corresponds to a perfectly calibrated vision system.
[0092] According to a particular embodiment, the threshold error is proportional to the first error; thus, the threshold error is equal to 1.5x, 2x, or 3x the first error. Therefore, the ratio between the threshold error and the first error is strictly greater than 1. Indeed, if the ratio were 1 or less, then the conformance indicator would represent a nonconformity, even when the vision system implements a newly learned depth prediction model.
[0093] The conformity indicator is, according to a first particular embodiment, equal to the ratio between the second error and the threshold error.
[0094] According to a second particular embodiment, the conformity indicator is equal to 1 when the second error is greater than the threshold error and equal to 0 otherwise.
[0095] In a step 36, the non-conformity rate is updated according to the conformity indicator.
[0096] For example, the nonconformity rate is equal to the average of the last N conformity indicators, including the conformity indicator determined in step 35, when this process is implemented iteratively, the number N being defined, for example, by a user or manufacturer. For example, the number N is equal to 10 or 100. Thus, the nonconformity rate represents the number of second errors most recently determined that exceed the threshold error. It should be noted that in this example, the last N conformity indicators are, for example, recorded or stored in the memory of the device implementing the process.
[0097] According to another particular embodiment, the non-conformity rate is equal to the average or a weighted average of the received non-conformity rate and the conformity indicator, for example determined from the following function: T' = 0.9 x T + 0.1 x Ic, with: • The updated non-compliance rate, • T is the non-conformity rate received in step 31, and • Ic is the compliance indicator determined in step 35.
[0098] According to the second particular embodiment, Ic = E2 / Es, where: • Here, the compliance indicator determined in step 35, • E2 the second error, and • Is the threshold error.
[0099] In a step 37, a calibration defect of the vision system is determined based on a result of a comparison of the non-conformity rate to a threshold value.
[0100] The threshold value is, for example, defined in such a way as to determine a defect in ca The vision system is deactivated when the non-conformity rate represents a number of conformity indicators exceeding a threshold. Indeed, the goal is not to identify a non-conformity from the first detected non-conformity, but rather when several non-conformities are identified over a certain number N of tests, for example.
[0101] According to a particular embodiment, the method 3 for determining a calibration fault in the vision system further comprises a step 32 for receiving meteorological data representative of the weather conditions of an environment 1, 2 of the vehicle 10, 20, the frequency of testing the vision system being determined based on the meteorological data. Indeed, a calibration fault occurs, for example, when a lens or objective of a camera is covered with transparent water droplets, which happens more frequently in rainy weather, for example. Thus, the meteorological data is obtained, for example, by analyzing an image acquired by the vision system, in which it is possible to perceive that the environment in which the vehicle 10, 20 is traveling is humid, or received from an on-board system of the vehicle 10, 20 when the windshield wipers are, for example, in operation.
[0102] According to another specific embodiment, the method 3 for determining a calibration fault in the vision system further comprises a step of checking the display of graphic content on a screen embedded in the vehicle 10, 20, the graphic content including information representative of the calibration fault when it is determined. Such a screen is, for example, part of an infotainment system embedded in the vehicle 10, 20 and is located in the passenger compartment of the vehicle 10, 20; for example, it is mounted on a dashboard or a center console. This screen is then connected to the device implementing the method 3 for determining a calibration fault in the vision system, for example, via a wired communication bus connection.
[0103] According to yet another specific embodiment, data is transmitted to an Advanced Driver-Assistance System (ADAS) which receives as input the depths predicted by the vehicle's onboard vision system 10, 20. This data is, for example, representative of a calibration fault when one is detected. The ADAS then switches, for example, to a degraded or safety mode to continue operating with unreliable input data, or is even deactivated. Similar to the previous example, the ADAS is connected to the device implementing method 3 for determining a calibration fault in the vision system, for example, via a wired communication bus connection.
[0104] Such a method for determining a calibration fault in the vision system thus makes it possible to test the quality or accuracy of depths predicted by the depth prediction system at a current time in comparison with the quality or accuracy of depths predicted by the depth prediction system at the end of the learning phase. In the event of several recurring non-conformities, a calibration fault in the vision system is then determined or detected, making it possible, for example, to alert the driver of the vehicle equipped with the vision system so that they can take into account a possible failure of the vision system. Similarly, an ADAS fed by depth data is, for example, switched to a degraded operating mode in order to maintain an adequate level of safety.A user of vehicle 10, 20 is then able to plan a maintenance operation, a calibration or simply a cleaning of the vision system on board vehicle 10, 20 in order to restore an operational vision system as it was at the end of the learning phase.
[0105] Figure 4 schematically illustrates a device 4 configured for determining a calibration fault in a vision system embedded in a vehicle, for example in the first vehicle 10 of Figure 1 or in the second vehicle 20 of Figure 2, according to a particular and non-limiting embodiment of the present invention. The device 4 corresponds, for example, to a device embedded in the first vehicle 10 or in the second vehicle 20, for example a computer.
[0106] Device 4 is, for example, configured to carry out the operations described opposite Figures 1 or 2 and / or the steps described opposite [Fig. 3]. Examples of such a device 4 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 4, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 4 may be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.
[0107] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41, corresponding, for example, to volatile and / or non-volatile memory, and / or includes a memory storage device which may include memory volatile and / or non-volatile, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.
[0108] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored on memory 41.
[0109] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.
[0110] According to a particular and non-limiting embodiment, the device 4 includes a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 include one or more of the following interfaces: - Radio frequency (RF) interface, for example, Wi-Fi® (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (from the English "Long-Term Evolution" or in French "Evolution à long terme"), LTE-Advanced (or in French LTE-avancé); - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HDMI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").
[0111] According to another particular and non-limiting embodiment, the device 4 includes a communication interface 43 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 430. The communication interface 43 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds, for example, to a wired network of the CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3) type.
[0112] According to a particular and non-limiting embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch or not, one or more speakers 450 and / or other peripherals 460 via the output interfaces 44, 45, 46 respectively. According to a variant, one or more of the external devices is integrated into the device 4.
[0113] Of course, the present invention is not limited to the embodiments described above but extends to a method for detecting a calibration fault in a vehicle-mounted vision system, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0114] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 4 of [Fig.4].
Claims
Demands
1. A method for determining a calibration fault in a vehicle-mounted vision system (10, 20), the vision system comprising an array of at least one camera, each camera of said array being arranged to acquire an image of a three-dimensional scene, said method being implemented by at least one processor, and being characterized in that it comprises the following steps: - receiving (31) data representative of a non-conformity rate and a first error of a depth prediction model associated with the vision system, the first error being determined by a loss function during a training phase of the depth prediction model; - receiving (33) image data representative of a set of images acquired by the vision system; - determining (34) a second error by applying the loss function to the image data;- determination (35) of a conformity indicator based on the result of a comparison of the second error to a threshold error, the threshold error being determined based on the first error; - updating (36) of the non-conformity rate based on the conformity indicator; - determination (37) of a vision system calibration defect based on the result of a comparison of the non-conformity rate to a threshold value.
2. A method according to claim 1, wherein the threshold error is proportional to the first error, a ratio between the threshold error and the first error being strictly greater than 1.
3. A method according to claim 1 or 2, further comprising a receiving step (32) of meteorological data representative of meteorological conditions of an environment (1, 2) of the vehicle (10, 20), a test frequency of the vision system being determined according to the meteorological data.
4. A method according to any one of claims 1 to 3, further comprising a step of checking the display of graphic content on a screen embedded in the vehicle (10, 20), said graphic content comprising information representative of the calibration fault.
5. A method according to any one of claims 1 to 4, wherein the vision system is a stereoscopic vision system comprising at least two cameras.
6. A method according to any one of claims 1 to 4, wherein the vision system is a monoscopic vision system comprising a single camera.
7. A method according to any one of claims 1 to 6, wherein a camera of the vision system is a wide-angle camera.
8. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by a processor.
9. Device (4) for detecting a calibration fault of a vision system on board a vehicle (10, 20), said device (4) comprising a memory (41) associated with at least one processor (40) configured for carrying out the steps of the method according to any one of claims 1 to 7.
10. Vehicle (10, 20) comprising the device (4) according to claim 9.