Method and device for generating representative image data to improve the reliability of depth determination by a vehicle-mounted vision system
The method enhances the reliability of stereoscopic vision systems in vehicles by training a depth prediction model using intentionally contaminated images, addressing issues like dirty lenses and glare to improve ADAS performance.
Patent Information
- Application Number
- FR2023010392
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-09-29
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-09-29
AI Technical Summary
Existing stereoscopic vision systems in vehicles are susceptible to contamination from issues like dirty lenses or glare, which affect the reliability of depth determination and compromise the functioning of Advanced Driver-Assistance Systems (ADAS).
A method and device for generating representative image data using a stereoscopic vision system with a first and second camera, involving image prediction and learning a depth prediction model through the use of intentionally 'contaminated' images to enhance robustness, allowing the system to predict reliable depths despite defects.
The method improves the reliability of ADAS systems by training the depth prediction model to handle image defects, ensuring accurate depth determination even when cameras are affected by contamination.
Smart Images

Figure 00000021_0000 
Figure 00000021_0001 
Figure 00000021_0002
Abstract
Description
Title of the invention: Method and device for generating representative image data to improve the reliability of depth determination by a vehicle-mounted vision system. Technical field
[0001] The present invention relates to methods and devices for generating image-representative data for a stereoscopic vision system embedded in a vehicle, for example, in a motor vehicle. The present invention also relates to a method and device for generating training data for a depth prediction model using a stereoscopic vision system embedded in a vehicle. The present invention further relates to a method and device for controlling one or more ADAS systems embedded in a vehicle based on a distance between an object and the vehicle determined by a stereoscopic vision system. Technological background
[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS systems use advanced technologies to assist the driver while driving and thus improve performance. ADAS systems use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.
[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.
[0004] The AD AS systems embedded in a vehicle are powered by data obtained one or more onboard sensors, such as cameras. These cameras make it possible to detect and locate other road users or potential obstacles around a vehicle in order to, for example: - to adapt the vehicle's lighting according to the presence of other road users; - to automatically regulate the vehicle's speed; - to act on the braking system in case of risk of impact with an object.
[0005] It sometimes happens that the data obtained from these sensors are defective, for example because of electronic problems disrupting their transmission, or because of Sensor problems, for example when a camera lens is dirty or the camera is dazzled, for example when it is pointed towards the sun or a car headlight shines directly on it.
[0006] The quality of the data emitted by a vision system therefore determines the proper functioning of the driving assistance devices using this data. Summary of the present invention
[0007] One object of the present invention is to solve at least one of the problems of the technological background described above.
[0008] Another object of the present invention is to propose a solution to improve the insensitivity of a stereoscopic vision system to contamination of an image acquired by this system, for example due to a dirty lens or glare.
[0009] Another object of the present invention is to improve road safety, in particular by improving the reliability of AD AS systems powered by data obtained from such a stereoscopic vision system.
[0010] According to a first aspect, the present invention relates to a method for generating representative image data for a stereoscopic vision system embedded in a vehicle, the stereoscopic vision system comprising a first camera and a second camera arranged so as to acquire an image of a three-dimensional scene from two different points of view, the process being characterized in that it comprises the following steps: - reception of representative data of a first pair of images, the pair of images comprising a first image acquired by a first camera at a time instant and a second image acquired by a second camera at the same time instant; - prediction by a prediction model of first depths associated with a first set of pixels of the first image as a function of optical flux values determined from the first set of pixels and a second set of pixels of the second image corresponding to the first set of pixels and prediction of second depths associated with the second set of pixels as a function of optical flux values determined from the second set of pixels and the first set of pixels; - generation of a second pair of images from the first pair of images, the second pair of images comprising the first image and a third image, the third image being generated from the second image for which a colorimetric value of each pixel of a third set of pixels is modified; - prediction by the third depth prediction model associated with a fourth set of pixels of the first image as a function of optical flux values determined from the fourth set of pixels and a fifth pixel set of the third image corresponding to the fourth pixel set and prediction of fourth depths associated with the fifth pixel set as a function of optical flux values determined from the fifth pixel set and the fourth pixel set; - learning the prediction model based on the result of a comparison of the third and fourth depths to the first and second depths.
[0011] Such a process thus makes it possible to generate "contaminated" images, that is, images intentionally containing defects, from images acquired by the stereoscopic vision system. The first and second predicted depths then serve as training data, and the depth prediction model is learned through the use of these "contaminated" images. The prediction model thus gains robustness and makes it possible to predict reliable depths for images containing defects.
[0012] According to a variant of the method, the third set of pixels comprises at least one group of adjacent pixels.
[0013] According to yet another variant of the process, at least one group of adjacent pixels forms a rectangle.
[0014] According to a further variant of the method, the pixels of the third set of pixels are modified into black pixels, a black pixel being a pixel whose set of colorimetric values is zero.
[0015] According to another variant of the process, the pixels of the third set of pixels are modified into white pixels, a white pixel being a pixel whose set of colorimetric values is maximal.
[0016] According to yet another variant of the method, the third set of pixels comprises a number of pixels less than a threshold, the threshold being defined in relation to the number of pixels of said second image.
[0017] According to a further variant of the method, the threshold is equal to 5% of the number of pixels of the second image.
[0018] According to one variant, the process comprises the following steps: - generation of a third pair of images from the first pair of images, the third pair of images comprising a fourth image and the second image, the fourth image being generated from the first image for which a colorimetric value of pixels from a sixth set of pixels is modified; - prediction of fifth depths associated with a seventh set of pixels from the fourth image based on optical flux values determined from the seventh set of pixels and an eighth set of pixels from the second image corresponding to the seventh set of pixels, and prediction of sixth depths associated with the eighth set of pixels based on flux values optical values determined from the eighth set of pixels and the seventh set of pixels, the prediction model is learned, moreover, based on the fifth and sixth depths by comparison with the first and second depths.
[0019] According to a second aspect, the present invention relates to a device for generating representative image data for a stereoscopic vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for implementing the steps of the process according to the first aspect of the present invention.
[0020] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0021] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0022] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0023] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.
[0024] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.
[0025] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.
[0026] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures
[0027] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 3, in which:
[0028] [Fig-1] schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting example of the present invention;
[0029] [Fig.2] illustrates a flowchart of the different stages of a process for generating representative image data for a stereoscopic vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention;
[0030] [Fig.3] schematically illustrates a device configured for generating representative image data for a stereoscopic vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention. Description of examples of achievements
[0031] A method and device for generating representative image data for a stereoscopic vision system embedded in a vehicle will now be described in what follows with joint reference to Figures 1 to 3. The same elements are identified with the same reference signs throughout the following description.
[0032] The terms "first," "second" (or "firsts," "seconds"), etc., are used in this document by arbitrary convention to allow for the identification and distinction of different elements (such as operations, means, etc.) implemented in the embodiments described below. Such elements may be distinct or correspond to a single element, depending on the embodiment.
[0033] According to a particular and non-limiting example of an embodiment of the present invention, a method for generating representative image data for a stereoscopic vision system embedded in a vehicle is, for example, implemented by a computer of the vehicle's embedded system controlling this stereoscopic vision system.
[0034] The vision system comprises a first camera and a second camera, arranged so as to each acquire an image of a three-dimensional scene from different points of view.
[0035] To this end, the method for generating representative image data for a stereoscopic vision system embedded in a vehicle includes receiving representative data from a first pair of images comprising a first image acquired by the first camera at a given time and a second image acquired by the second camera at the same time.
[0036] Initial depths associated with a first set of pixels in the first image are predicted by a prediction model, and second depths associated with a second set of pixels in the second image are also predicted by the same prediction model. These depths are further predicted as a function of optical flux values determined from the first and second sets of pixels.
[0037] A second pair of images is generated from the first pair of images, the second pair of images comprising the first image and a third image, the third image being generated from the second image for which a colorimetric value of each pixel of a third set of pixels is modified.
[0038] Third depths associated with a fourth set of pixels from the first image are then predicted by the prediction model, as well as fourth depths associated with a fifth set of pixels from the third image. These depths are further predicted as a function of optical flux values determined from the fourth and fifth sets of pixels.
[0039] The prediction model is then learned based on the result of a comparison of the third and fourth depths with the first and second depths.
[0040] Figure 1 schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.
[0041] Such an environment 1 corresponds, for example, to a road environment consisting of a network of roads accessible to the vehicle 10.
[0042] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.
[0043] The vehicle 10 advantageously comprises several on-board cameras 11, 12, each configured to acquire images of a scene in the environment of the vehicle 10. This set of cameras 11, 12 forms the stereoscopic vision system. Two cameras 11 and 12 are illustrated in [Fig. 1]. The present invention is not limited, however, to a stereoscopic vision system comprising two cameras but extends to any vision system comprising two or more cameras, for example, two, three, four, or five cameras.
[0044] The two cameras 11, 12 have known intrinsic parameters. These pa- Ramers consist in particular of: - the focal length fl of the first camera 11; - the focal length f2 of the second camera 12; - distortions which are due to imperfections in the optical system of each camera; - the direction Cl of the optical axis of the first camera 11; - the C2 direction of the optical axis of the second camera 12; and - the respective resolutions of cameras 11, 12.
[0045] The intrinsic parameters characterize the transformation that associates, for an image point, the camera coordinates to the pixel coordinates, in each camera. These parameters do not change if the camera is moved.
[0046] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in lens curvature, parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are either not taken into account or their correction is addressed during image acquisition.
[0047] These two cameras 11, 12 are arranged so that each acquires an image of a scene from a different viewpoint. The first viewpoint is, for example, located on or in the left-hand rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. The second viewpoint is, for example, located on or in the right-hand rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. If both cameras are located at the top of the windshield of the vehicle, they are then positioned at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, and the second camera 12 is located in the right-hand rearview mirror of the vehicle 10.
[0048] A first marker is associated with the first camera 11: - the direction of the y-axis is defined by the position of the second camera 12, so as to place the second camera 12 on the y-axis of the first camera 11. The distance B separating the two cameras 11,12 is called the reference base (in English "baseline") and the direction separating the two cameras 11,12 is that of the y-axis; - the direction of the x axis is defined orthogonal to that of the y axis and orthogonal to that of the optical axis Cl of the first camera 11; - The direction of the z-axis is defined as orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal coordinate system.
[0049] The extrinsic parameters related to the position of cameras 11, 12 are the following parameters: - 3 translations in the x, y, and z directions: Tx, Ty, and Tz forming the translation vector T; and - 3 rotations around the x, y and z axes: Rx, Ry and Rz, constituting the rotation matrix R.
[0050] Determining the extrinsic parameters constitutes the problem of calibrating a stereoscopic vision system.
[0051] A key constraint of stereoscopic vision systems used in automobiles is, for example, the large distance between the two cameras. Indeed, to cover a measurement range of 200 meters, the reference base must be 60 cm for cameras commonly used in this field.
[0052] The two cameras 11, 12 acquire images of a scene located in front of the vehicle 10, the first camera 11 alone covering a first acquisition field 13, the second camera 12 alone covering a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic view of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic view of the scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic view of the scene by the stereoscopic vision system composed of the two cameras 11, 12.
[0053] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.
[0054] Among these three fields, field 16 is visible from the second camera 12. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system composed of the second camera 12.
[0055] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable with the monoscopic vision system composed of the first camera 11.
[0056] Finally, field 19 is not visible from any of the cameras. The part of the scene present in this field 19 is therefore not observable.
[0057] According to one embodiment, the directions Cl, C2 of the optical axes representing an orientation of the field of vision of each camera are oriented non-parallel so as to obtain the third acquisition field 15 of the environment 1 as wide as possible.
[0058] According to another embodiment, the directions Cl, C2 of the optical axes represent Sensitive to an orientation of the field of vision of each camera are oriented parallel.
[0059] It is evident that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently.
[0060] The images acquired by cameras 11, 12 at a given acquisition time t1 are in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colors and brightness of objects in the observed scene, for example in the form of RGB (Red Green Blue) or HSL (Hint, Saturation, Luminosity) colorimetric coordinates.
[0061] The images acquired by cameras 11 and 12 represent views of the same scene taken from different viewpoints, the camera positions being distinct. For example, this scene includes: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0062] These images are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.
[0063] A process for generating representative image data for a stereoscopic vision system embedded in the vehicle 10 is advantageously implemented by the vehicle 10, i.e. by a computer or a combination of computers of the embedded system of the vehicle 10, for example by the computer or computers in charge of the stereoscopic vision system of the vehicle 10.
[0064] In a first operation, the computer receives data representative of a first pair of images, the pair of images comprising a first image acquired by the first camera 11 at a time instant and a second image acquired by the second camera 12 at the same time instant.
[0065] The two images received thus correspond to the same three-dimensional scene taking place around vehicle 10 seen from two different points of view.
[0066] In a second operation, to facilitate the analysis of the two received images, the first and second images are, for example, rectified using a method known to those skilled in the art. Such a method is described, for example, in "Projective Rectification of Uncalibrated Infrared Stereo Images with Global Consideration of Distortion Minimization" by Benoit Ducarouge, Thierry Sentenac, Florian Bugarin and Michel Devy, July 16, 2009.
[0067] The rectification method consists of reorienting the epipolar lines so that they are parallel with the horizontal axis of the image. This method is described by a transformation that projects the epipoles to infinity and whose corresponding points are necessarily on the same ordinate.
[0068] A rectification algorithm consists, for example, of 4 steps: - Rotate (virtually) the first camera 11 so that the epipole goes to infinity along the horizontal axis of the frame associated with it; - Apply the same rotation to the second camera to return to the initial geometric configuration; - Rotate the second camera by the rotation associated with the rotation matrix 'R', corresponding to the extrinsic parameter of the starting stereoscopic vision system; - Adjust the scale in both camera reference points.
[0069] It should be noted that rectification simplifies the matching of pixels in images obtained by a stereoscopic vision system. The pixel in the second image corresponding to a pixel in the first image (and vice versa) is positioned on the same line. Based on knowledge of the epipolar geometry and therefore of a fundamental matrix of the stereoscopic vision system, the objective is then to determine a pair of projective transformations, called homographies, which reorient the epipolar projections parallel to the image lines, and therefore to the horizontal axis of the rectified cameras.
[0070] It is also possible not to rectify the received images; disparities induced by the relative position of the optical axes of the first 11 and second 12 cameras must then be taken into account when predicting depths from optical flux values.
[0071] In a third operation, a second set of pixels from the second image corresponding to a first set of pixels from the first image is determined. This operation is commonly called "stereo matching" (in English, "feature matching").
[0072] Stereo matching consists of searching for pixels in stereoscopic views that correspond to the same 3D point in the three-dimensional scene. Rectified epipolar geometry simplifies this process of finding matches on the same epipolar line. It is not necessary to calculate the coordinates of the 3D point to find the corresponding pixel on the same line in the other image.
[0073] In a fourth operation, initial depths associated with the first set of pixels are predicted by a prediction model based on optical flux values determined from the first set of pixels and the second set of pixels from the second image.
[0074] The prediction model uses a method known as optical flow computation. Such a method is notably described in "UnOS: Unified Unsupervised Optical-flow and Stereo-depth Estimation by Watching Videos" by Yang Wang, Peng Wang, Zhenheng Yang, Chenxu Luo, Yi Yang and Wei Xu, June 2019.
[0075] The optical flow calculation method is, for example, implemented by a convolutional neural network (CNN). This type of tool is commonly used in image processing.
[0076] An optical flow is in particular representative of a displacement between each pixel of the first set of pixels of the first image and a pixel of the second set of pixels corresponding to each first pixel in the second image.
[0077] Similar to the fourth operation, in a fifth operation, second depths associated with the second set of pixels are predicted as a function of optical flux values determined from the second set of pixels and the first set of pixels.
[0078] The prediction model has, moreover, been previously learned by any learning method known to a person skilled in the art, whether using training data provided to the stereoscopic vision system or using data generated by the stereoscopic vision system itself, the prediction model being learned, for example, by minimizing a loss function comparing the predicted depths to actual depths. The predictions of the first and second depths are thus reliable.
[0079] In a sixth operation, a second pair of images is generated from the first pair of images. The second pair of images comprises the first image and a third image generated from the second image in which a colorimetric value of each pixel in a third set of pixels is modified.
[0080] Such a third image corresponds to a so-called "contaminated" image, that is to say that the pixels of the third set represent defects, due for example to problems in acquiring the second image or errors in transmitting the data representing the second image.
[0081] According to a particular embodiment, the third set of pixels comprises isolated pixels. The number of pixels in the third set is, in particular, less than a threshold defined relative to the number of pixels in the second image. This threshold represents, for example, 5%, 10%, or 15% of the total number of pixels in the second image.
[0082] According to another particular embodiment, the third set of pixels comprises at least one group of adjacent pixels, thus forming an “area contaminated”. A group of adjacent “contaminated” pixels represents, for example, an image acquired by the second camera 12 through a dirty lens.
[0083] According to a particular embodiment, at least one group of adjacent pixels forms a rectangle or a square. The width of the rectangle or square is, for example, less than a threshold defined relative to the width of the second image. For example, the width of the rectangle or square is less than 5% of the width of the second image, and is equal to 1%, 3%, or 5% of the width of the second image.
[0084] In the case of "contaminated areas", the total surface area of these areas is, for example, limited in relation to the total surface area of the second image, for example 3%, 7% or 10%.
[0085] The sixth operation consists of modifying the colorimetric value of each pixel in the third set of pixels.
[0086] According to a particular embodiment, the pixels of the third set of pixels are changed into black pixels, a black pixel being a pixel whose set of colorimetric values is zero.
[0087] Such a defect corresponds for example to an opaque object or mud present on the lens of the second camera 12 at the time of acquisition of the second image.
[0088] According to another particular embodiment, the pixels of the third set of pixels are changed into white pixels, a white pixel being a pixel whose set of colorimetric values is maximal.
[0089] This type of "contamination" corresponds, for example, to glare, occurring for example when vehicle 10 crosses a second vehicle and a headlight of the second vehicle illuminates the lens of the second camera 12 at the time of acquisition of the second image.
[0090] In a seventh operation similar to the third operation, a fifth set of pixels of the third image corresponding to a fourth set of pixels of the first image is determined.
[0091] In an eighth operation, third depths associated with the fourth set of pixels of the first image are predicted by the prediction model as a function of optical flux values determined from the fourth set of pixels and the fifth set of pixels of the third image.
[0092] Similarly, in a ninth operation, fourth depths associated with the fifth set of pixels are predicted by the prediction model as a function of optical flux values determined from the fifth set of pixels and the fourth set of pixels.
[0093] The prediction model is then learned in a tenth operation using any method known to a person skilled in the art. The prediction model is, for example learned based on the result of a comparison of the third and fourth depths to the first and second depths.
[0094] A first error is, for example, determined for the pixels of the first image by comparing the first and third depths. The first error is, for example, defined by the following function: [Math.l] 1^)= Il dAp)-d'sSp) Il 2
[0095] With: • L^p) the first error for a pixel in the first image, • D^p) the first depth of a pixel in the first image, and • D'^p} the third depth of a pixel of the first image.
[0096] A second error is, for example, determined for the pixels of the second image by comparing the second and fourth depths. The second error is, for example, defined by the following function: [Math.2] = He MP) -D'JP) He 2
[0097] With: • lAp) the second error for a pixel in the second image, • Dst{ p) the second depth of a pixel in the second image, and • D\tÇp) the fourth depth of a pixel of the second image.
[0098] According to a particular embodiment, learning is achieved by minimizing the following loss function: [Math.3] Ls = p), L2(p)^
[0099] With: * The loss due to contamination, • L](p) the first error for a pixel of the first image, and • L^p) the second error for a pixel of the second image.
[0100] According to a particular embodiment, in an eleventh operation, a third pair of images is generated from the first pair of images, the third pair of images comprising a fourth image and the second image.
[0101] Similar to the sixth operation, the fourth image is generated from the first image in which a colorimetric value of pixels from a sixth set of pixels is modified. The modifications are, for example, those presented previously.
[0102] In a twelfth operation similar to the third operation, an eighth set of pixels of the second image corresponding to a seventh set of pixels of the fourth image is determined.
[0103] In a thirteenth operation similar to the fourth operation, fifth depths associated with the seventh set of pixels are predicted by a prediction model based on optical flux values determined from the seventh and eighth sets of pixels.
[0104] In a fourteenth operation, sixth depths associated with the eighth set of pixels are predicted by a prediction model based on optical flux values determined from the eighth set of pixels and the seventh set of pixels.
[0105] In a manner similar to the previous learning, the prediction model is further learned as a function of the fifth and sixth depths by comparison with the first and second depths.
[0106] According to a particular embodiment, several successive learning operations are performed, with the first and second images being "contaminated" in turn. A complete learning sequence is obtained, for example, after the following successive "contaminations": - the first image is contaminated by individual black pixels, - the second image is contaminated by individual black pixels, - the first image is contaminated by individual white pixels, - the second image is contaminated by individual white pixels, - the first image is contaminated by squares of black pixels, - the second image is contaminated by squares of black pixels, - the first image is contaminated by squares of white pixels, - the second image is contaminated by squares of white pixels.
[0107] Thus, the "contamination" of an image allows the prediction model to learn to predict depths for pixels that are "contaminated" or have distorted colorimetric values. The training data is obtained from data acquired by the stereoscopic vision system; it is therefore easy to obtain and representative of real-world situations.
[0108] The prediction model is thus more robust and will be able to predict depths even if the input images have slight defects.
[0109] If the ADAS uses input data such as depths determined by the prediction model learned using the method described above to determine the distance between a part of the vehicle 10, for example the front bumper, and another road user, the ADAS is then able to accurately determine this distance even when one of the cameras has a defect such as a lens dirt or localized glare.
[0110] Fig. 2 illustrates a flowchart of the different steps of a process 2 for generating representative image data for a stereoscopic vision system embedded in the vehicle of Fig. 1.
[0111] The method is implemented for example by one or more processors of one or more computers embedded in the vehicle 10, for example by a computer controlling a stereoscopic vision system comprising a first 11 and a second 12 cameras.
[0112] In a step 21, the computer receives data representative of a first pair of images, the pair of images comprising a first image acquired by the first camera 11 at a time instant and a second image acquired by the second camera 12 at the same time instant.
[0113] In a step 22, first depths associated with a first set of pixels of the first image are predicted by a prediction model based on optical flux values determined from the first set of pixels and a second set of pixels of the second image corresponding to the first set of pixels.
[0114] Second depths associated with the second set of pixels are also predicted by the prediction model as a function of optical flux values determined from the second set of pixels and the first set of pixels.
[0115] In step 23, a second pair of images is generated from the first pair of images, the second pair of images comprising the first image and a third image. The third image is generated from the second image in which a colorimetric value of each pixel in a third set of pixels is modified.
[0116] In a step 24, third depths associated with a fourth set of pixels of the first image are predicted by the prediction model as a function of optical flux values determined from the fourth set of pixels and a fifth set of pixels of the third image corresponding to the fourth set of pixels and fourth depths associated with the fifth set of pixels are predicted as a function of optical flux values determined from the fifth set of pixels and the fourth set of pixels.
[0117] In a step 25, the prediction model is learned based on a result of a comparison of the third and fourth depths to the first and second depths.
[0118] Figure 3 schematically illustrates a device 4 configured to generate representative image data for a stereoscopic vision system mounted in a vehicle 10, according to a particular and non-limiting embodiment of the present invention. Device 4 corresponds for example to a device embedded in the first vehicle 10, for example a computer.
[0119] Device 4 is, for example, configured to carry out the operations described opposite [Fig. 1] and / or the steps described opposite [Fig. 2]. Examples of such a device 4 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 4, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 4 may be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.
[0120] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.
[0121] The computer code of the embedded software(s), including the instructions to be loaded and executed by the processor, is for example stored on memory 4L
[0122] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.
[0123] According to a particular and non-limiting embodiment, the device 4 includes a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 include one or more of the following interfaces: - Radio frequency (RF) interface, for example, Wi-Fi® (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox using UBN (Ultra Narrow Band) radio technology narrow), or LoRa in the 868 MHz frequency band, LTE (from the English "Long-Term Evolution" or in French "Evolution à long terme"), LTE-Advanced (or in French LTE-avancé); - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HD MI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").
[0124] According to another particular and non-limiting embodiment, the device 4 includes a communication interface 43 which enables communication with other devices (such as other computers in the embedded system) via a communication channel 430. The communication interface 43 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds, for example, to a wired network of the CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3) type.
[0125] According to a particular and non-limiting embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch or not, one or more loudspeakers 450 and / or other peripherals 460 (projection system) via the output interfaces 44, 45, 46 respectively. According to a variant, one or more of the external devices is integrated into the device 4.
[0126] Of course, the present invention is not limited to the embodiments described above but extends to a method for generating representative image data for a stereoscopic vision system embedded in a vehicle and for training a depth prediction model for a stereoscopic vision system using images generated from images acquired by this system, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0127] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 4 of [Fig.3].
Claims
1. Demands Method for generating representative image data for a stereoscopic vision system embedded in a vehicle (10), the stereoscopic vision system comprising a first camera (11) and a second camera (12) arranged to acquire an image of a three-dimensional scene from two different viewpoints, said method being characterized in that it comprises the following steps: - reception (21) of data representative of a first pair of images, said pair of images comprising a first image acquired by a first camera (11) at a time instant and a second image acquired by a second camera (12) at the same time instant; - prediction (22) by a prediction model of first depths associated with a first set of pixels of the first image as a function of optical flux values determined from said first set of pixels and a second set of pixels of the second image corresponding to said first set of pixels and prediction of second depths associated with said second set of pixels as a function of optical flux values determined from said second set of pixels and said first set of pixels; - generation (23) of a second pair of images from said first pair of images, said second pair of images comprising said first image and a third image, said third image being generated from said second image for which a colorimetric value of each pixel of a third set of pixels is modified; - prediction (24) by said prediction model of third depths associated with a fourth set of pixels of the first image as a function of optical flux values determined from said fourth set of pixels and a fifth set of pixels of the third image corresponding to said fourth set of pixels and prediction of fourth depths associated with said fifth set of pixels as a function of optical flux values determined from said fifth set of pixels and said fourth set of pixels; - learning (25) said prediction model based on a result from a comparison of the third and fourth depths to the first and second depths.
2. Method according to claim 1, wherein said third set of pixels comprises at least one group of adjacent pixels.
3. A method according to claim 2, wherein said at least one group of adjacent pixels forms a rectangle.
4. A method according to claim 1 to 3, wherein the pixels of said third set of pixels are modified into black pixels, a black pixel being a pixel whose set of colorimetric values is zero.
5. A method according to claim 1 to 3, wherein the pixels of said third set of pixels are modified into white pixels, a white pixel being a pixel whose set of colorimetric values is maximal.
6. A method according to any one of claims 1 to 5, wherein said third set of pixels comprises a number of pixels less than a threshold, said threshold being defined in relation to the number of pixels of said second image.
7. A method according to claim 6, wherein said threshold is equal to 5% of the number of pixels of said second image.
8. A method according to any one of claims 1 to 6, comprising the following steps: - generating a third pair of images from said first pair of images, said third pair of images comprising a fourth image and said second image, said fourth image being generated from said first image in which a colorimetric value of pixels from a sixth set of pixels is modified; - predicting by said prediction model fifth depths associated with a seventh set of pixels from the fourth image as a function of optical flux values determined from said seventh set of pixels and an eighth set of pixels from the second image corresponding to said seventh set of pixels and predicting sixth depths associated with said eighth set of pixels as a function of optical flux values determined from said eighth set of pixels and said seventh set of pixels;said prediction model being learned (25), furthermore, as a function of the fifth and sixth depths by comparison with the first and second depths.;
9. Device (4) for generating representative image data for a stereoscopic vision system embedded in a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for carrying out the steps of the method according to any one of claims 1 to 8.
10. Vehicle (10) comprising the device (4) according to claim 9.