Method and device for determining a visibility mask by a stereoscopic vision system on board a vehicle.

The method for determining a visibility mask in stereoscopic vision systems using image symmetry and neural networks addresses the complexity of existing algorithms, improving ADAS system data quality and safety by accurately identifying occluded areas.

FR3153682B1Active Publication Date: 2025-08-29STELLANTIS AUTO SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2023010303
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-09-28
Publication Date
2025-08-29
Estimated Expiration
2043-09-28

AI Technical Summary

Technical Problem

Existing stereoscopic vision systems in vehicles face challenges in accurately determining visibility masks, which are crucial for enhancing the performance of ADAS systems, as they rely on algorithms that require optical flow methods or multiple image comparisons, making them complex and resource-intensive.

Method used

A method for determining a visibility mask using a stereoscopic vision system with at least two cameras, involving image symmetry, rectification, disparity calculation, and error comparison to define a visibility mask, utilizing a convolutional neural network for self-supervised learning without annotated data.

Benefits of technology

Improves the quality of data for ADAS systems by accurately identifying occluded areas, enhancing operational safety and reducing computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000023_0000
    Figure 00000023_0000
  • Figure 00000024_0000
    Figure 00000024_0000
  • Figure 00000024_0001
    Figure 00000024_0001
Patent Text Reader

Abstract

Method or device implementing a method for determining a visibility mask by a stereoscopic vision system on board a vehicle, the vision system comprising at least two cameras arranged so as to acquire images of a scene from a different point of view. Indeed, the method comprises the reception of first and second images (31a, 32a), a determination of symmetrical images (31b, 32b) of the first and second images, the determination of disparities and the reconstruction of images from the received and symmetrical images and from the determined disparities and the determination of reconstruction errors by comparing the reconstructed images to the received images. A visibility mask associated with the pixels of the first image is determined from the reconstruction errors determined by comparing the errors to predefined threshold values. Figure for the abstract: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method and device for determining a visibility mask by a stereoscopic vision system on board a vehicle. Technical field

[0001] The present invention relates to methods and devices for determining a visibility mask by a stereoscopic vision system on board a vehicle, for example in a motor vehicle. The present invention also relates to a method and a device for controlling one or more AD AS systems on board a vehicle from a stereoscopic vision system on board a vehicle. Technological background

[0002] Many modern vehicles are equipped with so-called AD AS (Advanced Driver-Assistance System). Such AD AS systems are passive and active safety systems designed to eliminate the element of human error in the driving of vehicles of all types. AD AS use advanced technologies to assist the driver while driving and thus improve their performance. AD AS use a combination of sensor technologies to perceive the environment around a vehicle, then provide information to the driver or act on certain vehicle systems.

[0003] There are several levels of ADAS, such as rearview cameras and blind spot sensors, lane departure warning systems, adaptive cruise control and automatic parking systems.

[0004] The AD AS embedded in a vehicle are supplied with data obtained from one or more embedded sensors such as, for example, cameras. These cameras make it possible in particular to detect and locate other road users or possible obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle speed; - to act on the braking system in the event of a risk of impact with an object.

[0005] The proper functioning of the driving assistance peripherals using this data therefore depends on the quality of the data emitted by a vision system.

[0006] Many vision systems perceive an environment around a vehicle from several images acquired by one or more cameras. When exploiting the images, occluded areas of the images which correspond to areas of the environment that are not present on all the acquired images are defined. A visibility mask associated with an image then defines, for example, a filter making it possible to determine pixels associated with areas not found in the other images.

[0007] Solutions exist for detecting an occlusion, that is to say for determining a visibility mask.

[0008] A first solution presented by "Occlusion Aware Unsupervised Learning of Optical Flow" by Yang Wang, Yi Yang, Zhenheng Yang, Liang Zhao, Peng Wang and Wei Xu published on April 4, 2018 is based on reverse optical flow. For each pixel of a first image represented by its coordinates, the algorithm checks whether a pixel of a second image arrives at this pixel of the first image with reverse optical flow by scanning all the pixels of the second image. This method can be used for both directions of optical flow to identify the occluded areas of the two images. Such a method, however, requires the use of an algorithm based on the optical flow method.

[0009] A second solution described in "Digging Into Self-Supervised Monocular Depth Estimation" by Clément Godard, Oisin Mac, Michael Firman and Gabriel Brostow uses a loss function with an algorithm avoiding occluded areas without explicitly identifying them. Two reconstructions of a first image acquired at a given time instant are made on shooting times directly before and after the time instant of acquisition of the first image and are compared to two other images acquired at the time instant directly before and at the time instant directly after. The occluded areas in one image may be present in another image to be reconstructed. A loss function is calculated for the images and the smallest error for each pixel is added to a total error. This solution requires the use of three images. Summary of the present invention

[0010] An object of the present invention is to solve at least one of the problems of the technological background described above.

[0011] Another object of the present invention is to propose a solution for determining a visibility mask for a stereoscopic vision system in order to improve the quality of the data from the cameras of the stereoscopic vision system.

[0012] Another object of the present invention is to improve road safety, in particular by improving the operational safety of AD AS systems supplied by data obtained from at least one camera.

[0013] According to a first aspect, the present invention relates to a method for determining a visibility mask by a stereoscopic vision system on board a vehicle, the stereoscopic vision system comprising at least two cameras each arranged so as to acquire an image of a scene from a different point of view, the method being characterized in that it comprises the following steps: - reception of first and second data respectively representative of a first and second image acquired by respectively a first and second camera of the set of cameras at the same acquisition time instant; - determination of a third image by symmetry of the second image with respect to a vertical axis of the second image and determination of a fourth image by symmetry of the first image with respect to a vertical axis of the first image; - rectification of the first and second images, respectively third and fourth images according to extrinsic parameters of the stereoscopic vision system, so as to render horizontal epipolar lines in the first and second rectified images, respectively in the third and fourth rectified images; - determining a second set of pixels of said second rectified image corresponding to a first set of pixels of the first rectified image, and determining first disparities associating the pixels of the first set of pixels with the pixels of the second set of pixels, and determining a fourth set of pixels of the fourth rectified image corresponding to a third set of pixels of the third rectified image and determining second disparities associating the pixels of the third set of pixels with the pixels of the fourth set of pixels, - reconstruction of a fifth image from the first rectified image and the first disparities and reconstruction of a sixth image from the third rectified image and the second disparities; - determining a first reconstruction error by comparing the fifth image and the second rectified image and determining a second reconstruction error by comparing the sixth image and the fourth rectified image; - associating the second error with a fifth set of pixels of the first rectified image, the fifth set of pixels corresponding to the fourth set of pixels obtained by the symmetry of the first image; - determining a visibility mask associated with the first set of pixels of the first image from the first and second errors by respective comparison of the first and second errors to predefined threshold values.

[0014] The visibility mask representative of the pixels of the first image associated with an object of the scene not visible in the second image is thus determined.

[0015] According to a variant of the method, the first and second disparities are determined by implementing a convolutional neural network.

[0016] According to an additional variant, the method comprises a step of determining a third error from an average of the first and second errors and adjustment of input parameters of the convolutional neural network by minimizing the third error.

[0017] The convolutional neural network is thus learned in a self-supervised manner by minimizing a determined error for images acquired by the stereoscopic vision system, i.e. without requiring annotated data.

[0018] According to yet another variant of the method, the third error is obtained by the following loss function: L* - Epavg(L4p), L Jp)) With: • the third error, • an average of arguments, • L.^p) the first reconstruction error for a pixel P of the second rectified image, • L^^p) the second reconstruction error for a pixel P of the first rectified image.

[0019] According to a variant of the additional method, the first and second errors respectively are obtained by the following loss function: l(p)= E p [(la)'|Z(p)-î(p)|+a-(14SS™( / ( 7) ),î(p)))]With : • L^p) the first, respectively second error, • I(p) a value of the pixel P in the second rectified image, respectively the fourth rectified image, * l^p^ a value of the pixel P in the reconstructed image: the fifth image, respectively the sixth image' • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.

[0020] According to another variant of the method, the visibility mask is determined by the following function: ^(P) =1( / \LdP)<b)> c)With: * V*t(.P) 'visibility of a pixel P of the first image, • 1 a function returning 0 or 1, • (J the union of pixels, • L.^p) said first reconstruction error for a pixel P of the second rectified image, • L^p) said second reconstruction error for a pixel P of the first rectified image, • an AND operator, and • a, b and c are determined parameters.

[0021] According to another variant of the method, the reconstruction of the fifth and respectively sixth images is obtained by the following formula: ^ = ^-d(^)With: • p* the abscissa of a pixel of the fifth image, respectively sixth image, • the abscissa of a pixel of the first rectified image, respectively third image, and • d( p ) a first disparity determined for a pixel P; of the first image, respectively a second disparity determined for a pixel Pt of the third image

[0022] According to a second aspect, the present invention relates to a device for determining a visibility mask by a stereoscopic vision system on board a vehicle, the device comprising a memory associated with at least one processor configured for implementing the steps of the method according to the first aspect of the present invention.

[0023] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0024] According to a fourth aspect, the present invention relates to a computer program which comprises instructions adapted for executing the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0025] Such a computer program may use any programming language and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0026] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to the first aspect of the present invention.

[0027] On the one hand, the recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM memory, a CD-ROM or a microelectronic circuit type ROM memory, or a magnetic recording means or a hard disk.

[0028] On the other hand, this recording medium may also be a transmissible medium such as an electrical or optical signal, such a signal being able to be conveyed via an electrical or optical cable, by conventional or hertzian radio or by self-directed laser beam or by other means. The computer program according to the present invention can in particular be downloaded from an Internet-type network.

[0029] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to perform or to be used in performing the method in question. Brief description of the figures

[0030] Other characteristics and advantages of the present invention will emerge from the description of the particular and non-limiting exemplary embodiments of the present invention below, with reference to the appended figures 1 to 4, in which:

[0031] [Fig-1] schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention;

[0032] [Fig.2] illustrates a flowchart of the different stages of a determination process mination of a visibility mask by vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;

[0033] [Fig.3] schematically illustrates images obtained from a vision system stereoscopic on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention;

[0034] [Fig.4] schematically illustrates a device configured for the determination of a Visibility mask by on-board vision system in the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention. Description of the exemplary embodiments

[0035] A method and a device for determining a visibility mask by a vision system on board a vehicle will now be described in what follows with joint reference to Figures 1 to 4. The same elements are identified with the same reference signs throughout the description which follows.

[0036] The terms "first(s)", "second(s)" (or "first(s)", "second(s)"), etc. are used in this document by arbitrary convention to enable different elements (such as operations, means, etc.) implemented in the embodiments described below to be identified and distinguished. Such elements may be distinct or correspond to a single element, depending on the embodiment.

[0037] According to a particular and non-limiting example of embodiment of the present invention, a method for determining a visibility mask by a stereoscopic vision system on board a vehicle is for example implemented by a computer of the on-board system of the vehicle controlling this stereoscopic vision system.

[0038] The vision system comprises at least two cameras, each arranged to acquire an image of a scene from a different point of view.

[0039] For this purpose, the method for determining a visibility mask by a vision system on board a vehicle comprises receiving first and second data respectively representative of a first and second image acquired by respectively a first and second camera of the set of cameras at the same acquisition time instant.

[0040] The method also comprises determining a third image by symmetry of the second image, determining a fourth image by symmetry of the first image and rectifying the first, second, third and fourth images according to extrinsic parameters of the stereoscopic vision system.

[0041] First disparities associated with pixels of the first image are determined from the first and second images and disparities associated with pixels of the third image are determined from the third and fourth images.

[0042] A first reconstruction error is then determined by comparing the second image to a fifth image reconstructed from the first rectified image and the first disparities.

[0043] Similarly, a second reconstruction error is determined by comparing the fourth image to a sixth image reconstructed from the third rectified image and the second disparities.

[0044] A visibility mask associated with pixels of the first image is then determined from the first and second errors by respective comparison of the first and second errors to predefined threshold values.

[0045] [Fig. 1] schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.

[0046] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, with an electric motor(s) or even a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as an automobile, a truck, a bus, a motorcycle. Finally, the vehicle 10 corresponds to an autonomous vehicle or not, that is to say a vehicle traveling according to a determined level of autonomy or under the total supervision of the driver.

[0047] The vehicle 10 advantageously comprises several on-board cameras 11, each configured to acquire images of a scene in the environment of the vehicle 10. This set of cameras 11 forms the stereoscopic vision system. Two cameras 11 are illustrated in [Fig.l]. The present invention is however not limited to a stereoscopic vision system comprising two cameras but extends to any vision system comprising 1 or more cameras, for example 1, 2, or 5 cameras. A system comprising a single camera 11 then forms a monoscopic vision system.

[0048] The cameras 11 have known intrinsic parameters. These parameters consist in particular of: - focal length f of the camera 11; - distortions which are due to imperfections in the optical system of each camera; - direction C of the optical axis of the camera 11; - respective resolutions of the cameras 11.

[0049] The intrinsic parameters characterize the transformation which associates, for an image point, the camera coordinates with the pixel coordinates, in each camera. These parameters do not change if the camera is moved.

[0050] The distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions which generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, that is to say that the distortions are not taken into account or that their correction is processed at the time of image acquisition.

[0051] These cameras 11 are arranged so as to each acquire an image of a scene from a different point of view, the first point of view is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10, the second point of view is for example located on or in the right rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. In the case where two cameras are located at the top of the windshield of the vehicle, they are then placed at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, the second camera 11 is located in the right rearview mirror of the vehicle 10.

[0052] A first reference point is associated with the first camera 11: - the direction of the y axis is defined by the position of the second camera 11, so as to place the second camera 11 on the y axis of the first camera 11. The distance B separating the two cameras 11 is called the reference base (in English “baseline”) and the direction separating the two cameras 11 is that of the y axis; - the direction of the x axis is defined orthogonal to that of the y axis and orthogonal to that of the optical axis Cl of the first camera 11; - the direction of the z axis is defined orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal reference frame.

[0053] The extrinsic parameters linked to the position of the cameras 11 are the following parameters: - 3 translations in the x, y and z directions: Tx, Ty and Tz constituting the translation vector T; and - 3 rotations around the x, y and z axes: Rx, Ry and Rz, constituting the rotation matrix R.

[0054] According to a particular embodiment, the first 11 and second 12 cameras are arranged relative to each other in a symmetrical manner relative to a plane, the plane being for example normal to the axis connecting the first 11 and second 12 cameras and located midway between the two cameras 11 and 12, that is to say at a distance from one of the cameras 11, 12 equal to half of the reference base, B / 2.

[0055] A main constraint of the stereoscopic vision system used in automobiles is, for example, the large distance between the two cameras. Indeed, to be able to cover a measurement range of 200 meters, the reference base must reach 60cm for the cameras commonly used in this field.

[0056] The two cameras 11 acquire images of a scene located in front of the vehicle 10, the first camera covering only a first acquisition field 13, the second camera covering only a second acquisition field 14 and the two cameras 11 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic vision of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic vision of the scene by the second camera 11 and the third acquisition field 15 allows a stereoscopic vision of the scene by the stereoscopic vision system composed of the two cameras 11.

[0057] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.

[0058] Among these three fields, field 16 is visible from the second camera 11. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system composed of the second camera 11.

[0059] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable using the monoscopic vision system composed of the second camera 11.

[0060] Finally, field 19 is not visible from any of the cameras. The part of the scene present in this field 19 is therefore not observable.

[0061] The directions C1, C2 of the optical axes representative of an orientation of the field of vision of each camera are oriented non-parallel so as to obtain the third acquisition field 15 of the environment 1 as wide as possible.

[0062] It is obvious that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with differently placed and oriented cameras.

[0063] The images acquired by the cameras 11 at a given acquisition time instant tl are presented in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colors and brightness of objects in the observed scene in the form, for example, of RGB colorimetric coordinates (from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness).

[0064] The images acquired by the cameras 11 represent views of the same scene taken from different viewpoints, the positions of the cameras being distinct. On this scene are found for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.

[0065] These images are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.

[0066] A process for determining a visibility mask by the stereoscopic vision system on board the vehicle 10 is advantageously implemented by the vehicle 10, that is to say by a computer or a combination of computers of the on-board system of the vehicle 10, for example by the computer(s) in charge of the stereoscopic vision system of the vehicle 10.

[0067] In a first operation, the computer receives first data representative of a first image 31a acquired by the first camera 11 at an acquisition time instant.

[0068] In a second operation, the computer receives second data representative of a second image 32a acquired by the second camera 12 at the same acquisition time instant.

[0069] The two images 31a and 32a received correspond to two views of the same scene taking place around the vehicle 10.

[0070] In a third operation, a third image 31b is determined by symmetry of the second image 32a with respect to a vertical axis 320 of the second image 32a (such an operation being called “flip” in English). The vertical axis 320 is the central vertical axis of the second image 32a, that is to say the axis which divides the second image 32a in two equal parts, for example, each comprising the same number of pixels.

[0071] Similarly, in a fourth operation, a fourth image 32b is determined by symmetry of the first image 31a with respect to a vertical axis 310 of the first image 31a. The vertical axis 310 is the central vertical axis of the first image 31a, that is to say the axis which divides the first image 31a into two equal parts, for example which each comprise the same number of pixels.

[0072] In this way, the third 31b and fourth 32b images are similar to images acquired by the first 11 and second 12 cameras. In other words, the third image 31b is comparable to an image acquired by the first camera 11 at an acquisition time instant and the fourth image 32b is comparable to an image acquired by the second camera 12 at this same acquisition time instant.

[0073] In a fifth operation, the first 31a and second 32a images are rectified according to extrinsic parameters of the stereoscopic vision system, so as to render horizontal epipolar lines in the first and second rectified images.

[0074] Similarly, in a sixth operation, the third 31b and fourth 32b images are rectified according to extrinsic parameters of the stereoscopic vision system, so as to render horizontal epipolar lines in the third and fourth rectified images.

[0075] The first and second images are rectified according to a method known to those skilled in the art. Such a method is described, for example, in “Projective Rectification of Non-Calibrated Infrared Stereo Images with Global Consideration of Distortion Minimization” by Benoit Ducarouge, Thierry Sentenac, Florian Bugarin and Michel Devy of July 16, 2009.

[0076] The rectification method consists of reorienting the epipolar lines so that they are parallel with the horizontal axis of the image. This method is described by a transformation which projects the epipoles to infinity and whose corresponding points are necessarily on the same ordinate.

[0077] A rectification algorithm consists, for example, of 4 steps: - Rotate (virtually) the first camera 11 so that the epipole goes to infinity along the horizontal axis of the reference frame associated with it; - Apply the same rotation to the second camera to end up in the initial geometric configuration; - Rotate the second camera 12 by the rotation associated with the rotation matrix 'R', corresponding to the extrinsic parameter of the starting stereoscopic vision system; - Adjust the scale in both camera markers.

[0078] It should be noted that rectification simplifies the matching of pixels in stereo images, i.e. images obtained by a stereoscopic vision system. The corresponding pixel in the second image to a pixel in the first image (and vice versa) is positioned on the same line. From the knowledge of the epipolar geometry and therefore of a fundamental matrix of the stereo system, the objective is then to determine a pair of projective transformations, called homographies, which reorient the epipolar projections parallel to the lines of the images, therefore to the horizontal axis of the rectified cameras.

[0079] In a seventh operation, a second set of pixels of the second rectified image corresponding to a first set of pixels of the first rectified image is determined and first disparities associating the pixels of the first set of pixels with the pixels of the second set of pixels are obtained, for example by implementing a convolutional neural network, this operation is called “stereo matching” (in English “feature matching”).

[0080] Stereo matching or disparity estimation is the operation of finding pixels in stereoscopic views that correspond to the same 3D point in the scene. Rectified epipolar geometry simplifies this operation of finding correspondence on the same epipolar line. It is not necessary to calculate the coordinates of the 3D point to find the corresponding pixel on the same line of the other image. Disparity is the distance d between a pixel and its correspondence in the other image. This disparity is horizontal when the images are rectified.

[0081] The output data of this operation is a disparity representative of a displacement between each pixel of the first set of pixels of the first rectified image and the second pixel corresponding to each first pixel in the second image.

[0082] In an eighth operation, a fourth set of pixels of the fourth rectified image corresponding to a third set of pixels of the third rectified image is determined and second disparities associating the pixels of the third set of pixels with the pixels of the fourth set of pixels are obtained, for example by implementing the method previously described.

[0083] In a ninth operation, a fifth image is reconstructed from the first rectified image and the first disparities.

[0084] Similarly, in a tenth operation, a sixth image is reconstructed from the third rectified image and the second disparities.

[0085] According to a particular exemplary embodiment, the reconstruction of the fifth and respectively sixth images is obtained by the following formula:

[0086] [Math.l]

[0087] With: • / f* the abscissa of a pixel of the fifth image, respectively sixth image, • P* the abscissa of a pixel of the first rectified image, respectively third image, and * d(p ) a first disparity determined for a pixel Pt of the first image, respectively a second disparity determined for a pixel Pt of the third image.

[0088] In an eleventh operation, a first reconstruction error of the fifth image is determined by comparing the fifth image and the second rectified image and in a twelfth operation, a second reconstruction error of the sixth image is determined by comparing the sixth image and the fourth rectified image. The first error is determined for a pixel of the second set of pixels and the second error is determined for a pixel of the fourth set of pixels.

[0089] According to a particular exemplary embodiment, the first and respectively second errors are obtained by the following loss function:

[0090] [Math.2] U(p) = ■ UM'lMi+a - (l-±SSIM{l(p)Mp)))]

[0091] With: • L-ip) the first, respectively second error, • I(.p) a value of the pixel P in the second rectified image, respectively the fourth rectified image, • a value of the pixel P in the reconstructed image: the fifth image, respec tively the sixth image' • SSIM a function that takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.

[0092] In a thirteenth operation, the second error is associated with a fifth set of pixels of the first rectified image, the fifth set of pixels corresponding to the fourth set of pixels by the symmetry of the first image.

[0093] Thus, the second error is defined for pixels of the first rectified image, the pixels of the fifth set of pixels corresponding to the pixels of the first rectified image having as image the pixels of the fourth set of pixels obtained by the symmetry carried out during the fourth operation.

[0094] According to a particular exemplary embodiment, in a fourteenth operation, a third error is determined from an average of the first and second errors. The input parameters of the convolutional neural network are then adjusted by minimization of the third error.

[0095] The third error is, for example, obtained by the following loss function:

[0096] [Math.3] = L*s(p))

[0097] With: • L* the third error, • avg an average of arguments, • the first reconstruction error for a pixel P of the second rectified image, • L* / p) the second reconstruction error for a pixel P of the first rectified image.

[0098] The input parameters of the convolutional neural network are then adjusted by minimizing the third error. The convolutional neural network is then said to be self-supervised because it is able to autonomously adjust its input parameters in order to improve the output data, here the first and second geometric parameters. No annotated data is necessary for this learning.

[0099] The use of the average function (in English "average" or "avg") allows, for its part, to take into account the first and second errors in an equivalent manner in order to improve the reliability of the predictions of the convolutional neural network.

[0100] In a fifteenth operation, a visibility mask associated with the first set of pixels of the first image is determined from the first and second errors.

[0101] This visibility mask is determined by respective comparison of the first and second errors with predefined threshold values ​​and is obtained, for example, by the following function:

[0102] [Math.4] =1( UpCWp) >a<b)> c)

[0103] With: * the visibility of a pixel P of the first image, • 1 a function returning 0 or 1, • (J the union of pixels, • L^p} said first reconstruction error for a pixel P of the second rectified image, • L^p) said second reconstruction error for a pixel P of the first rectified image, • an AND operator, and • a, b and c are determined parameters.

[0104] The error of the pixel not visible in the target image, but visible in the source image, must exceed a certain level defined by the value a, when the error of the reconstruction of the other direction must be lower than a certain level defined by the value b. The notable difference between a and b is linked to the maturity of the training of the learning model.

[0105] The symbol U signifies the union of pixels, provided that the number of pixels in this union is greater than the criterion c. From experience the value of c is, for example, greater than 4 pixels. The symbol 1 indicates the visibility mask resulting from a logical operator, either 1 or 0.

[0106] [Fig.2] illustrates a flowchart of the different steps of a method 2 for determining a visibility mask by a stereoscopic vision system on board a vehicle, for example the vehicle of [Fig.l].

[0107] The vision system comprises at least two cameras 11, 12, each arranged so as to acquire an image of a scene from a different point of view at the same time instant, according to a particular and non-limiting exemplary embodiment of the present invention.

[0108] The method 2 is for example implemented by one or more processors of one or more computers on board the vehicle 10, for example by a computer controlling the stereoscopic vision system.

[0109] In a first step 21, the computer receives first data representative of a first image acquired by the first camera 11 at an acquisition time instant.

[0110] In a second step 22, the computer receives second data representative of a second image acquired by the second camera 12 at the same acquisition time instant.

[0111] The two images received correspond to two views of the same scene taking place around the vehicle 10.

[0112] In a step 23, a third image 31b is determined by symmetry of the second image 32a relative to a vertical axis 320 of the second image 32a and a fourth image 32b is determined by symmetry of the first image 31a relative to a vertical axis 310 of the first image 31a.

[0113] In a step 24a, the first 31a and second 32a images are rectified according to extrinsic parameters of the stereoscopic vision system so as to render horizontal epipolar lines in the first and second rectified images.

[0114] In a step 24b, the third 31b and fourth 32b images are rectified according to extrinsic parameters of the stereoscopic vision system so as to render horizontal epipolar lines in the first and second rectified images.

[0115] In a step 25a, a second set of pixels of the second rectified image corresponding to a first set of pixels of the first rectified image is determined and first disparities associating the pixels of the first set of pixels with the pixels of the second set of pixels are determined.

[0116] In a step 25b, a fourth set of pixels of the fourth rectified image corresponding to a third set of pixels of the third rectified image is determined and second disparities associating the pixels of the third set of pixels with the pixels of the fourth set of pixels are determined.

[0117] In a step 26a, a fifth image is reconstructed from the first rectified image and the first disparities.

[0118] In a step 26b, a sixth image is reconstructed from the third rectified image and the second disparities.

[0119] In a step 27a, a first reconstruction error is determined by comparing the fifth image and the second rectified image, the first error being determined for a pixel of the second set of pixels.

[0120] In a step 27b, a second reconstruction error is determined by comparing the sixth image and the fourth rectified image, the second error being determined for a pixel of the fourth set of pixels.

[0121] In a step 28, the second error is associated with a fifth set of pixels of the first rectified image, the fifth set of pixels corresponding to the fourth set of pixels obtained by the symmetry of the first image

[0122] In a step 29, a visibility mask associated with the first set of pixels of the first image is determined from the first and second errors by respective comparison of the first and second errors to predefined threshold values.

[0123] Thus, the definition of a visibility mask makes it possible to identify the pixels visible in the first image and not visible in the second image via one of the embodiments presented.

[0124] If the ADAS uses input data such as the depths determined by the monoscopic vision system to determine the distance between a part of the vehicle 10, for example the front bumper, and another user present on the road, the ADAS is then able to determine whether the predicted depth is reliable when the pixel is clearly visible in the first and second images.

[0125] [Fig. 4] schematically illustrates a device 4 configured for determining a visibility mask by a vision system on board a vehicle 10, according to a particular and non-limiting exemplary embodiment of the present invention. The device 4 corresponds for example to a device on board the first vehicle 10, for example a calculator.

[0126] The device 4 is for example configured for the implementation of the operations and / or steps described with regard to figures 1, 2 and 3. Examples of such a device 4 include, but are not limited to, on-board electronic equipment such as an on-board computer of a vehicle, an electronic calculator such as an ECU (“Electronic Control Unit”), a smartphone, a tablet, a laptop. The elements of the device 4, individually or in combination, can be integrated in a single integrated circuit, in several integrated circuits, and / or in discrete components. The device 4 can be produced in the form of electronic circuits or software (or computer) modules or even a combination of electronic circuits and software modules.

[0127] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the method and / or for executing the instructions of the software(s) embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41 corresponding for example to a volatile and / or non-volatile memory and / or comprises a memory storage device which may comprise volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.

[0128] The computer code of the embedded software(s) comprising the instructions to be loaded and executed by the processor is for example stored in the 4L memory.

[0129] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Telematic Control Unit”), for example via a communication bus or through dedicated input / output ports.

[0130] According to a particular and non-limiting exemplary embodiment, the device 4 comprises a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 comprise one or more of the following interfaces: - RF radio frequency interface, for example Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced (or in French LTE-advanced); - USB interface (from the English “Universal Serial Bus” or “Universal Serial Bus” in French); HD MI interface (from the English “High Definition Multimedia Interface” or “High Definition Multimedia Interface” in French); - LIN interface (from the English “Local Interconnect Network”).

[0131] According to another particular and non-limiting exemplary embodiment, the device 4 comprises a communication interface 43 which makes it possible to establish communication with other devices (such as other computers of the on-board system) via a communication channel 430. The communication interface 43 corresponds for example to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds for example to a wired network of the CAN (from the English “Controller Area Network” or in French “Réseau de contrôles”) type, CAN FD (from the English “Controller Area Network Flexible Data-Rate” or in French “Réseau de contrôles à débit de données flexible”), FlexRay (standardized by the ISO 17458 standard) or Ethernet (standardized by the ISO / IEC 802-3 standard).

[0132] According to a particular and non-limiting exemplary embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch-sensitive or not, one or more speakers 450 and / or other peripherals 460 (projection system) via the output interfaces 44, 45, 46 respectively. According to a variant, one or other of the external devices is integrated into the device 4.

[0133] Of course, the present invention is not limited to the exemplary embodiments described above but extends to a method for determining a visibility mask by a stereoscopic vision system on board a vehicle, which would include secondary steps without thereby departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0134] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-based motor vehicle, comprising the device 4 of [Fig.4].

Claims

1. Claims Method for determining a visibility mask by a stereoscopic vision system on board a vehicle (10), the stereoscopic vision system comprising at least two cameras (11, 12) each arranged so as to acquire an image of a scene from a different point of view, said method being characterized in that it comprises the following steps: - reception (21, 22) of first and second data respectively representative of a first (31a) and second (32a) images acquired by respectively a first (11) and second (12) cameras of said set of cameras at the same acquisition time instant; - determination (23) of a third image (31b) by symmetry of the second image (32a) relative to a vertical axis of the second image (32a) and determination of a fourth image (32b) by symmetry of the first image (31a) relative to a vertical axis of the first image (31a); - rectification (24a, 24b) of the first (31a) and second (32a) images, respectively third (31b) and fourth (32b) images as a function of extrinsic parameters of the stereoscopic vision system, so as to render horizontal epipolar lines in the first and second rectified images, respectively in the third and fourth rectified images; - determining (25a) a second set of pixels of said second rectified image corresponding to a first set of pixels of said first rectified image, and determining first disparities associating the pixels of said first set of pixels with the pixels of said second set of pixels, and determining a fourth set of pixels of said fourth rectified image corresponding to a third set of pixels of said third rectified image and determining (25b) second disparities associating the pixels of said third set of pixels with the pixels of said fourth set of pixels, - reconstruction (26a) of a fifth image from said first rectified image and said first disparities and reconstruction (26b) of a sixth image from said third rectified image and said second disparities; - determining (27a) a first reconstruction error by comparing said fifth image and second rectified image and determining (27b) a second reconstruction error by comparing said sixth image and fourth rectified image, - associating (28) said second error with a fifth set of pixels of said first rectified image, said fifth set of pixels corresponding to said fourth set of pixels obtained by said symmetry of the first image; - determining (29) a visibility mask associated with said first set of pixels of said first image from said first and second errors by respective comparison of said first and second errors with predefined threshold values.

2. The method of claim 1, wherein said first and second disparities are determined by implementing a convolutional neural network.

3. A method according to claim 2, comprising a step of determining (28) a third error from an average of said first and second errors and adjusting input parameters of said convolutional neural network by minimizing said third error.

4. Method according to claim 3, for which said third error is obtained by the following loss function: L.^'Lp^g(L.^p\ L^p))^^ : • L* the third error, • avS an average of arguments, • L» / p) the first reconstruction error for a pixel P of the second rectified image, • L^p) the second reconstruction error for a pixel P of the first rectified image.

5. Method according to one of claims 1 to 4, for which said first and respectively second errors are obtained by the following loss function: L^p) = Ep[(l-«) ■ \l(p)-I(p)\+a-(^-±SSIM(i(p)J(p)])] With: • L / p) the first, respectively second error, • I(p) a value of the pixel P in the second rectified image, respectively the fourth rectified image, * a value of the pixel P in the reconstructed image: the fifth image, respectively the sixth image' • SSIM a function which takes into account a local structure, and • has a weighting factor depending in particular on the type of environment.

6. Method according to claim 5, for which the visibility mask is determined by the following function: =1( UP(Mr)>« AMp) <&) >c')With: * P) the visibility of a pixel P of the first image, • 1 a function making 0 or 1, • (J the union of the pixels, • L»^p) said first reconstruction error for a pixel P of the second rectified image, • L^p) said second reconstruction error for a pixel P of the first rectified image, • an AND operator, and • a, b ande determined parameters.

7. Method according to one of claims 1 to 6, for which the reconstruction of said fifth and respectively sixth images is obtained by the following formula: pl = p^-d(pt)^c- • p* the abscissa of a pixel of the fifth image, respectively sixth image, • p] the abscissa of a pixel of the first rectified image, respectively third image, and * d( p a first disparity determined for a pixel Pt of the first image, respectively a second disparity determined for a pixel Pt of the third image.

8. Computer program comprising instructions for implementing the method according to any one of the preceding claims, when these instructions are executed by a processor.

9. Device (4) for determining a visibility mask by a stereoscopic vision system on board a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for implementing the steps of the method

10. according to any one of claims 1 to 7. Vehicle (10) comprising the device (4) according to claim 9.