Method and device for determining a visibility mask for a vision system on board a vehicle

EP4706002A1Pending Publication Date: 2026-03-11STELLANTIS AUTO SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-09
Publication Date
2026-03-11

AI Technical Summary

Technical Problem

Current methods for determining visibility masks in vision systems for vehicles are resource-intensive and may not accurately identify occluded areas, affecting the quality of data used by ADAS systems.

Method used

A method using a stereoscopic vision system with two cameras to predict depths and reproject pixels into a 3D scene, generating visibility masks by identifying occluded pixels based on field of view and focal distances, without requiring optical flow calculations.

Benefits of technology

This approach improves the efficiency and accuracy of determining visibility masks, reducing computational resources and enhancing the reliability of data for ADAS systems by directly associating occluded pixels with their 3D scene coordinates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FR2024050462_14112024_PF_FP_ABST
    Figure FR2024050462_14112024_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method and a device for determining a visibility mask for a stereoscopic vision system on board a vehicle (10). The vision system comprises at least two cameras (11, 12) for acquiring images of the same three-dimensional scene from determined viewpoints. For this purpose, first and second images are received, depths associated with a set of pixels in the first image are predicted, this set of pixels is reprojected into the three-dimensional scene as a set of points according to the predicted depths and a third image is generated by projecting the set of points. A visibility mask is then determined from the spatial co-ordinates of the points of the set of points and from the co-ordinates of the pixels projected into the third image.
Need to check novelty before this filing date? Find Prior Art

Description

DESCRIPTION Title: Method and device for determining a visibility mask for a vision system on board a vehicle. Technical field

[0001] The present invention claims priority from French application 2304517 filed on 05.05.2023, the content of which (text, drawings and claims) is incorporated herein by reference. The present invention relates to methods and devices for determining a visibility mask for an on-board vision system in a vehicle, for example in a motor vehicle. The present invention also relates to a method and device for controlling one or more ADAS systems on-board a vehicle from a determined visibility mask. Technological background

[0002] Many modern vehicles are equipped with so-called ADAS (Advanced Driver Assistance System). ADAS are passive and active safety systems designed to eliminate human error in the operation of all types of vehicles. ADAS uses advanced technologies to assist the driver while driving and thus improve their performance. ADAS uses a combination of sensor technologies to perceive the environment around a vehicle, then provides information to the driver or influences certain vehicle systems.

[0003] There are several levels of ADAS, such as rearview cameras and blind spot sensors, lane departure warning systems, adaptive cruise control, and automatic parking systems.

[0004] ADAS systems embedded in a vehicle are powered by data obtained from one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or possible obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other users; - to automatically regulate the vehicle's speed; - to act on the braking system in the event of a risk of impact with an object.

[0005] The proper functioning of the driving assistance devices using this data therefore depends on the quality of the data emitted by a vision system.

[0006] Many vision systems perceive the environment around a vehicle from multiple images acquired by one or more cameras. When processing the images, occluded areas of the images are defined, corresponding to areas of the environment that are not present in all the acquired images. A visibility mask associated with an image then defines, for example, a filter to determine pixels associated with areas not found in the other images.

[0007] Solutions exist for detecting occlusion, i.e. for determining a visibility mask.

[0008] A first solution presented by "Occlusion Aware Unsupervised Learning of Optical Flow" by Yang Wang, Yi Yang, Zhenheng Yang, Liang Zhao, Peng Wang and Wei Xu published on April 4, 2018 is based on reverse optical flow. For each pixel of a first image represented by its coordinates, the algorithm checks whether a pixel of a second image arrives at this pixel of the first image with reverse optical flow by scanning all the pixels of the second image. This method can be used for both directions of optical flow to identify occluded areas of both images.

[0009] A second solution is described in the document "Geometry-based Occlusion-Aware Unsupervised Stereo Matching for Autonomous Driving algorithm". The detection of occluded areas is based on a geometric constraint: the occluded pixel and another pixel that hides it are projected into the same pixel of a reconstructed image. Summary of the present invention

[0010] An object of the present invention is to solve at least one of the problems of the technological background described above.

[0011] Another object of the present invention is to propose an alternative solution for determining a visibility mask for any vision system in order to improve the quality of the data from the camera(s) of the vision system.

[0012] Another object of the present invention is to reduce the resources required for determining a visibility mask.

[0013] According to a first aspect, the present invention relates to a method for determining visibility masks by a stereoscopic vision system on board a vehicle, the stereoscopic vision system comprising a set of cameras of at least two cameras arranged so as to each acquire an image of a three-dimensional scene from a different point of view, said second camera being located to the right of the first camera from a point of view of the first camera, the method being characterized in that it comprises the following steps: - reception of first and second data respectively representative of a first and second image acquired by respectively a first and second camera of the set of cameras at the same acquisition time instant;- prediction of depths associated with a set of pixels of the first image by the stereoscopic vision system from a learned prediction model, each pixel of the first image having principal coordinates in the first image; - reprojection in the three-dimensional scene of the set of pixels in the form of a set of points as a function of the depths, of an intrinsic matrix of the first camera and of extrinsic parameters of the stereoscopic vision system; - determination of a first visibility mask associated with the pixels of the set of pixels as a function of the coordinates of points of the set of points, a pixel of the set of pixels being not visible in the second image if the coordinates of a point of said set of points associated with said pixel locate it outside a field of; vision of the second camera determined as a function of a width of the second image and a focal length of the second camera; - generation of a third image by projection of the set of points as a function of an intrinsic matrix of the second camera, secondary coordinates in the third image being associated with each pixel of the set of pixels;- rectification of the first image from the intrinsic and extrinsic parameters of the two cameras, for each row of the first rectified image, scanning from left to right, according to a point of view of the first camera, arrival column index values ​​and detection of a set of irregular pixels whose arrival column index does not follow a monotonic function representative of an evolution of arrival column indices as a function of a column index in the first rectified image, and for each irregular pixel of the set, identification of a set of occluded pixels in the row to the left of each irregular pixel whose arrival column index is greater than or equal to an arrival column index of each irregular pixel, a second visibility mask being determined as a union of the sets of occluded pixels;and - determining a third visibility mask by associating the first visibility mask and the second visibility mask.;

[0014] According to a variant of the process, the reprojection of a pixel into the three-dimensional scene is carried out using the following formula: ^^ 3 ^^ ( ^^ ^^ ) = ^^. ^^( ^^ ^^ | ^^ −1 , ^^( ^^ ^^ )) with : the three-dimensional scene of the point resulting from the reprojection of the pixel ^^ ^^ of the first image, - ^^ a displacement matrix between a position of the first camera and a position of the second camera, - ^^ the intrinsic matrix of the first camera associated with the projection of a point of the three-dimensional scene into an image acquired by the first camera, - ∅ a reprojection function in the three-dimensional scene of a pixel as a function from its depth, - ^^( ^^ ^^ ) is a pixel depth ^^ ^^predicted by the stereoscopic vision system.

[0015] According to another variant of the method, the projection of a point of the three-dimensional scene is carried out using the following formula: ^ ^ ^^ = ^^ ( ^^′ [ ^^3 ^^ ( ^^ ^^ )]) with: - ^^ a function to go from homogeneous coordinates in three-dimensional space to pixel coordinates in two dimensions by removing one dimension from a vector, - ^^′ the intrinsic matrix of the second camera associated with the projection of a point of the three-dimensional scene into an image acquired by the second camera, - ^^ 3 ^^ ( ^^ ^^ ) the coordinates in the three-dimensional scene of the point resulting from the reprojection of the pixel ^^ ^^ of the first image.

[0016] According to yet another variant of the method, the first visibility mask is obtained using the following formula: ^^ ⋃ ^^^^ ^^ ^^ ^^ ^^ ^^ ^^^^3 ^^( ^^ ^^) ^^ / 2 ^^ ^^( ^^ ^^) > ^^ resulting from the reprojection of the pixel ^^ ^^ of the first image, an abscissa axis being defined parallel to an axis along which the first and second cameras are located, - ^^( ^^ ^^ ) the depth of the pixel ^^ ^^ of the first image predicted by the stereoscopic vision system, - ^^ the width of the second image, and - ^^ the focal length of the second camera.

[0017] In yet another method variant, depths are predicted by a convolutional neural network.

[0018] According to an additional method variant, the convolutional neural network is trained to minimize a photometric error defined by the following loss function: ^^ ( ^^ ) = ∑ [ ( 1 − ^^ ) ⋅ | ^^ (^^ ) − ^^ ( ^^ ) | + ^^ ⋅ (1 − 1 ^^ ^^ ^^ ^^ ( ^^ ( ^^ ) , ^^ ^^ ))] ^^ 2 ( ) - ^^ ( ^^) a value of the pixel ^^ in the third image; - SSIM a function that takes into account a local structure; and - ^^ a weighting factor depending on a type of road environment.

[0019] According to a second aspect, the present invention relates to a device for determining a visibility mask for a vision system on board a vehicle, the device comprising a memory associated with at least one processor configured for implementing the steps of the method according to the first aspect of the present invention.

[0020] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.

[0021] According to a fourth aspect, the present invention relates to a computer program which comprises instructions adapted for executing the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.

[0022] Such a computer program may use any programming language and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0023] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to the first aspect of the present invention.

[0024] On the one hand, the recording medium may be any entity or device capable of storing the program. For example, the medium may include a means storage media, such as ROM, CD-ROM, or microelectronic circuit-type ROM, or magnetic recording media or hard disk.

[0025] Furthermore, this recording medium may also be a transmissible medium such as an electrical or optical signal, such a signal being able to be conveyed via an electrical or optical cable, by conventional or hertzian radio or by self-directed laser beam or by other means. The computer program according to the present invention may in particular be downloaded from a network such as the Internet.

[0026] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to perform or to be used in performing the method in question.

[0027] Brief description of the figures

[0028] Other characteristics and advantages of the present invention will emerge from the description of the particular and non-limiting exemplary embodiments of the present invention below, with reference to the appended figures 1 to 5, in which:

[0029] [Fig.1] schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention;

[0030] [Fig.2] schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention;

[0031] [Fig.3] schematically illustrates a device configured for determining a visibility mask by a vision system on board the vehicle of FIG. 1, according to a particular and non-limiting exemplary embodiment of the present invention;

[0032] [Fig.4] illustrates a flowchart of the different steps of a method for determining a visibility mask by a vision system on board the vehicle of Figure 1, according to a particular and non-limiting exemplary embodiment of the present invention.

[0033] [Fig.5] illustrates a matrix having different column indices for pixels of a row and a visibility criterion for an on-board vision system in the vehicle of Figure 1, according to a particular and non-limiting exemplary embodiment of the present invention.

[0034] Description of examples of implementation

[0035] A method and a device for determining a visibility mask for a vision system on board a vehicle will now be described in the following with joint reference to Figures 1 to 5. The same elements are identified with the same reference signs throughout the description which follows.

[0036] According to a particular and non-limiting example of embodiment of the present invention, a method for determining a visibility mask for a stereoscopic vision system on board a vehicle is for example implemented by a computer of the on-board system of the vehicle controlling this vision system.

[0037] The stereoscopic vision system comprises a set of cameras of at least two cameras arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the second camera being located to the right of the point of view of said first camera.

[0038] For this purpose, the method for determining a visibility mask by a stereoscopic vision system on board a vehicle comprises the reception of first and second data respectively representative of a first and second image acquired from a different point of view by the first and second cameras at the same acquisition time instant.

[0039] The method also includes predicting depths associated with a set of pixels of the first image by the stereoscopic vision system from a learned prediction model, each pixel of the first image having principal coordinates in the first image, reprojecting the pixels of the first image into the three-dimensional scene as a set of points based on the depths, an intrinsic matrix of the first camera and extrinsic parameters of the stereoscopic vision system.

[0040] The method then determines a first visibility mask associated with the pixels of the set of pixels as a function of the coordinates of points of the set of points, a pixel of the set of pixels being invisible in the second image if the coordinates of a point of the set of points associated with the pixel locate it outside a field of vision of the second camera determined as a function of a width of the second image and a focal length of the second camera.

[0041] The method then comprises generating a third image by projecting the set of points according to an intrinsic matrix of the second camera, secondary coordinates in the third image being associated with each pixel of the set of pixels. The generation of this third image then allows the determination of a second visibility mask associated with the pixels of the set of pixels. For each row of the rectified first image, arrival column index values ​​are scanned and a set of irregular pixels is detected. For each irregular pixel of the set, a set of occluded pixels is identified in the row to the left of each detected irregular pixel. A second visibility mask is then determined as a union of the sets of occluded pixels.

[0042] A third visibility mask is then determined by associating the first visibility mask and the second visibility mask.

[0043] Figure 1 schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.

[0044] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, an electric motor(s) or a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as an automobile, a truck, a bus, a motorcycle. Finally, the vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say a vehicle circulating according to a determined level of autonomy or under the total supervision of the driver.

[0045] The vehicle 10 advantageously comprises several on-board cameras 11, 12, each configured to acquire images of a three-dimensional scene in the environment of the vehicle 10. This set of cameras 11, 12 forms the stereoscopic vision system. Two cameras 11, 12 are illustrated in FIG. 1. The present invention is however not limited to a stereoscopic vision system comprising two cameras but extends to any vision system comprising two or more cameras, for example 3, 4 or 5 cameras.

[0046] The cameras 11, 12 have known intrinsic parameters. These parameters consist in particular of: - focal length f1 of the first camera 11; - focal length f2 of the second camera 12; - distortions which are due to imperfections in the optical system of each camera; - direction C1 of the optical axis of the first camera 11; - direction C2 of the optical axis of the second camera 12; - respective resolutions of the cameras 11, 12.

[0047] Intrinsic parameters characterize the transformation that associates, for an image point, the camera coordinates with the pixel coordinates, in each camera. These parameters do not change if the camera is moved.

[0048] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, that is to say that the distortions are not taken into account or that their correction is processed at the time of image acquisition.

[0049] These cameras 11, 12 are arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the second camera 12 being located to the right of the point of view of said first camera 11, the first point of view is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10, the second point of view is for example located on or in the right rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. In the case where two cameras are located at the top of the windshield of the vehicle, they are then placed at a certain distance.

[0050] A first reference frame is associated with the first camera 11: - the direction of the y axis is defined by the position of the second camera 11, so as to place the second camera 12 on the y axis of the first camera 11. The distance B separating the two cameras 11, 12 is called the reference base (in English "baseline") and the direction separating the two cameras 11, 12 is that of the y axis; - the direction of the x axis is defined orthogonal to that of the y axis and orthogonal to that of the optical axis C1 of the first camera 11; - the direction of the z axis is defined orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal reference frame.

[0051] The extrinsic parameters related to the position of the cameras 11, 12 are the following parameters: - 3 translations in the x, y and z directions: Tx, Ty and Tz constituting the translation vector T; and - 3 rotations around the x, y and z axes: Rx, Ry and Rz, constituting the rotation matrix R.

[0052] A major constraint of the stereoscopic vision system used in automotive applications, for example, is the large distance between the two cameras. In fact, to be able to cover a measuring range of 200 meters, the baseline must reach 60 cm for the cameras commonly used in this field.

[0053] The two cameras 11, 12 acquire images of a three-dimensional scene located in front of the vehicle 10, the first camera 11 covering only a first acquisition field 13, the second camera 12 covering only a second field acquisition 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic vision of the three-dimensional scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic vision of the three-dimensional scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic vision of the three-dimensional scene by the stereoscopic vision system composed of the two cameras 11, 12.

[0054] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.

[0055] Among these three fields, field 16 is visible from the second camera 12. The part of the three-dimensional scene present in this field 16 is therefore observable using the monoscopic vision system composed of the second camera 12.

[0056] Field 17 is visible from the first camera 11. The part of the three-dimensional scene present in this field 17 is therefore observable using the monoscopic vision system composed of the second camera 12.

[0057] Finally, field 19 is not visible from any of the cameras. The part of the three-dimensional scene present in this field 19 is therefore not observable.

[0058] The directions C1, C2 of the optical axes are representative of an orientation of the field of vision of each camera 11, 12.

[0059] It is obvious that it is possible to use such a stereoscopic vision system to take images of three-dimensional scenes located on the sides or behind the vehicle 10 by equipping it with differently placed and oriented cameras.

[0060] The images acquired by the cameras 11, 12 at a given acquisition time instant are presented in the form of data representing pixels characterized by: - ​​coordinates in each image; and - data relating to the colors and brightness of the objects in the observed three-dimensional scene in the form, for example, of colorimetric coordinates. RGB (from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness).

[0061] The images acquired by the cameras 11, 12 represent views of the same three-dimensional scene taken from different viewpoints, the positions of the cameras being distinct. On this three-dimensional scene are for example: - buildings; - road infrastructures; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.

[0062] These images are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.

[0063] Figure 2 schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.

[0064] Points 20, 21, 22 of the three-dimensional scene are visible from the point of view of the first camera 11.

[0065] Points 21 are also visible from the viewpoint of the second camera 12.

[0066] Points 20 are occluded from the point of view of the second camera 12 because they are masked by points 21 located on the same axes in the field of vision of the second camera 12.

[0067] Points 22 are located outside the field of view of the second camera 12.

[0068] Thus, during the acquisition of the first and second images by the first camera 11 and the second camera 12 respectively, pixels associated with the visible points 20, 21, 22 from the point of view of the first camera 11 will be present in the first image, whereas only pixels associated with the points 21 visible from the point of view of the second camera will be present in the second image.

[0069] The pixels associated with the points 20 visible from the point of view of the first camera 11 and not visible from the point of view of the second camera 12 are subsequently called occluded pixels.

[0070] Figure 3 schematically illustrates a device 4 configured for determining a visibility mask for a vision system embedded in a vehicle 10, according to a particular and non-limiting exemplary embodiment of the present invention. The device 4 corresponds for example to a device embedded in the first vehicle 10, for example a computer.

[0071] The device 4 is for example configured for the implementation of the operations and / or steps described with regard to figures 1, 2 and 4. Examples of such a device 4 include, but are not limited to, on-board electronic equipment such as an on-board computer of a vehicle, an electronic calculator such as an ECU (“Electronic Control Unit”), a smartphone, a tablet, a laptop. The elements of the device 4, individually or in combination, can be integrated in a single integrated circuit, in several integrated circuits, and / or in discrete components. The device 4 can be produced in the form of electronic circuits or software (or computer) modules or even a combination of electronic circuits and software modules.

[0072] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the method and / or for executing the instructions of the software(s) embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41 corresponding for example to a volatile and / or non-volatile memory and / or comprises a memory storage device which may comprise volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.

[0073] The computer code of the embedded software(s) including the instructions to be loaded and executed by the processor is for example stored in memory 41.

[0074] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Telematic Control Unit”), for example via a communication bus or through dedicated input / output ports.

[0075] According to a particular and non-limiting exemplary embodiment, the device 4 comprises a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 comprise one or more of the following interfaces: - RF radio frequency interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN radio technology (Ultra Narrow Band), or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (Universal Serial Bus); HDMI interface (High Definition Multimedia Interface); - LIN interface (Local Interconnect Network).

[0076] According to another particular and non-limiting exemplary embodiment, the device 4 comprises a communication interface 43 which makes it possible to establish communication with other devices (such as other computers of the on-board system) via a communication channel 430. The communication interface 43 corresponds for example to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds for example to a wired network of the CAN type (from the English “Controller Area Network” or in French “Réseau de contrôles”), CAN FD (from the English “Controller Area Network”) or CAN FD (from the English “Controller Area Network”). Flexible Data-Rate” or in French “Flexible Data Rate Controller Network”), FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3).

[0077] According to a particular and non-limiting exemplary embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch-sensitive or not, one or more speakers 450 and / or other peripherals 460 (projection system) via the output interfaces 44, 45, 46 respectively. According to a variant, one or other of the external devices is integrated into the device 4.

[0078] Figure 4 illustrates a flowchart of the different steps of a method 2 for determining a visibility mask for a vision system on board the vehicle of Figure 1, the vision system comprising at least one camera 11 arranged so as to acquire an image of a three-dimensional scene from a determined point of view, according to a particular and non-limiting exemplary embodiment of the present invention.

[0079] The method is for example implemented by one or more processors of one or more computers on board the vehicle 10, for example by a computer controlling the vision system.

[0080] In a first step 31, the computer receives first data representative of a first image acquired by a first camera 11 at a given time instant.

[0081] In a second step 32, the computer receives second data representative of a second image acquired by a second camera 12 at the same given time instant.

[0082] The two images received correspond to two views of the same three-dimensional scene taking place around the vehicle 10 taken from two different points of view at the same given time instant.

[0083] In order to facilitate the analysis of the two received images, the first and second images are rectified according to a method known to those skilled in the art. The rectification of the images is made using the intrinsic and extrinsic parameters of the cameras. Such a method is described, for example, in "Projective Rectification of Uncalibrated Infrared Stereo Images with Global Consideration of Distortion Minimization" by Benoit Ducarouge, Thierry Sentenac, Florian Bugarin and Michel Devy of July 16, 2009.

[0084] The rectification method consists of reorienting the epipolar lines so that they are parallel with the horizontal axis of the image. This method is described by a transformation that projects the epipoles to infinity and whose corresponding points are necessarily on the same ordinate.

[0085] A rectification algorithm consists, for example, of 4 steps: - Rotate (virtually) the first camera 11 so that the epipole goes to infinity along the horizontal axis of the reference frame associated with it; - Apply the same rotation to the second camera 12 to end up in the initial geometric configuration; - Rotate the second camera by the rotation associated with the rotation matrix 'R', corresponding to the extrinsic parameter of the initial stereoscopic vision system; - Adjust the scale in the two camera reference frames.

[0086] It should be noted that rectification simplifies the matching of pixels in stereo images, i.e. those obtained by a stereoscopic vision system. The corresponding pixel in the second image to a pixel in the first image (and vice versa) is positioned on the same line. Based on knowledge of the epipolar geometry and therefore of a fundamental matrix of the stereo system, the objective is then to determine a pair of projective transformations, called homographies, which reorient the epipolar projections parallel to the lines of the images, therefore to the horizontal axis of the rectified cameras.

[0087] In a step 33, depths associated with a set of pixels of the first image are predicted by the stereoscopic vision system from a learned prediction model.

[0088] Such self-supervised learning, i.e. not requiring external intervention or the use of annotated data, is for example achieved by minimizing the photometric error calculated during image reconstructions.

[0089] Disparities are determined from obtaining first and second images from the first camera 11 and the second camera 12 at the same time instant, the disparities being defined by the following function:

[0090] [Math 1]

[0091] ^^ ^ ^ ^ ^ ^^ = ^^ ^ ^ ^ ^ − ^^( ^^ ^^ )

[0092] with: - ^^ ^ ^ ^ ^ ^^ the abscissa of a pixel in the second image, - ^^ ^ ^ ^ ^ the abscissa of a pixel in the first image, and - ^^( ^^ ^^) a disparity determined for a pixel ^^ ^^ of the first image.

[0093] from previously determined disparities, depths are calculated for the pixels of the first image:

[0094] [Math 2]

[0095] ^^ ( ^^ ) = ^^ × ^^ ^^ ^^ ^^ 1 ^^( ^^ ^^)

[0096] - ^^ ^^ of the pixel ^^ ^^ of the first image predicted by the stereoscopic vision system, - ^^( ^^ ^^ ) a disparity determined for a pixel ^^ ^^ of the first image, and - ^^1the focal length of the first camera 11.

[0097] A third image is reconstructed from the first image and the previously calculated depths using the following formula:

[0098] [Math 3]

[0099] ^^ ^^ = ^^( ^^′[ ^^. ^^( ^^ ^^ | ^^ −1 , ^^( ^^ ^^ ))]) with : homogeneous coordinates in three-dimensional space dimensions to two-dimensional pixel coordinates by removing one dimension of a vector, - ^^ the intrinsic matrix of the first camera 11 associated with the projection of a point of the three-dimensional scene into an image acquired by the first camera 11, - ^^′ the intrinsic matrix of the second camera 12 associated with the projection of a point of the three-dimensional scene into an image obtained by the second camera 12, - ^^ a displacement matrix between a position of the first camera and a position of the second camera, - ∅ a reprojection function in the three-dimensional scene of a pixel as a function of its depth, and - ^^( ^^ ^^ ) is a pixel depth ^^ ^^ of the first image predicted by the stereoscopic vision system.

[0100] The reconstructed image is then compared to the second image to determine a photometric error:

[0101] [Math 4]

[0102] ^^ ( ^^ ) = ∑ ^^ [ ( 1 − ^^ ) ⋅ | ^^ ( ^^ ) − ^^ ( ^^ ) | + ^^ ⋅ ( 1 − 1 2^^ ^^ ^^ ^^ ( ^^ ( ^^ ) , ^^ ( ^^ ) ) ) ] - - ^^ ( ^^) a value of the pixel ^^ in the third image, - SSIM (from the English "structural similarity index measure") a function which takes into account a local structure, and - ^^ a weighting factor depending on a type of road environment.

[0104] The convolutional neural network is then learned to minimize the previously defined photometric error.

[0105] Thus, at the output of step 33, each pixel is defined according to principal coordinates (x,y) in the first image and a predicted depth for this pixel.

[0106] In a step 34, the pixels of the set of pixels are reprojected into the three-dimensional scene in the form of a set of points as a function of the depths predicted during step 33, of the intrinsic matrix K of the first camera 11 and of extrinsic parameters T of the stereoscopic vision system.

[0107] Such a reprojection is done for example from the following formula:

[0108] [Math 5]

[0109] ^^ 3 ^^ ( ^^ ^^ ) = ^^. ^^ ( ^^ ^^ | ^^ −1 , ^^( ^^ ^^ ) )

[0110] with: - ^^ 3 ^^ ( ^^ ^^ ) the coordinates in the three-dimensional scene of the point resulting from the reprojection of the pixel ^^ ^^of the first image, - ^^ a displacement matrix between a position of the first camera 11 and a position of the second camera 12, - ^^ the intrinsic matrix of the first camera 11 associated with the projection of a point of the three-dimensional scene into an image acquired by the first camera 11, - ∅ a reprojection function in the three-dimensional scene of a pixel as a function of its depth, - ^^( ^^ ^^ ) is a pixel depth ^^ ^^ of the first image predicted by the stereoscopic vision system.

[0111] Thus, the set of points is located in the three-dimensional scene at positions such that the second camera 12 could see them.

[0112] It is however possible that some of the points 22 projected into the three-dimensional scene are located outside the field of vision of the second camera 12. Indeed, certain parts of the three-dimensional scene are not visible to both cameras 11, 12 at the same time.

[0113] In a step 35, a first visibility mask associated with the pixels of the set of pixels is determined based on the coordinates of points of the set of points.

[0114] A pixel in the pixel set is defined as not visible in the second image if the coordinates of point 22 of the point set associated with the pixel locate it in outside a field of vision of the second camera 12 determined as a function of a width of the second image and a focal length of the second camera 12.

[0115] For example, it is possible that a pixel is the projection of a point 22 of the three-dimensional scene which is located in the first acquisition field 13 which only the first camera 11 perceives. In this case, the point 22 of the three-dimensional scene is outside the field of vision of the second camera 12. It is then necessary to detect this point 22 which is not visible to the second camera 12 because it cannot have an associated pixel in the second image.

[0116] The coordinate system of the points in the three-dimensional scene is defined as a function of the orientation of the cameras 11, 12. Thus, the x axis of the system associated with the three-dimensional scene is parallel to an axis defined by the positions of the cameras 11, 12, the cameras being placed on this axis, the z axis of the system is the focal axis of the second camera 12.

[0117] The principle is to compare the ratio between an abscissa ^^ ^^3 ^^of a point in the three-dimensional scene and the depth of the point in the three-dimensional scene to the ratio of the half-width of the second image to the focal length of the second camera.

[0118] The first visibility mask is obtained for example using the following formula:

[0119] [Math 6] ^^

[0120] ⋃ ^^ ^^^^3 ^^( ^^ ^^) ^^ / 2 ^ ^ ^^ ^^ ^^ ^^ ^^ ^^^^( ^^ ^^)> ^^ in the three-dimensional scene resulting from the reprojection of an abscissa axis being defined parallel to an axis the first and second cameras 11, 12, - ^^( ^^ ^^ ) said pixel depth ^^ ^^ of the first image predicted by the stereoscopic vision system, - ^^ the width of the second image, and - ^^ the focal length of the second camera 12.

[0122] Thus, the first visibility mask makes it possible to identify the pixels of the first image for which the points 22 reprojected into the three-dimensional scene are located outside the field of vision of the second camera 12.

[0123] In a step 36, a third image is generated by projection of the set of points according to an intrinsic matrix K' of the second camera 12.

[0124] Secondary coordinates (i,j) in the third image are associated with each pixel in the pixel set.

[0125] The projection of a point of the three-dimensional scene is carried out for example using the following formula: ^ ^ ^^ = ^^ ( ^^′ [ ^^3 ^^ ( ^^ ^^ )])with: - ^^ a function for going from homogeneous coordinates in three-dimensional space to pixel coordinates in two dimensions by removing one dimension from a vector, - ^^′ the intrinsic matrix of the second camera 12 associated with the projection of a point of the three-dimensional scene into an image acquired by the second camera 12, - ^^ 3 ^^ ( ^^ ^^ ) the coordinates in the three-dimensional scene of the point resulting from the reprojection of the pixel ^^ ^^ of the first image.

[0126] When this operation is performed, each pixel of the first image is thus defined by its (x,y) coordinates in the first image, by an arrival column index i and by an arrival row index j in said third image.

[0127] In a step 37, a second visibility mask associated with the pixels of the set of pixels is determined.

[0128] The principle is to determine the points of the three-dimensional scene which are masked by other points closer to the second camera 12, that is to say the occluded pixels associated with these points.

[0129] The abscissa axis of the first image is oriented positively from left to right.

[0130] For each line of the first rectified image, a scan of the pixels is carried out from left to right according to a point of view of the first camera 11. A matrix, such that presented in Figure 5, is constituted, presenting arrival column index values ​​i for each pixel of the line whose departure column is defined by 'x'.

[0131] In the absence of occluded pixels, i.e. if all pixels in a row 'y' of the first rectified image have a different arrival pixel on row 'j' of the third rectified image, then the arrival column indices 'i' are distributed in the matrix according to a monotonically increasing function.

[0132] An irregular pixel is a pixel in the first image whose arrival column index i' does not follow the monotonic function described above. Such an irregular pixel, placed at rank n of the matrix, is then detected when i'n < in-1.

[0133] Following the detection at position n of an irregular pixel, all the pixels of the scanned line masked by this irregular pixel are then identified, a masked or occluded pixel being a pixel whose arrival column index ik is greater than or equal to the arrival column index i'n of the irregular pixel. Thus, a pixel to the left of the irregular pixel is occluded by the irregular pixel if ik≥i'n.

[0134] The second visibility mask is thus defined as the union of the previously identified occluded pixels.

[0135] For example, in Figure 5, the pixel in sixth position (x(p) = 6) is an irregular pixel. Indeed, its arrival column index is equal to 3 while the arrival column index of the pixel preceding it (x(p) = 5) is equal to 5.

[0136] The hidden or occluded pixels located to the left of the irregular pixel are then identified; these are the pixels in the third, fourth, and fifth positions. Indeed, their arrival column indices in the third image are respectively greater than or equal to the arrival column index of the pixel in the sixth position in the matrix. Thus, if V(p) represents the visibility of a pixel p in the second image, then V(p)=0 for the previously identified hidden pixels.

[0137] Conversely V(p)=1 for the pixels p of the line visible in the second image.

[0138] In a step 38, a third visibility mask is determined as being the association of the first visibility mask and the second visibility mask.

[0139] Thus, this third visibility mask takes into consideration all of the pixels of the first image associated with points of the three-dimensional scene which are located outside the field of vision of the second camera 12 and all of the occluded pixels of the first image.

[0140] The advantage of such a definition of a visibility mask is to determine it without additional calculation, reprojections and projections being already used for learning and depths being already predicted by the stereoscopic vision system. This solution also makes it possible to do without optical flow calculation often used for this type of application and requiring a lot of resources for calculations.

[0141] This definition of a visibility mask allows you to identify the pixels visible in the first image and not visible in the second image.

[0142] Using this third visibility mask when training the convolutional neural network makes it possible to improve the relevance of the definition of input parameters of this convolutional neural network, thus making training more efficient.

[0143] If the ADAS uses input data such as depths determined by the stereoscopic vision system to determine the distance between a part of the vehicle 10, for example the front bumper, and another user present on the road, the ADAS is then able to determine whether the predicted depth is reliable when the pixel is clearly visible in the first and second images.

[0144] Of course, the present invention is not limited to the exemplary embodiments described above but extends to a method for determining a visibility mask for a vision system embedded in a vehicle, which would include secondary steps without thereby departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.

[0145] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-based motor vehicle, comprising the device 4 of figure 3.

Claims

CLAIMS 1. Method for determining visibility masks by a stereoscopic vision system on board a vehicle (10), the stereoscopic vision system comprising a set of cameras of at least two cameras (11, 12) arranged so as to each acquire an image of a three-dimensional scene from a different point of view, said second camera (12) being located to the right of said first camera (11) from a point of view of said first camera (11), said method being characterized in that it comprises the following steps: - reception (31, 32) of first and second data respectively representative of a first and second image acquired by respectively a first and second camera (11, 12) of said set of cameras at the same acquisition time instant;- prediction (33) of depths associated with a set of pixels of the first image by said stereoscopic vision system from a learned prediction model, each pixel of the first image having principal coordinates (x,y) in the first image; - reprojection (34) in the three-dimensional scene of said set of pixels in the form of a set of points as a function of said depths, of an intrinsic matrix of the first camera (11) and of extrinsic parameters of said stereoscopic vision system;- determining (35) a first visibility mask associated with the pixels of said set of pixels as a function of the coordinates of points of said set of points, a pixel of said set of pixels being not visible in said second image if the coordinates of a point of said set of points associated with said pixel locate it outside a field of vision of said second camera (12) determined as a function of a width of the second image and a focal length of the second camera (12); - generating (36) a third image by projection of said set of points as a function of an intrinsic matrix of said second camera (12), the arrival column indices (i) and the arrival line indices (j) in said third image; being associated with each pixel of said set of pixels;- rectification of the first image from the intrinsic and extrinsic parameters of the two cameras, for each line of said first rectified image, scanning from left to right, according to a point of view of the first camera (11), arrival column index values (i), constitution of a matrix having arrival column index values (i) as a function of a column index (x) in the first rectified image, and, from said matrix, detection of a set of irregular pixels and for each irregular pixel of its arrival column index (i'), and for each irregular pixel of said set, identification of a set of occluded pixels in said line to the left of said each irregular pixel of which an arrival column index (i) is greater than or equal to an arrival column index (i') of said each irregular pixel, a second visibility mask being determined (37) as a union of said sets of pixels occluded;and - determination (38) of a third visibility mask by associating said first visibility mask and said second visibility mask.

2. Method according to claim 1, for which the reprojection of a pixel in the three-dimensional scene is carried out using the following formula: ^^; 3 ^^ ( ^^ ^^ ) = ^^. ^^( ^^ ^^ | ^^ −1 , ^^( ^^ ^^ )) with: in the three-dimensional scene of the point resulting from the reprojection of the pixel ^^ ^^ of the first image, - ^^ a displacement matrix between a position of the first camera (11) and a position of the second camera (12), - ^^ the intrinsic matrix of the first camera (11) associated with a projection of a point of the three-dimensional scene in an image acquired by the first camera (11), - ^^ a reprojection function in the three-dimensional scene of a pixel as a function of its depth, - ^^( ^^ ^^ ) is a pixel depth ^^ ^^of the first image predicted by the stereoscopic vision system.

3. Method according to one of claims 1 to 2, for which the projection of a point of the three-dimensional scene is carried out using the following formula: ^ ^ ^^ = ^^ ( ^^′ [ ^^3 ^^ ( ^^ ^^ )]) with: - ^^ a function for going from homogeneous coordinates in three-dimensional space to pixel coordinates in two dimensions by removing one dimension from a vector, - ^^′ the intrinsic matrix of the second camera (12) associated with a projection of a point of the three-dimensional scene into an image acquired by the second camera (12), - ^^ 3 ^^ ( ^^ ^^ ) the coordinates in the three-dimensional scene of the point resulting from the reprojection of the pixel ^^ ^^of the first image.

4. Method according to one of claims 1 to 3, for which said first visibility mask is obtained using the following formula: ^^ ⋃ ^^^^3 ^^( ^^ ^^) ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ / 2 ^^ ^^( ^^ ^^) > ^^ resulting from the reprojection of the pixel ^^ ^^ of the first image, an abscissa axis being defined parallel to an axis along which the first and second cameras (11, 12) are located, - ^^( ^^ ^^ ) said pixel depth ^^ ^^of the first image predicted by the stereoscopic vision system, - ^^ the width of the second image, and - ^^ the focal length of the second camera (12).

5. Method according to one of claims 1 to 4, for which said depths are predicted by a convolutional neural network.

6. Method according to claim 5, for which the convolutional neural network is trained to minimize a photometric error defined by the following loss function: ^^ ( ^^ ) = ∑ [ ( 1 − ^^ ) ⋅ | ^^ ( ^^ − ^^ ^^ | + ^^ ⋅ (1 − 1 ^^ ^^ ^^ ^^ ( ^^ ^^ , ^^ ^^ ))] ^^ ) ( ) 2 ( ) ( ) - ^^ ( ^^) a value of the pixel ^^ in the third image; - SSIM a function which takes into account a local structure; and - ^^ a weighting factor depending on a type of road environment.

7. Computer program comprising instructions for implementing the method according to any one of the preceding claims, when these instructions are executed by a processor.

8. Device (4) for determining a visibility mask for a vision system embedded in a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for implementing the steps of the method according to any one of claims 1 to 6.

9. System for determining a visibility mask for a vision system embedded in a vehicle (10) comprising at least two cameras (11, 12) and a device according to claim 8. 10.Vehicle (10) comprising the device (4) according to claim 8 or the system according to claim 9.