Method and device for determining a visibility mask for a vision system embedded in a vehicle.
The stereoscopic vision system method efficiently determines visibility masks for vehicles, addressing occlusion challenges in ADAS systems by reducing resource needs and improving data quality.
Patent Information
- Application Number
- FR2023004517
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-05-05
- Publication Date
- 2026-02-20
- Estimated Expiration
- 2043-05-05
AI Technical Summary
Existing vision systems in vehicles face challenges in determining visibility masks to accurately identify occluded areas, which affects the quality of data used by ADAS systems, and requires significant computational resources.
A method for determining visibility masks using a stereoscopic vision system with multiple cameras, involving image acquisition, depth prediction, and reprojection to identify occluded pixels, reducing the need for optical flow calculations.
Improves the quality of data for ADAS systems by efficiently identifying occluded areas without requiring additional computational resources, enhancing the reliability of depth predictions.
Smart Images

Figure 00000024_0000 
Figure 00000024_0001 
Figure 00000024_0002
Abstract
Description
Title of the invention: Method and device for determining a visibility mask for a vision system embedded in a vehicle. technical field
[0001] The present invention relates to methods and devices for determining a visibility mask for a vision system installed in a vehicle, for example, in a motor vehicle. The present invention also relates to a method and device for controlling one or more ADAS systems installed in a vehicle based on a determined visibility mask. Technological background
[0002] Many modern vehicles are equipped with Advanced Driver-Assistance Systems (ADAS). Such ADAS systems are passive and active safety systems designed to eliminate human error in driving all types of vehicles. ADAS systems use advanced technologies to assist the driver while driving and thus improve performance. ADAS systems use a combination of sensor technologies to perceive the environment around a vehicle and then provide information to the driver or act on certain vehicle systems.
[0003] There are several levels of ADAS, such as reversing cameras and blind spot sensors, lane departure warning systems, adaptive cruise control or automatic parking systems.
[0004] Vehicle-mounted ADAS systems are powered by data obtained from one or more on-board sensors such as, for example, cameras. These cameras make it possible, in particular, to detect and locate other road users or any obstacles present around a vehicle in order, for example: - to adapt the vehicle's lighting according to the presence of other road users; - to automatically regulate the vehicle's speed; - to act on the braking system in case of risk of impact with an object.
[0005] The quality of the data emitted by a vision system therefore determines the proper functioning of the driving assistance devices using this data.
[0006] Many vision systems perceive the environment around a vehicle from several images acquired by one or more cameras. During image processing, occluded areas of the images, corresponding to areas of the environment not present in all the acquired images, are defined. A visibility mask associated with an image then defines, for example, a filter allowing the determination of pixels associated with areas not found in other images.
[0007] Solutions exist for detecting an occlusion, that is to say for determining a visibility mask.
[0008] A first solution presented in "Occlusion Aware Unsupervised Leaming of Optical Flow" by Yang Wang, Yi Yang, Zhenheng Yang, Liang Zhao, Peng Wang, and Wei Xu, published on April 4, 2018, is based on reverse optical flow. For each pixel of a first image represented by its coordinates, the algorithm checks whether a pixel of a second image reaches that pixel of the first image with the reverse optical flow by scanning all the pixels of the second image. This method can be used for both directions of optical flow to identify the occluded areas of the two images.
[0009] A second solution is described in the document "Geometry-based Occlusion-Aware Unsupervised Stereo Matching for Autonomous Driving algorithm". The detection of occluded areas is based on a geometric constraint: the occluded pixel and another pixel that hides it are projected into the same pixel of a reconstructed image. Summary of the present invention
[0010] One object of the present invention is to solve at least one of the problems of the technological background described above.
[0011] Another object of the present invention is to propose an alternative solution for determining a visibility mask for any vision system in order to improve the quality of the data from the camera(s) of the vision system.
[0012] Another object of the present invention is to reduce the resources needed for determining a visibility mask.
[0013] According to a first aspect, the present invention relates to a method for determining visibility masks by a stereoscopic vision system mounted in a vehicle, the stereoscopic vision system comprising a set of cameras of at least two cameras arranged so as to each acquire an image of a three-dimensional scene from a different point of view, said second camera being located to the right of the first camera from the point of view of the first camera, the method being characterized in that it comprises the following steps: - reception of first and second data respectively representative of a first and second image acquired by respectively a first and second camera of the camera set at the same time instant of acquisition; - Prediction of depths associated with a set of pixels in the first image by the stereoscopic vision system based on a learned prediction model, each pixel of the first image having principal coordinates in the first image; - reprojection into the three-dimensional scene of the pixel set as a set of points as a function of depths, an intrinsic matrix of the first camera and extrinsic parameters of the stereoscopic vision system; - determination of a first visibility mask associated with the pixels of the pixel set as a function of the coordinates of points of the point set, a pixel of the pixel set being not visible in the second image if the coordinates of a point of said point set associated with said pixel place it outside a field of view of the second camera determined as a function of a width of the second image and a focal length of the second camera; - generation of a third image by projection of the set of points according to an intrinsic matrix of the second camera, secondary coordinates in the third image being associated with each pixel of the pixel set; - for each line of the first rectified image, scanning from left to right, according to a viewpoint of the first camera, of destination column index values and detection of a set of irregular pixels whose destination column index does not follow a monotonic function representing an evolution of destination column indices as a function of a column index in the first rectified image, and for each irregular pixel of the set, identification of a set of occluded pixels in the line to the left of each irregular pixel whose destination column index is greater than or equal to a destination column index of each irregular pixel, a second visibility mask being determined as a union of the sets of occluded pixels; and - determination of a third visibility mask by combining the first visibility mask and the second visibility mask.
[0014] According to an alternative method, the reprojection of a pixel in the three-dimensional scene is carried out using the following formula: PJj>, )=Tf( - n ( n Wes coordinates in the three-dimensional scene of the point from the re-7 3Z>V t / projection of pixel Pt from the first image, - T is a displacement matrix between a position of the first camera and a position of the second camera, - K, the intrinsic matrix of the first camera associated with the projection of a point from the three-dimensional scene into an image acquired by the first camera, - 0 a reprojection function in the three-dimensional scene of a pixel in depending on its depth, ■ D^p ) is a depth of the pixel Pt predicted by the stereoscopic vision system.
[0015] According to another variant of the method, the projection of a point of the three-dimensional scene is carried out using the following formula: ] )with 77 a function to convert coordinates homogeneous in three-dimensional space to pixel coordinates in two dimensions by removing one dimension from a vector, - K' the intrinsic matrix of the second camera associated with the projection of a point from the three-dimensional scene into an image acquired by the second camera, - n I n Ves coordinates in the three-dimensional scene of the point originating from the re- * 3D\ptJ projection of pixel P{ of the first image.
[0016] According to yet another variant of the process, the first visibility mask is obtained at using the following formula: An x with: such that ' ~y - the abscissa of a point resulting from the reprojection of pixel Pt from the first image, an x-axis being defined parallel to an axis along which the first and second cameras are located, ■ IP(p) is the depth of pixel Pt of the first image predicted by the stereoscopic vision system, - W the width of the second image, and - / the focal length of the second camera.
[0017] According to yet another variant of the process, the depths are predicted by a convolutional neural network.
[0018] According to a further variant of the method, the convolutional neural network is trained to minimize a photometric error defined by the following loss function: With : - Kp) a value of pixel P in the second image; - a value of pixel P in the third image; - SSIM is a function that takes into account a local structure; and - has a weighting factor dependent on a type of road environment.
[0019] According to a second aspect, the present invention relates to a device for determining a visibility mask for a vision system embedded in a vehicle, the device comprising a memory associated with at least one processor configured for the implementation of the steps of the process according to the first aspect of the present invention.
[0020] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0021] According to a fourth aspect, the present invention relates to a computer program which includes instructions adapted for carrying out the steps of the process according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0022] Such a computer program may use any programming language and be in the form of source code, object code, or an intermediate form between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0023] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the process according to the first aspect of the present invention.
[0024] On the one hand, the recording medium can be any entity or device capable of storing the program. For example, the medium can include a storage means, such as a ROM, a CD-ROM or a microelectronic circuit-type ROM, or a magnetic recording means or a hard disk drive.
[0025] On the other hand, this recording medium can also be a transmissible medium such as an electrical or optical signal, such a signal being able to be transmitted via an electrical or optical cable, by conventional or radio frequency, by self-directing laser beam, or by other means. The computer program according to the present invention can, in particular, be downloaded from an Internet-type network.
[0026] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to execute or to be used in the execution of the process in question. Brief description of the figures
[0027] Other features and advantages of the present invention will become apparent from the description of the particular and non-limiting embodiments of the present invention below, with reference to the attached Figures 1 to 5, in which:
[0028] [Fig-1] schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting example of the present invention;
[0029] [Fig.2] schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention;
[0030] [Fig.3] schematically illustrates a device configured for determining a visibility mask by a vision system onboard in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention;
[0031] [Fig.4] illustrates a flowchart of the different stages of a process for determining a visibility mask by a vision system onboard in the vehicle of [Fig.1], according to a particular and non-limiting example of the present invention.
[0032] [Fig.5] illustrates a matrix presenting different column indices for pixels of a row and a visibility criterion for a vision system embedded in the vehicle of [Fig.1], according to a particular and non-limiting embodiment of the present invention.
[0033] Description of examples of implementation
[0034] A method and device for determining a visibility mask for a vision system on board a vehicle will now be described in what follows with joint reference to Figures 1 to 5. The same elements are identified with the same reference signs throughout the following description.
[0035] According to a particular and non-limiting example of an embodiment of the present invention, a method for determining a visibility mask for a stereoscopic vision system embedded in a vehicle is for example implemented by a computer of the vehicle's embedded system controlling this vision system.
[0036] The stereoscopic vision system comprises a set of cameras of at least two cameras arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the second camera being located to the right of the point of view of said first camera.
[0037] To this end, the method of determining a visibility mask by a stereoscopic vision system embedded in a vehicle includes the reception of first and second data respectively representative of a first and second image acquired from a different point of view by the first and second cameras at the same time instant of acquisition.
[0038] The method also includes the prediction of depths associated with a set of pixels from the first image by the stereoscopic vision system from a learned prediction model, each pixel of the first image having principal coordinates in the first image, the reprojection into the three-dimensional scene dimensionality of the pixels of the first image in the form of a set of points as a function of depths, an intrinsic matrix of the first camera and extrinsic parameters of the stereoscopic vision system.
[0039] The method then determines a first visibility mask associated with the pixels of the pixel set as a function of the coordinates of points in the point set, a pixel of the pixel set being not visible in the second image if the coordinates of a point in the point set associated with the pixel place it outside a field of view of the second camera determined as a function of a width of the second image and a focal length of the second camera.
[0040] The method then comprises generating a third image by projecting the point set onto an intrinsic matrix of the second camera, with secondary coordinates in the third image associated with each pixel in the pixel set. Generating this third image then allows the determination of a second visibility mask associated with the pixels in the pixel set. For each row of the first rectified image, values of the arrival column indices are scanned, and a set of irregular pixels is detected. For each irregular pixel in the set, a set of occluded pixels is identified in the row to the left of each detected irregular pixel. A second visibility mask is then determined as a union of the sets of occluded pixels.
[0041] A third visibility mask is then determined by combining the first visibility mask and the second visibility mask.
[0042] Fig. 1 schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.
[0043] In this example, vehicle 10 corresponds to a vehicle with an internal combustion engine, an electric motor(s), or a hybrid vehicle with an internal combustion engine and one or more electric motors. Vehicle 10 thus corresponds, for example, to a land vehicle such as a car, a truck, a bus, or a motorcycle. Finally, vehicle 10 corresponds to an autonomous or non-autonomous vehicle, that is to say, a vehicle operating according to a predetermined level of autonomy or under the total supervision of the driver.
[0044] The vehicle 10 advantageously comprises several onboard cameras 11, 12, each configured to acquire images of a three-dimensional scene in the environment of the vehicle 10. This set of cameras 11, 12 forms the stereoscopic vision system. Two cameras 11, 12 are illustrated in [Fig. 1]. The present invention is not limited, however, to a stereoscopic vision system comprising two cameras but extends to any vision system comprising two or more cameras, for example, 3, 4, or 5 cameras.
[0045] The cameras 11, 12 have known intrinsic parameters. These parameters are include, in particular: - focal length fl of the first camera 11; - focal length f2 of the second camera 12; - distortions which are due to imperfections in the optical system of each camera; - direction Cl of the optical axis of the first camera 11; - direction C2 of the optical axis of the second camera 12; - respective resolutions of cameras 11, 12.
[0046] The intrinsic parameters characterize the transformation that associates, for an image point, the camera coordinates to the pixel coordinates, in each camera. These parameters do not change if the camera is moved.
[0047] Distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of camera lenses, will deflect the light beams and thus induce a positioning error for the projected point relative to an ideal model. It is then possible to complete the camera model by introducing the three distortions that generate the most significant effects, namely radial, decentering, and prismatic distortions, induced by defects in lens curvature, parallelism, and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, meaning that distortions are either not taken into account or their correction is addressed during image acquisition.
[0048] These cameras 11, 12 are arranged so as to each acquire an image of a three-dimensional scene from a different point of view, the second camera 12 being located to the right of the point of view of said first camera 11, the first point of view is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10, the second point of view is for example located on or in the right rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. In the case where two cameras are located at the top of the windshield of the vehicle, they are then placed at a certain distance.
[0049] A first marker is associated with the first camera 11: - the direction of the y-axis is defined by the position of the second camera 11, so as to place the second camera 12 on the y-axis of the first camera 11. The distance B separating the two cameras 11,12 is called the reference base (in English "baseline") and the direction separating the two cameras 11,12 is that of the y-axis; - the direction of the x axis is defined orthogonal to that of the y axis and orthogonal to that of the optical axis Cl of the first camera 11; - The direction of the z-axis is defined as orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal coordinate system.
[0050] The extrinsic parameters related to the position of cameras 11, 12 are the following parameters: - 3 translations in the x, y, and z directions: Tx, Ty, and Tz forming the translation vector T; and - 3 rotations around the x, y and z axes: Rx, Ry and Rz, constituting the rotation matrix R.
[0051] A key constraint of stereoscopic vision systems used in automobiles is, for example, the large distance between the two cameras. Indeed, to cover a measurement range of 200 meters, the baseline must reach 60 cm for cameras commonly used in this field.
[0052] The two cameras 11, 12 acquire images of a three-dimensional scene located in front of the vehicle 10, the first camera 11 alone covering a first acquisition field 13, the second camera 12 alone covering a second acquisition field 14 and the two cameras 11, 12 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic view of the three-dimensional scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic view of the three-dimensional scene by the second camera 12 and the third acquisition field 15 allows a stereoscopic view of the three-dimensional scene by the stereoscopic vision system composed of the two cameras 11, 12.
[0053] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.
[0054] Among these three fields, field 16 is visible from the second camera 12. The part of the three-dimensional scene present in this field 16 is therefore observable using the monoscopic vision system composed of the second camera 12.
[0055] The field 17 is visible from the first camera 11. The part of the three-dimensional scene present in this field 17 is therefore observable with the monoscopic vision system composed of the second camera 12.
[0056] Finally, field 19 is not visible from any of the cameras. The part of the three-dimensional scene present in this field 19 is therefore not observable.
[0057] The directions Cl, C2 of the optical axes are representative of an orientation of the field of vision of each camera 11, 12.
[0058] It is evident that it is possible to use such a stereoscopic vision system to take images of three-dimensional scenes located on the sides or behind the vehicle 10 by equipping it with cameras placed and oriented differently.
[0059] The images acquired by cameras 11, 12 at a given acquisition time are in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colours and brightness of objects in the observed three-dimensional scene in the form of, for example, RGB colourimetric coordinates (from the English "Red Green Blue", in French "Rouge Vert Bleu") or HSL (Tone, Saturation, Luminosity).
[0060] The images acquired by cameras 11, 12 represent views of the same three-dimensional scene taken from different viewpoints, the camera positions being distinct. This three-dimensional scene includes, for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0061] These images are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.
[0062] Figure 2 schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting embodiment of the present invention.
[0063] Points 20, 21, 22 of the three-dimensional scene are visible from the point of view of the first camera 11.
[0064] Points 21 are also visible from the point of view of the second camera 12.
[0065] Points 20 are, however, occluded from the point of view of the second camera 12 because they are masked by points 21 located on the same axes in the field of vision of the second camera 12.
[0066] Points 22 are located outside the field of vision of the second camera 12.
[0067] Thus, during the acquisition of the first and second images by respectively the first camera 11 and second camera 12, pixels associated with points visible 20, 21, 22 from the viewpoint of first camera 11 will be present in the first image, while only pixels associated with points 21 visible from the viewpoint of second camera will be present in the second image.
[0068] The pixels associated with the points 20 visible from the viewpoint of the first camera 11 and not visible from the viewpoint of the second camera 12 are hereafter called occluded pixels.
[0069] Figure 3 schematically illustrates a device 4 configured for determining a visibility mask for a vision system embedded in a vehicle 10, according to a particular and non-limiting embodiment of the present invention. The device 4 corresponds, for example, to a device embedded in the first vehicle 10, for example, a computer.
[0070] Device 4 is, for example, configured to carry out the operations and / or steps described opposite Figures 1, 2, and 4. Examples of such a device 4 include, but are not limited to, embedded electronic equipment such as a vehicle's on-board computer, an electronic control unit such as an ECU (Electronic Control Unit), a smartphone, a tablet, or a laptop computer. The elements of device 4, individually or in combination, may be integrated into a single integrated circuit, into several integrated circuits, and / or into discrete components. Device 4 may be implemented in the form of electronic circuits or software (or computer) modules, or a combination of electronic circuits and software modules.
[0071] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the process and / or for executing instructions from the software embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41, for example, volatile and / or non-volatile memory, and / or includes a memory storage device that may include volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk, or optical disk.
[0072] The computer code of the embedded software(s), including the instructions to be loaded and executed by the processor, is for example stored in memory 4L
[0073] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (Telematic Control Unit), for example via a communication bus or through dedicated input / output ports.
[0074] According to a particular and non-limiting embodiment, the device 4 includes a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 include one or more of the following interfaces: - radio frequency RF interface, for example of the Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or of the Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or of the Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English "Universal Serial Bus" or "Universal Serial Bus" in French); HD MI interface (from the English "High Definition Multimedia Interface", or "High Definition Multimedia Interface" in French); - LIN interface (from the English "Local Interconnect Network", or in French "Réseau interconnecté local").
[0075] According to another particular and non-limiting embodiment, the device 4 includes a communication interface 43 which allows communication to be established with other devices (such as other computers in the embedded system) via a communication channel 430. The communication interface 43 corresponds, for example, to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds, for example, to a wired network of the CAN (Controller Area Network) type, CAN FD (Controller Area Network Flexible Data-Rate) type, FlexRay (standardized by ISO 17458) or Ethernet (standardized by ISO / IEC 802-3).
[0076] According to a particular and non-limiting embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch or not, one or more loudspeakers 450 and / or other peripherals 460 (projection system) via the output interfaces 44, 45, 46 respectively. According to a variant, one or more of the external devices is integrated into the device 4.
[0077] Figure 4 illustrates a flowchart of the different steps of a method 2 for determining a visibility mask for a vision system embedded in the vehicle of Figure 1, the vision system comprising at least one camera 11 arranged to acquire an image of a three-dimensional scene from a determined point of view, according to a particular and non-limiting embodiment of the present invention.
[0078] The process is implemented for example by one or more processors of one or more computers embedded in the vehicle 10, for example by a computer controlling the vision system.
[0079] In a first step 31, the computer receives first data representative of a first image acquired by a first camera 11 at a given time instant.
[0080] In a second step 32, the computer receives second data representing a second image acquired by a second camera 12 at the same given time instant.
[0081] The two images received correspond to two views of the same three-dimensional scene taking place around the vehicle 10 taken from two different viewpoints at the same given time instant.
[0082] To facilitate the analysis of the two received images, the first and second images are rectified using a method known to those skilled in the art. Such a method is described, for example, in "Projective Rectification of Uncalibrated Infrared Stereo Images with Global Consideration of Distortion Minimization" by Benoit Ducarouge, Thierry Sentenac, Florian Bugarin and Michel Devy, dated July 16, 2009.
[0083] The rectification method consists of reorienting the epipolar lines so that they are parallel with the horizontal axis of the image. This method is described by a transformation that projects the epipoles to infinity and whose corresponding points are necessarily on the same ordinate.
[0084] A rectification algorithm consists, for example, of 4 steps: - Rotate (virtually) the first camera 11 so that the epipole goes to infinity along the horizontal axis of the frame associated with it; - Apply the same rotation to the second camera 12 to return to the initial geometric configuration; - Rotate the second camera by the rotation associated with the rotation matrix 'R', corresponding to the extrinsic parameter of the starting stereoscopic vision system; - Adjust the scale in both camera reference points.
[0085] It should be noted that rectification simplifies the matching of pixels in stereo images, i.e., images obtained by a stereoscopic vision system. The pixel in the second image corresponding to a pixel in the first image (and vice versa) is positioned on the same line. Based on knowledge of the epipolar geometry and thus of a fundamental matrix of the stereo system, the objective is then to determine a pair of projective transformations, called homographies, which reorient the epipolar projections parallel to the image lines, and therefore to the horizontal axis of the rectified cameras.
[0086] In a step 33, depths associated with a set of pixels of the first image are predicted by the stereoscopic vision system from a learned prediction model.
[0087] Such self-supervised learning, i.e. not requiring external intervention or the use of annotated data, is for example carried out by minimizing the photometric error calculated during image reconstructions.
[0088] Disparities are determined from obtaining the first and second images from the first camera 11 and the second camera 12 at the same time temporal, the disparities being defined by the following function:
[0089] [Math.l] p^p^d(pt)
[0090] with: - p', the x-coordinate of a pixel in the second image, - p^ the x-coordinate of a pixel in the first image, and ■ d( p J a disparity determined for a pixel Pt of the first image.
[0091] Using the previously determined disparities, depths are calculated for the pixels of the first image:
[0092] [Math.2]
[0093] with: ■ Dt^p ) the depth of pixel Pt of the first image predicted by the stereoscopic vision system, ■ d(p) a disparity determined for a pixel Pt of the first image, and -j the focal length of the first camera 11.
[0094] A third image is reconstructed from the first image and the depths previously calculated using the following formula:
[0095] [Math.3] P s = M / " [ ( P^' D(p) ) ] ) with [Math.3] 7T a function to convert homogeneous coordinates in three-dimensional space to pixel coordinates in two dimensions by removing one dimension from a vector, [Math.3] K the intrinsic matrix of the first camera 11 associated with the projection of a point of the three-dimensional scene into an image acquired by the first camera 11, [Math.3] K' the intrinsic matrix of the second camera 12 associated with the projection of a point of the three-dimensional scene in an image obtained by the second camera 12, [Math.3] T a displacement matrix between a position of the first camera and a position of the second camera, [Math.3] 0 a reprojection function in the three-dimensional scene of a pixel based on its depth, and [Math.3] dp,) is a pixel depth [Math.3] Pt of the first image predicted by the stereoscopic vision system.
[0096] The reconstructed image is then compared to the second image in order to determine a photometric error:
[0097] [Math.4] L(p) = Ep[ ( l-«) ■ \I(p) -î(p)|+a- (l-^SSIM(l(p)J(p)))]
[0098] with: - I(p) a value of pixel P in the second image, 'I^pj a value of pixel P in the third image, - SSIM (from the English "structural similarity index measure") is a function that takes into account a local structure, and - has a weighting factor dependent on a type of road environment.
[0099] The convolutional neural network is then learned to minimize the previously defined photometric error.
[0100] Thus, at the output of step 33, each pixel is defined according to principal coordinates (x,y) in the first image and a predicted depth for that pixel.
[0101] In a step 34, the pixels of the pixel set are reprojected into the three-dimensional scene as a set of points according to the depths predicted in step 33, the intrinsic matrix K of the first camera 11 and extrinsic parameters T of the stereoscopic vision system.
[0102] Such a reprojection is done, for example, using the following formula:
[0103] [Math.5] PjdM =MpF'' riù)
[0104] with: - nin 'i the coordinates in the three-dimensional scene of the point resulting from the re- projection of pixel P( of the first image, - T is a displacement matrix between a position of the first camera 11 and a position of the second camera 12, - K the intrinsic matrix of the first camera 11 associated with the projection of a point from the three-dimensional scene into an image acquired by the first camera 11, - 0 a reprojection function in the three-dimensional scene of a pixel as a function of its depth, ■ I^p ) is a depth of the pixel Pt of the first image predicted by the stereoscopic vision system.
[0105] Thus, the set of points is located in the three-dimensional scene at positions such that the second camera 12 could see them.
[0106] However, it is possible that some of the points 22 projected into the three-dimensional scene are located outside the field of vision of the second camera 12. Indeed, some parts of the three-dimensional scene are not visible by both cameras 11, 12 at the same time.
[0107] In a step 35, a first visibility mask associated with the pixels of the pixel set is determined as a function of the point coordinates of the point set.
[0108] A pixel in the pixel set is defined as not visible in the second image if the coordinates of the point 22 in the point set associated with the pixel place it outside a field of view of the second camera 12 determined as a function of a width of the second image and a focal length of the second camera 12.
[0109] For example, it is possible that a pixel is the projection of a point 22 of the three-dimensional scene located in the first acquisition field 13, which only the first camera 11 perceives. In this case, the point 22 of the three-dimensional scene is outside the field of view of the second camera 12. It is then necessary to detect this point 22, which is not visible to the second camera 12, because it cannot have an associated pixel in the second image.
[0110] The coordinate system of the points in the three-dimensional scene is defined according to the orientation of the cameras 11, 12. Thus, the x-axis of the coordinate system associated with the scene three-dimensional is parallel to an axis defined by the positions of cameras 11, 12, the cameras being placed on this axis, the z-axis of the frame is the focal axis of the second camera 12.
[0111] The principle is to compare the ratio between an XP3D abscissa of a point in the three-dimensional scene and the depth of the point in the three-dimensional scene to the ratio between the half-width of the second image and the focal length of the second camera.
[0112] The first visibility mask is obtained, for example, using the following formula:
[0113] [Math.6]
[0114] with: 'XpJP') the abscissa of the point in the three-dimensional scene resulting from the reprojection of pixel Pt from the first image, an x-axis being defined parallel to an axis along which the first and second cameras 11, 12 are located, ■ D^pJ said pixel depth Pt of the first image predicted by the stereoscopic vision system, - It is the width of the second image, and - f the focal length of the second camera 12.
[0115] Thus, the first visibility mask makes it possible to identify the pixels of the first image for which the points 22 reprojected in the three-dimensional scene are located outside the field of vision of the second camera 12.
[0116] In a step 36, a third image is generated by projection of the set of points according to an intrinsic matrix K' of the second camera 12.
[0117] Secondary coordinates (i,j) in the third image are associated with each pixel in the pixel set.
[0118] The projection of a point in the three-dimensional scene is carried out, for example, using the following formula: Ps JVioW])™' 77 a function to convert coordinates homogeneous in three-dimensional space to pixel coordinates in two dimensions by removing one dimension from a vector, - K' the intrinsic matrix of the second camera 12 associated with the projection of a point of the three-dimensional scene into an image acquired by the second camera 12, - n ( n Ves coordinates in the three-dimensional scene of the point from the re-^30 VJ projection of pixel Pt from the first image.
[0119] When this operation is performed, each pixel of the first image is thus defined by its coordinates (x,y) in the first image, by an arrival column index i and by an arrival row index j in said third image.
[0120] In a step 37, a second visibility mask associated with the pixels of the pixel set is determined.
[0121] The principle is to determine the points of the three-dimensional scene which are masked by other points closer to the second camera 12, i.e. the occluded pixels associated with these points.
[0122] The x-axis of the first image is oriented positively from left to right.
[0123] For each line of the first rectified image, a scan of the pixels is carried out from left to right according to a viewpoint of the first camera 11. A matrix, as shown in [Fig.5], is constituted, presenting values of arrival column index i for each pixel of the line whose starting column is defined by 'x'.
[0124] In the absence of occluded pixels, that is, if all the pixels of a row 'y' of the first rectified image have a different arrival pixel on the row 'j' of the third rectified image, then the arrival column indices 'i' are distributed in the matrix according to a monotonically increasing function.
[0125] An irregular pixel is defined as a pixel in the first image whose index i' of the arrival column does not follow the monotonic function described above. Such an irregular pixel, placed at rank n of the matrix, is then detected when i'n < in_i.
[0126] Following the detection of an irregular pixel at position n, all pixels in the scanned line masked by this irregular pixel are then identified. A masked or occluded pixel is a pixel whose arrival column index ik is greater than or equal to the arrival column index i'n of the irregular pixel. Thus, a pixel to the left of the irregular pixel is occluded by the irregular pixel if ik > i'n.
[0127] The second visibility mask is thus defined as the union of the previously identified occluded pixels.
[0128] For example, in [Fig.5], the pixel in the sixth position (x(p) = 6) is an irregular pixel. Indeed, its destination column index is equal to 3 while the destination column index of the pixel preceding it (x(p) = 5) is equal to 5.
[0129] The masked or occluded pixels located to the left of the irregular pixel are then identified; these are the pixels in the third, fourth, and fifth positions. Indeed, their destination column indices in the third image are respectively greater than or equal to the destination column index of the pixel in the sixth position in the matrix. Thus, if V(p) represents the visibility of a pixel p in the second image, then V(p) = 0 for the previously identified masked pixels.
[0130] Conversely, V(p)=l for the pixels p of the line visible in the second image.
[0131] In a step 38, a third visibility mask is determined as being the combination of the first visibility mask and the second visibility mask.
[0132] Thus, this third visibility mask takes into consideration all the pixels of the first image associated with points of the three-dimensional scene which are outside the field of vision of the second camera 12 and all the pixels of the first occluded image.
[0133] The advantage of such a definition of a visibility mask is that it can be determined without additional calculation, as reprojections and projections are already used for training and depths are already predicted by the stereoscopic vision system. This solution also eliminates the need for optical flow calculations, which are often used for this type of application and require significant computational resources.
[0134] This definition of a visibility mask makes it possible to identify the pixels visible in the first image and not visible in the second image.
[0135] The use of this third visibility mask during the training of the convolutional neural network improves the relevance of the definition of input parameters of this convolutional neural network, thus making learning more efficient.
[0136] If ADAS uses input data such as depths determined by the stereoscopic vision system to determine the distance between a part of the vehicle 10, for example the front bumper, and another road user, ADAS is then able to determine whether the predicted depth is reliable when the pixel is clearly visible in the first and second images.
[0137] Of course, the present invention is not limited to the embodiments described above but extends to a method for determining a visibility mask for a vision system embedded in a vehicle, which would include secondary steps without departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0138] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-powered vehicle, comprising the device 4 of [Fig.3].
Claims
Demands
1. Method for determining visibility masks by a stereoscopic vision system mounted in a vehicle (10), the stereoscopic vision system comprising a set of cameras of at least two cameras (11, 12) arranged so as to each acquire an image of a three-dimensional scene from a different point of view, said second camera (12) being located to the right of said first camera (11) from a point of view of said first camera (11), said method being characterized in that it comprises the following steps: - reception (31, 32) of first and second data respectively representing a first and second image acquired by respectively a first and second camera (11, 12) of said set of cameras at the same time instant of acquisition; - prediction (33) of depths associated with a set of pixels of the first image by said stereoscopic vision system from a learned prediction model, each pixel of the first image having principal coordinates (x,y) in the first image; - reprojection (34) into the three-dimensional scene of said pixel set in the form of a set of points as a function of said depths, of an intrinsic matrix of the first camera (11) and of extrinsic parameters of said stereoscopic vision system; - determination (35) of a first visibility mask associated with the pixels of said pixel set as a function of the coordinates of points of said point set, a pixel of said pixel set being not visible in said second image if the coordinates of a point of said point set associated with said pixel place it outside a field of view of said second camera (12) determined as a function of a width of the second image and a focal length of the second camera (12); - generation (36) of a third image by projection of said set of points according to an intrinsic matrix of said second camera (12), the arrival column indices (i) and the arrival row indices (j) in said third image being associated with each pixel of said set of pixels; - for each line of said first rectified image, scanning from left to right, according to a viewpoint of the first camera (11), of values of arrival column indices (i) and detection of a set of irregular pixels whose arrival column index (i') does not follow a monotonic function representing an evolution of arrival column indices (i) as a function of a column index (x) in the first rectified image, and for each irregular pixel of said set, identification of a set of occluded pixels in said row to the left of said each irregular pixel whose arrival column index (i) is greater than or equal to an arrival column index (i') of said each irregular pixel, a second visibility mask being determined (37) as a union of said sets of occluded pixels; and - determination (38) of a third visibility mask by association of said first visibility mask and said second visibility mask.
2. A method according to claim 1, wherein the reprojection of a pixel in the three-dimensional scene is carried out using the following formula: pM = - d ( n ï 'cs coordinates in the three-dimensional scene of the point 1 3D\} t' resulting from the reprojection of the pixel Pt of the first image, - T a displacement matrix between a position of the first camera (11) and a position of the second camera (12), - K the intrinsic matrix of the first camera (11) associated with a projection of a point of the three-dimensional scene into an image acquired by the first camera (11), - 0 a reprojection function in the three-dimensional scene of a pixel as a function of its depth, ■ I)(p ) is a depth of the pixel Pt of the first image predicted by the stereoscopic vision system.
3. A method according to any one of claims 1 to 2, wherein the projection of a point in the three-dimensional scene is carried out using the following formula: p — jj ^ with 71 a function for going from homogeneous coordinates in three-dimensional space to pixel coordinates in two dimensions by removing one dimension of a vector, - K' the intrinsic matrix of the second camera (12) associated with a projection of a point of the three-dimensional scene into an image acquired by the second camera (12), - n ( nj the coordinates in the three-dimensional scene of the point PM) from the reprojection of the pixel Pt of the first image.
4. A method according to any one of claims 1 to 3, wherein said first visibility mask is obtained using the following formula: 1 1 n tel nyo SW!2 aveC : MP,el que-^y — - xp^pt} the abscissa of a point resulting from the reprojection of the pixel Pt of the first image, an x-axis being defined parallel to an axis along which the first and second cameras (11, 12) are located, - I^ p^ l^ite depth of the pixel Pt of the first image predicted by the stereoscopic vision system, - W the width of the second image, and - f the focal length of the second camera (12).
5. A method according to any one of claims 1 to 4, wherein said depths are predicted by a convolutional neural network.
6. A method according to claim 5, wherein the convolutional neural network is trained to minimize a photometric error defined by the following loss function: L(p) = EJ(1-«) ■ \l(p)-I(p')\+a'(l-jSSIM[l(p), / ( / ;)))] With: - I(p) a value of pixel P in the second image; - a value of pixel P in the third image; - SSIM a function that takes into account a local structure; and - a a weighting factor depending on a type of road environment.
7. A computer program comprising instructions for carrying out the method according to any one of the preceding claims, when such instructions are executed by a processor.
8. Device (4) for determining a visibility mask for a vision system embedded in a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for carrying out the steps of the method according to any one of claims 1 to 6.
9. System for determining a visibility mask for a system of vision embedded in a vehicle (10) comprising at least two cameras (11, 12) and a device according to claim 8.
10. Vehicle (10) comprising the device (4) according to claim 8 or the system according to claim 9.