Method and device for determining a visibility mask by a vision system on board a vehicle.
The method uses a convolutional neural network to determine visibility masks for vehicle cameras, addressing the challenge of occlusion in monoscopic systems and enhancing ADAS system performance and safety.
Patent Information
- Application Number
- FR2023003337
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-04-04
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-04-04
AI Technical Summary
Existing ADAS systems in vehicles face challenges in determining visibility masks for monoscopic vision systems, which are crucial for improving data quality and road safety, as existing methods are not applicable to these systems.
A method using a convolutional neural network to determine a visibility mask by comparing images from different viewpoints, adjusting input parameters to minimize reconstruction errors, and applying predefined threshold values to define the visibility mask, suitable for both monoscopic and stereoscopic vision systems.
Improves the quality of data from vehicle cameras by accurately identifying occluded areas, enhancing the performance of ADAS systems and improving road safety.
Smart Images

Figure 00000030_0000 
Figure 00000031_0000 
Figure 00000032_0000
Abstract
Description
Title of the invention: Method and device for determining a visibility mask by a vision system on board a vehicle. Technical field
[0001] The present invention relates to methods and devices for determining a visibility mask by a vision system on board a vehicle, for example in a motor vehicle. The present invention also relates to a method and a device for controlling one or more AD AS systems on board a vehicle from a determined visibility mask. Technological background
[0002] Many modern vehicles are equipped with so-called AD AS (Advanced Driver-Assistance System). Such AD AS systems are passive and active safety systems designed to eliminate the element of human error in the driving of vehicles of all types. AD AS use advanced technologies to assist the driver while driving and thus improve their performance. AD AS use a combination of sensor technologies to perceive the environment around a vehicle, then provide information to the driver or act on certain vehicle systems.
[0003] There are several levels of ADAS, such as rearview cameras and blind spot sensors, lane departure warning systems, adaptive cruise control and automatic parking systems.
[0004] The AD AS embedded in a vehicle are supplied with data obtained from one or more embedded sensors such as, for example, cameras. These cameras make it possible in particular to detect and locate other road users or possible obstacles present around a vehicle in order, for example: - to adapt the lighting of the vehicle according to the presence of other users; - to automatically regulate the speed of the vehicle; - to act on the braking system in the event of a risk of impact with an object.
[0005] The proper functioning of the driving assistance peripherals using this data therefore depends on the quality of the data emitted by a vision system.
[0006] Many vision systems perceive an environment around a vehicle from several images acquired by one or more cameras. When exploiting the images, occluded areas of the images which correspond to areas of the environment which are not present on all of the acquired images are defined. A visibility mask associated with an image then defines, for example, a filter making it possible to determine pixels associated with areas not found in other images.
[0007] Solutions exist for detecting an occlusion, that is to say for determining a visibility mask.
[0008] A first solution, described in "Dense point trajectories by gpu-accelerated large displacement optical flow" by Narayanan Sundaram, Thomas Brox, and Kurt Keutzer, consists of checking the coherence between optical flows from a first image to a second image and vice versa. An optical flow indicates the displacement of the position of a pixel in a first image to a corresponding pixel in a second image. This pixel in the second image must have the opposite optical flow to return to the position of the initial pixel in the first image. The sum of the two optical flows is an indicator of occlusion. This solution is used in an optical flow algorithm and it can also be directly used for a stereoscopic vision system with at least two cameras. This solution is not used for a monoscopic vision system.
[0009] A second solution presented by "Occlusion Aware Unsupervised Learning of Optical Flow" by Yang Wang, Yi Yang, Zhenheng Yang, Liang Zhao, Peng Wang and Wei Xu published on April 4, 2018 is based on reverse optical flow. For each pixel of a first image represented by its coordinates, the algorithm checks whether a pixel of a second image arrives at this pixel of the first image with reverse optical flow by scanning all the pixels of the second image. This method can be used for both directions of optical flow to identify occluded areas of both images. As before, this method is not used for monoscopic vision systems.
[0010] A third solution described in "Digging Into Self-Supervised Monocular Depth Estimation" by Clément Godard, Oisin Mac, Michael Firman and Gabriel Brostow uses a loss function with an algorithm avoiding occluded areas without explicitly identifying them. Two reconstructions of a first image acquired at a given time instant are made on shooting times directly before and after the time instant of acquisition of the first image and are compared to two other images acquired at the time instant directly before and at the time instant directly after. The occluded areas in one image may be present in another image to be reconstructed. A loss function is calculated for the images and the smallest error for each pixel is added to a total error. This solution requires the use of three images. Summary of the present invention
[0011] An object of the present invention is to solve at least one of the problems of the technological background described above.
[0012] Another object of the present invention is to propose a solution for determining a visibility mask for any vision system in order to improve the quality of the data from the camera(s) of the vision system.
[0013] Another object of the present invention is to improve road safety.
[0014] According to a first aspect, the present invention relates to a method for determining a visibility mask by a vision system on board a vehicle, the vision system comprising at least one camera arranged so as to acquire an image of a scene from a determined point of view, the method being characterized in that it comprises the following steps: - reception of first and second data respectively representative of a first and second image acquired from a different point of view by said at least one camera; - determining a second set of pixels of the second image corresponding to a first set of pixels of the first image and obtaining first geometric data associating the pixels of the first set of pixels with the pixels of the second set of pixels by implementing a convolutional neural network; - reconstruction of a third image from the first image and the first geometric data; - determination of a first reconstruction error by comparing the third and second images; - determining a fourth set of pixels of the first image corresponding to a third set of pixels of the second image and obtaining second geometric data associating the pixels of the third set of pixels with the pixels of the fourth set of pixels by implementing the convolutional neural network; - reconstructing a fourth image from the second image and the second geometric data; - determination of a second reconstruction error by comparing the fourth and first images; - determining a third error from the first and second errors and adjusting input parameters of the convolutional neural network by minimizing the third error; - determining a visibility mask associated with the first set of pixels of the first image from the first and second errors by respective comparison of the first and second errors to predefined threshold values.
[0015] According to a method variant, the first and second errors respectively are obtained by the following loss function: L*(p) = Up[(l-«) + î(p)))] With : ■ l(p) a value of the pixel P in the second image, respectively the first image; " î(p) a value of the pixel P in the reconstructed image: the third image, respectively the fourth image
[0016]
[0017] - SSIM (from the English “structural similarity index measure”) a function which takes into account a local structure; and - has a weighting factor depending in particular on the type of environment. According to another method variant, the third error is obtained by the following loss function: ù=2 p ^(l4pï Mp)) avec : ■ 'a First reconstruction error for a pixel P of the second image; " L* (p) 'a second reconstruction error for a pixel P of the first image. According to yet another method variant, the visibility mask is determined by the following function: V n (P) = l(U p (^*t(p) > a / \L* s (p)<b)> c) with : - 1 a function returning 0 or 1; - (J the union of pixels; " L^p) 'a first reconstruction error for a pixel P of the second image; " L^p) 'a second reconstruction error for a pixel P of the first image; - an AND operator; and - a, b and c are determined parameters.
[0018] According to an additional method variant, the vision system is a monoscopic vision system formed from a single moving camera, the first image being acquired at a first time instant and the second image being acquired at a second time instant prior to the first time instant.
[0019] According to yet another method variant, the first and second geometric data are each representative of a movement of the camera between the second and first time instants, the reconstruction of the third and respectively fourth images being obtained by the following formula: ]) with : - n a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector; - K an intrinsic matrix of the camera associated with the projection of a point in space at 3-dimensional coordinates into the image at 2-dimensional coordinates; - T a displacement matrix between the camera positions at the first acquisition time instant and at the second acquisition time instant; - 0 a backprojection function in the scene of a pixel according to its depth; D / ' 1 is a depth of the pixel P, predicted by the vision system tnAPt) monoscopic.
[0020] According to another method variant, the vision system is a stereoscopic vision system comprising two cameras arranged so as to each acquire an image of a scene at the same time instant from a different point of view, optical axes representative of an orientation of the field of vision of each camera being oriented in a non-parallel manner, the first and second images being acquired by a first and a second camera respectively.
[0021] According to an additional method variant, the first and second geometric data are each representative of an optical flow determined from the first and second images by an optical flow calculation method, the reconstruction of the third and respectively fourth images being obtained by the following formula: pfs=ps+Ft^- with : - Pfs a pixel of the third image, respectively fourth image; - Fle optical flow associated with a source pixel Ps of the first image, respectively second image.
[0022] According to yet another method variant, the first and second geometric data are each representative of disparities determined from the first and second images, the reconstruction of the third and respectively fourth images being obtained by the following formula: P*ss = tf-d(Pt]âKC: - the abscissa of a pixel of the third image, respectively fourth image; - the abscissa of a pixel of the first image, respectively second image; and 'd(Pt) a disparity determined for a pixel Pt of the first image, respectively second image.
[0023] According to a second aspect, the present invention relates to a device for determining a visibility mask by a vision system on board a vehicle, the device comprising a memory associated with at least one processor configured for implementing the steps of the method according to the first aspect of the present invention.
[0024] According to a third aspect, the present invention relates to a vehicle, for example of the automobile type, comprising a device as described above according to the second aspect of the present invention.
[0025] According to a fourth aspect, the present invention relates to a computer program which comprises instructions adapted for executing the steps of the method according to the first aspect of the present invention, in particular when the computer program is executed by at least one processor.
[0026] Such a computer program may use any programming language and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.
[0027] According to a fifth aspect, the present invention relates to a computer-readable recording medium on which is recorded a computer program comprising instructions for carrying out the steps of the method according to the first aspect of the present invention.
[0028] On the one hand, the recording medium may be any entity or device capable of storing the program. For example, the medium may comprise a storage means, such as a ROM memory, a CD-ROM or a microelectronic circuit type ROM memory, or a magnetic recording means or a hard disk.
[0029] On the other hand, this recording medium may also be a transmissible medium such as an electrical or optical signal, such a signal being able to be conveyed via an electrical or optical cable, by conventional or terrestrial radio or by self-directed laser beam or by other means. The computer program according to the present invention can in particular be downloaded from an Internet-type network.
[0030] Alternatively, the recording medium may be an integrated circuit in which the computer program is incorporated, the integrated circuit being adapted to perform or to be used in performing the method in question. Brief description of the figures
[0031] Other characteristics and advantages of the present invention will emerge from the description of the particular and non-limiting exemplary embodiments of the present invention below, with reference to the appended figures 1 to 4, in which:
[0032] [Fig-1] schematically illustrates a monoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention;
[0033] [Fig.2] schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention;
[0034] [Fig.3] schematically illustrates a device configured for determining a visibility mask by a vision system on board the vehicle of [Fig.1], according to a particular and non-limiting exemplary embodiment of the present invention;
[0035] [Fig.4] illustrates a flowchart of the different steps of a method for determining a visibility mask by a vision system on board the vehicle of [Fig.l], according to a particular and non-limiting exemplary embodiment of the present invention. Description of examples of implementation
[0036] A method and a device for determining a visibility mask by a vision system on board a vehicle will now be described in what follows with joint reference to Figures 1 to 4. The same elements are identified with the same reference signs throughout the description which follows.
[0037] According to a particular and non-limiting example of embodiment of the present invention, a method for determining a visibility mask by a vision system on board a vehicle is for example implemented by a computer of the on-board system of the vehicle controlling this vision system.
[0038] The vision system comprises at least one camera arranged so as to acquire an image of a scene according to a determined point.
[0039] For this purpose, the method for determining a visibility mask by a vision system on board a vehicle comprises receiving first and second data respectively representative of a first and second image acquired from a different point of view by the at least one camera.
[0040] The method also includes determining a second set of pixels of the second image corresponding to a first set of pixels of the first image and obtaining first geometric data associating the pixels of the first set of pixels with the pixels of the second set of pixels by implementing a convolutional neural network. A third image is reconstructed from the first image and the first geometric data and a first reconstruction error is determined by comparing the third and second images.
[0041] The method also comprises determining a fourth set of pixels of the first image corresponding to a third set of pixels of the second image and obtaining second geometric data associating the pixels of the third set of pixels with the pixels of the fourth set of pixels by implementing the convolutional neural network. A fourth image is reconstructed from the second image and the second geometric data and a second reconstruction error is determined by comparing the fourth and first images.
[0042] The method then determines a third error from the first and second errors and adjusts input parameters of the convolutional neural network by minimizing the third error.
[0043] Finally, a visibility mask associated with the first set of pixels of the first image is determined from the first and second errors by respective comparison of the first and second errors to predefined threshold values.
[0044] [Fig. 1] schematically illustrates a monoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.
[0045] Such an environment 1 corresponds, for example, to a road environment formed of a network of roads accessible to the vehicle 10.
[0046] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, with an electric motor(s) or even a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as an automobile, a truck, a bus, a motorcycle. Finally, the vehicle 10 corresponds to an autonomous vehicle or not, that is to say a vehicle traveling according to a determined level of autonomy or under the total supervision of the driver.
[0047] The vehicle 10 advantageously comprises at least one on-board camera 11 configured to acquire images of a scene in the environment of the vehicle 10. This camera 11 forms the monoscopic vision system. A camera 11 is illustrated in [Fig.l]. The present invention is however not limited to a vision system monoscopic comprising a single camera but extends to any vision system comprising 1 or more cameras, for example 1, 2, 3, 4 or 5 cameras.
[0048] The camera 11 has known intrinsic parameters. These parameters consist in particular of: - the focal length f of the camera 11; - distortions which are due to imperfections in the camera's optical system; - the direction Cl of the optical axis of the camera 11; - the resolution of the camera 11.
[0049] The intrinsic parameters characterize the transformation which associates, for an image point, the camera coordinates with the pixel coordinates, in each camera. These parameters do not change if the camera is moved.
[0050] The distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions which generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the camera 11 is assumed to be perfect, that is to say that the distortions are not taken into account or that their correction is processed at the time of image acquisition.
[0051] This camera 11 is arranged so as to acquire an image of a scene according to a defined point of view, the point of view is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10.
[0052] A marker is associated with the first camera 11: - the direction of the x axis is defined as the longitudinal axis of the vehicle 10; - the direction of the y axis is defined as being the transverse axis of the vehicle 10, therefore orthonormal to the x axis; - the direction of the z axis is defined orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal reference frame.
[0053] The camera 11 acquires images of a scene located in front of the vehicle 10, the camera 11 covers a first acquisition field 13.
[0054] An obstacle 18 is placed in the acquisition field 13 of the camera. The presence of the obstacle 18 defines an occlusion field for the monoscopic vision system composed here of the field 16.
[0055] The direction Cl of the optical axis representative of an orientation of the field of vision of the camera 11 is oriented so as to obtain the widest possible acquisition field 13 of the environment 1.
[0056] It is obvious that it is possible to use such a monoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with differently placed and oriented cameras.
[0057] An image acquired by the camera 11 at a given acquisition time instant is presented in the form of data representing pixels characterized by: - coordinates in the image; and - data relating to the colors and brightness of objects in the observed scene in the form, for example, of RGB colorimetric coordinates (from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness).
[0058] When the vehicle 10 is in motion, that is to say when the camera 11 is in motion, the images acquired by the camera 11 at different time instants represent views of the same scene taken from different viewpoints, the positions of the camera 11 being distinct. On this scene are for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0059] These images are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.
[0060] [Fig. 2] schematically illustrates a stereoscopic vision system equipping a vehicle, according to a particular and non-limiting exemplary embodiment of the present invention.
[0061] In this example, the vehicle 10 corresponds to a vehicle with a thermal engine, with an electric motor(s) or even a hybrid vehicle with a thermal engine and one or more electric motors. The vehicle 10 thus corresponds, for example, to a land vehicle such as an automobile, a truck, a bus, a motorcycle. Finally, the vehicle 10 corresponds to an autonomous vehicle or not, that is to say a vehicle traveling according to a determined level of autonomy or under the total supervision of the driver.
[0062] The vehicle 10 advantageously comprises several on-board cameras 11, each configured to acquire images of a scene in the environment of the vehicle 10. This set of cameras 11 forms the stereoscopic vision system. Two cameras 11 are illustrated in [Fig. 1]. The present invention is however not limited to a stereoscopic vision system comprising two cameras but extends to any vision system comprising 1 or more cameras, for example 1, 2, or 5 cameras. A system comprising a single camera 11 then forms a monoscopic vision system.
[0063] The cameras 11 have known intrinsic parameters. These parameters consist in particular of: - focal length f of the camera 11; - distortions which are due to imperfections in the optical system of each camera; - direction C of the optical axis of the camera 11; - respective resolutions of the cameras 11.
[0064] The intrinsic parameters characterize the transformation which associates, for an image point, the camera coordinates with the pixel coordinates, in each camera. These parameters do not change if the camera is moved.
[0065] The distortions, which are due to imperfections in the optical system such as defects in the shape and positioning of the camera lenses, will deflect the light beams and therefore induce a positioning deviation for the projected point compared to an ideal model. It is then possible to complete the camera model by introducing the three distortions which generate the most effects, namely radial, decentering and prismatic distortions, induced by defects in curvature, parallelism of the lenses and coaxiality of the optical axes. In this example, the cameras are assumed to be perfect, that is to say that the distortions are not taken into account or that their correction is processed at the time of image acquisition.
[0066] These cameras 11 are arranged so as to each acquire an image of a scene from a different point of view, the first point of view is for example located on or in the left rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10, the second point of view is for example located on or in the right rearview mirror of the vehicle 10 or at the top of the windshield of the vehicle 10. In the case where two cameras are located at the top of the windshield of the vehicle, they are then placed at a certain distance. In this example, the first camera 11 is located at the top of the windshield of the vehicle 10, the second camera 11 is located in the right rearview mirror of the vehicle 10.
[0067] A first marker is associated with the first camera 11: - the direction of the y axis is defined by the position of the second camera 11, so as to place the second camera 11 on the y axis of the first camera 11. The distance B separating the two cameras 11 is called the reference base (in English “baseline”) and the direction separating the two cameras 11 is that of the y axis; - the direction of the x axis is defined orthogonal to that of the y axis and orthogonal to that of the optical axis Cl of the first camera 11; - the direction of the z axis is defined orthogonal to the directions of the x and y axes. The three axes x, y and z thus form an orthonormal reference frame.
[0068] The extrinsic parameters linked to the position of the cameras 11 are the following parameters: - 3 translations in the x, y and z directions: Tx, Ty and Tz constituting the translation vector T; and - 3 rotations around the x, y and z axes: Rx, Ry and Rz, constituting the rotation matrix R.
[0069] A main constraint of the stereoscopic vision system used in automobiles is, for example, the large distance between the two cameras. Indeed, to be able to cover a measurement range of 200 meters, the "baseline" must reach 60cm for the cameras commonly used in this field.
[0070] The two cameras 11 acquire images of a scene located in front of the vehicle 10, the first camera covering only a first acquisition field 13, the second camera covering only a second acquisition field 14 and the two cameras 11 both covering a third acquisition field 15. The first and third acquisition fields 13, 15 thus allow a monoscopic vision of the scene by the first camera 11, the second and third acquisition fields 14, 15 allow a monoscopic vision of the scene by the second camera 11 and the third acquisition field 15 allows a stereoscopic vision of the scene by the stereoscopic vision system composed of the two cameras 11.
[0071] An obstacle 18 is placed in the acquisition field of the cameras, for example in the third acquisition field 15. The presence of the obstacle 18 defines an occlusion field for the stereoscopic vision system composed here of the three fields 16, 17 and 19.
[0072] Among these three fields, field 16 is visible from the second camera 11. The part of the scene present in this field 16 is therefore observable using the monoscopic vision system composed of the second camera 11.
[0073] The field 17 is visible from the first camera 11. The part of the scene present in this field 17 is therefore observable using the monoscopic vision system composed of the second camera 11.
[0074] Finally, field 19 is not visible from any of the cameras. The part of the scene present in this field 19 is therefore not observable.
[0075] The directions C1, C2 of the optical axes representative of an orientation of the field of vision of each camera are oriented non-parallel so as to obtain the third acquisition field 15 of the environment 1 as wide as possible.
[0076] It is obvious that it is possible to use such a stereoscopic vision system to take images of scenes located on the sides or behind the vehicle 10 by equipping it with differently placed and oriented cameras.
[0077] The images acquired by the cameras 11 at a given acquisition time instant tl are presented in the form of data representing pixels characterized by: - coordinates in each image; and - data relating to the colors and brightness of objects in the observed scene in the form, for example, of RGB colorimetric coordinates (from the English “Red Green Blue”) or TSL (Tone, Saturation, Brightness).
[0078] The images acquired by the cameras 11 represent views of the same scene taken from different viewpoints, the positions of the cameras being distinct. On this scene are for example: - buildings; - road infrastructure; - other stationary users, for example a parked vehicle; and / or - other mobile users, for example another vehicle, a cyclist or a moving pedestrian.
[0079] These images are sent to a computer of a device equipping the vehicle 10 or stored in a memory of a device accessible to a computer of a device equipping the vehicle 10.
[0080] [Fig. 3] schematically illustrates a device 4 configured for determining a visibility mask by a vision system on board a vehicle 10, according to a particular and non-limiting exemplary embodiment of the present invention. The device 4 corresponds for example to a device on board the first vehicle 10, for example a computer.
[0081] The device 4 is for example configured for the implementation of the operations and / or steps described with regard to figures 1, 2 and 4. Examples of such a device 4 include, but are not limited to, on-board electronic equipment such as an on-board computer of a vehicle, an electronic calculator such as an ECU (“Electronic Control Unit”), a smartphone, a tablet, a laptop. The elements of the device 4, individually or in combination, can be integrated in a single integrated circuit, in several integrated circuits, and / or in discrete components. The device 4 can be produced in the form of electronic circuits or software (or computer) modules or even a combination of electronic circuits and software modules.
[0082] The device 4 comprises one (or more) processor(s) 40 configured to execute instructions for carrying out the steps of the method and / or for executing the instructions of the software(s) embedded in the device 4. The processor 40 may include integrated memory, an input / output interface, and various circuits known to those skilled in the art. The device 4 further comprises at least one memory 41 corresponding for example to a volatile and / or non-volatile memory and / or comprises a memory storage device which may comprise volatile and / or non-volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic or optical disk.
[0083] The computer code of the embedded software(s) comprising the instructions to be loaded and executed by the processor is for example stored in the memory 41.
[0084] According to various particular and non-limiting embodiments, the device 4 is coupled in communication with other similar devices or systems (for example other computers) and / or with communication devices, for example a TCU (from the English “Telematic Control Unit” or in French “Telematic Control Unit”), for example via a communication bus or through dedicated input / output ports.
[0085] According to a particular and non-limiting exemplary embodiment, the device 4 comprises a block 42 of interface elements for communicating with external devices. The interface elements of the block 42 comprise one or more of the following interfaces: - RF radio frequency interface, for example Wi-Fi® type (according to IEEE 802.11), for example in the 2.4 or 5 GHz frequency bands, or Bluetooth® type (according to IEEE 802.15.1), in the 2.4 GHz frequency band, or Sigfox type using UBN (Ultra Narrow Band) radio technology, or LoRa in the 868 MHz frequency band, LTE (Long-Term Evolution), LTE-Advanced; - USB interface (from the English “Universal Serial Bus” or “Universal Serial Bus” in French); HD MI interface (from the English “High Definition Multimedia Interface” or “High Definition Multimedia Interface” in French); - LIN interface (from the English “Local Interconnect Network”).
[0086] According to another particular and non-limiting exemplary embodiment, the device 4 comprises a communication interface 43 which makes it possible to establish communication with other devices (such as other computers of the on-board system) via a communication channel 430. The communication interface 43 corresponds for example to a transmitter configured to transmit and receive information and / or data via the communication channel 430. The communication interface 43 corresponds for example to a wired network of the CAN (Controller Area Network) type, CAN FD (Controller Area Network Flexible Data-Rate), FlexRay (standardized by the ISO 17458 standard) or Ethernet (standardized by the ISO / IEC 802-3 standard).
[0087] According to a particular and non-limiting exemplary embodiment, the device 4 can provide output signals to one or more external devices, such as a display screen 440, touch-sensitive or not, one or more speakers 450 and / or other peripherals 460 (projection system) via the output interfaces 44, 45, 46 respectively. According to a variant, one or other of the external devices is integrated into the device 4.
[0088] [Fig.4] illustrates a flowchart of the different steps of a method 2 for determining a visibility mask by a vision system on board the vehicle of [Fig.1], the vision system comprising at least one camera 11 arranged so as to acquire an image of a scene from a determined point of view, according to a particular and non-limiting exemplary embodiment of the present invention.
[0089] The method is for example implemented by one or more processors of one or more computers on board the vehicle 10, for example by a computer controlling the vision system.
[0090] In a first step 21, the computer receives first data representative of a first image acquired by a camera 11 at a first acquisition time instant t1.
[0091] In a second step 31, the computer receives second data representative of a second image acquired by a camera 11 at a second acquisition time instant t2.
[0092] The two images received correspond to two views of the same scene taking place around the vehicle 10.
[0093] Three embodiments of this method are then defined. These different modes can be implemented independently depending on the type of vision system but can also be combined. A combination of the different embodiments can for example allow supervision of a vision system using one embodiment by another vision system using another embodiment.
[0094] According to a first embodiment, the first and second images are taken by the same camera 11 in motion, the vehicle 10 itself being in motion. The first and second images are taken at two distinct time instants t1, t2, the second time instant t2 being prior to t1. The vision system is then a monoscopic vision system as presented in [Fig.l].
[0095] In a step 22, a second set of pixels of the second image corresponding to a first set of pixels of the first image is determined and first geometric data associating the pixels of the first set of pixels with the pixels of the second set of pixels are obtained by implementing a convolutional neural network.
[0096] According to this first embodiment, the first geometric data correspond to obtaining the displacement of the camera 11 between the two time instants t1 and t2 as well as the prediction of depths associated with the pixels of the first image.
[0097] In this first embodiment, the extrinsic parameters of the monoscopic vision system are, for example, determined by a computer associated with this same monoscopic vision system. The determination of the extrinsic parameters of the monoscopic vision system, including the movement of the camera 11 between the two time instants t1 and t2, is known to those skilled in the art and presented, for example, in the document Unsupervised Learning of Depth and Ego-Motion from Video by Tinghui Zhou, Matthew Brown, Noah Snavely and David G. Lowe published on August 1, 2017.
[0098] Depths associated with the first pixels of the first image are predicted by the monoscopic vision system from a depth prediction model, learned during a training phase, applied to the first image. Such a method is described in the document: HR-Depth: High Resolution Self-Supervised Monocular Depth Estimation by Xiaoyang Lyu, Liang Liu, Mengmeng Wang, Xin Kong, Lina Liu, Yong Liu, Xinxin Chen, and Yi Yuan published on December 14, 2020.
[0099] In a step 32, a fourth set of pixels of the first image corresponding to a third set of pixels of the second image is determined and second geometric data associating the pixels of the third set of pixels with the pixels of the fourth set of pixels are obtained by implementing the convolutional neural network.
[0100] In this first embodiment, the second geometric data correspond to obtaining the displacement of the camera 11 between the two time instants t2 and t1 as well as the prediction of depths associated with the pixels of the second image. The displacement of the camera 11 between the instants t2 and t1 is the opposite of the displacement obtained during step 22. The depths associated with the pixels of the second image are obtained using the same method as that used in step 22.
[0101] In a step 23, a third image is reconstructed from the first image and the first geometric data.
[0102] In a step 33, a fourth image is reconstructed from the second image and the second geometric data.
[0103] The third and fourth images respectively are reconstructed using the following formula: [Math 1]
[0104]
[0105]
[0106] with : - n a function to go from homogeneous coordinates to pixel coordinates by removing a dimension from a vector; - K an intrinsic matrix of the camera 11 associated with the projection of a point in space with 3-dimensional coordinates into the image with 2-dimensional coordinates; - T a displacement matrix between the positions of the camera 11 at the second acquisition time instant t2 and at the first acquisition time instant tl for the reconstruction of the third image, respectively a displacement matrix between the positions of the camera 11 at the first acquisition time instant tl and at the second acquisition time instant t2 for the reconstruction of the fourth image; - 0 a backprojection function in the scene of a pixel according to its depth; - D j is a depth of the pixel P[ of the first image predicted by the monoscopic vision system for the reconstruction of the third image, respectively a depth of the pixel Pt of the second image predicted by the monoscopic vision system for the reconstruction of the fourth image. In a step 24, a first reconstruction error is determined for each pixel of the second image by comparing the second image to the third image. Similarly, in a step 34, a second reconstruction error is determined for each pixel of the first image by comparing the first image to the fourth image. The first and second errors respectively are determined by the following loss function: [Math 2] L n Ap)= U p [(l-«)-^(P)-î(P)l + a-(l-^S / M(l(p) z î(p)))] With : ■ l(p) is a value of the pixel P in the second image, respectively the first image; " j(p) is a value of the pixel P in the reconstructed image: the third image, respectively the fourth image the first reconstruction error for a pixel P for a system of - SSIM (from the English “structural similarity index measure”) is a function which takes into account a local structure; and - a is a weighting factor depending in particular on the type of environment.
[0107] In a step 25, a third error is determined from the first and second errors. This third error is defined by the following loss function: [Math 3] Lm = 2pmin ( LnM Lm£p] )with: monoscopic vision; ' Lms(p) the second reconstruction error for a pixel P for a monoscopic vision system.
[0108] The input parameters of the convolutional neural network are then adjusted by minimizing the third error. The convolutional neural network is then said to be self-supervised because it is able to autonomously adjust its input parameters in order to improve the output data, here the first and second geometric parameters.
[0109] In a step 26, a visibility mask associated with the first set of pixels of the first image is determined. This visibility mask is obtained by respective comparison of the first and second errors with predefined threshold values and is obtained for example by the following function: [Math 4] ^rntiP) = 1( Up(^nt(p)> 3 / \Lms(p)<b)> c)
[0110] with : - 1 a function returning 0 or 1; - the union of pixels; - 'a Prcm'^rc reconstruction error for a pixel P of the second image for a monoscopic vision system; ' the second reconstruction error for a pixel P of the first image for a monoscopic vision system; - an AND operator; and - a, b and c are determined parameters. The error of the pixel not visible in the target image, but visible in the source image, must exceed a certain level defined by the value a, when the error of the reconstruction of the other direction must be less than a certain level defined by the value b. The notable difference between a and b is related to the maturity of the learning model training.
[0111] The symbol U signifies the union of pixels, provided that the number of pixels in this union is greater than the criterion c. From experience the value of c is, for example, greater than 4 pixels. The symbol 1 indicates the visibility mask resulting from a logical operator, either 1 or 0.
[0112] According to a second embodiment, the first and second images are taken by two separate cameras 11 at the same time instant t1. The vision system is then a stereoscopic vision system as presented in [Fig.2].
[0113] In a step 22, a second set of pixels of the second image corresponding to a first set of pixels of the first image is determined and first geometric data associating the pixels of the first set of pixels with the pixels of the second set of pixels are obtained by implementing a convolutional neural network.
[0114] According to this second embodiment, the first geometric data correspond to an optical flow.
[0115] Step 22 is done by implementing a method called optical flow calculation. Such a method is notably described in “UnOS: Unified Unsupervised Optical-flow and Stereo-depth Estimation by Watching Videos” by Yang Wang, Peng Wang, Zhenheng Yang, Chenxu Luo, Yi Yang and Wei Xu from June 2019.
[0116] The optical flow calculation method is performed by a convolutional neural network (CNN). This type of tool is commonly used in image processing.
[0117] The output data of this operation 22 is an optical flow representative of a displacement between each pixel of the first set of pixels of the first image and the second pixel corresponding to each first pixel in the second image.
[0118] Similarly, in a step 32, a fourth set of pixels of the first image corresponding to a third set of pixels of the second image is determined and second geometric data associating the pixels of the third set of pixels with the pixels of the fourth set of pixels are obtained by implementing the convolutional neural network.
[0119] The output data of this operation 32 is an optical flow representative of a displacement between each pixel of the third set of pixels of the second image and the fourth pixel corresponding to each third pixel in the first image.
[0120] In a step 23, a third image is reconstructed from the first image and the first geometric data.
[0121] In a step 33, a fourth image is reconstructed from the second image and the second geometric data.
[0122] The third and fourth images respectively are reconstructed using the following formula: [Math 5] p[s=ps+Ft~s- with : - Pfs one pixel of the third image, respectively fourth image; - Ft-_^s the optical flow associated with a source pixel Ps of the first image, respectively second image.
[0123] In a step 24, a first reconstruction error is determined for each pixel of the second image by comparing the second image to the third image.
[0124] Similarly, in a step 34, a second reconstruction error is determined for each pixel of the first image by comparing the first image to the fourth image.
[0125] The first and second errors respectively are determined by the following loss function: [Math 6] L f (p)= Up[(i-«)-U(p)-î(p)l + ct-(i4ss / M(i(p)7î(p)))] With : ■ l(p) is a value of the pixel P in the second image, respectively the first image; " is a value of the pixel P in the reconstructed image: the third image, respectively the fourth image - SSIM (from the English “structural similarity index measure”) is a function which takes into account a local structure; and - a is a weighting factor depending in particular on the type of environment.
[0126] In a step 25, a third error is determined from the first and second errors. This third error is defined by the following loss function: [Math 7] Û = y mm(LJdI Ldnb avec : ' Lf^p) 'a first reconstruction error for a pixel P for a stereoscopic vision system using an optical flow calculation method; ' Lf (p) ' is the second reconstruction error for a pixel P for a stereoscopic vision system using an optical flow calculation method.
[0127] The input parameters of the convolutional neural network are then adjusted by minimizing the third error. The convolutional neural network is then said to be self-supervised because it is able to autonomously adjust its input parameters in order to improve the output data, here the first and second geometric parameters.
[0128] In a step 26, a visibility mask associated with the first set of pixels of the first image is determined. This visibility mask is obtained by respective comparison of the first and second errors with predefined threshold values and is obtained for example by the following function: [Math 8] V n (p) = l(U p (Mp)>a / \L fs (p)<b)> c] with :
[0129] - 1 a function returning 0 or 1; - [J the union of pixels; " Lf^p] 'a first reconstruction error for a pixel P of the second image for a stereoscopic vision system using an optical flow calculation method; "'a second reconstruction error for a pixel P of the first image for a stereoscopic vision system using an optical flow calculation method; - an AND operator; and - a, b and c are determined parameters. The error of the pixel not visible in the target image, but visible in the source image, must exceed a certain level defined by the value a, while the error of the reconstruction of the other direction must be less than a certain level defined by the value b. The notable difference between a and b is related to the maturity of the training of the learning model.
[0130] The symbol U signifies the union of pixels, provided that the number of pixels in this union is greater than the criterion c. From experience the value of c is, for example, greater than 4 pixels. The symbol 1 indicates the visibility mask resulting from a logical operator, either 1 or 0.
[0131] According to a third embodiment, the first and second images are taken by two separate cameras 11 at the same time instant t1. The vision system is then a stereoscopic vision system as presented in [Fig.2].
[0132] According to this second embodiment, the first geometric data correspond to disparities.
[0133] In order to facilitate the analysis of the two received images, the first and second images are rectified according to a method known to those skilled in the art. Such a method is described, for example, in “Projective Rectification of Non-Calibrated Infrared Stereo Images with Global Consideration of Distortion Minimization” by Benoit Ducarouge, Thierry Sentenac, Florian Bugarin and Michel Devy of July 16, 2009.
[0134] The rectification method consists of reorienting the epipolar lines so that they are parallel with the horizontal axis of the image. This method is described by a transformation which projects the epipoles to infinity and whose corresponding points are necessarily on the same ordinate.
[0135] A rectification algorithm consists, for example, of 4 steps: - Rotate (virtually) the first camera 11 so that the epipole goes to infinity along the horizontal axis of the reference frame associated with it; - Apply the same rotation to the second camera to end up in the initial geometric configuration; - Rotate the second camera by the rotation associated with the rotation matrix 'R', corresponding to the extrinsic parameter of the starting stereoscopic vision system; - Adjust the scale in the two camera frames.
[0136] It should be noted that rectification simplifies the matching of pixels in stereo images, i.e. images obtained by a stereoscopic vision system. The corresponding pixel in the second image to a pixel in the first image (and vice versa) is positioned on the same line. From the knowledge of the epipolar geometry and therefore of a fundamental matrix of the stereo system, the objective is then to determine a pair of projective transformations, called homographies, which reorient the epipolar projections parallel to the lines of the images, therefore to the horizontal axis of the rectified cameras.
[0137] In a step 22, a second set of pixels of the second image corresponding to a first set of pixels of the first image is determined and first geometric data associating the pixels of the first set of pixels with the pixels of the second set of pixels are obtained by implementing a convolutional neural network, this step is called “stereo matching” (in English “feature matching”).
[0138] Stereo matching or disparity estimation is the process of finding pixels in stereoscopic views that correspond to the same 3D point in the scene. Rectified epipolar geometry simplifies this process of finding correspondences on the same epipolar line. There is no need to calculate the coordinates of the 3D point to find the corresponding pixel on the same line in the other image. Disparity is the distance d between a pixel and its correspondence in the other image. This disparity is horizontal when the images are rectified, with horizontality defined along the x-direction of the first image.
[0139] The output data of this operation 22 is a disparity representative of a displacement between each pixel of the first set of pixels of the first image and the second pixel corresponding to each first pixel in the second image.
[0140] Similarly, in a step 32, a fourth set of pixels of the first image corresponding to a third set of pixels of the second image is determined and second geometric data associating the pixels of the third set of pixels with the pixels of the fourth set of pixels are obtained by implementing the convolutional neural network.
[0141] The output data of this operation 32 is a disparity representative of a displacement between each pixel of the third set of pixels of the second image and the fourth pixel corresponding to each third pixel in the first image.
[0142] In a step 23, a third image is reconstructed from the first image and the first geometric data.
[0143] In a step 33, a fourth image is reconstructed from the second image and the second geometric data.
[0144] The third and fourth images respectively are reconstructed using the following formula: [Math 9] - p^s the abscissa of a pixel of the third image, respectively fourth image; - the abscissa of a pixel of the first image, respectively second image; and 'Pt) a disparity determined for a pixel of the first image, respectively second image.
[0145] In a step 24, a first reconstruction error is determined for each pixel of the second image by comparing the second image to the third image.
[0146] Similarly, in a step 34, a second reconstruction error is determined for each pixel of the first image by comparing the first image to the fourth image.
[0147] The first and second errors respectively are determined by the following loss function: [Math 10] L s (p)= U p[d-«)-k(P)-î(P)l + a-(l-iSSlM(l(p),î(p)))] With : ■ l(p) is a value of the pixel P in the second image, respectively the first image; " l(p) is a value of the pixel P in the reconstructed image: the third image, respectively the fourth image
[0148] - SSIM (from the English “structural similarity index measure”) is a function which takes into account a local structure; and - a is a weighting factor depending in particular on the type of environment. In a step 25, a third error is determined from the first and second errors. This third error is defined by the following loss function: [Math 11] Ls = ^pmin(Lst(p), LsJ,p))^c: " Lst(p) 'a first reconstruction error for a pixel P for a system of stereoscopic vision using a disparity calculation method; ■ Lss(p) the second reconstruction error for a pixel P for a stereoscopic vision system using a disparity calculation method.
[0149] The input parameters of the convolutional neural network are then adjusted by minimizing the third error. The convolutional neural network is then said to be self- supervised because it is able to autonomously adjust its input parameters in order to improve the output data, here the first and second geometric parameters.
[0150] In a step 26, a visibility mask associated with the first set of pixels of the first image is determined. This visibility mask is obtained by respective comparison of the first and second errors with predefined threshold values and is obtained for example by the following function: [Math 12] V s t(P) = 1( U p (bjp)> a A^SS (P) < b) > c) with : - 1 a function returning 0 or 1; - [J the union of pixels; " Lst(p) 'a first reconstruction error for a pixel P of the second image for a stereoscopic vision system using a disparity calculation method; ■ Lss(p) the second reconstruction error for a pixel P of the first image for a stereoscopic vision system using a disparity calculation method; - an AND operator; and - a, b and c are determined parameters.
[0151] The error of the pixel not visible in the target image, but visible in the source image, must exceed a certain level defined by the value a, when the error of the reconstruction of the other direction must be less than a certain level defined by the value b. The notable difference between a and b is linked to the maturity of the training of the learning model.
[0152] The symbol U signifies the union of pixels, provided that the number of pixels in this union is greater than the criterion c. From experience the value of c is, for example, greater than 4 pixels. The symbol 1 indicates the visibility mask resulting from a logical operator, either 1 or 0.
[0153] Thus, the definition of a visibility mask makes it possible to identify the pixels visible in the first image and not visible in the second image via one of the embodiments presented.
[0154] If the ADAS uses input data such as the depths determined by the monoscopic vision system to determine the distance between a part of the vehicle 10, for example the front bumper, and another user present on the road, the ADAS is then able to determine whether the predicted depth is reliable when the pixel is clearly visible in the first and second images.
[0155] Of course, the present invention is not limited to the exemplary embodiments described above but extends to a method for determining a visibility mask by a vision system embedded in a vehicle, which would include secondary steps without thereby departing from the scope of the present invention. The same would apply to a device configured for implementing such a method.
[0156] The present invention also relates to a vehicle, for example an automobile or more generally an autonomous land-based motor vehicle, comprising the device 4 of [Fig.2].
Claims
1. Claims Method for determining a visibility mask by a vision system on board a vehicle (10), the vision system comprising at least one camera (11) arranged so as to acquire an image of a scene from a determined point of view, said method being characterized in that it comprises the following steps: - reception (21, 31) of first and second data respectively representative of a first and second image acquired from a different point of view by said at least one camera (11); - determining (22) a second set of pixels of said second image corresponding to a first set of pixels of said first image and obtaining first geometric data associating the pixels of said first set of pixels with the pixels of said second set of pixels by implementing a convolutional neural network; - reconstruction (23) of a third image from said first image and said first geometric data; - determination (24) of a first reconstruction error by comparing said third and second images; - determining (32) a fourth set of pixels of said first image corresponding to a third set of pixels of said second image and obtaining second geometric data associating the pixels of said third set of pixels with the pixels of said fourth set of pixels by implementing said convolutional neural network; - reconstruction (33) of a fourth image from said second image and said second geometric data; - determination (34) of a second reconstruction error by comparing said fourth and first images; - determining (25) a third error from said first and second errors and adjusting input parameters of said convolutional neural network by minimizing said third error; - determining (26) a visibility mask associated with said first set of pixels of said first image from said first and second errors by respective comparison of said first and second errors with predefined threshold values, said first and respectively second errors being obtained by the following loss function: L*(p) = Up[(l-cr)-k(P)-î(P) | + a-(l-iSSTM(î{p)J(p)))] With: - l(p) a value of pixel P in the second image, respectively the first image; " î(p) a value of pixel P in the reconstructed image: the third image, respectively the fourth image - SSIM (from the English "structural similarity index measure") a function which takes into account a local structure; and - a weighting factor depending in particular on the type of environment, said third error being obtained by the following loss function: L* = 2pmm(^ L*£p))™eC : " L.^(p) said first reconstruction error for a pixel P of the second image; " L*^p) said second reconstruction error for a pixel P of the first image.
2. Method according to claim 1, for which the visibility mask is determined by the following function: Vn(p) = l(Up(Mp)>a / \L*s(p)<b)> c] with: - 1 a function returning 0 or 1; - the union of the pixels; " L*^p) said first reconstruction error for a pixel P of the second image; " L ^p) said second reconstruction error for a pixel P of the first image; - an AND operator; and - a, b and c are determined parameters.
3. Method according to one of claims 1 or 2, for which said vision system is a monoscopic vision system formed of a single moving camera (11), said first image being acquired at a first time instant and said second image being acquired at a second time instant prior to said first time instant.
4. Method according to claim 3, for which said first and second geometric data are each representative of a displacement of said camera (11) between the second and first time instants, the reconstruction of said third and respectively fourth images being obtained by the following formula: Pms= D <j]) avec : - n une fonction pour passer de coordonnées homogènes à des pixels en supprimant dimension d’un vecteur ; k matrice intrinsèque la caméra (11) associée projection point l’espace aux 3 dimensions dans l’image 2 t déplacement entre les positions au premier instant temporel d’acquisition et deuxième 0 rétroprojection scène pixel sa profondeur ■ (n cst du pf prédite par le système vision monoscopique.
5. Method according to one of claims 1 to 3, for which said vision system is a stereoscopic vision system comprising two cameras (11) arranged so as to each acquire an image of a scene at the same time instant from a different point of view, optical axes representative of an orientation of the field of vision of each camera being oriented in a non-parallel manner, said first and second images being acquired by a first and a second camera respectively (H).
6. Method according to claim 5, for which said first and second geometric data are each representative of an optical flow determined from the first and second images by an optical flow calculation method, the reconstruction of said third and respectively fourth images being obtained by the following formula: Pfs = Ps + Ft^ with: ' Pfs a pixel of the third image, respectively fourth image; - Fle optical flow associated with a source pixel Ps of the first image, respectively second image.
7. Method according to claim 5, for which said first and second geometric data are each representative of disparities determined from the first and second images, the reconstruction of said third and respectively fourth images being obtained by the following formula: Pfs=P?-d(Pf)with: - p^s the abscissa of a pixel of the third image, respectively fourth image; - the abscissa of a pixel of the first image, respectively second image; and - d( a disparity determined for a pixel of the first image, respectively second image.
8. Device (4) for determining a visibility mask by a vision system on board a vehicle (10), said device (4) comprising a memory (41) associated with at least one processor (40) configured for implementing the steps of the method according to any one of claims 1 to 7.