Method for estimating distance, computer element, computer program and embedded system for implementing the method
The method uses an image processing module with a lighting device and camera to project and analyze patterns for precise distance estimation in low light, addressing cost and performance issues in existing technologies, enhancing autonomous vehicle perception.
Patent Information
- Application Number
- FR2024001991
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-28
- Publication Date
- 2025-08-29
AI Technical Summary
Existing technologies for estimating distance in low light conditions, such as night or adverse weather, face challenges including high costs, performance degradation, and scale inconsistencies, particularly in autonomous vehicle applications.
A method using an image processing module with a lighting device to project a light pattern and a camera to acquire images, trained on a plurality of images to learn distance estimation parameters, enabling precise and robust distance mapping in low light conditions without additional vehicle equipment.
The method provides accurate and cost-effective distance estimation in low light conditions, enhancing scene perception for autonomous vehicles by leveraging existing vehicle lighting and camera systems.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method for estimating distance, computer element, computer program and embedded system for implementing the method
[0001] The present invention relates to the field of motor vehicles, in particular automobiles, and more particularly to the perception and estimation (evaluation) of the distance or depth of elements present in the field of vision of the vehicle, despite low light conditions, such as at night for example.
[0002] A known problem in the field of motor vehicles and in particular automobiles concerns visibility at night or in low light conditions and in particular the perception of distances (or depth). Adverse conditions, such as night (or dawn or dusk) or difficult weather conditions (snow, rain, dark sky, storm, etc.), represent a significant challenge for various computer vision applications. Indeed, despite significant advances in autonomous driving, the challenges of night navigation persist, in particular because of the complexity of the scenes to be analyzed. Accurate perception of distance or depth is particularly important, especially in low visibility conditions, for example to provide assistance or assistance to driving or for autonomous vehicles.This perception goes beyond immediate applications and has a profound impact on the overall perception of the scene, especially at night. Information about distance (also called "depth") is a crucial aspect of scene perception.
[0003] Various solutions are known from the prior art but they all have various drawbacks. For example, some solutions use LiDAR (Light Detecting and Ranging) which offers high depth accuracy both day and night, but its widespread adoption is hampered by significant costs. On the other hand, thermal images have demonstrated their effectiveness but they introduce new challenges such as contrast and resolution. Furthermore, camera-based methods excel in daylight but fail in low-light conditions, leading to the development of dedicated approaches.
[0004] Also known are solutions known as "structure from motion" (SfM) which exploit or predict the relative positions of cameras in a video along with depth estimation. However, a significant challenge associated with SfM is the inconsistency scale and its deviation from reality. Attempts to address scale inconsistencies have so far introduced a loss in scene geometry and / or required the use of other physical properties such as vehicle speed.
[0005] Depth information can also be obtained with RGB-D cameras. While some rely solely on stereovision, others additionally use an infrared laser-projected pattern, called active stereovision. The main challenge of active stereovision is the pattern matching between the projection and the reference, traditionally addressed by mathematical models. Recent advances incorporate machine learning and deep learning. Unlike monocular depth estimation methods, active stereovision provides scale information derived from the distance between the camera and the projector. However, studies have indicated considerable performance degradation when applied outdoors due to ambient lighting conditions and projector performance.In such scenarios, the estimation of the maximum distance and the accuracy of the reconstruction are reduced. In addition, the use of infrared projectors also presents an additional cost, especially for use on a vehicle that is not necessarily equipped with them.
[0006] The present invention therefore aims to provide solutions for estimating the distance for a motorized vehicle in low light conditions which are precise and robust, while limiting additional costs, particularly in terms of additional vehicle equipment.
[0007] This aim is achieved by a method for estimating the distance, or depth, in a field of vision of a vehicle, at night or in low light conditions, by an image processing module, said vehicle comprising, on the one hand, at least one lighting device for projecting at least one light pattern in said field of vision and, on the other hand, at least one camera for acquiring images of said field of vision, the method being characterized in that it comprises: - A prior training of said module for estimating the distance from data corresponding to a plurality of images of said field of vision, called training images, in which a lighting device projects at least one contrasting design, called pattern, in a projection zone, during a prior training phase allowing said module to generate a plurality of learned parameters corresponding to characteristics of elements belonging to said pattern and of elements belonging to the content of the scene captured in these training images, as a function of the distance at which these elements are located, then an establishment of a mapping the distance of these elements in at least one area of interest within these images, then, - Acquisition and processing, by said module, when said vehicle is in circulation at night or in low light conditions, of images provided by said camera, to provide a mapping of the distance of the elements contained in at least one surveillance zone within these images acquired in circulation, using said parameters learned during the training phase.
[0008] This aim is also achieved by a computer element comprising means for implementing the steps of a method according to the invention.
[0009] This object is also achieved by a computer program comprising instructions which, when the program is executed by an image processing module, cause this module to execute the steps of a method according to the invention.
[0010] This aim is also achieved by a vehicle assistance system comprising: - an image processing module capable of carrying out the steps of the method according to the invention; - at least one lighting device capable of projecting at least one pattern in the field of vision; - at least one camera to provide images acquired in traffic.
[0011] According to another particularity, said learned parameters correspond to characteristics representative of modifications of elements present in the projected pattern with respect to a reference pattern, these modifications being due to the distance, the shape and the orientation of the visible surfaces of elements belonging to the content of the scene in the training images and corresponding to at least one parameter among the disparity, the shape and the size of the elements belonging to the projected pattern.
[0012] According to another feature, the reference pattern is learned implicitly by said module thanks to its repeated presence in the training images during prior learning.
[0013] According to another feature, the reference pattern is provided to said module at least during prior learning.
[0014] According to another feature, said area of interest is restricted relative to said projection area within the training images.
[0015] According to another particularity, said learned parameters which are used to provide the mapping of the distance of the elements contained in the surveillance zone, within the images acquired in circulation, correspond at least to characteristics of shape and / or size of elements belonging to the content of the scene captured in the images.
[0016] According to another feature, said pattern is also projected into a projection zone in said field of vision, by said lighting device, during the acquisition and processing, by said module, of images acquired in circulation by the camera.
[0017] According to another particularity, said surveillance zone has dimensions greater than or equal to those of said zone of interest within the training images and / or to those of the projection zone within said images acquired in circulation by the camera.
[0018] According to another feature, said projection zone during the acquisition and processing of the images acquired in circulation by the camera is restricted to a part of the field of vision which is determined by a detection module as a function of elements contained in the scene captured by said camera.
[0019] According to another feature, advantage is taken of a projection, by at least one lighting device, of at least one second pattern in a second projection zone.
[0020] According to another feature, the second pattern has identical content to the first pattern but with an identical or different resolution.
[0021] According to another feature, the second pattern has a different content from that of the first pattern.
[0022] According to another feature, the first pattern and the second pattern are projected into the field of vision at different distances from the vehicle.
[0023] According to another particular feature, said lighting device is capable of projecting the pattern thanks to the fact that it comprises a plurality of light sources each capable of illuminating a restricted position of the field of vision with a variable intensity, the combined control of these intensities as a function of these positions making it possible to obtain said pattern.
[0024] Other features and advantages of the present invention will appear more clearly on reading the description of various embodiments below, made with reference to the appended drawings. Indeed, to complete the description and to allow a better understanding of the invention, a set of drawings is provided. These drawings illustrate an embodiment of the invention and form an integral part of the description, which should not be interpreted as limiting the scope of the invention, but simply as an example of the manner in which the invention can be implemented. The drawings include the following figures:
[0025] [Fig.l] [Fig.l] represents a schematic top view of a vehicle equipped with a system implementing the method according to various embodiments of the invention;
[0026] [Fig.2] [Fig.2] represents a schematic view of a system implementing the method according to certain embodiments of the invention;
[0027] [Fig.3] [Fig.3] represents a schematic view of a system implementing the method according to other embodiments of the invention than [Fig.2];
[0028] [Fig.4] [Fig.4] represents, at the top, an example of a reference pattern and, at the bottom, an example of a pattern projected onto a surface perpendicular to the image acquisition camera with its deformation due to the components of the projector, according to certain embodiments of the invention;
[0029] [Fig.5] [Fig.5] represents, on the left, an example of a part of a pattern projected onto a surface which is perpendicular to the image acquisition camera and located at a first distance and, on the right, the same part of the pattern projected onto the same surface but at a second distance greater than the first distance;
[0030] [Fig.6] [Fig.6] represents a view of an image acquired by a camera and in which a pattern is projected, with elements of the pattern and elements of the content of the scene in this image, according to certain embodiments of the invention;
[0031] [Fig.7] [Fig.7] represents, on the left, an image acquired by a black and white camera and, on the right, the distance mapping obtained on this image by implementing the method according to certain embodiments of the invention;
[0032] [Fig.8] [Fig.8] schematically represents, in black and white, real distance maps of three different scenes, with elements of the content of these scenes identified for comparison with figures 9 and 10;
[0033] [Fig.9] [Fig.9] schematically represents, in black and white, distance maps of the three scenes of [Fig.8] obtained by implementing the method according to certain embodiments of the invention using the projection of the pattern during training and during the acquisition of circulating images of these scenes;
[0034] [Fig. 10] [Fig. 10] schematically represents, in black and white, distance maps of the three scenes of [Fig.8] obtained by implementing the method according to other embodiments of the invention than [Fig.9], using the projection of the pattern only during training but not during the acquisition of circulating images of these scenes.
[0035] The present invention relates to a method for estimating distance in the field of vision of a vehicle, a computer element, a computer program and a system (embedded in the vehicle) for implementing the method. This estimation of the distance, or depth, is carried out by an image processing module (1), in a field of vision of a vehicle, at night or in low light conditions (generally less than 20 lux, or even 10 lux). In general, the vehicle comprises, on the one hand, at least one lighting device (5) for projecting at least one light pattern in said field of vision and, on the other hand, at least one camera (3) for acquiring images of said field of vision. A single "lighting" or "pattern projection" device is sufficient for the model to operate. In a vehicle, this projection can be made by one of the lighting devices (i.e., the headlights for example, in particular the dipped beam headlights) with which the vehicle is equipped, but it is possible to use another one which is dedicated to the pattern. However, the present invention makes it possible to take advantage of the equipment already present in many vehicles, as detailed below.
[0036] The present application refers to an image processing module (1) which is trained during at least one learning process and which performs an estimation of the distance within images acquired by a camera (3) thanks to the projection of a pattern by at least one lighting device (5), but also to a control unit (2) which can for example control the lighting devices of the vehicle and which can integrate driving assistance functions, as known from the prior art. Indeed, in general a control unit (2) controls the lighting devices of the vehicle to adapt the lighting to the traffic conditions and the brightness, with or without manual action by the driver and modern control units increasingly integrate advanced ("intelligent") functions, in particular for driving assistance or autonomous driving.The skilled person will of course understand that the expression “image processing module (1)” is a functional definition and in fact refers to a model, in particular of neural networks, for example of the CNN type (for “Convolutional Neural Network”). Such a model may for example be based on transformers (or “self-attentive models”) designed to manage sequential data. It has been noted by the applicant of the present application that the invention can operate with a conventional, or even basic, model, such as for example the model known in the prior art under the name “U-net” or with models dedicated to distance estimation (such as for example the models known in the prior art under the names “Adabins” or “depthformer”) and the invention is not limited to the type of model used in the module (1).Concretely, this model (or module) can be implemented by the control unit or other data processing resources (computers in particular). In addition, it may in fact be several modules and / or control units dedicated to the different tasks or functions and cooperating to implement the invention, or a single unit performing all of the tasks or functions described in the present application. The figures therefore illustrate a control unit (2) for controlling the lighting devices (5) and distinct from the image processing module (1), but it is understood that this configuration is only one non-limiting example among others. The same applies to the detection module (not illustrated) described in the present application as being capable of detecting obstacles, objects or characteristics of the environment or the scene, because such a module (known from the prior art) can be implemented in the same way.The invention also fits advantageously into the context. vehicles already equipped with driver assistance or even an autonomous driving system by providing a method and a system complementary to existing systems (often more expensive by using other sensor technologies). This complementarity makes it possible to provide redundancy in the measurements and estimations carried out, which is essential in the case of autonomous vehicles. In addition, the person skilled in the art naturally understands, by the nature of the tasks carried out, that such units, for example comprising modules, will generally comprise at least one processor executing instructions and that various forms of implementation are possible and that it is therefore not necessary to provide details on the types of hardware that can be used or used.The drawings presented on the figure pages are provided solely as examples and in a schematic and functional manner, so that the person skilled in the art will be able to imagine any type of variants from the content of this application.
[0037] Unless otherwise defined, all terms (including technical and scientific terms) used in this document must be interpreted in accordance with the practices of the profession, in particular in the field of lighting and signaling of motorized vehicles, in particular automobiles. For example, the term "field of view" (corresponding to the Anglo-Saxon term "field of view", often expressed in degrees) may be used without implying any particular limitation in the present application. It is also understood that terms in common use must be interpreted as being customary in the relevant art and not in an idealized or overly formal sense and must not be interpreted in a limiting manner, unless they are expressly defined as such in this document.The term "in circulation" ("images acquired in circulation") is used here to designate the fact that these are images of real scenes during the implementation of the method "in the field" in a vehicle ready to circulate, but the model does not require knowing the speed of the vehicle unlike certain models of the prior art, this term is not limiting and especially not as to the fact that the vehicle is moving or not. The term circulation therefore covers for example also stopping or parking.
[0038] In the present application, as generally accepted in the field of patent applications, the terms "comprises", "comprises" and "includes" as well as their derivatives (such as "comprising", "comprising", etc.) should not be understood in an exclusive sense, that is to say, these terms should not be interpreted as excluding the possibility that what is described and defined may include other elements, steps, etc.
[0039] The term “lighting device (5)” is understood in the present application more particularly to mean devices which make it possible to illuminate the environment of the vehicle: - either to be able to see, such as low beam headlights (LB) or high beam headlights (HB) even if the latter are generally prohibited in certain conditions, particularly in built-up areas, - either to be able to be seen, such as for example the position lights (PL, according to the Anglo-Saxon terminology "position light"), or the daytime running lights (DRL, according to the Anglo-Saxon terminology "day running light") but given the distances at which the distance estimation is generally desired, the pattern projection will preferably concern the dipped beam headlights (or possibly the main beam headlights).
[0040] The present invention uses the projection of a "pattern" (4) or "contrasted drawing" within the images of the scene in which the estimation of the distance is carried out by an image processing module (1) or "model". The pattern can be a drawing, a diagram, a motif, an image or a shape, preferably contrasted, that is to say preferably having contrasts (generally of brightness) sufficient for their detection by said camera despite the low brightness. The contrast can be obtained by an alternation of dark parts and lighter parts and it is not necessary to use a contrast as high as the black and white shown in the figures which are not limiting on this point either. Preferably, the transition (the edges) between the dark parts and the light parts will on the other hand be as clear as possible to facilitate the detection of deformations.In addition, the pattern may include a recurring motif or a repeated shape, such as the vehicle brand logo. Preferably, the pattern includes a high concentration of discontinuities and contrasts, particularly with edges and corners. Thus, a checkerboard is particularly advantageous since it has a high repetition of contrasting patterns with edges and corners. However, the invention is not limited to this example and the module is capable of learning the deformations of various types of pattern, even if learning is faster with such repeated patterns. The invention is therefore not limited to the example of the checkerboard illustrated in the figures because other patterns can of course be used, preferably with easily identifiable contrast zones. Indeed, the checkerboard was chosen as an example because of its ability to have high discontinuities and strong contrast.The dense concentration of corners and transitions in the pattern serves as easily recognizable and detectable features by the model, but other less regular patterns are usable. On the other hand, the invention is not limited to the use of a single pattern at a time nor to the use of a single identical pattern in all scenarios.
[0041] It will be noted that the term "light pattern" is used in the present application to designate the fact that the lighting devices project light beams of various shapes (or photometry) depending on the devices and configurations (and regulations). Indeed, the various devices emit a light beam which is specific to them and which has a particular shape, defined by photometry. For example, dipped headlights often have a beam illuminating a symmetrical zone close to the vehicle (called "fiat" in English) and an asymmetrical zone further away (called "kink" in English) thus providing a particular light pattern. However, this term pattern is not limiting of a pattern in the proper sense (which can be wrongly interpreted for example as a diagram) and therefore does not necessarily imply a variation of the photometry.Furthermore, it is known to modify the light patterns, for example, to save energy or to illuminate certain parts of the field of vision more or less in order to better identify objects or traffic conditions (for example, by retaining only the "fiat" or only the lighting necessary to indicate the size of the vehicle). The pattern (4) of the invention may therefore be projected within the light pattern of the lighting device (5) which projects this pattern or outside this pattern, in particular when this pattern has been restricted for other reasons. In general, and preferably, the pattern is projected by the lighting device within its light pattern, but it is possible for the pattern to be projected outside the projected pattern, in particular in the case where this pattern has been restricted by the control unit for other reasons (energy saving or anti-glare, for example).Furthermore, since the pattern is only needed for training, it may not be projected during the processing of online (circulating) images that are usable thanks to the illumination provided by the light pattern of at least one of the illumination devices.
[0042] Furthermore, it is understood that the pattern (4) can advantageously be projected by lighting devices (5) known from the prior art, such as “pixelated headlights” which comprise a plurality of light sources each capable of illuminating a restricted position of the field of vision with a variable intensity, the combined control of these intensities as a function of these positions making it possible to obtain said pattern (4). Indeed, various types of headlights with multiple individually controllable light sources (therefore forming pixels) are known and which easily make it possible to project a pattern (4) usable for the implementation of the present invention. It is known for example, in particular from document EP4251473, luminous lighting devices emitting light by individually controllable pixels (with in particular the advantage of an anti-glare function in certain areas of the field of vision, thanks to advanced control by a control unit).
[0043] Thus, preferably, instead of requiring that an additional specific projector be provided in the vehicle to project the pattern, the invention makes it possible to take advantage of the fact that the vehicle comprises lighting devices (5) of which at least one is capable of projecting the pattern, which represents an interesting saving, in particular compared to other methods of the prior art using other technologies.
[0044] The model is capable of learning with low resolution and provides satisfactory results with a resolution of 320x320 (in its area of interest (ROI) or its monitoring area (ZS) described in the present application), but can operate with even lower resolutions. Thus, the present invention is capable of operating on certain vehicles without requiring modification of their camera or lighting device, the resolutions of which are generally higher than that required by the model. Indeed, cameras generally have resolutions in the order of megapixels and modern HD headlights have resolutions of several hundred or thousands of pixels, or even tens of thousands of pixels. On the other hand, a pattern for example in the form of a checkerboard whose boxes (or cells) measure 2x2 pixels proves sufficient for the operation of the model, which allows it to distinguish the pattern even at lower resolutions.The implementation of the invention is therefore advantageously not limiting with regard to the technical specifications of the devices used. In certain embodiments, at least one of the lighting devices (5) of the vehicle comprises a matrix arrangement of light pixels (2) in rows and columns. For example, the matrix can comprise at least 1000 to 2000 light sources and so-called HD (high definition) devices are now known which have more than 4000 sources, or even 25000 or even 50000 sources.
[0045] Generally speaking, the invention is based on a method using at least: - A prior learning of said module (1) for estimating the distance from data corresponding to a plurality of images of said field of vision, called training images, in which a lighting device (5) projects at least one contrasting design, called pattern (4), in a projection zone (ZP), during a prior training phase allowing said module (1) to generate a plurality of learned parameters corresponding to characteristics of elements (EP) belonging to said pattern (4) and of elements (EC) belonging to the content of the scene captured in these training images, as a function of the distance at which these elements (EP, EC) are located, then an establishment of a mapping of the distance of these elements in at least one zone of interest (ROI) within these images, then, - An acquisition and a processing, by said module (1),when said vehicle is in circulation at night or in low light conditions, images provided by said camera (3), to provide a mapping of the distance of the elements contained, in at least one surveillance zone (ZS) within these images acquired in circulation, using said parameters learned during the training phase.
[0046] [Fig.7] illustrates, in the left part, an image acquired at night where most elements of the scene are very difficult to distinguish and their distances almost impossible to estimate, whereas in the left part, the mapping obtained using the process shows that the elements of the scene are better recognized and that the distances are estimated correctly.
[0047] Certain embodiments therefore relate to a method based at least on this method and certain embodiments relate to a computer element comprising means for implementing this method. Such means may for example be a computer program. Thus, certain embodiments relate to a computer program comprising instructions which, when the program is executed by an image processing module (1), cause this module (1) to execute the method described in the present application. It is therefore understood that by integrating such an image processing module (1) or such a model into the data processing means present in a vehicle, a driving assistance system for a motorized vehicle is obtained. Thus, certain embodiments relate to a vehicle assistance system comprising: - an image processing module (1) capable of carrying out the steps of the method; - at least one lighting device (5) capable of projecting at least one pattern (4) into the field of vision; - at least one camera (3) for providing images acquired in traffic.
[0048] In some of these embodiments of the system, said lighting device (5) is capable of projecting the pattern thanks to the fact that it comprises a plurality of light sources capable of illuminating, each, a restricted position of the field of vision with a variable intensity, the combined control of these intensities as a function of these positions making it possible to obtain said pattern (4). On the other hand, the system can in fact be integrated into the vehicle itself and certain embodiments therefore relate to a vehicle comprising the same means as those of the system described above.
[0049] [Fig.l] schematically and non-limitingly illustrates an example of such a vehicle or system projecting a pattern (4) in a projection zone (ZP) in which a region of interest (ROI) is used for training while a wider monitoring zone (ZS) is used for distance estimation according to certain embodiments, as explained below. [Fig.2] schematically and non-limitingly illustrates an example of such a system comprising a camera (3) for providing the images of the field of vision to the image processing module (1) and a device (5), controlled in this example by a control unit (2) for projecting the pattern (4) in the field of the camera. [Fig.3] illustrates another non-limiting example limiting other embodiments of the system which comprise two devices (5) controlled by a control unit comprising an image processing module (1) for projecting patterns allowing a camera (3) to acquire images of the scene and patterns to provide them to the module (1) allowing the estimation of distances in the scene. Based on these examples, it is understood that numerous variants are possible.
[0050] On the other hand, it is understood that the training actually allows the model to learn parameters for recognizing features due to distance. These learned parameters are referred to here as being generated because they are then stored in memory for use when processing the images acquired in circulation.
[0051] In various embodiments, the learning by the control unit (2) comprises the use of an automatic computer learning algorithm, for example by neural networks. The learning is called "preliminary" because it precedes the steps of the method which are implemented during the circulation of the vehicle (acquisition, calculation, detection, comparison, etc.), but it will be noted that this learning can either be carried out before installation on the vehicle, or be carried out after this installation so that the learning is carried out in real conditions. This learning can therefore be done with images from simulation or real images acquired, either before putting into circulation ("offline" or "offline" in English), or for example by the vehicle in circulation ("online" or "online" in English).Once the corresponding results are validated, the parameters (and values) obtained during training are used to implement various embodiments of the “online” method when the vehicle is in circulation. On the other hand, for this “online” method, the steps implemented by the module (or the control unit) when the vehicle is in circulation can be carried out continuously (for example as soon as the brightness falls below a threshold, for example 20 lux or 10 lux) or only following the detection of particular conditions by a driver assistance system.
[0052] In some embodiments, said learned parameters correspond to characteristics representative of modifications of elements (EP) present in the projected pattern (4) with respect to a reference pattern (4r), these modifications being due to the distance, the shape and the orientation of the visible surfaces of elements (EC) belonging to the content of the scene in the training images and corresponding to at least one parameter among the disparity, the shape and the size of the elements (EP) belonging to the projected pattern (4). In some of these embodiments, the reference pattern (4r) is implicitly learned by said module (1) thanks to its repeated presence in the training images during the prior learning. In others embodiments, the reference pattern (4r) is provided to said module (1) at least during the prior learning, while during the processing of the acquired real images, this reference pattern (4r) is no longer necessary and may or may not be provided to the module (1) or model.
[0053] The term "disparity" here designates a difference in position deduced using the known distance between the projector (5) and the camera (3). This term is known in the field but has never been used for a pattern projected by a vehicle headlight and used by a processing module processing the images from a vehicle camera. In addition, the pattern (4) allows the module to recognize the shapes and / or sizes of the elements of the projected pattern, which is particularly advantageous for estimating the distance. [Fig.4] illustrates, in its upper part, a reference pattern (4r) and, in its lower part, a pattern as projected in a projection zone (ZP), with the deformations generally observed (due to the optical properties of the lighting device) and in particular the degradation of the pattern on the edges. Thus, the use of a restricted area of interest (ROI) within the projection zone (ZP) is preferred to optimize learning.Thus, in some embodiments, said area of interest (ROI) is restricted relative to said projection area (ZP) within the training images. In addition, the projected pattern also deforms as a function of distance as illustrated in [Fig.5] in which the left part shows a part of the pattern projected onto a surface perpendicular to the camera axis which is 10 meters away from the camera, while the right part shows this same part of the pattern when the surface is 100 meters away from the camera. It is therefore understood that the size of the elements (EP) of the pattern, such as the size of the boxes or cells in this example of the checkerboard in [Fig.5], provides useful information to the model for estimating the distance from the surface onto which a part of the pattern is projected. On the other hand, [Fig.6] illustrates the projection of the pattern into a more complex real scene and shows the deformations of the elements (EP) depending on the orientation of the surfaces onto which the pattern is projected. We see that if the surface is horizontal, the squares of the checkerboard are deformed into rectangles or trapezoids and if the surface is curved, the squares are deformed accordingly. These elements (EP) of the pattern therefore provide information on the elements (EC) belonging to the scene content, including the orientation, shape and distance of their surfaces. Thus, [Fig.6] also illustrates the elements (EC) of the scene content that the model can recognize outside the pattern projection area (e.g. the tree on the left or the wall on the right).
[0054] In some embodiments, said learned parameters which are used to provide the mapping of the distance of the elements contained in the surveillance zone (ZS), within the images acquired in circulation, correspond at least to shape and / or size characteristics of elements (EC) belonging to the content of the scene captured in the images. Indeed, even in the absence of a pattern, the module is able to use the size and shape of the elements to estimate their distance, as explained above and detailed below.
[0055] In certain embodiments, said pattern (4) is also projected into a projection zone (ZP) in said field of vision, by said lighting device (5), during the acquisition and processing, by said module (1), of the images acquired in circulation by the camera (3). Indeed, it has been observed that the module trained with the pattern was capable of estimating distances even in the absence of the pattern, even if its performance is sometimes degraded depending on the content of the field of vision.Thus, it is possible not to use the pattern at all or to limit its use, for example by projecting it for a certain time (e.g. to avoid dazzling an oncoming vehicle) and / or projecting it only in a part of the field (to also avoid dazzling in a part of the field while estimating the distance in the rest of the field or to improve the distance estimation only in a part of the field, for example defined following an obstacle detection by another module. Figures 8, 9 and 10 illustrate this capability. [Fig.8] corresponds to a real mapping of distances in a scene, while [Fig.9] corresponds to a mapping obtained with the projection of the pattern both during training and during the acquisition of this scene. It can be seen that all the elements (EC) of the scene content are recognized with a correct estimation of the distance. [Fig.10] illustrates a mapping obtained with the projection of the pattern only during training but not during the “online” acquisition of the image of this scene. We see that most of the elements (EC) of the scene are recognized but that the particular elements EC1 and EC6 are particularly blurred and that the element EC5 is not recognized at all. We therefore understand that the model works even in the absence of the pattern for the “in circulation” estimation but that it is more efficient if the pattern remains present during the “in circulation” acquisition.
[0056] On the other hand, in certain embodiments, said monitoring zone (ZS) has dimensions greater than or equal to those of said area of interest (ROI) within the training images and / or to those of the projection zone (ZP) within said images acquired in circulation by the camera (3). Such a configuration is represented for example in [Fig.l] and an example of a result is illustrated in [Fig.6]. Indeed, it has been observed that even by training the module on a restricted region of interest, it was capable of processing a wider field after the training phase. Thus, it is possible to limit the projection of the pattern to a zone of interest (ROI) during training but to use a larger monitoring zone. large (ZS) for use in traffic. Moreover, as demonstrated by [Fig.8], 9 and 10, the model correctly trained with a pattern can then operate in its absence, which removes a number of constraints on the extent of the surveillance zone (ZS), even if the performance is reduced in the absence of the pattern.
[0057] The embodiments projecting the pattern also for the images acquired in circulation are therefore preferred, in particular because the characteristics identified by the model can then also correspond to the disparity as explained above. In addition, thanks to the presence of the pattern, the performance of the model is significantly improved thanks to the information provided by the modifications of shape and / or size of the pattern (4) which are due to its projection onto the elements (EC) belonging to the content of the scene. These embodiments are therefore clearly preferred for effective distance estimation despite the low ambient light.The model is then in fact capable of also distinguishing characteristics representative of modifications of elements (EP) present in the projected pattern (4) compared to the reference pattern (4r), these modifications being due to the distance, the shape and the orientation of the visible surfaces of elements (EC) belonging to the content of the scene in the training images and corresponding to at least one parameter among the disparity, the shape and the size of the elements (EP) belonging to the projected pattern (4). These parameters concerning the elements (EP) of the pattern facilitate the identification of the elements (EC) of the content of the scene and improve the result (in terms of speed and precision).
[0058] In some of these advantageous embodiments, said projection zone (ZP) during the acquisition and processing of the images acquired in traffic by the camera (3) is restricted to a part of the field of vision which is determined by a detection module as a function of elements (EC) contained in the scene captured by said camera (3). The detection module can for example detect vehicles traveling in the opposite direction and thus prevent the projection of the pattern onto these vehicles to avoid dazzling or detect an unforeseen obstacle and on the contrary specifically project the pattern onto this obstacle.This makes it possible to improve the estimation of the distance over this projection zone (ZP) and the monitoring zone (ZS) can then either be limited to this projection zone (or even restricted to an area inside this projection zone), but the model remains capable of continuing to estimate the distance outside this limited projection zone and it is therefore not necessary to limit the monitoring zone (ZS), except when necessary to speed up the calculation time for example. Such embodiments make it possible for example to target areas in the field of view where the estimation becomes crucial or to limit the use of computing resources for example in favor of other functions or to save energy.
[0059] Finally, in certain embodiments, the method comprises a projection, by at least one lighting device (5), of at least a second pattern, for example in a second projection zone (or in the same zone as illustrated in [Fig.3]). The same zone allows a superposition of the patterns and allows for example to benefit from the presence of the pattern at certain locations of the scene by one of the devices while an obstacle prevents the projection of the pattern on these locations by the other device. Different patterns make it possible to combine the advantages of the two patterns on the recognition of the characteristics of the content of the scene. In some of these embodiments, the second pattern has identical content to the first pattern (4) but with an identical or different resolution. The use of different resolutions also presents advantages in terms of estimation accuracy and processing time.In other such embodiments, the second pattern has a different content than the first pattern (4). This makes it possible, in particular, to double the amount of information that can be extracted from the pattern. For example, the first pattern may comprise vertical stripes while the second comprises horizontal stripes. Their combination provides a checkerboard grid that is excellent for distance estimation, while ensuring good distance mapping even if an obstacle obscures one of the two patterns over a portion of the field of view. In some embodiments, the first pattern (4) and the second pattern are projected into the field of view at different distances from the vehicle. Thus, optimization of the distance estimation by the presence of "in circulation" patterns can be achieved over a larger portion of the field of view.
[0060] It is understood from the present application that the pattern is defined at the time of training because the model learns it implicitly but it can also be provided at the time of training and / or in traffic (eg, real conditions and processing of acquired images). In addition, certain embodiments allow the selective application of the pattern to specific regions of interest within a scene. This targeted application could prove particularly useful in driver assistance or autonomous driving, in scenarios requiring obtaining precise information on the distance of specific objects, such as for example the detection of lost goods on highways.
[0061] Since edge sharpness is generally important regardless of distance, the model is trained to learn how to detect edges. In addition, the model enables accurate representation of surfaces in depth maps, thus enabling the identification of planar areas such as roads and buildings, but also facilitating the reconstruction of object surfaces, for example for possible classification tasks.
[0062] It is understood from the present application that the present invention makes it possible to take advantage of the pixelated headlights of modern vehicles capable of projecting patterns easily (although other means of projecting the pattern are conceivable and therefore within the scope of the present application, for example with masks in front of the headlights). Numerous tests have been able to demonstrate the effectiveness of the method, highlighting notable and robust improvements in depth perception inside and beyond the illuminated areas. The versatility of the method has further been confirmed by its implementation in a classic model such as the U-net model but also in more complex models with cutting-edge architectures such as the Adabins model and the DepthFormer model. It is therefore understood that the present invention is not limited to these example models and may be used with other types of past, present and future models.
[0063] The present application describes various technical features and advantages with reference to the figures and / or to various embodiments. Those skilled in the art will understand that the technical features of a given embodiment may in fact be combined with features of another embodiment unless the opposite is explicitly mentioned or it is obvious that these features are incompatible or that the combination does not provide a solution to at least one of the technical problems mentioned in the present application. Furthermore, the technical features described in a given embodiment may be isolated from the other features of this embodiment unless the opposite is explicitly mentioned.
[0064] Detailed list of references in the figures: 1 image processing module 2 control unit 3 camera 4 pattern 4r reference pattern 5 lighting device
Claims
Claims
1. Method for estimating the distance, or depth, in a field of vision of a vehicle, at night or in low light conditions, by an image processing module (1), said vehicle comprising, on the one hand, at least one lighting device (5) for projecting at least one light pattern in said field of vision and, on the other hand, at least one camera (3) for acquiring images of said field of vision, the method being characterized in that it comprises: - A prior training of said module (1) for estimating the distance from data corresponding to a plurality of images of said field of vision, called training images, in which a lighting device (5) projects at least one contrasting design, called pattern (4), in a projection zone (ZP),during a preliminary training phase allowing said module (1) to generate a plurality of learned parameters corresponding to characteristics of elements (EP) belonging to said pattern (4) and of elements (EC) belonging to the content of the scene captured in these training images, as a function of the distance at which these elements (EP, EC) are located, then an establishment of a mapping of the distance of these elements in at least one zone of interest (ROI) within these images, then, - An acquisition and processing, by said module (1), when said vehicle is in circulation at night or in low light conditions, of images provided by said camera (3), to provide a mapping of the distance of the elements contained in at least one surveillance zone (ZS) within these images acquired in circulation, thanks to said parameters learned during the training phase.,
2. Method according to claim 1, characterized in that said learned parameters correspond to characteristics representative of modifications of elements (EP) present in the projected pattern (4) with respect to a reference pattern (4r), these modifications being due to the distance, the shape and the orientation of the visible surfaces of elements (EC) belonging to the content of the scene in the training images and corresponding to at least one parameter among the disparity, shape and size of the elements (EP) belonging to the projected pattern (4).
3. Method according to claim 2, characterized in that the reference pattern (4r) is learned implicitly by said module (1) thanks to its repeated presence in the training images during prior learning.
4. Method according to claim 2 or 3, characterized in that the reference pattern (4r) is provided to said module (1) at least during prior learning.
5. Method according to one of the preceding claims, characterized in that said area of interest (ROI) is restricted relative to said projection area (ZP) within the training images.
6. Method according to one of the preceding claims, characterized in that said learned parameters which are used to provide the mapping of the distance of the elements contained in the surveillance zone (ZS), within the images acquired in circulation, correspond at least to characteristics of shape and / or size of elements (EC) belonging to the content of the scene captured in the images.
7. Method according to one of the preceding claims, characterized in that said pattern (4) is also projected into a projection zone (ZP) in said field of vision, by said lighting device (5), during the acquisition and processing, by said module (1), of the images acquired in circulation by the camera (3).
8. Method according to one of the preceding claims, characterized in that said monitoring zone (ZS) has dimensions greater than or equal to those of said zone of interest (ROI) within the training images and / or to those of the projection zone (ZP) within said images acquired in circulation by the camera (3).
9. Method according to one of claims 7 or 8, characterized in that said projection zone (ZP) during the acquisition and processing of the images acquired in circulation by the camera (3) is restricted to a part of the field of vision which is determined by a detection module as a function of elements (EC) contained in the scene captured by said camera (3).
10. Method according to one of the preceding claims, characterized in that it comprises a projection, by at least one device (5) lighting at least a second pattern in a second projection zone.
11. Method according to claim 10, characterized in that the second pattern has a content identical to the first pattern (4) but with an identical or different resolution.
12. Method according to claim 10, characterized in that the second pattern has a different content from that of the first pattern (4).
13. Method according to one of claims 10 to 12, characterized in that the first pattern (4) and the second pattern are projected into the field of vision at different distances from the vehicle.
14. A computer element comprising means for implementing the steps of a method according to any one of the preceding claims.
15. Computer program comprising instructions which, when the program is executed by an image processing module (1), cause this module (1) to execute the steps of a method according to any one of claims 1 to 13.
16. Vehicle assistance system comprising: - an image processing module (1) capable of carrying out the steps of the method according to any one of claims 1 to 13; - at least one lighting device (5) capable of projecting at least one pattern (4) in the field of vision; - at least one camera (3) for providing images acquired in traffic.
17. System (1) according to the preceding claim, in which said lighting device (5) is capable of projecting the pattern thanks to the fact that it comprises a plurality of light sources capable of illuminating, each, a restricted position of the field of vision with a variable intensity, the combined control of these intensities as a function of these positions making it possible to obtain said pattern (4).
Citation Information
Patent Citations
Method for controlling a lighting system using a non-glare lighting function
EP4251473A1
Accelerating speckle image block matching using convolution techniques
US20230072702A1
Automotive system with controllable headlamp
US20240029282A1