Active stereo depth camera for operation in optically contaminated environments

The stereo depth camera with event sensors and active beam steering collects multiple depth points and applies statistical methods to enhance reliability and accuracy in optically contaminated environments, addressing the limitations of current depth cameras.

WO2026012578A1PCT designated stage Publication Date: 2026-01-15JABIL OPTICS GERMANY GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/069387
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Current depth cameras, particularly time-of-flight and active stereo solutions, are susceptible to errors in optically contaminated environments due to single depth readings being corrupted by contaminants like rain, dust, and snow, and increasing frame rates leads to increased processing load and power consumption.

Method used

A stereo depth camera using multiple detectors configured as event sensors and an illumination source with active beam steering, collecting multiple depth points per pixel during a predetermined time interval, and applying statistical analysis, filters, deep learning, or machine learning algorithms to determine the most probable distance, while using bandpass filters to restrict wavelengths and employing multiple illumination sources with different paths to enhance reliability.

Benefits of technology

The solution provides a stereo depth camera that operates reliably in optically clear and contaminated environments by reducing noise and improving accuracy through event-based pixel detection and statistical methods, enhancing reliability and speed in harsh conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024069387_15012026_PF_FP_ABST
    Figure EP2024069387_15012026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates a stereo depth camera and in particular to a stereo depth camera based on multiple detectors configured as event sensors and at least one illumination source oriented by 1-axis or 2-axis high-speed scanning for operation in optically clear to contaminated environments. A stereo depth camera (100) according to the present invention comprises an illumination unit (10) including a first illumination source (12-1) for emitting a first light beam, wherein the illumination unit (10) is configured for actively steering the first light beam to a first measurement position (P) in a surrounding of the stereo depth camera (100) within the field of view of the illumination unit (10); an imaging unit (20) including at least two detectors (22-1, 22-2) for detecting the first light beam reflected by a first object point at the first measurement position (P), wherein the detectors (22-1, 22-2) have different alignments and are configured to image a common field of view to enable stereoscopic imaging of the first object point; wherein the detectors (22-1, 22-2) include multiple pixels (S1, S2) that are configured as event sensors sharing a predefined pixel threshold trigger setting and which are adapted to collect multiple depth points from the first object point at the first measurement position (P) during a predetermined time interval Δt, wherein statistical analysis, filters, deep learning or machine learning algorithms, and / or a trained AI-model are applied to the collected depth points for determining and outputting the most probable distance of the first object point as pixel depth value for the first measurement position (P).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Active Stereo Depth Camera for Operation in Optically Contaminated Environments

[0002] The present invention relates a stereo depth camera and in particular to a stereo depth camera based on multiple detectors configured as event sensors and at least one illumination source oriented by 1-axis or 2-axis high-speed scanning for operation in optically clear to optically contaminated environments.

[0003] Technological Background

[0004] There exist different depth camera technologies for calculating depth of objects in a scene such as light detection and ranging (LiDAR), time-of-flight (direct and indirect TOF), frequency modulated continuous wave (FMCW), structured light, stereo, active stereo, etc. Such technologies are used to support applications for collision avoidance, obstacle detection, object identification, path monitoring, etc. for robots, automobiles, forklifts, drones, and many other platforms that operate in optically hostile conditions (snow, rain, dust, etc.).

[0005] A key limitation particularly of current state-of-the-art time-of-flight and active stereo solutions is the collection of only a single depth point per pixel and per frame. In optically clear environments this typically works fine, however, in harsh visual environments contaminated with rain, dust, hail, snow, etc. the single depth readings can be easily corrupted by contaminations along the measurement path.

[0006] If the depth camera is frame based, then the image capture frame rate is a critical element for operation in harsh environments. Traditional depth cameras based on time-of-flight, structured light, or active stereo are based on frame-based cameras, typically working with 15-30 fps depending on the detector resolution and pixel bit depth. If particles are in the measurement path between an object of interest and the illumination source or receiving detector, then an errant depth measurement could result. An obvious approach is to increase the frame rate of the depth camera to capture more depth measurements for each pixel per time interval, but this can lead to an inordinate increase in pixel processing load and required laser power for actively illuminated depth cameras.

[0007] It is therefore an object of the present invention to avoid or at least reduce the problems of current state-of-the-art time-of-flight and active stereo solutions and to provide a depth camera that can operate in optically clear and optically contaminated environments, and which thus is widely insensitive to rain, dust, hail, snow, contaminants, etc. along the measurement path.

[0008] Summary of Invention

[0009] The present invention solves the objective problem by providing a stereo depth camera as defined in independent claim 1. Further preferred embodiments of the invention result from features mentioned in the dependent claims.

[0010] A stereo depth camera according to the present invention comprises an illumination unit including a first illumination source for emitting a first light beam, wherein the illumination unit is configured for actively steering the first light beam to a first measurement position in a surrounding of the stereo depth camera within the field of view of the illumination unit; an imaging unit including at least two detectors (also referred to as image sensors or pixelated image detectors) for detecting the first light beam reflected by a first object point at the first measurement position, wherein the detectors have different alignments and are configured to image a common field of view to enable stereoscopic imaging of the first object point. The detectors include multiple pixels that are configured as event sensors (events are preferably triggered per pixel) sharing a predefined pixel threshold trigger setting and which are adapted to collect multiple depth points from the first object point at the first measurement position during a predetermined time interval At, wherein statistical analysis, filters (e.g., Kalman), deep learning or machine learning algorithms, and / or a trained artificial intelligence model (Al-model) are applied to the collected depth points for determining and outputting the most probable distance of the first object point as pixel depth value for the first measurement position.

[0011] At least one bandpass filter may be placed in the optical path before the detector(s) (i.e., before at least one event sensor) restricting the incoming light to wavelengths near the wavelength of the illumination source (EELs, VCSELS, LEDs, etc.), wherein all or substantially all other wavelengths are blocked.

[0012] The statistical analysis, filter, deep learning or machine learning algorithms, and / or the trained Al-model may be performed by or running on a suitable processor which may be configured for determining and outputting the most probable distance of the object point as pixel depth value according to the first measurement position after the statistical analysis, filter, deep learning or machine learning algorithms, and / or the trained Al-model have been applied to the collected depth points. A stereo depth camera uses at least two detectors to calculate depth of an object in the surrounding. The cameras are configured to image a common field of view (“observable object space”) but have different alignments with respect to said common field of view. Therefore, an object point of an object in the surrounding can be observed by two or more detectors under slightly different viewing angles. Based on the known respective alignment of the detectors and their internal spacing, the position of the object point (at which illumination light from the illumination unit is reflected) can be calculated. The illumination unit typically emits a pulsed or continuous light beam. The light beam may be raster scanned across the field of view of the illumination unit by actively steering to various measurement positions. The active steering may be performed by a galvanometer-based scanning motor with an optical mirror (a so-called galvo scanner), a MEMS structure, polygon scanner, liquid lens, etc. The illumination sources may be laser sources, wherein the emitted laser radiation may have a typical frequency (application dependent) in a spectral range reaching from the visible (VIS) to short wave infrared (SWIR).

[0013] The detectors may be two-dimensional (2D) matrix detectors including a planar or curved detector surface comprising individual pixels. Such detectors are well-known in the prior art. However, the pixels (detectors) are configured as event sensors, which means that a predefined pixel threshold trigger setting may define a lower and / or higher threshold for the detection of an event at a pixel. Pixel triggering, or event sensing, may be set by adjusting the magnitude of change of the light intensity, which may be an increase or a decrease, on a pixel. Weak signals may thus not be able to trigger such an detection event. This allows to reduce detection noise possibly caused by contaminations along the beam path. For example, shortterm reflections from rain or snow which temporarily passes the light beam may be too weak to be detected by the detectors and are then excluded from the statistics.

[0014] The benefits of using an event sensor for stereo imaging are:

[0015] 1. Path planning of the laser(s) along the image sensor provides inherent filtering of noisy pixels, defined as triggered pixels or events outside of the path of illumination. In spite of being triggered, these pixels may be removed from the stereo matching process.

[0016] 2. Event sensors simplify the correspondence or stereo matching of pixels between the image sensors. At any timestamp only a few pixels may be triggered at a predefined time, and thus the stereo-matching of pixels can be dramatically simplified.

[0017] 3. A common weakness of stereo in depth sensing is the reliance on texture in a scene for stereo pairs matching. In scenes with limited texture, stereo pairs matching may fail to occur despite the use of complex algorithms. In contrast, stereo matching in the present invention may be accomplished through matching of the pixels of the independent event sensors, by an associated timestamp assigned to the pixel(s) triggered by the active illumination.

[0018] 4. Pixels, can trigger and reset in nanoseconds allowing for a rapid accumulation of depth values at a prescribed measurement position.

[0019] The detectors are adapted to collect multiple depth points from the object point at a measurement position during a predetermined time interval At. That means in the case of pulsed illumination, depth points may be repeatedly collected by detecting several pulses reflected at the object point at a specific measurement position during the time interval At. In the prior art, typically only one depth point is collected at a specific measurement position, i.e., only a single reflected pulse is measured at each measurement position. The time interval At may be a continuous single time interval (repeated measurement without breaks in between) or a composite time interval based on a sequence of non-continuous sub-intervals (repeated measurement with short breaks in between as sub-frame repetition).

[0020] By collecting multiple depth points for a specific measurement position, statistical methods can be applied to the individual detection events. Therefore, statistical analysis, filters, deep learning or machine learning algorithms, and / or a trained Al-model are applied to the collected depth points for determining and outputting the most probable distance of the object point as pixel depth value according to the first measurement position. In contaminated environments, the statistic approach may highly increase the reliability of the outputted depth values. However, in clear environments the number of collected depth points may be reduced or only a single depth points may be collected. This may increase the effective frame rate of the depth camera.

[0021] The present invention thus provides a stereo depth camera with improved reliability for working in contaminated environments by defining a pixel-based detection threshold (or thresholds) for noise reduction and by collecting multiple depth points at a specific measurement position during a predetermined time interval At to allow statistical methods on the collected multiple depth points to be applied (also learning algorithms are based on statistical methods). The use of deep learning or machine learning algorithms, and / or a trained Al-model for analysing large amounts data is well-known in the prior art and widely used for image enhancement and data processing. In the present case, they are applied to the collected multiple depth points for determining and outputting the most probable distance of the object point as pixel depth value according to the first measurement position.

[0022] The objective problem is thus solved by collecting multiple depth points per pixel, preferably in an extremely short time period, and applying statistical analysis, filtering, and / or machine learning algorithms to output a final pixel depth reading. For increasing the detection reliably and to reduce detection noise, an active stereo solution employing multiple detectors with multiple pixels configured as event sensors is applied. Unlike frame-based cameras, event sensors outputs pixel data only if a pixel experiences a predefined threshold triggering event. The individual pixels in an array detector, which may be obscured by contaminants in the field of view, can output repeated depth values without the processing penalty typically associated with conventional frame-based cameras.

[0023] If two or more detectors observing a common field of view are present and the pixels share a common or similar pixel threshold trigger setting, then an event, typically a laser beam, in the environment can trigger pixel events (and corresponding timestamps) in each of the event sensors. Stereo depth can then be determined by correlating the triggered pixels of the individual detectors. Triangulation of the pixels provides the relevant baseline distance and angles necessary for the determination of stereo depth values. The use of an illumination source with active beam steering provides real-time control of the laser path trajectory, speed, and pixel repetition.

[0024] Preferably, the illumination unit includes a second illumination source emitting a second light beam independent of and separable from the first light beam and the illumination unit is configured for actively steering the second light beam to a second measurement position different from the first measurement position of the first light beam. However, more than two further illumination sources may be used. In particular, several independent light beams can be used for a parallel illumination (“stereo” illumination). Multiple detectors and illumination sources can thus be simultaneously combined with advantage for faster and more accurate depth sensing.

[0025] Preferably, the detectors are adapted to collect multiple depth points from a second object point at the second measurement position during a predetermined time interval At’ and wherein statistical analysis, filters (e.g., Kalman), deep learning or machine learning algorithms, and / or a trained Al-model are applied to the collected multiple depth points for determining and outputting the most probable distance of the second object point as pixel depth value for the second measurement position. The “stereo” illumination unit is thus configured for actively steering at least a first light beam to a first measurement position and a second light beam to a second measurement position in a surrounding of the stereo depth camera within in the field of view of the “stereo” illumination unit.

[0026] The detectors are further adapted to collect multiple depth points from a second object point at the second measurement position (in addition to the collection of multiple depth points from a first object point at the first measurement position). In other words, a depth camera according to the present invention is configured as a stereo depth camera able to stereo image several object points simultaneously. In some embodiments, the number of illumination sources and detectors may be the same (e.g., two of each) or in the same order while in other embodiments their number may be different (e.g., three illumination sources and two detectors). The predetermined time interval At’ and the predetermined time interval At may be different time intervals, however, preferably they refer to a common time interval AT.

[0027] Preferably, the illumination unit includes a second illumination source emitting a second light beam independent of and separable from the first light beam and the illumination unit is configured for actively steering the second light beam to a second measurement position coincident with the first measurement position. In contrast to the previously described embodiment, here the first and second measurement position are identical and thus the light beams are overlapping at a certain distance from the illumination source. This allows to collect depth points of an object point along different measurement paths and thus individual contaminants in only one of the illumination directions may be detected. A further advantage of such a configuration is the increase in the total intensity of the illumination light, which is, however, split between different paths. This way safety requirements may still be fulfilled even with a high total illumination intensity. The use of multiple illumination sources for illuminating a single object point allows to increase the detection reliability in harsh environments.

[0028] Preferably, the wavelength of the first light beam and the second light beam are different. With different wavelengths, the first light beam and the second light beam can easily be distinguished. However, another advantage may be different scattering properties along a contaminated measurement path. According to the difference in wavelength, the scattering may be different for the first light beam and the second light beam and thus also the type of contamination may be considered in the applied statistic method.

[0029] Preferably, the triggered pixels of the at least two detectors are matched by related event timestamps, laser trajectory analysis or beam progression analysis. The collection of multiple depth points from an object point at a measurement position requires that the right pair of triggered pixels on the various detectors can be correlated for depth calculation. Only pixels that are triggered corresponding to a single illumination event and from the same object point can be used for depth calculation. Therefore, correlation between related triggered pixels on the different detectors is required. This matching may be based on related event timestamps (prerequisite is an appropriate temporal separability between the events), laser trajectory analysis or beam progression analysis.

[0030] Preferably, the imaging unit comprises at least one polygon scanner for 2D depth capture in a first direction, or at least one polygon scanner for 2D depth capture in a first direction in combination with another beam steering technology for 3D depth capture in a second direction different from the first direction. In cases with more than one illumination sources, one polygon scanner per light beam may be used. A polygon scanner may be a rotating element with a polygonal structure comprising a variety of planar reflecting surfaces or facets which have one axis parallel to the rotation axis. However, a polygon scanner may have a different alignment.

[0031] For example, a polygon scanner may be an octagonal structure with eight reflecting surfaces that are arranged to form a graduated ring (see FIG. 9). When the polygonal structure is rotated, a light beam incident to a side surface will be reflected under constantly varying conditions and thus, depending on the alignment, it is scanned in at least one dimension. Polygon scanners are known from material processing where they are used for separating high frequency pulses due to their high deflection rate. In the present invention they allow a fast scanning of light beams in one dimension. A galvanometer-based scanning motor with an optical mirror is another known beam steering technology which is typically slower than polygon scanners but may provide a higher precision of deflection.

[0032] Preferably, for 3D depth capture, the measurement position is scanned in the first direction by the at least one polygon scanner to collect multiple depth points from the object point for a predetermined time interval before the measurement position is varied in a second direction to implement a raster scan pattern on the detectors. Specifically, the combination of two scanning technologies with different properties for 3D depth capture (i.e. , two scanning dimensions and depth sensing) allows optimized raster scanning schemes to sequentially collect multiple depth points from multiple object points at measurement positions along a first dimension before the light beam is varied in the second dimension. Preferably, the raster scan pattern on the detectors are adapted to parameters of the stereo depth camera and / or environmental conditions. Parameters of the stereo depth camera can be operation parameters or specific modes of operation. For example, the raster scan may be adapted to specific imaging requirements in which only a part of the observable field of view is to be observed and thus a discontinuous scanning can be performed in which some raster elements may be skipped. In other embodiments, some of the raster elements may not be skipped but selected to be scanned with lower resolution or less accuracy as compared to other (e.g., more relevant or important) raster elements.

[0033] The adaption to environmental conditions may include the consideration of a degree of contamination. In cases with a low or vanishing contamination, a raster scan pattern with faster scan speeds across the FOV may be used while in cases with high contamination, the raster scan pattern may be adapted to provide more reliability but which may be slower, to cover the entire FOV, due to an increased number of measurement and subsequent total measurement time. Preferably, the skipping or selecting of raster elements for scanning may be adapted to the environmental conditions. The environmental conditions may be determined independent from the optical system of the stereo depth camera, e.g., taken from another measurement device or loaded from an external data source.

[0034] Preferably, the parameters of the raster scan are varied in real-time based on an analysis of the depth points. The parameters of the raster scan may relate to any variable, function or selection in relation to the raster scan and in particular also to the definition of the raster scan pattern itself. The analysis of the depth points may be based on the statistical analysis, filtering (e.g., Kalman), deep learning or machine learning algorithms, and / or a trained Al-model which is applied to the multiple collected depth points. The parameters of the raster scan may thus be vary based on data derived from the optical system of the stereo depth camera.

[0035] Preferably, the stereo depth camera further includes a raster scan controller adapted to control the measurement position of the light beams. The raster scan controller enables the actively steering of a light beam to a specific measurement position in a surrounding.

[0036] Preferably, the collection of multiple depth points from the object at the measurement position during a predetermined time interval is determined by a feedback loop of the pixel depth values in which a predetermined number of readout values is applied or where the number of depth values to be applied is determined in real-time based on a statistical analysis of the likelihood of the quality of depth values. In other words, the collection of multiple depth points may be based on a predetermined number of readout values or the number of readout values may depend on the quality of the measured depth values. In cases with a low or vanishing contamination, a low number of readout values may be sufficient, while in cases with high contamination, the number of readout values may have to be increased to improve the quality of the pixel depth value.

[0037] Specifically, repeated illumination of a pixel and / or illumination trajectory may be determined by a feedback loop of the pixel depth values. A predetermined number of readout values or scans can be applied to develop a histogram or a body of depth data for an analysis of pixels in at least one row. Another implementation may be to employ a feedback loop where the number of depth values is not predetermined but determined in real-time based on a statistical analysis of the likelihood of the quality of depth values.

[0038] In lieu of a predetermined number of sweeps on an individual row, the number of pixels on a row, or the number of rows scanned, the present invention may thus use a feedback loop that adjusts these parameters in real-time based on a statistical analysis of the depth values. The range of the depth camera, the contaminant in the field of view (dust, wind, snow, fog, rain, metal shavings, grain, etc.), the density and / or size of the contaminants, ambient lighting, etc. may all impact the quality of the depth data. Integrating feedback with real-time control can improve the ability to capture higher quality data across a field of view, albeit with a lower spatial resolution or number of pixels with effective depth data.

[0039] Preferably, pixel depth data from neighbouring pixels (i.e., pixels around a pixel related to a specific measurement position) is applied into filtering algorithms to identify depth outliers, smooth depth data, filter depth data, and / or create intra-pixel data. In contrast to determining pixel depth data individually per pixel, the additional consideration of pixel depth data also from neighbouring pixels may allow to improve the statistical analysis of depth values and in particular to increase and estimate the quality of depth values. Furthermore, by additionally considering pixel depth data also from neighbouring pixels, the influence of rapidly moving contaminants can be estimated over a larger object space volume. Identify depth outliers, smooth depth data, filter depth data, and / or create intra-pixel data further increases the reliability of the depth sensing.

[0040] Preferably, pixel depth data from an individual pixel (e.g., a pixel related to a specific measurement position) is collected over a period of time and applied into filtering algorithms to identify depth outliers, smooth depth data, filter depth data, and / or create intra-pixel data. In contrast to the previous embodiment, pixel depth data acquired over a certain time interval (temporal development of pixel depth values) is used for the filtering algorithms per pixel.

[0041] Preferably, filters are applied to the detection of light beams reflected by an object point at the measurement position. For, example, the filters may comprise region of interest (ROI) areas on the detectors, software filters employing spatial and temporal return filters, etc.

[0042] In contaminated environments, the presence of dust, organic particulates, rain drops, etc. may alter or block either the outbound or the inbound path of a light beam (e.g., laser pulse) or both. As discussed above, the use of multiple light beams (e.g., laser pulses), in time, can increase the probability of an uninterrupted light beam on both outbound and inbound paths. The use of multiple illumination sources (two or more), with different positional vectors to the measurement position (towards an object point or target), but pointing at a common object point in object space, may increase the probability of an uninterrupted light beam.

[0043] In contrast to the use of multiple illumination sources, in standard operation, that traverse different paths in object space simultaneously (increasing the coverage time), the light beams may all follow the same path in object space. During operation, when particulates, dust, rain, etc. appear, the light paths can be changed dynamically from independent trajectory to a coincident path. Additional light sources could be resident on a platform and only used when the object space becomes contaminated, much like the use of fog lights on an automobile.

[0044] In the time domain, the light beams may traverse across object space in a coincident manner. As the positioning of the light beams on the detectors are unique, each illumination source may follow a unique control algorithm to ensure the light beams, from the multiple illumination sources, follow the same path in both time and object space.

[0045] Preferably, the scan pattern on the detectors employ alternative patterns (such as Lissajous, raster or spiral) or at least one region of interest (ROI) based on environmental conditions or scene information. For example, in the case of real-time tracking of objects in the scene, an ROI or bounding of an object may be adjusted based on the size of the object, speed of the object, and / or trajectory of the object of interest that moves within the FOV. In cases where more than one illumination source is used, individual laser sources may employ different patterns to optimize gathering of key scene information.

[0046] Further preferred embodiments of the invention result from features mentioned in the dependent claims. The various embodiments and aspects of the invention mentioned in this application can be combined with each other to advantage, unless otherwise specified in the particular case.

[0047] Brief Description of the Drawings

[0048] In the following, the invention will be described in further detail by figures. The examples given are adapted to describe the invention. The figures show:

[0049] Fig. 1 a schematic illustration of a first embodiment of a stereo depth camera according to the present invention;

[0050] Fig. 2 a schematic illustration of a second embodiment of a stereo depth camera according to the present invention;

[0051] Fig. 3 a schematic illustration demonstrating how dust or other contaminants may obscure scene information of a stereo depth camera;

[0052] Fig. 4 schematic illustrations of outcomes for a transmitted light beam in an environment with contaminants in the path to an object point;

[0053] Fig. 5 schematic illustrations of capturing repeated pulses in a contaminated environment;

[0054] Fig. 6 a schematic illustration of multiple light beams having coincident measurement positions;

[0055] Fig. 7 an example of histogram plots according to a detector comprising nine pixels;

[0056] Fig. 8 a schematic illustration for improved depth estimation by a consolidation of depth values from neighbouring pixels;

[0057] Fig. 9 a schematic illustration of an embodiment of an illumination unit comprising a polygon scanner according to the present invention;

[0058] Fig. 10 an example of row capture in a raster scan according to the present invention;

[0059] Fig. 11 an example of filtering depth by combining information from the depth data distributions of the neighbouring pixels;

[0060] Fig.12 a schematic illustration of an illumination unit comprising three illumination source configured to provide three light beams having coincident measurement positions; and

[0061] Fig. 13a schematic illustration of a scan pattern on the detectors based on environmental conditions or scene information.

[0062] Detailed Description of the Invention

[0063] Figure 1 shows a schematic illustration of a first embodiment of a stereo depth camera 100 according to the present invention. The stereo depth camera 100 comprises an illumination unit 10 including a first illumination source 12-1 for emitting a first light beam, wherein the illumination unit 10 is configured for actively steering the first light beam to a first measurement position P in a surrounding of the stereo depth camera 100 within the field of view of the illumination unit 10; an imaging unit 20 including at least two detectors 22-1 , 22-2 for detecting the first light beam reflected by a first object point at the first measurement position P, wherein the detectors 22-1 , 22-2 have different alignments and are configured to image a common field of view to enable stereoscopic imaging of the first object point. The detectors 22-1 , 22-2 include multiple pixels S1 , S2 that are configured as event sensors sharing a predefined pixel threshold trigger setting and which are adapted to collect multiple depth points from the first object point at the first measurement position P during a predetermined time interval At, wherein statistical analysis, filters (e.g., Kalman), deep learning or machine learning algorithms, and / or a trained Al-model are applied to the collected depth points for determining and outputting the most probable distance of the first object point as pixel depth value for the first measurement position P.

[0064] Multiple depth points from the object point at the first measurement position P are therefore collected during a predetermined time interval At. Since the detectors 22-1 , 22-2 are configured to image a common field of view to enable stereoscopic imaging of the object point, there is a direct relation between the position of object points in the surrounding (object position) and the respective positions of the image of the object points on the pixels S1 , S2 of the detectors 22- 1 , 22-2 (image pixel position). This relates to a kind of “spherical” imaging in two-dimensions (X, Y). The depth of an object point is related to the third dimension (Z) and can be determined by the stereoscopic imaging of the object point for a specific time (T).

[0065] In the present invention, multiple depth points are collected from an object point and thus from single pixels S1 , S2 such that statistical analysis, deep learning or machine learning algorithms, and / or a trained Al-model can be applied to the collected depth points for determining and outputting the most probable distance of the object point as pixel depth value for the first measurement position P. For matching the triggered pixels of the at least two detectors 22-1 , 22-2 that are related to the same object point, event timestamps, laser trajectory analysis or beam progression analysis may be applied.

[0066] Figure 2 shows a schematic illustration of a second embodiment of a stereo depth camera 100 according to the present invention. In this embodiment, the illumination unit 10 further includes a second illumination source 12-2 emitting a second light beam independent of and separable from the first light beam and the illumination unit 10 is configured for actively steering the second light beam to a second measurement position P’ different from the first measurement position P of the first light beam. Thus, a “stereo” illumination unit 10 is provided which allows to illuminated two separate object points at different measurement positions P, P’ simultaneously. However, this is only an exemplarily embodiment and three or more illumination source can be used to illuminated multiple object points at different measurement positions simultaneously.

[0067] In agreement with the embodiment of FIG. 1 , all detectors 22-1 , 22-2 may be adapted to collect multiple depth points from a second object point at the second measurement position P’ during the predetermined time interval At’ (which may differ from or coincide with the predetermined time interval At) and wherein statistical analysis, filters (e.g., Kalman), deep learning or machine learning algorithms, and / or a trained Al-model are applied to the collected multiple depth points for determining and outputting the most probable distance of the second object point as pixel depth value for the second measurement position P’. By using a “stereo” illumination source 10 including at least two independent illumination sources 12-1 , 12-2, the frame rate of the stereo depth camera 100 can be increased due to parallel processing (increased depth measuring speed). However, the different light beams which are emitted into the surrounding may have to be distinguished from one another. This may be achieved by imposing an individual modulation to the different light beams, for example, using pulse width, amplitude, frequency or polarization modulation schemes. Preferably, the wavelength of the first light beam and the second light beam are different.

[0068] The second embodiment is further distinguished from the embodiment of FIG.1 in that the imaging unit 20 comprises a third detector 22-3 include multiple pixels S3, S3’ that are configured as event sensors sharing a predefined pixel threshold trigger setting (may be the same as for the corresponding pixels S1 , ST and S2, S2’) and which are adapted to collect multiple depth points from the object point at the first measurement position P and the second measurement position P’ during a predetermined time interval At” (which may differ from or coincide with the other predetermined time intervals At, At’), respectively. However, the number of detectors is not limited and an integer number n of such pixelated array detectors may be used in parallel. Increasing the number of detectors allows to observe an image point under various imagining conditions. This expanses the database usable for determining and outputting the most probable distance of the object point by statistical analysis, deep learning or machine learning algorithms, and / or a trained Al-model, which may be performed by or running on a suitable processor.

[0069] The present invention thus uses stereoscopic imaging by two or more two-dimensional detectors having individual pixels for selectively determing the most probable distance of an object point as resulting pixel depth value by applying advanced statistical methods based on collecting multiple depth points from the object point under a number of different imaging conditions. For increasing the speed of depth measurement, also the number of illumination sources may be increased for parallel processing. The illumination sources can also be used for increasing the reliability of the pixel depth value determination.

[0070] Figure 3 shows a schematic illustration demonstrating how dust or other contaminants C may obscure scene information of a stereo depth camera 100. The figure highlights the problem of contamination impeding the light beam or the field of view of the detectors. The field of contaminants C can reflect, block, and / or impede both a transmit pulse and / or an reflected return pulse. In the present case, a dust cloud is located between the stereo depth camera 100 and an object point at the corresponding measurement position P (e.g., a head of a human). Depending on the degree of contamination, strong time-dependent fluctuations in the pixel depth value may be observed or the object point may be invisible for the stereo depth camera 100 at certain times. Such situations are typical examples of a contained environment in which conventional depth sensing approaches become highly unreliable.

[0071] Figure 4 shows schematic illustrations of outcomes for a transmitted light beam in an environment with contaminants C in the path to an object point. The different scenarios are similar to the situation shown in FIG. 3. Specifically, the figure highlights three outcomes for a transmitted laser pulse (i.e., a pulsed laser beam). The ideal scenario is a transmitted pulse that is reflected from an object of interest (i.e., from an object point) and returned to the detector unobstructed as shown under Scenario A. Scenario B illustrates a transmitted pulse reflected from an object point but blocked or diverted from its return path to the detector. Scenario C is a common failure mode typically experienced by time-of-flight systems. The pulse is reflected by a contaminant C resulting in a closer and errant distance value.

[0072] With a single reading (i.e., acquiring only a single depth point from an object point at a measurement position P), scenarios B and C become more pronounced in dense snow, hail, and dusty environments. For increasing the probability of achieving Scenario A, the present invention captures multiple readings (i.e., collects multiple depth points from an object point at a measurement position P during a predetermined time interval At) at a pixel and applies a statistics to determine the most probable distance value for the corresponding object point.

[0073] Figure 5 shows schematic illustrations of capturing repeated pulses in a contaminated environment. The figure provides an example of benefits of capturing multiple data points for a pixel (i.e., collecting multiple depth points from the object point at a measurement position P during a predetermined time interval At). Over time, contaminants C move through the space between the stereo depth camera 100 and the object point and obscure the illumination source. By capturing multiple depth readings at a pixel, the likelihood of capturing the distances of objects of interest within the scene highly increases.

[0074] Figure 6 shows a schematic illustration of multiple light beams having coincident measurement positions. Two separate illumination sources of a stereo depth camera 100 are aligned such that the light beams can be coincidently directed to an object point at a common measurement position P. As the range increases, the spatial volume of the pixel in space increases. Using multiple illumination sources pointing in the volume can increase the amount of acquired data. By adding a second illumination source, additional data is added in both the temporal domain and the spatial domain with respect to the pixel.

[0075] Figure 7 shows an example of histogram plots according to a detector comprising nine pixels. For determining the depth of an object of interest, the pixel data can be input into a statistical model for determining the most probable distance Z of an object point or the data sets can be input into deep learning or machine learning algorithms, and / or an Al-model trained to determine the most probable distance. Specifically, the probability W for a certain depth value Z may be determined per pixel. For providing a reliable database from which an informative histogram can be derived, multiple depth points from the object point are collected during a predetermined time interval At (i.e., the number of triggered events in said time interval At should be sufficient to allow statistical methods to be applied with some certainty).

[0076] Figure 8 shows a schematic illustration for improved depth estimation by a consolidation of depth values from neighbouring pixels. In particular, pixel depth data from neighbouring pixels may be applied into filtering algorithms to identify depth outliers, smooth depth data, filter depth data, and / or create intra-pixel data. If contaminants C in the field of view obstruct the collection of accurate pixel depth data, pixel depth data from neighbouring pixels may be applied into filtering algorithms to improve depth estimation. For example, in the 2D detector example of FIG. 7, the centre pixel may be analysed by using also data from its neighbouring eight pixels for an extended statistical analysis. Related algorithms may use row data, column data, or other array sizes. Different methods of statistical analysis may be applied to the depth values based on environmental factors such as nature of contamination, density, etc.

[0077] Figure 9 shows a schematic illustration of an embodiment of an illumination unit comprising a polygon scanner 13 according to the present invention. The imaging unit 20 comprises a first illumination source 12-1 and one polygon scanner 13 for 2D depth capture in a first direction in combination with another beam steering technology 14 (e.g., a “galvo” scanner) for 3D depth capture in a second direction different from the first direction. For collimating the light beam, a projection lens (e.g., an F-theta lens) may be used.

[0078] For 3D depth capture, the measurement position P, P’ may be scanned in the first direction (e.g., the horizontal direction) by the at least one polygon scanner 13 to collect multiple depth points from the object point for a predetermined time interval (e.g., At, At’, At”, AT) before the measurement position P, P’ is varied in the second direction (e.g., the vertical direction) to implement a raster scan pattern on the various detectors. The collection of multiple depth points for one pixel is in this case a compound time interval having a total length according to the respective predetermined time interval. The shown stereo depth camera 100 further includes a raster scan controller 16 to control the measurement positions P, P’ of the light beams. Preferably, the raster scan pattern for the detectors are adapted to parameters of the stereo depth camera 100 and / or environmental conditions. Furthermore, the parameters of the raster scan may be varied in real-time based on an analysis of the depth points.

[0079] In contrast to prior art based on time-of-flight technologies, the present invention is based on using an active illumination source with event-based stereo cameras (i.e., at least two detectors including pixels configured as event sensors) to determine depth values. Stereo imaging is the foundation of determining depth in contrast to the time-of-flight measurement approach of pulsed light. Furthermore, the present invention can combine an innovative fast scanning solution using polygon scanners for 2D depth capture (e.g., depth measurement only in horizontal directions) or polygon scanners in combination with another beam steering technology for 3D capture (e.g., depth measurement in horizontal and vertical directions). The another beam steering technology may be based on a MEMS scanner, galvo, gimbal, metasurface, OPA, or other.

[0080] In the prior art, the use of Lissajous patterns in capturing depth measurements with event sensors is known. In contrast, the present invention may apply a raster scan pattern based on a polygon scanner for beam steering. Polygon scanners, rotating at extremely high speeds, with repeatable precision, allows for the capture of extremely fast depth measurements within a row (i.e., a row in the field of view or of the pixels of the detectors). For example, a corresponding row capture may start in the middle of a vertical pixel array and is not required to start at the top or bottom of the pixel array in the raster scan pattern. Additionally, raster scanning is not limited to sequential rows, but can skip to any row as determined by the system and environmental conditions.

[0081] It is clear for the skilled person that multiple illumination sources can be used simultaneously with a single or multiple polygon scanners of a illumination unit 10 according to the present invention. Related to the field of view, scanning may be performed with one illumination source working from the middle of and moving vertically upward in the field of view and an additional illumination source may be scanned from the middle of and moving vertically downward in the field of view. The orientation of the raster scanning may also be implemented with the primary axis being in the vertical direction. Therefore, the polygon scanner may be oriented in the vertical direction and the another beam steering technology may move the light beam in the horizonal direction.

[0082] Figure 10 shows an example of row capture in a raster scan according to the present invention. Multiple passes or ‘n’ sweeps of the bottom row of the detector may be accomplished through the use of a polygon scanner 13. After ‘n’ sweeps, the another beam steering technology 14 may be varied in the vertical direction, and the process is repeated at the second row. The numbers of sweeps and indexing of rows (sequential, skipping rows, etc.) may be adaptively varied depending on specific parameters.

[0083] Figure 11 shows an example of filtering depth by collecting depth data from a pixel over a period of time. In some environments, contaminants C in the field of view may obstruct the collection of accurate pixel data. However, the concentration of contaminants C, direction of travel, speed of travel, etc. may change over time. For overcoming this issue, pixel depth data from an individual pixel collected over a period time can be applied into several filtering algorithms to identify depth outliers, smooth depth data, filter depth data, create intra-pixel data, etc. The figure highlights the collection of pixel data over time. The period of time should be long enough to demonstrate the temporal development of the pixel depth value of an respective pixel. The period of time shown in the figure is much larger than the predetermined time interval used to generate the individual histograms in each of the shown time frames. According to the figure, the shown period of time may be 12-times the predetermined time interval.

[0084] Figure 12 shows a schematic illustration of an illumination unit 10 comprising three different illumination sources configured to provide three independent light beams having coincident measurement positions. In particular, the illumination unit 10 may additionally include a second and a third illumination source 12-2, 12-3 emitting light beams independent of and separable from one another and the illumination unit 10 may be configured for actively steering all light beams to a common measurement position (e.g., a contaminant C). An individual modulation may be imposed to the different light beams to distinguish the different light beams from one another, for example, using pulse width, amplitude, frequency or polarization modulation schemes. Preferably, the wavelength of the first light beam and the second light beam are different.

[0085] Figure 13 shows a schematic illustration of a scan pattern on the detectors based on environmental conditions or scene information. In this example, the scan pattern on the detectors employs a region of interest (ROI) based on actual scene information. Specifically, a ROI may be defined by one or more objects of interest (in this case a single person) that moves within the FOV. For example, in the case of real-time tracking of objects in the scene, the ROI or bounding of an object may be adjusted based on the size of the object, speed of the object, and / or trajectory of the object of interest that moves within the FOV. In cases where more than one illumination source (only a first illumination source 12-1 is exemplarily shown in the illustration) is used, individual laser sources may employ different patterns to optimize gathering of key scene information.

[0086] In the shown exemplary embodiment, raster scanning may first be used to identify a person walking through the FOV of the camera as shown under a). After identifying the person, an alternative scanning pattern different from raster scanning may be applied. In particular, as shown under b), the illumination pattern may be changed to a ROI that follows the identified person. However, while the person is moving within the FOV, the position of the ROI has to be adapted in real-time. This can be done, for example, by applying known tracking algorithms, path estimation and prediction methods, and / or known or received object data. The ROI may then be moved through the FOV to follow the person (see under c)). When the followed object has finally moved outside the FOV and cannot be tracked anymore, the employed illumination pattern may automatically return to a standard raster scan as it is shown under d). However, a ROI is only an example for an alternative scan pattern. Based on the environmental conditions or scene information other scan patterns may be employed.

[0087] Reference List

[0088] 10 illumination unit

[0089] 12-1 first illumination source

[0090] 12-2 second illumination source

[0091] 12-3 third illumination source

[0092] 13 polygon scanner

[0093] 14 another beam steering technology

[0094] 15 projection lens (e.g., F-theta lens)

[0095] 16 raster scan controller

[0096] 20 imaging unit

[0097] 22-1 first detector

[0098] 22-2 second detector

[0099] 22-3 third detector

[0100] 100 stereo depth camera

[0101] 51 , Si’ pixels of the first detector

[0102] 52, S2’ pixels of the second detector

[0103] 53, S3’ pixels of the third detector

[0104] At, At’, At” time intervals

[0105] AT common time interval

Claims

1. Claims1. Stereo depth camera (100), comprising: an illumination unit (10) including a first illumination source (12-1) for emitting a first light beam, wherein the illumination unit (10) is configured for actively steering the first light beam to a first measurement position (P) in a surrounding of the stereo depth camera (100) within the field of view of the illumination unit (10); an imaging unit (20) including at least two detectors (22-1 , 22-2) for detecting the first light beam reflected by a first object point at the first measurement position (P), wherein the detectors (22-1 , 22-2) have different alignments and are configured to image a common field of view to enable stereoscopic imaging of the first object point; characterized in that the detectors (22-1 , 22-2) include multiple pixels (Si, S2) that are configured as event sensors sharing a predefined pixel threshold trigger setting and which are adapted to collect multiple depth points from the first object point at the first measurement position (P) during a predetermined time interval At, wherein statistical analysis, filters, deep learning or machine learning algorithms, and / or a trained Al-model are applied to the collected depth points for determining and outputting the most probable distance of the first object point as pixel depth value for the first measurement position (P).

2. Stereo depth camera (100) of claim 1 , wherein the illumination unit (10) includes a second illumination source (12-2) emitting a second light beam independent of and separable from the first light beam and the illumination unit (10) is configured for actively steering the second light beam to a second measurement position (P’) different from the first measurement position (P) of the first light beam.

3. Stereo depth camera (100) of claim 2, wherein the detectors (22-1 , 22-2) are adapted to collect multiple depth points from a second object point at the second measurement position (P’) during a predetermined time interval At’ and wherein statistical analysis, filters, deep learning or machine learning algorithms, and / or a trained Al-model are applied to the collected multiple depth points for determining and outputting the most probable distanceof the second object point at the second measurement position (P’) as pixel depth value for the second measurement position (P’).

4. Stereo depth camera (100) of claim 1 , wherein the illumination unit (10) includes a second illumination source (12-2) emitting a second light beam independent of and separable from the first light beam and the illumination unit (10) is configured for actively steering the second light beam to a second measurement position (P’) coincident with the first measurement position (P).

5. Stereo depth camera (100) of any one of claims 2 to 4, wherein the wavelength of the first light beam and the second light beam are different.

6. Stereo depth camera (100) of any one of the preceding claims, wherein triggered pixels of the at least two detectors (22-1 , 22-2) are matched by related event timestamps, laser trajectory analysis or beam progression analysis.

7. Stereo depth camera (100) of any one of the preceding claims, wherein the imaging unit (20) comprises at least one polygon scanner (13) for 2D depth capture in a first direction, or at least one polygon scanner (13) for 2D depth capture in a first direction in combination with another beam steering technology (14) for 3D depth capture in a second direction different from the first direction.

8. Stereo depth camera (100) of claim 7, wherein for 3D depth capture, the measurement position (P, P’) is scanned in the first direction by the at least one polygon scanner (13) to collect multiple depth points from the object point for a predetermined time interval before the measurement position (P, P’) is varied in a second direction to implement a raster scan pattern on the detectors (22-1 , 22-2).

9. Stereo depth camera (100) of claim 8, wherein the raster scan pattern on the detectors (22-1 , 22-2) are adapted to parameters of the stereo depth camera (100) and / or environmental conditions.

10. Stereo depth camera (100) of any one of claims 8 or 9, wherein the parameters of the raster scan are varied in real-time based on an analysis of the depth points.11 . Stereo depth camera (100) of claim 8 to 10, wherein the stereo depth camera (100) further includes a raster scan controller (16) adapted to control the measurement position (P, P’) of the light beams.

12. Stereo depth camera (100) of any one of the preceding claims, wherein the collection of multiple depth points from the object point at the measurement position (P, P’) during a predetermined time interval is determined by a feedback loop of the pixel depth values in which a predetermined number of readout values is applied, or wherein the number of depth values to be applied is determined in real-time based on a statistical analysis of the likelihood of the quality of depth values.

13. Stereo depth camera (100) of any one of the preceding claims, wherein pixel depth data from neighbouring pixels is applied into filtering algorithms to identify depth outliers, smooth depth data, filter depth data, and / or create intra-pixel data.

14. Stereo depth camera(100) of any one of the preceding claims, wherein pixel depth data from an individual pixel (Si , S2) is collected over a period of time and applied into filtering algorithms to identify depth outliers, smooth depth data, filter depth data, and / or create intra-pixel data.

15. Stereo depth camera (100) of any one of the preceding claims, wherein filters are applied to the detection of light beams reflected by the object point at the measurement position (P, P’).