In-cabin monitoring circuitry and in-cabin monitoring method
A machine-learned model transforms 2D infrared data into depth data, addressing the challenges of monocular depth estimation in vehicle cabins by training on cabin-specific geometry and illumination cues, ensuring accurate depth measurements for safety and spatial audio applications.
Patent Information
- Application Number
- PCT/EP2025/051669
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-30
- Filing Date
- 2025-01-23
- Publication Date
- 2025-08-07
AI Technical Summary
Existing monocular depth estimation techniques struggle to provide reliable depth measurements at a millimeter scale in diverse in-cabin environments due to varying cabin geometries, lighting conditions, and sensor parameters, making it challenging to accurately monitor occupant safety and spatial audio in vehicles.
Utilizing a machine-learned model that transforms 2D infrared sensor data into depth data by training on cabin-specific geometry and illumination cues, leveraging a convolutional neural network or Vision Transformer, and incorporating time-of-flight data as ground truth to generate monocular depth maps.
Enables accurate and reliable millimeter-scale depth estimation in vehicle cabins, supporting occupant safety monitoring, seat dimensioning, and spatial audio calibration by compensating for diverse cabin conditions and lighting variations.
Smart Images

Figure EP2025051669_07082025_PF_FP_ABST
Abstract
Description
[0001] IN-CABIN MONITORING CIRCUITRY AND IN-CABIN MONITORING
[0002] METHOD
[0003] TECHNICAL FIELD
[0004] The present disclosure generally pertains to in-cabin monitoring circuitry and an in-cabin monitoring method for monitoring an in-cabin of a vehicle.
[0005] TECHNICAL BACKGROUND
[0006] Generally, monocular depth estimation from large-scale RGB (red, green, blue) datasets and / or in generic indoor or outdoor environments is known. Such estimation may be generated by an artificial intelligence (Al) trained on a mix of synthetic or real datasets, for example.
[0007] It is also known to determine a depth of a scene. For example, time-of-flight (ToF) systems are generally known. In ToF, a depth of a scene or a distance to the scene (e.g., object) is determined based on a roundtrip delay of emitted light. The delay may be directly measured by measuring a time from emission to reception, which is known as dToF (direct ToF). The delay may be indirectly measured by measuring correlating a phase of the emitted (modulated) light with the received (reflected) light, which is known as iToF (indirect ToF).
[0008] Although there exist techniques for monocular depth estimation, it is generally desirable to provide in-cabin monitoring circuitry and an in-cabin monitoring method.
[0009] SUMMARY
[0010] According to a first aspect, the disclosure provides in-cabin monitoring circuitry for monitoring an in-cabin environment of a vehicle, the circuitry being configured to: obtain infrared sensor data; and generate monocular depth data with a machine-learned model, wherein the machine-learned model has depth data as ground truth and uses the obtained infrared sensor data as input.
[0011] According to a second aspect, the disclosure provides an in-cabin monitoring method for monitoring an in-cabin environment of a vehicle, the method comprising: obtaining infrared sensor data; and generating monocular depth data with a machine-learned model, wherein the machine- learned model has depth data as ground truth and uses the obtained infrared sensor data as input.
[0012] Further aspects are set forth in the dependent claims, the drawings and the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Embodiments are explained by way of example with respect to the accompanying drawings, in which:
[0014] Fig. la depicts an embodiment of in-cabin monitoring circuitry using single-exposure IR;
[0015] Fig. lb depicts an embodiment of in-cabin monitoring circuitry using multi-exposure IR;
[0016] Fig. 2 depicts a data collection procedure and evaluation framework according to the present disclosure;
[0017] Fig. 3 depicts obtained IR images which are transformed into depth data according to the present disclosure;
[0018] Fig. 4 depicts an embodiment of a method for training a CNN, which is fed with IR data at its input, and at its output with a pre-calibrated depth map;
[0019] Fig. 5 depicts an embodiment of a method for training a CNN, which is fed with IR data and a precalibrated depth map at its input.
[0020] Fig. 6 depicts an embodiment of a method for training a CNN, which is fed with IR data, precalibrated camera rays, and pre-calibrated cabin keypoints at its input;
[0021] Fig. 7 depicts an embodiment of a method for training a CNN based on a differentiable Tenderer using pre-calibrated camera rays and illumination;
[0022] Fig. 8 depicts an embodiment of a method for in-cabin environment monitoring according to the present disclosure;
[0023] Fig. 9 depicts an embodiment of a method for in-cabin monitoring environment according to the present disclosure, wherein RGB data is additionally used;
[0024] Fig. 10 depicts an embodiment of a training method according to the present disclosure;
[0025] Fig. 11 depicts a server for providing data that is used according to the present disclosure;
[0026] Fig. 12 depicts an occupant status monitoring system according to the present disclosure;
[0027] Fig. 13 depicts a schematic block diagram of a method for generating body feature information according to the present disclosure;
[0028] Fig. 14 depicts an illustrational result of a method according to Fig. 13;
[0029] Fig. 15 depicts a skeleton in a filtered depth image and the result of a segmentation in a confidence image; and Fig. 16 depicts a depth image and a confidence image depicting points on an occupant’s body.
[0030] DETAILED DESCRIPTION OF EMBODIMENTS
[0031] Before a detailed description of the embodiments starting with Fig. 1 is given, general explanations are made.
[0032] As mentioned in the outset, monocular depth estimation is generally known. However, known setups of such datasets are rather generic or for other environments (e.g., outdoor) and thus, reliable depth (at scale of millimeters or meters) can often not be obtained based on monocular depth estimation by training with such datasets. Reasons for this may be changes of scene geometry, lens and sensor parameters (e.g., camera intrinsics), illumination conditions (e.g., shading, illumination intensity) .
[0033] Hence, it has been recognized that reliable depth at millimeter scale may be obtained by exploiting a controlled environment, such as in an in-cabin environment sensing setup under one or more active infrared (IR) illuminators or exposures captured by a 2D (two-dimensional) infrared (IR) sensor.
[0034] It has been recognized that dense volumetric perception (e.g., millimeter scale depth) may be a key to ensure reliable automotive in-cabin safety functions, e.g., monitoring of driver / occupant size, posture, behavior, and the like. For example, relative distances between a driver and one or more incabin points (e.g., steering wheel) may be determined for passive safety monitoring, e.g., for crash protection in vehicles with corresponding adaptive restraint systems. It should be noted that the present disclosure is not limited to the use-case of vehicle safety as it may be used in any case in which volumetric data is important. For example, spatial in-cabin audio may be such a use case. In such a case, it may be important to know where exactly an occupant’s ears are in order to calibrate / beamform speaker’s output onto the user’s ears. The present disclosure may also be used in the context of functions such as seat dimensioning / tuning, occupant volume and weight estimation, or the like.
[0035] It has been recognized that, instead of using depth maps / point clouds produced by active depth sensors with complex calibration procedures (e.g., stereo camera, structured light, ToF), 2D IR for obtaining an image of the in-cabin environment (wherein, generally, the present disclosure is not limited to that case) and a trained machine-learning model that transforms the 2D IR image into depth data may be used. In the context of the present disclosure, such an approach may be referred to as “monocular depth estimation” or “in-cabin depth estimation” (ICDE). In the context of the present invention, ICDE may refer to an estimation of a depth (e.g., metric, such as meters or millimeters, at scale) from a single (monocular) IR sensor (e.g., camera) and a single (optional) active illuminator placed in a static position inside a vehicle, e.g., for monitoring front seat(s), back seat(s), or both (e.g., placed on a rear-view mirror with a field of view covering the respective seat(s), and / or placed on a side-mirror for monitoring a driver, or the like) .
[0036] This is made possible by learning cabin geometry and illumination cues during a training of a machine-learning model inside the cabin, e.g., based on a CNN (convolutional neural network), a Vision Transformer (ViT) model, or the like.
[0037] In ICDE, it has been recognized that it may be challenging to map input 2D frames to a depth since no explicit geometry may be inferred from the 2D frames. Moreover, there may be diverse cabin geometries and a model may not be able to generalize to different cabins, for example having different dimensions or layouts, such that retraining may be required. Also, diverse cabin content may be present, such as different occupants at different places, objects of different sizes, different textures or different reflectivity garments (which may be more or less challenging to distinguish based on the used sensor), retroreflective elements, and the like. Furthermore, diverse camera extrinsics and intrinsics may be present (e.g., positioning and geometry may vary depending on camera lens and cabin design or layout). Also, different illumination conditions may be present, such as low or high ambient light, active or passive IR illumination, or the like. All these conditions may be challenging to train for, but it has been recognized that the present disclosure may compensate for such conditions and more.
[0038] Therefore, some embodiments pertain to in-cabin monitoring circuitry for monitoring an in-cabin environment of a vehicle, the circuitry being configured to: obtain infrared sensor data; and generate monocular depth data with a machine-learned model, wherein the machine-learned model has depth data as ground truth and uses the obtained infrared sensor data as input.
[0039] Circuitry may include any entity or multitude of entities which are configurable to carry generate monocular depth data according to the present disclosure, such as a CPU (central processing unit), GPU (graphics processing unit), and FPGA (field-programmable gate array), or the like. The circuitry may be (partly or fully) included in a camera, (partly or fully) included in a vehicle, or it may communicate with a camera and / or a vehicle. Also, the circuitry may be included in a mobile phone, or the like.
[0040] In some embodiments, training of the machine-learned method may (partially or fully) be carried out at a camera manufacturer, a car manufacturer, a phone manufacturer, or the like. In some embodiments, training may (partially or fully) be carried out in a mobile phone of a driver or an occupant of a vehicle. In some embodiments, training may (partially or fully) be carried out on onboard computational resources of a vehicle. As will be described further below, depth data (such as time-of-flight (ToF) data) may be obtained as ground truth (i.e., dense depth ground truth data) for the machine-learned model by a ToF camera, such as a ToF camera of a manufacturer (of a vehicle, portable device such as phone handset, head mounted display, camera, or the like), but there are also embodiments in which the ground truth is collected by a ToF camera of a user’s portable device, such as a mobile phone. For example, there may be a predetermined position in a vehicle in which the user may put his or her portable device and an integrated ToF sensor may collect the ground truth data for the training. Also, arbitrary positions of the portable may be envisaged in such embodiments. In some embodiments as head mounted display may co-operate with the portable device to relay ground truth. It should be noted that, in some embodiments, other (dense) depth data may be used as ground truth, instead ToF data or in combination with ToF data. The depth data (e.g., ToF data) may be used as ground truth depth maps at a training state, for example.
[0041] In some embodiments, the circuitry is configured to obtain infrared (IR) sensor data. The IR sensor may include a camera, a monochrome (i.e., one-band) pixel array, a polychromatic pixel array (e.g., RGB-infrared), or the like. IR sensor data may refer to any type of infrared sensor data, such as near-infrared (NIR), super-wide infrared (SWIR, e.g., 1 pm to 2.5 pm), or any other spectral band that includes or is related to infrared.
[0042] Based on the IR sensor data, monocular depth data are generated, in some embodiments. Since the IR sensor may only be a single camera, the generated depth data may be referred to as “monocular”.
[0043] In some embodiments, the monocular depth data is generated by mapping via a neural network the 2D infrared sensor data to a depth map of the in-cabin environment, where the neural network is trained with the depth (map) data as ground truth. Mapping may refer to assigning parameters to a function of a neural network.
[0044] In some embodiments, the infrared sensor data is obtained based on an active illumination. For example, light may be emitted (such as structured light, modulated light, or the like) by a light source and the IR sensor may detect a reflection of the emitted light, which may be the basis of the IR data. In some embodiments, the active illumination may be synchronized with the IR sensor shutter to provide improved signal-to-noise ratio, and to capture one or more frames under different exposure conditions, from high active illumination intensity to no active illumination,
[0045] In some embodiments, the active illumination is an infrared illumination.
[0046] In some embodiments, the machine-learned algorithm is based on a convolutional neural network. However, the present disclosure is not limited to that case as also other types of neural networks may be envisaged, such as a deep neural network (DNN), a vision transformer network (ViT), or the like. In some embodiments, the in-cabin monitoring circuitry is further configured to: obtain RGB sensor data as an auxiliary input to the machine-learned model.
[0047] The RGB sensor data may be aligned to IR data by co-located or stacked sensor designs, and / or by using a color filter array, in some embodiments. The RGB sensor data may be combined with the IR sensor data, e.g., in order to improve the IR data. For example, with an RGB frame (single or multiexposure), illumination variations may be compensated, edges may be refined, estimates for low- reflectivity objects may be refined, or the like. Also, an RGB-IR sensor / camera may be envisaged as a data acquisition source. RGB is used herein to generally describe image data. The skilled person will appreciate that image data may for example be RGB data, luminance and / or chrominance data, or other image data formats / types. It should be noted that the present disclosure is not limited to any type of RGB sensor data. Any way of constructing an RGB frame may be envisaged according to the present disclosure, such as mosaics, or the like.
[0048] In some embodiments, the machine-learned model is based on at least one time-of-flight measurement (for obtaining depth data) of the in-cabin environment, as discussed above. It should be noted that the present disclosure is not limited to the case of a time-of-flight measurement since any other technology may be utilized that makes it possible to generate 3D (or depth) data, such as a 3D scanner, photogrammetry, or the like. Accordingly, other data than ToF data or equivalents of ToF data may be used, such as 3D scanner data, photogrammetry data, or the like.
[0049] In some embodiments the machine-learned model is based on synthetic depth data of an in-cabin environment. Thereby, acquisition and / or computation costs of real data may be reduced and potential ground truth training sets may be increased.
[0050] In some embodiments, the machine-learned model is based on an augmentation of real time-of- flight data with synthetic depth data. Thereby, potential ground truth training sets may be increased, e.g., with scenarios that are hard to simulate in a factory, or the like. It should be noted that the synthetic depth data may be synthetic ToF data, but the present disclosure is not limited in that regard.
[0051] For example, a wide range of vehicles, of lighting conditions, of (out-of-distribution) objects, of (dangerous) conditions (e.g., crashes, violent theft or assault, and such harmful situations for the occupants), of contents, of occupants, of textures, of materials (e.g., for seats and / or clothing), dashboards, and the like may be simulated with synthetic or augmented (hybrid) data. It should be noted that, for augmenting the data, first the ground truth may be needed, in some embodiments.
[0052] The augmentation may be based on at least one of the following (without limiting the present disclosure in that regard): Using a segmentation algorithm to compose real occupants from collected data sequences with synthetic cabins in both (RGB+)IR and depth
[0053] Generating assets from real occupants and compositing them with synthetic cabins in both (RGB+)IR and depth
[0054] Generating assets from synthetic occupants and compositing them with real cabins in both (RGB+)IR and depth
[0055] All of the above, for different objects, textures, and materials interchangeably added to synthetic elements in the rendered frames
[0056] In some embodiments the machine-learned model uses illumination information to estimate depth as according to the ground truth depth maps provided during training. This includes leveraging the inverse square law relation of an active illuminator co-located with the IR sensor, by which the IR signal intensity recorded at a photodetector simplifies to: where p is an albedo (for a Lambertian target); ? is a slant angle between the illumination ray and the camera ray; Aois a IR signal amplitude that one may calibrate for a unit-albedo Lambertian target (e.g., a flat white panel) at one meter and with slant angle fl ■= 0, and which depends on the illumination power. In some embodiments, the above equation may be used partially or fully to obtain the depth. In other embodiments, the network may leam to regress from the IR signal intensity part of the above equation. Accordingly, this requires the network to learn different lighting conditions (that, in the above equation, would set the term Ao) and different albedos per in-cabin element.
[0057] Some embodiments pertain to a data collection method to obtain real data for ICDE.
[0058] The method may include performing at least one static measurement. For example, in a static measurement, a scanning device or a 3D scanning technology may be used, such as 3D reconstruction from photogrammetry, laser scanning, neural radiance field, ToF, or the like, thereby obtaining data of the empty cabin as ground truth 3D model (in the form of a point cloud, mesh, neural radiance field, or the like). Optionally, the cabin may be additionally rigged or annotated a posteriori with markers to register the ground truth 3D model points to the camera, and thus obtain camera extrinsics of the IR or RGB-IR camera that is acting as a main data source for the monocular depth estimation in ICDE. The method may include (additionally or alternatively) performing at least one dynamic measurement. For example, frames from one or more synchronized ground truth depth sensor may be streamed along with the IR (or RGB-IR) camera (see also Fig. 2 as a configuration example of the cameras in the cabin). The ground truth depth sensors may include ToF sensors / cameras, passive / active stereo infrared sensors / cameras, passive / active RGB sensors / cameras, structured light sensors / cameras, or the like (or any combination of the sensors / cameras). The sensors may be different for the static measurement than for the dynamic measurement because in a dynamic measurement, a real-time stream may be needed for obtaining for obtaining the ground truth depth and the stream may be synchronized with an IR sensor stream for training the network.
[0059] Some embodiments pertain to an in-cabin monitoring method for monitoring an in-cabin environment of a vehicle, the method including: obtaining infrared sensor data; and generating monocular depth data with a machine-learned model, wherein the machine-learned model has depth data as ground truth and uses the obtained infrared sensor data as input, as discussed herein.
[0060] The method may be carried out by in-cabin monitoring circuitry, as discussed herein.
[0061] In some embodiments, the monocular depth data is generated by mapping the infrared sensor data to a depth map of the in-cabin environment learned based on training with the depth data as ground truth, as discussed herein. In some embodiments, the infrared sensor data is obtained based on an active illumination, as discussed herein. In some embodiments, the active illumination is an infrared illumination, as discussed herein. In some embodiments, the machine-learned model is based on a convolutional neural network, as discussed herein. In some embodiments, the method further includes: obtaining RGB sensor data as an auxiliary input to the machine-learned model, as discussed herein. In some embodiments, the machine-learned model is based on at least one time- of-flight measurement of the in-cabin environment used as ground truth, as discussed herein. In some embodiments, the machine-learned model is based on synthetic IR depth data of an in-cabin environment, as discussed herein. In some embodiments, the machine-learned model is based on an augmentation of real time-of-flight data with synthetic depth data, as discussed herein. In some embodiments, the machine-learned model has illumination cues as ground truth, as discussed herein.
[0062] Some embodiments also pertain to a training method for the machine-learned model.
[0063] The training may optionally include, for example, at least one of the following techniques to improve quality:
[0064] Filtering and segmentation of reliable ground truth depth data with arbitrary complexity models at training time. Additional input from a segmentation model that receives IR images and outputs salient elements in the scene, e.g., only the driver and occupant, to focus the monocular depth estimation model on reconstruction of humans in the scene.
[0065] Providing a calibration frame containing metric depth of the ground truth cabin at the chosen camera extrinsics and intrinsics; the calibration frame can be assumed to be captured once per cabin model, and the cabin model may be updated dynamically during the lifetime of the system.
[0066] Providing fiducial points containing metric depth of ground truth keypoints in the cabin at the chosen camera extrinsics and intrinsics; same considerations as the previous point hold.
[0067] Retraining with parts of the monocular depth estimation model being fixed, e.g., after synthetic pretraining (note: this is undetectable)
[0068] Using a differentiable rendering equation or differentiable Tenderer to obtain a single-view reconstruction of the reflectance (albedo), shading, normals, depth (altogether known as geometry or G-buffer).
[0069] Using a calibrated, known active illuminator model and position, so as to leverage the illumination decay model (inverse square law) to reconstruct the measured input IR image as a form of self-supervision.
[0070] The methods as described herein are also implemented in some embodiments as a computer program causing a computer and / or a processor to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer-readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed.
[0071] Returning to Fig. 1, there is depicted, in Fig. la, in-cabin monitoring circuitry 1 according to the present disclosure, wherein single-exposure IR is used. An IR frame 2 and an RGB frame 3 are obtained and fed into a monocular depth estimation model 4. In this embodiment, the IR frame and RGB frame are obtained based on a mosaicked IR / RGB sensor, but it should be noted that the present disclosure is not limited to that case. For example, a stacked sensor may be used or different sensors whose respective location may be calibrated in the depth estimation model 4.
[0072] The monocular depth estimation model 4 outputs estimated depth data in millimeters.
[0073] Fig. lb depicts in-cabin monitoring circuitry 4 according to the present disclosure, wherein multiexposure acquisitions are carried out herein, thereby achieving a high dynamic range. For the IR acquisition a Hi IR frame 6 (high exposure, high integration time) and a Low IR frame 7 (low exposure, low integration time) is obtained. For the RGB acquisition, a Hi (exposure) RGB frame 8 and a Low (exposure) RGB frame 9 are obtained. The data obtained based on the frames is fed into a monocular depth estimation model 10 which outputs an estimated depth in millimeters.
[0074] Fig. 2 depicts a data collection procedure and evaluation framework 20 according to the present disclosure according to which depth is estimated. Also, it should be noted that the ToF camera 22 is used as ground truth and the IR camera 22 is used as an input to inference model during monitoring stage of the in-cabin environment 23.
[0075] On the top of Fig. 2, there is shown a IR camera 21 and a ToF camera 22 which obtain image data of an in-cabin environment 23 of a vehicle.
[0076] On the bottom of Fig. 2, as ground truth (lower path), and IR image and a depth map are acquired, which are filtered at 24. Thereby, ground truth depth frames are obtained which provide dense depth map labels that are the target of the monocular depth model’s 25 output (i.e., the depth estimate). By labels, one obtains ground truth depth values in millimeters. The filtering (or processing) performed on the depth map may involve any arbitrarily complex processing (e.g., denoising via neural networks, multipath correction, non-causal temporal filtering, or the like).
[0077] The ground truth depth (GT depth in the figure) is fed into a point cloud alignment model 26 together with the estimated depth from the depth estimation model 25.
[0078] Also, based on a segmentation model 27, ground truth of salient segments, such as human body parts and / or static cabin elements are obtained, which provides general geometric / scale cues. Even if elements of the cabin change (e.g., seat positions, orientation, and the like), a global scale estimate is simplified by at least partial knowledge of the ground truth cabin and its content. This information can be used as input as it is calibrated. This is either dense, in the form of a depth map, or sparse, in the form of fiducial points inside the cabin, i.e., always-in-view positions that act as “anchors” for obtaining metric depth.
[0079] On the upper path of the method 20, an IR image is fed into the monocular depth estimation model 25, which outputs an estimated depth, which is fed into a point cloud alignment model 26, as discussed herein.
[0080] The IR image or frame (single or multi-exposure) is the main input of the monocular depth model and is indicative of a mix of cabin elements, occupants, objects and is recorded under different conditions, such as day / night, low / high ambient light, and the like. As mentioned above, active illumination is used for enforcing specific illumination conditions. Multiple exposures may be used to capture a scene under high-intensity illumination mode and / or one or more lower-intensity illumination modes in a rapid temporal sequence, so as to provide diversity in the recorded inputs.
[0081] Also, an RGB frame is used as an auxiliary input to the depth estimation model 25, as discussed above.
[0082] The output of the segmentation model 27 and the point cloud alignment model 26 are fed into an metrics evaluation model 28 that is configured to output an error between the estimated depth and the ground truth. It should be noted that such a method may be carried out during validation time of the model, e.g., in a factory, to evaluate how reliable the depth model is. However, the present disclosure is not limited in that regard.
[0083] Fig. 3 depicts, on the top, IR images which, after processing by a monocular depth estimation model 30, are transformed into depth data, as discussed herein.
[0084] Figs. 4 to 7 depict different method for training the depth estimation model (machine-learned model):
[0085] Fig. 4 depicts a method 40, wherein a convolutional neural network (CNN) 41 is fed, at its input, with IR data 42, and, at its output, with a pre-calibrated cabin depth map 43. Thereby, the CNN 41 leams a difference depth map 44 including only objects, humans, moving elements that were not present in the pre-calibrated cabin depth map.
[0086] Fig. 5 depicts a method 50, wherein IR data 51, a pre-calibrated cabin depth map 52, and (optionally) pre-calibrated cabin segments 53 are fed, at an input, into a CNN 54. It should be noted that also other metadata may be fed into the input of the CNN, in some embodiments.
[0087] The CNN 54 learns an at-scale depth map 55 including humans and moving elements that were not present in the pre-calibrated cabin. The scale source is the pre-calibrated cabin depth map. Optional labels and segments may steer the attention of the CNN 54 onto specific cabin parts to extract more reliable scale information.
[0088] Fig. 6 depicts a method 60, wherein IR data 61, pre-calibrated camera rays 62 and pre-calibrated cabin keypoints (sparse XYZ points) are fed into an input of a CNN 64.
[0089] The CNN leams an at-scale depth map 65 using the pre-calibrated cabin keypoints. These can be either all in-view, or partly in-view. Moreover, the pre-calibrated camera rays obtained from intrinsics / extrinsics are used to generalize over different lenses (e.g., low tier vehicle model could have cheaper lens — models higher in same range could have expensive, wide lenses, or the like).
[0090] Fig. 7 depicts a method 70 which uses a differentiable Tenderer using pre-calibrated camera rays and illumination. It should be noted that the pre-calibration may be partial, there may be a selfcalibration or semi-calibration aspect considered. For example, a module / car manufacturer will provide at least partial information on the camera and illuminator based on which it may be possible to carry out a training.
[0091] The input IR data 71 is fed into a first CNN at a training step which is configured to generate an Albedo map. The IR data 71 is also fed in parallel into a second CNN at a training step which is configured to generate a depth map. The Albedo map and the depth map are input into a differentiable rendering module 72.
[0092] Based on the depth map, normal estimation 73 is carried out and a normals map is generated. The differentiable rendering module is also fed with pre-calibrated camera rays 74, illuminator parameters, and camera parameters.
[0093] By using a differentiable tenderer, the CNN leams to reproduce the input IR image from the predicted depth map, albedo map, and normals (from depth); internally, the render creates a 3D scene representation and reproduces the IR input view. Learning to reproduce the input results in a more accurate depth map. At inference, only the depth network, i.e., the second CNN may be used to produce a depth map, as the training has concluded.
[0094] Hence, the machine-learned model input, i.e., the IR image, may be forced to be free of ambient light and / or sunlight.
[0095] For this, two frames may be captured, one with active light and one without active light, and taking difference. In some embodiments, the result may be denoised because the difference may have two times the noise variance of each capture as it is the sum of two independent noise samples.lt has been recognized that samples may expose a pitfall of “unfiltered” IR light that might be problematic for monocular depth estimation. Hence, in some embodiments, light conditions are controlled such that ambient light (e.g., sun or any other light sources) is removed from the data. However, in order to further improve the inference, the embodiment of Fig. 7 is envisaged so that one enforces the assumption that the sole source of illumination is an active illuminator whose parameters (intensity, position, etc.) are controlled and can be included in the system design. On the other hand, other light sources may be considered, as well, such as additional lights in the incabin environment or passive illumination as sunlight that generates shadows in the cabin and enters the cabin from different angles. However, such light sources may lead to a bad quality of the monocular depth estimation. Hence, in some embodiments, ambient light is cancelled.
[0096] This may be achieved as follows:
[0097] Assuming a first frame I’ is captured: and assuming a second frame
[0098] I” = Isun+ N” without active illumination is captured, the difference signal
[0099] I := F - I” = Iactive +N’ - N” may be obtained (where Isunmay be cancelled) based on the assumption that a scene has remained static, sunlight has remained roughly constant during the capture of a first (one prime) and a second frame (two primes). This contains active light only and noise contributions, so that I can be used as input to the machine-learned model, as if no passive light was recorded.
[0100] According to such an assumption, an IR camera may be provided with an external trigger signal (e.g., 30 Hertz) that also drives an IR LED periodically with a square wave (on / off). On the rising edge of the trigger signal, the capture of the frame I’ may be triggered with active light. On the falling edge, the frame I” may be captured. Thereby, the camera captures at 60 Hz (twice per LED period).
[0101] The timing requirements of the waveform may be realized on a PCB (printed circuit board), or between cameras and illumination drivers using synchronization signals for example. Thereby, ambient light cancellation may be included in the system that performs monocular depth estimation so that, during inference, one verifies the assumption used during training that active IR illumination is the only illumination contribution.
[0102] Fig. 8 depicts an in-cabin monitoring method 80 according to the present disclosure.
[0103] At 81, IR data are obtained by an IR camera with an active illumination, as discussed herein.
[0104] At 82, monocular depth data are generated with a machine-learned model, as discussed herein.
[0105] Fig. 9 depicts an in-cabin monitoring method 90 according to the present disclosure.
[0106] At 91, IR data are obtained by an IR camera with an active illumination, as discussed herein. At 93, RGB data are obtained an input to the machine-learned model as auxiliary data, as discussed herein.
[0107] At 93, monocular depth data are generated by the machine-learned model, as discussed herein.
[0108] Fig. 10 depicts a training and inference method 100 according to the present disclosure. Empty precalibrated cabin 3D data 101 are fed into a Tenderer 102. This allows the pipeline to query the tenderer for any number of views of this static 3D model, including the current view as long as its pose is known. The pipeline in Fig. 10 describes a means to obtain and use the precalibrated cabin depth map of Fig. 4 and Fig. 5 by solving the problem of estimating the pose between the current 2D IR view and the 2D IR view of the precalibrated empty cabin, of which a 3D model is fully known.
[0109] In order to solve this pose estimation problem, from a reference view of empty cabin 2D IR data 103, 2D image features are extracted (104). Examples of these 2D features (also referred to as keypoints) are SIFT features, ORB features, DISK features, or the like. The feature extraction process 104 is repeated also based on a current 2D IR view 105 obtained with an IR camera.
[0110] Based on the extracted features 104, a perspective -n-point (PnP) algorithm 106 is applied to find matches between empty cabin features and the current view 2D features. As the empty cabin features correspond to empty cabin 3D data for each of their coordinates, it is possible to find via PnP a pose that explains 2D-3D cabin correspondences, where the 3D points are sourced from the known empty cabin 3D data at the 2D features.
[0111] The perspective -n-point algorithm 106 based on empty cabin features therefore yields an at-scale pose between the current view and the reference view of empty cabin. This can be fed into the tenderer 102; finally, this yields an empty cabin depth map view from the current IR view pose.
[0112] This monocular depth (at-scale, in meters) for the empty cabin may be partially providing the monocular depth for the current IR view and can be used as the precalibrated cabin depth in Fig. 4 and Fig. 5. We then proceed to explain how to handle mismatches between the empty cabin and the current view.
[0113] Hereafter, the current IR view 105 is always fed into a neural network 109 as its main input.
[0114] In one embodiment, an output of the tenderer 102 and the current IR view 105 is fed into a merging module 107; in its simplest embodiment, this module may be a sum between the precalibrated depth map and the machine-learned model output (Fig. 4), which leams to estimate the non-empty cabin depth accordingly. In other embodiments, the merging module may include another operator that combines pixelwise elements between the precalibrated depth map and the machine-learned model output, such as a weighted sum, weighted average, product, difference, or the like.
[0115] In a further embodiment, the output of the Tenderer 102 is fed into a cabin consistency module 108 before the merging module 107. The cabin consistency module checks what is the similarity between the current IR view and the precalibrated empty cabin view, and marks the invalid pixels in the precalibrated empty cabin (due to movement of cabin content, movable cabin elements such as seats, or the like) that are afterwards discarded.
[0116] A final output of the neural network 109 and an output of the cabin consistency module 108 is fed into a merge module 110.
[0117] It should be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding. For example, the ordering of 24 and 25 in the embodiment of Fig. 2 may be exchanged. Also, the ordering of 91 and 92 in the embodiment of Fig. 9 may be exchanged. Other changes of the ordering of method steps may be apparent to the skilled person.
[0118] Please note that the division of the circuitry 1 or 5 into units 2 to 10 is only made for illustration purposes and that the present disclosure is not limited to any specific division of functions in specific units. For instance, the circuitry 1 or 5 could be implemented by a respective programmed processor, field programmable gate array (FPGA) and the like.
[0119] Some embodiments pertain to a method for controlling an electronic device, such as mobile terminal including in-cabin monitoring circuitry according to the present disclosure. Also, some embodiments pertain to a mobile terminal including in-cabin monitoring circuitry according to the present disclosure.
[0120] Some embodiments pertain to a system for ICDE using active illumination (such as one or more (non-modulated), flood IR illuminators), wherein the system uses IR single-exposure or multi exposure, or optionally (instead) RGB and IR single-exposure or multi-exposure.
[0121] Some embodiments pertain to a data collection method including performing a static measurement of the in-cabin environment and / or a dynamic measurement of objects / content and the cabin.
[0122] Some embodiments pertain to a data generation method including using synthetic data to obtain monocular depth supervision for IR cameras and / or hybrid (real and synthetic) data (such as real vehicle, synthetic driver or vice versa; or a mix of real and synthetic objects and / or occupants in the vehicle; or the like) Some embodiments pertain to a training method of an artificial intelligence (e.g., a CNN, as discussed herein) including training for IR to depth (optionally with RGB and / or HDR); and / or training measured cabin ground truth to depth; and / or training measured and / or observable cabin keypoints to depth; and / or training cabin segments to depth; and / or training intrinsics and / or camera rays to depth; and / or training differentiable in-cabin rendering to depth.
[0123] Some embodiments pertain to a system that links the estimated monocular depth to in-cabin passive safety applications.
[0124] Fig. 11 depicts a server 110 that includes a first memory 111 configured to store a plurality of instances of ground truth of in cabin measurements relative to a predetermined position in a vehicle cabin, each of the plurality of instances relating to one or more vehicle types. The server further includes a second memory 112 configured to store a first machine learned model with first data as ground truth data. Furthermore, the server includes processing circuitry 113 configured to process the first machine learned model into a second machine learned model using at least one of the plurality of instances of ground truth obtained from the first memory. The server further includes communication circuitry 114 configured to provide, based on a selected vehicle type the second machined learned model to an interface of in cabin monitoring circuitry which uses a monocular depth infrared sensor
[0125] Accordingly, some embodiments pertain to a server comprising: first memory configured to store a plurality of instances of ground truth of in cabin measurements relative to a predetermined position in a vehicle cabin, each of the plurality of instances relating to one or more vehicle types; second memory configured to store a first machine learned model with first data as ground truth data; circuitry configured to process the first machine learned model into a second machine learned model using at least one of the plurality of instances of ground truth obtained from the first memory; communication circuitry configured to provide, based on a selected vehicle type the second machined learned model to an interface of in cabin monitoring circuitry which uses a monocular depth infrared sensor. The first data may comprise ground truth data pertaining to average vehicle sizing. The first machine learned model may a basic or default machine learned model and may be suitable for use with a first range of vehicle types. The second machine learned model may be relatively more adapted to a particular vehicle type and provide different, possibly better results than if a the first machine learned model were used. The second machine learned model may be communicated to in cabin monitoring circuitry at manufacture, at installation or as an upgrade. Further embodiments provide for adapting the ground truth when a different mounting position for the sensor is chosen for example by a manufacturer. Ground truth data may be determined based on a default position is at the top mid point of the front windshield in the vehicle interior. The ground truth data may be adapted based on providing a different mounting position. Transforms of the ground truth data, for example ground truth data to which a geometric transform is applied may be the basis for adapted ground truth data. As such the above mentioned server may include: circuitry configured to process an instance of ground truth of in cabin measurements based on data indicating a different predetermined position in a vehicle cabin.
[0126] As mentioned in the outset, the present disclosure may be carried out based on measured relative distances between a driver and one or more in-cabin points (e.g., steering wheel). In the following, non-limiting examples are discussed on how to determine such distances, without limiting the present disclosure in that regard.
[0127] Fig. 12, there is depicted a Occupant Status Monitoring (OSM) 120 as an embodiment of a vehicle safety system.
[0128] The OSM 1 includes a ToF sensor 121, in this embodiment an iToF camera configured to acquire time-of-flight data indicative of a driver of a vehicle, without limiting the present disclosure in that regard since any other occupant may be monitored accordingly. The ToF data includes depth and confidence information as discussed herein.
[0129] The OSM 110 further includes ToF circuitry 122 according to the present disclosure.
[0130] Based on obtained ToF data, the ToF circuitry 122 is further configured to generate body segmentation information and to provide the body segmentation information to a safety function 123 included in the OSM 120.
[0131] In this embodiment, the safety function 123 also obtains an RGB camera output for the surveillance of the driver. It should be noted that the present disclosure is not limited to the case described herein since in some embodiments, an IR camera or RGB-IR camera may be used accordingly.
[0132] Fig. 13 depicts a schematic block diagram of a method 130 for generating body feature information 131.
[0133] ToF data including confidence 132 and depth 133 are fed into a confidence-based depth filtering algorithm 134 configured to generated different outputs, based on a coarse determination whether the obtained result has a sufficient quality: One output is statistics for gain control 135 which is used for refining the confidence 132, i.e., a gain of the iToF camera is adapted based on this information. The confidence 132 is again used for refining the statistics for gain control 135.
[0134] Another output is statistics for modulation frequency 136 which is used to refine the depth measurement. For example, the modulation frequency may be adapted such that the depth measurement becomes more precise, as generally known to the skilled person.
[0135] Also, based on the confidence 132, segmentation labels 137 are generated based on a machinelearning algorithm. Based on the labels 137, the confidence 132, as well as the confidence filtered depth 134, the body feature information is generated. The labels, the confidence, and the depth may allow the system to estimate the 3D positions and orientations of predetermined body parts. Thereby, a corresponding skeleton fitting may be carried out based on which the body features are determined.
[0136] Fig. 14 depicts an illustrational result of a method according to Fig. 13. An interior of a vehicle is shown captured with an iToF camera, wherein the ToF data are only pre-processed at that point. In the interior, a driver 141 and a passenger 142 are shown.
[0137] Based on the pre-processed depth, certain points 143 (white rectangles) on the occupants are identified which are indicative of their respective postures (indicated with white lines in Fig. 14).
[0138] Based on the identified points, a distance to a predetermined location in the vehicle (or with respect to the vehicle) is determined. In this embodiment, a distance between the driver’s 141 head and a steering wheel 144 is determined. In some embodiments predetermined locations may be fixed positions, such as fixed armrests, comers of windows, or positions on the roof lining, where adaptive restraint system might be placed It should be noted that any distance between identified points and predetermined locations may be determined, in some embodiments. Also, multiple distances may be determined in some embodiments.
[0139] It should also be noted that a distance to any point (or location) of the vehicle may be determined and the present disclosure is not limited to points of the in-cabin. For example, a distance of a body feature to a windshield (inside or outside or intermediate point) may be determined. Also, a distance of a body feature between an intermediate layer of a car roof (e.g., a stabilizing metal layer) may be determined, in some embodiments.
[0140] Fig. 15 depicts a skeleton in a filtered depth image (top image) and the result of a segmentation is shown in a confidence image (bottom image). As in Fig. 14, certain points and a skeleton of the driver are identified in the depth image. In the confidence image, body parts 151 are identified.
[0141] Similarly, Fig. 16 depicts a depth image (top left) and a confidence image (bottom left), wherein in the confidence image, points on the occupant’s body are identified based on the depth image with a high precision. Higher confidence values can be indicative of a front of the occupant’s body parts used for segmentation with depth information, while low confidence values indicate areas of depth with lower reliability to be discarded for further processing. For example, fingers of the driver can be recognized in this manner (indicated with white circles).
[0142] On the right of Fig. 16, it is shown that distances of the driver’s head and back are determined based on a method discussed herein and are used as an indicator of the driver’s position or posture. For example, in case it is determined, based on the determined distances, that the driver is bent forward, an alert is generated that he is out of position. Additionally or alternatively, an airbag restraint system is adapted or suppressed based on the driver’s position.
[0143] In more general terms,
[0144] The methods discussed herein can also be implemented as a computer program causing a computer and / or a processor, to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer-readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the method described to be performed.
[0145] All units and entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software.
[0146] In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure.
[0147] Note that the present technology can also be configured as described below.
[0148] (1) In-cabin monitoring circuitry for monitoring an in-cabin environment of a vehicle, the circuitry being configured to: obtain infrared sensor data; and generate monocular depth data with a machine-learned model, wherein the machine-learned model has depth data as ground truth and uses the obtained infrared sensor data as input. (2) The in-cabin monitoring circuitry of (1), wherein the monocular depth data is generated by mapping the infrared sensor data to a depth map of the in-cabin environment learned based on training with the depth data as ground truth.
[0149] (3) The in-cabin monitoring circuitry of (1) or (2), wherein the infrared sensor data is obtained based on an active illumination.
[0150] (4) The in-cabin monitoring circuitry of (3), wherein the active illumination is an infrared illumination.
[0151] (5) The in-cabin monitoring circuitry of anyone of (1) to (4), wherein the machine-learned model is based on a convolutional neural network.
[0152] (6) The in-cabin monitoring circuitry of anyone of (1) to (5), further configured to: obtain RGB sensor data as an auxiliary input to the machine-learned model.
[0153] (7) The in-cabin monitoring circuitry of anyone of (1) to (6), wherein the machine-learned model is based on at least one time-of-flight measurement of the in-cabin environment used as ground truth.
[0154] (8) The in-cabin monitoring circuitry of anyone of (1) to (7), wherein the machine-learned model is based on synthetic depth data of an in-cabin environment.
[0155] (9) The in-cabin monitoring circuitry of (8), wherein the machine-learned model is based on an augmentation of real time-of-flight data with the synthetic depth data.
[0156] (10) The in-cabin monitoring circuitry of anyone of (1) to (9), wherein the machine-learned model has illumination cues as ground truth.
[0157] (11) An in-cabin monitoring method for monitoring an in-cabin of a vehicle, the method comprising: obtaining infrared sensor data; and generating monocular depth data with a machine-learned model, wherein the machine- learned model has depth data as ground truth and uses the obtained infrared sensor data as input.
[0158] (12) The in-cabin monitoring method of (11), wherein the monocular depth data is generated by mapping the infrared sensor data to a depth map of the in-cabin environment learned based on training with the depth data as ground truth.
[0159] (13) The in-cabin monitoring method of (11) or (12), wherein the infrared sensor data is obtained based on an active illumination. (14) The in-cabin monitoring method of (13), wherein the active illumination is an infrared illumination.
[0160] (15) The in-cabin monitoring method of anyone of (11) to (14), wherein the machine-learned model is based on a convolutional neural network.
[0161] (16) The in-cabin monitoring method of anyone of (11) to (15), further comprising: obtaining RGB sensor data as an auxiliary input to the machine-learned model.
[0162] (17) The in-cabin monitoring method of anyone of (11) to (16), wherein the machine-learned model is based on at least one time-of-flight measurement of the in-cabin environment used as ground truth.
[0163] (18) The in-cabin monitoring method of anyone of (11) to (17), wherein the machine-learned model is based on synthetic time-of-flight data of an in-cabin environment.
[0164] (19) The in-cabin monitoring method of (18), wherein the machine-learned model is based on an augmentation of real time-of-flight data with the synthetic depth data.
[0165] (20) The in-cabin monitoring method of anyone of (11) to (19), wherein the machine-learned model has illumination cues as ground truth.
[0166] (21) A computer program comprising program code causing a computer to perform the method according to anyone of (11) to (20), when being carried out on a computer.
[0167] (22) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (11) to (20) to be performed.
Claims
CLAIMS1. In-cabin monitoring circuitry for monitoring an in-cabin environment of a vehicle, the circuitry being configured to: obtain infrared sensor data; and generate monocular depth data with a machine-learned model, wherein the machine-learned model has depth data as ground truth and uses the obtained infrared sensor data as input.
2. The in-cabin monitoring circuitry of claim 1, wherein the monocular depth data is generated by mapping the infrared sensor data to a depth map of the in-cabin environment learned based on training with the depth data as ground truth.
3. The in-cabin monitoring circuitry of claim 1, wherein the infrared sensor data is obtained based on an active illumination.
4. The in-cabin monitoring circuitry of claim 3, wherein the active illumination is an infrared illumination.
5. The in-cabin monitoring circuitry of claim 1, wherein the machine-learned model is based on a convolutional neural network.
6. The in-cabin monitoring circuitry of claim 1, further configured to: obtain RGB sensor data as an auxiliary input to the machine-learned model.
7. The in-cabin monitoring circuitry of claim 1, wherein the machine-learned model is based on at least one time-of-flight measurement of the in-cabin environment used as ground truth.
8. The in-cabin monitoring circuitry of claim 1, wherein the machine-learned model is based on synthetic depth data of an in-cabin environment.
9. The in-cabin monitoring circuitry of claim 8, wherein the machine-learned model is based on an augmentation of real time-of-flight data with the synthetic depth data.
10. The in-cabin monitoring method of claim 1, wherein the near-infrared sensor data is used as input for at least one of a training stage and an inference stage.
11. An in-cabin monitoring method for monitoring an in-cabin environment of a vehicle, the method comprising: obtaining infrared sensor data; and generating monocular depth data with a machine-learned model, wherein the machine- learned model has depth data as ground truth and uses the obtained infrared sensor data as input.
12. The in-cabin monitoring method of claim 11, wherein the monocular depth data is generated by mapping the infrared sensor data to a depth map of the in-cabin environment learned based on training with the depth data as ground truth.
13. The in-cabin monitoring method of claim 11, wherein the infrared sensor data is obtained based on an active illumination.
14. The in-cabin monitoring method of claim 13, wherein the active illumination is an infrared illumination.
15. The in-cabin monitoring method of claim 11, wherein the machine-learned model is based on a convolutional neural network.
16. The in-cabin monitoring method of claim 11, further comprising: obtaining RGB sensor data as an auxiliary input to the machine-learned model.
17. The in-cabin monitoring method of claim 11, wherein the machine-learned model is based on at least one time-of-flight measurement of the in-cabin environment used as ground truth.
18. The in-cabin monitoring method of claim 11, wherein the machine-learned model is based on synthetic time-of-flight data of an in-cabin environment.
19. The in-cabin monitoring method of claim 18, wherein the machine-learned model is based on an augmentation of real time-of-flight data with the synthetic depth data.
20. The in-cabin monitoring method of claim 11, wherein the machine-learned model has illumination cues as ground truth.
Citation Information
Patent Citations
Cab scene semantic and visual depth joint analysis method
CN110781717A