Multi-Modal In-Cabin Occupant Evaluation for 3D Pose Scaling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing occupant monitoring systems (OMS) struggle to accurately determine a three-dimensional (3D) pose and size of vehicle occupants due to limitations in sensor resolution and susceptibility to occlusions, especially when occupants are not in an upright posture, and rely heavily on expensive depth-sensing cameras or limited RADAR data.

Innovation Solution

A multi-modal sensor fusion approach combining optical image sensors and point cloud generating depth sensors, such as RADAR or LIDAR, to generate a 3D representation of vehicle occupants by processing optical image frames and depth data, using machine learning models to derive scale-normalized 3D poses and anchor them to absolute scales for precise 3D estimates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If depth-sensing cameras or RADAR sensors are used to determine 3D pose and size, then measurement precision is improved, but device complexity and cost increase

Engineering Contradiction:
Improve3D pose and size estimation accuracyVSAvoidsensor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines optical image sensors (cameras) with point cloud generating depth sensors (RADAR or LIDAR) to create a multi-modal sensor fusion system. This merging allows the system to achieve accurate 3D pose and size estimation by integrating the high-resolution visual data from cameras with the depth information from RADAR/LIDAR, resolving the contradiction between measurement precision and device complexity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an occupant evaluation function as an intermediary processing layer that receives data from multiple sensors and generates 3D representations. This intermediary component fuses the sensor data, correlates optical image frames with point cloud data, and produces accurate 3D pose and size estimates without requiring direct complex hardware integration

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multi-modal sensor fusion is implemented, then 3D representation accuracy is improved, but processing complexity increases

Engineering Contradiction:
Improve3D representation accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the processing into distinct functional components: an optical image sensor processing path that generates feature representations, and a depth sensor processing path that determines depth values. This segmentation allows each path to be optimized independently and simplifies the overall processing complexity while maintaining high 3D representation accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms 2D optical image data into 3D representations by integrating depth information from point cloud sensors. This dimensionality change approach allows the system to generate accurate 3D occupant representations by adding the depth dimension to the existing 2D visual data, improving accuracy without requiring complete reprocessing of all sensor data

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Adaptability or versatility

If scale-normalized 3D pose estimation is used, then adaptability to different distances is improved, but absolute scale accuracy deteriorates

Engineering Contradiction:
Improveviewpoint invarianceVSAvoidabsolute scale accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent uses depth information from point cloud sensors as feedback to correct and scale the normalized 3D pose estimates. The system first generates scale-invariant 3D representations for viewpoint adaptability, then uses measured depth values from RADAR/LIDAR sensors to feedback-correct the absolute scale, achieving both adaptability and accuracy

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Provides accurate 3D pose and size estimates of vehicle occupants invariant to viewpoint, enabling effective child presence detection and enhancing vehicle safety features like airbag deployment and posture classification, while reducing reliance on precise sensor calibration.

Implementation Method 1

Depth-perception sensors may use radio waves, laser light, and/or sound waves, for example, to detect the presence or movements of living beings within a vehicle interior

Methodology Applied
Scientific EffectRadar: Radar

Implementation Method 2

a point cloud generating depth sensor (e.g., a RADAR sensor and/or a LIDAR sensor)

Methodology Applied
Scientific EffectLIDAR: LIDAR

Data Source

PatentUS12462586B2Occupant evaluation using multi-modal sensor fusion for in-cabin monitoring systems and applications
Publication Date: 2025.11.04 NVIDIA CORP
  • US12462586B2 patent drawing
  • US12462586B2 patent drawing
  • US12462586B2 patent drawing

AI summary

In various examples, occupant assessment using multi-modal sensor fusion for monitoring systems and applications are provided. In some embodiments, an occupant monitoring system comprises an occupant evaluation function that may predict at least one characteristic representative of a size of the occupant. The occupant evaluation function may include a first processing path that generates a representation of features corresponding to the occupant based on optical image data, and a second processing path that performs operations to determine a depth corresponding to the one or more features based on depth data derived from the optical image data and the point cloud depth data. In some embodiments, a three-dimensional pose detection model generates a three-dimensional pose estimate of the occupant using the optical image data, and the three-dimensional pose estimate is scaled to an absolute pose based on the point cloud depth data.