In-Cabin Audio-Visual Monitoring for Multi-Occupant Mental State Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current in-cab monitoring systems in automotive vehicles face challenges such as significant visual and audio noise, suboptimal camera angles, and multi-occupancy, which limit their accuracy and capability in monitoring driver and passenger behavior effectively.

Innovation Solution

A confidence-aware stochastic process regression-based audio-visual fusion approach that assesses occupant mental state by determining expressed face, voice, and body behaviors, and provides a short list of potential causes for these behaviors, improving accuracy and enabling new capabilities beyond visual-only monitoring.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If visual-only monitoring via cameras is used, then the system is simpler to implement, but the accuracy and capability in monitoring occupant behavior are limited

Engineering Contradiction:
Improvemonitoring accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines visual monitoring (cameras) with audio monitoring (microphones and audio processing) into a unified audio-visual monitoring system. This fusion of multiple sensing modalities improves measurement precision by cross-validating signals and reducing false positives, while the integrated architecture manages complexity through coordinated processing of both audio and visual data streams

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If audio-visual fusion is implemented, then false positives are reduced, but computational requirements and processing complexity increase

Engineering Contradiction:
Improvefalse positive rateVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system employs feedback mechanisms where audio and visual signals continuously cross-validate each other. When visual detection suggests a certain state (e.g., occupant alertness), audio signals provide feedback to confirm or refute this assessment. This mutual verification reduces false positives by requiring consistent evidence from multiple modalities before triggering alerts

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces an intermediary processing layer that fuses audio and visual data before final interpretation. This intermediary stage reconciles conflicting signals, weighs evidence from both modalities, and produces a unified assessment of occupant state, thereby reducing false positives while managing computational complexity through structured integration

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If monitoring covers multiple occupants, then comprehensive behavior tracking is achieved, but confusion about audio signal sources and attention focus increases

Engineering Contradiction:
Improvemulti-occupant monitoring capabilityVSAvoidaudio source attribution difficulty
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The system segments the monitoring space into distinct zones (e.g., driver area, passenger area) with dedicated cameras and audio sensors for each zone. Visual tracking establishes which occupant is in which zone, and this spatial segmentation is used to attribute audio signals to specific occupants based on their location, thereby resolving confusion about audio source origins in multi-occupant scenarios

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If visual noise from lighting conditions is addressed, then detection accuracy improves, but system complexity and calibration requirements increase

Engineering Contradiction:
Improvedetection accuracy under varying lightingVSAvoidlighting adaptation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines visual and audio modalities to compensate for lighting-related visual noise. When lighting conditions degrade visual detection accuracy, the system relies more heavily on audio signals (which are unaffected by lighting) to assess occupant state. This fusion approach maintains detection accuracy across varying lighting conditions without requiring complex lighting adaptation mechanisms

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20240054794A1Multistage Audio-Visual Automotive Cab Monitoring
Publication Date: 2024.02.15 BLUESKEYE AI LTD
  • US20240054794A1 patent drawing
  • US20240054794A1 patent drawing
  • US20240054794A1 patent drawing

AI summary

Described is a task for an automobile interior having at least one subject that creates a video input, an audio input, and a context descriptor input. The video input relates to the at least one subject and is processed by a face detection module and a facial point registration module to produce a first output. The first output is further processed by at least one of: a facial point tracking module, a head orientation tracking module, a body tracking module, a social gaze tracking module, and an action unit intensity tracking module. The audio input relating to the at least one subject is processed by a valence and arousal affect states tracking module to produce a second output and to produce a valence and arousal scores output. A temporal behavior primitives buffer produce a temporal behavior output. Based on the foregoing, a mental state prediction module predicts the mental state of at least one subject in the automobile interior.