Adaptive Driving HMI Policy Training for Personalized Safety Alerts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing vehicle safety systems face challenges in effectively communicating potentially unsafe situations to drivers, as some drivers benefit from earlier and continuous alerts while others find them annoying, leading to reduced system trust and effectiveness.

Innovation Solution

A method and system for training policies using a framework that encodes human behaviors and preferences in a driving environment, utilizing a Markov Decision Process (MDP) to model interactions between a simulated human driver and an adaptive Human-Machine Interface (HMI) system, allowing for personalized interventions based on driver-specific traits and preferences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If earlier and continual alerts are provided to drivers, then driving safety is improved for distracted drivers, but system trust and usefulness deteriorate for less distracted drivers

Engineering Contradiction:
Improvedriving safetyVSAvoidsystem adaptability to different driver types
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The HMI system dynamically adapts its alert behavior based on the driver's distraction level. The system transitions from static, one-size-fits-all alerting to dynamic, personalized alerting where the timing, frequency, and content of alerts are adjusted in real-time according to the detected driver state. This resolves the contradiction by making the system reliable for distracted drivers while avoiding over-alerting for attentive drivers.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes key parameters of the HMI interface based on driver characteristics. Specifically, it modifies alert timing parameters, notification frequency parameters, and intervention thresholds according to the driver's distraction level. This allows the same system to provide appropriate levels of assistance to different driver types, improving safety for distracted drivers without annoying attentive ones.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If personalized interventions are implemented, then system usefulness is improved, but device complexity increases

Engineering Contradiction:
Improvepersonalization capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments drivers into different categories based on their distraction levels (e.g., highly distracted, moderately distracted, less distracted). This segmentation allows the complex personalization logic to be organized into manageable segments or modules, where each segment has its own simplified intervention strategy. The overall system complexity is reduced by dividing it into discrete, handleable segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary component (the distraction detection and classification module) that sits between the driver monitoring sensors and the HMI intervention system. This intermediary processes raw driver state data, classifies distraction levels, and translates them into appropriate HMI parameters. This intermediary layer simplifies the overall system architecture by creating a clear separation of concerns and standardized interfaces.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20230331240A1System and method for training at least one policy using a framework for encoding human behaviors and preferences in a driving environmet
Publication Date: 2023.10.19 TOYOTA RESEARCH INSTITUTE INC
  • US20230331240A1 patent drawing
  • US20230331240A1 patent drawing
  • US20230331240A1 patent drawing

AI summary

Disclosed are systems and methods for training at least one policy using a framework for encoding human behaviors and preferences in a driving environment. In one example, the method includes the steps of setting parameters of rewards and a Markov Decision Process (MDP) of the at least one policy that models a simulated human driver of a simulated vehicle and an adaptive human-machine interface (HMI) system configured to interact with each other and training the at least one policy to maximize a total reward based on the parameters of the rewards of the at least one policy.