Shared-Encoder Estimator Training for Accurate Driver State Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for estimating a target person's state from face image data, such as driver drowsiness or fatigue, often converge to less accurate local solutions during machine learning, leading to ineffective estimation.

Innovation Solution

An estimator generation apparatus that combines face image data with physiological data from sensors to train a shared encoder, allowing for more accurate estimation by converging to higher-accuracy local solutions, using a common encoder for both face image and physiological data processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If a trained model with neural network is used to estimate driver state from face image data, then the driver state can be estimated automatically, but the model may converge to less accurate local solutions during machine learning

Engineering Contradiction:
Improveautomatic driver state estimationVSAvoidaccuracy of driver state estimation
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent combines face image data and physiological data into a unified learning dataset, allowing the neural network to learn from multiple data sources simultaneously. This merging enables the model to converge to more accurate global solutions by integrating complementary information from both visual and physiological measurements, resolving the contradiction between automation and precision.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a composite learning dataset that integrates heterogeneous data types (face images and physiological signals) into a unified training framework. This composite approach allows the neural network to leverage diverse data characteristics, improving estimation accuracy while maintaining automated operation.

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If physiological data from sensors is collected to improve estimation accuracy, then higher-order information becomes available, but the device complexity and operational costs increase

Engineering Contradiction:
Improveaccuracy of state estimationVSAvoidcomplexity of data collection system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent designs a shared encoder architecture that processes both face image data and physiological data through a common neural network pathway. This universal processing framework allows the system to handle multiple data types with a single model structure, reducing overall system complexity while maintaining high estimation accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs a self-supervised learning approach where the model learns to reconstruct physiological data from face images alone, using this reconstruction capability to pre-train the shared encoder. This self-service mechanism enables the system to leverage easily obtainable face image data to improve processing of more complex physiological data, reducing the burden of complex sensor integration.

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP3876191B1Estimator generation device, monitoring device, estimator generation method, estimator generation program
Publication Date: 2024.01.24 OMRON CORP
  • EP3876191B1 patent drawingFigure 1~2
  • EP3876191B1 patent drawingFigure 3~4A
  • EP3876191B1 patent drawingFigure 4B

AI summary

A technique allows generation of an estimator that can estimate a target person's state more accurately. An estimator generation apparatus includes a first estimator and a second estimator sharing a common encoder. The first estimator is trained to determine a target person's state from face image data. The second estimator is trained to reconstruct physiological data from face image data. The machine learning allows the common encoder to have its parameters converging toward higher-accuracy local solutions for estimating the target person's state, thus generating the estimator that can estimate the target person's state more accurately.