Shared-Encoder Estimator Training for Accurate Driver State Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for estimating a target person's state from face image data, such as driver drowsiness or fatigue, often converge to less accurate local solutions during machine learning, leading to ineffective estimation.
Innovation Solution
An estimator generation apparatus that combines face image data with physiological data from sensors to train a shared encoder, allowing for more accurate estimation by converging to higher-accuracy local solutions, using a common encoder for both face image and physiological data processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If a trained model with neural network is used to estimate driver state from face image data, then the driver state can be estimated automatically, but the model may converge to less accurate local solutions during machine learning
Solution Approach 1:
The patent combines face image data and physiological data into a unified learning dataset, allowing the neural network to learn from multiple data sources simultaneously. This merging enables the model to converge to more accurate global solutions by integrating complementary information from both visual and physiological measurements, resolving the contradiction between automation and precision.
Solution Approach 2:
The patent creates a composite learning dataset that integrates heterogeneous data types (face images and physiological signals) into a unified training framework. This composite approach allows the neural network to leverage diverse data characteristics, improving estimation accuracy while maintaining automated operation.
2Measurement precision
If physiological data from sensors is collected to improve estimation accuracy, then higher-order information becomes available, but the device complexity and operational costs increase
Solution Approach 1:
The patent designs a shared encoder architecture that processes both face image data and physiological data through a common neural network pathway. This universal processing framework allows the system to handle multiple data types with a single model structure, reducing overall system complexity while maintaining high estimation accuracy.
Solution Approach 2:
The patent employs a self-supervised learning approach where the model learns to reconstruct physiological data from face images alone, using this reconstruction capability to pre-train the shared encoder. This self-service mechanism enables the system to leverage easily obtainable face image data to improve processing of more complex physiological data, reducing the burden of complex sensor integration.
Data Source
Figure 1~2
Figure 3~4A
Figure 4B
AI summary
A technique allows generation of an estimator that can estimate a target person's state more accurately. An estimator generation apparatus includes a first estimator and a second estimator sharing a common encoder. The first estimator is trained to determine a target person's state from face image data. The second estimator is trained to reconstruct physiological data from face image data. The machine learning allows the common encoder to have its parameters converging toward higher-accuracy local solutions for estimating the target person's state, thus generating the estimator that can estimate the target person's state more accurately.