Shared-Encoder Estimator Training for Accurate Face State Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for estimating a target person's state from face image data, such as those using human-designed feature point extraction, may converge to less accurate local solutions during machine learning, leading to inaccurate state estimation.

Innovation Solution

An estimator generation apparatus that includes a learning processor to construct a first estimator trained on face image data and a second estimator trained on physiological data, sharing a common encoder to automatically design feature quantities, allowing for more accurate state estimation by converging to higher-accuracy local solutions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a learning model is trained to directly estimate the target person's state from face image data, then the feature quantity is automatically designed and estimation accuracy is improved, but the model may converge to less accurate local solutions

Engineering Contradiction:
Improvestate estimation accuracyVSAvoidconvergence to accurate solution
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The learning model is segmented into two distinct estimators: a first estimator that estimates the target person's state from face image data, and a second estimator that reconstructs physiological data from face image data. This segmentation allows each estimator to specialize in different tasks while sharing the common encoder, thereby improving convergence reliability without sacrificing estimation accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The second estimator acts as an intermediary that bridges the face image data and the target state by first reconstructing physiological data. This intermediary process guides the common encoder to learn more accurate feature representations, preventing convergence to poor local solutions while maintaining high estimation accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If human-designed feature point extraction is used, then the system structure is simpler, but the feature quantity may not accurately reflect the target person's state

Engineering Contradiction:
Improvesystem structure complexityVSAvoidstate estimation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The common encoder automatically designs the feature quantity by itself through the dual-estimator training process, eliminating the need for manual feature point extraction. The encoder learns optimal feature representations directly from the data by simultaneously serving both estimation and reconstruction tasks, achieving high accuracy without increasing system complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the parameters of the common encoder through iterative training with both estimators. The encoder's internal parameters are automatically optimized to extract meaningful features from face image data, replacing fixed human-designed features with dynamically learned parameters that accurately reflect the target person's state.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11834052B2Estimator generation apparatus, monitoring apparatus, estimator generation method, and computer-readable storage medium storing estimator generation program
Publication Date: 2023.12.05 OMRON CORP
  • US11834052B2 patent drawing
  • US11834052B2 patent drawing
  • US11834052B2 patent drawing

AI summary

An estimator generation apparatus may include a first estimator and a second estimator sharing a common encoder. The first estimator may be trained to determine a target person's state from face image data. The second estimator may be trained to reconstruct physiological data from face image data. The machine learning may allow the common encoder to have its parameters converging toward higher-accuracy local solutions for estimating the target person's state, thus generating the estimator that may estimate the target person's state more accurately.