Facial Marker Tracking for Accurate Action Unit Data Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current techniques face challenges in accurately judging Action Units (AUs) from facial images, particularly due to the difficulty in generating large amounts of annotated teacher data and accurately measuring small changes in facial regions, which hinders efficient AU estimation.
Innovation Solution
A system utilizing RGB and IR cameras to capture images with markers, where the generating device specifies marker positions and judges occurrence intensity based on predefined criteria, enabling the creation of learning data for machine learning models to estimate AU occurrence intensities without manual annotation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation by specialists is used to create teacher data for AU estimation, then the accuracy of AU estimation is improved, but the time and cost required for data preparation increases significantly
Solution Approach 1:
The patent uses markers attached to facial landmarks as physical copies/reference points that can be automatically detected and tracked. These markers serve as surrogates for manual annotation, allowing the system to automatically generate teacher data by tracking marker positions rather than requiring specialists to manually annotate each frame, thus resolving the contradiction between accuracy and time consumption
Solution Approach 2:
The system performs self-annotation by automatically detecting marker positions on the face and generating AU estimation teacher data without human intervention. The automated tracking and classification processes enable the system to serve itself in creating training data, eliminating the need for time-consuming manual annotation while maintaining consistent quality
2Measurement precision
If image processing is performed on facial images to measure facial region changes, then AU detection is attempted, but the accuracy is insufficient due to difficulty in detecting small changes
Solution Approach 1:
The patent attaches markers to specific local regions of the face corresponding to key facial landmarks. By concentrating measurement effort on these localized marker positions rather than attempting to measure subtle changes across the entire facial image, the system achieves high precision in detecting small facial movements and expressions
Solution Approach 2:
The markers used in the patent have distinct visual characteristics (such as color or reflectivity properties) that make them easily distinguishable from the surrounding facial features. This allows the detection system to clearly identify marker positions and track their movements with high precision, overcoming the difficulty of detecting subtle facial changes in natural images
3Measurement precision
If a large amount of teacher data is generated for machine learning, then the AU estimation model accuracy is improved, but the manual annotation workload increases
Solution Approach 1:
The marker-based approach creates easily detectable visual copies of facial landmarks that can be automatically tracked. This enables the system to generate large volumes of annotated training data automatically by detecting marker positions across many images, rather than requiring manual annotation for each sample, thus resolving the contradiction between model accuracy and data generation efficiency
Solution Approach 2:
The automated system performs self-annotation by detecting marker positions and generating corresponding AU labels without human intervention. This self-service capability enables rapid generation of large training datasets, allowing the system to produce the large amount of teacher data needed for high-accuracy machine learning models without incurring proportional increases in manual workload
Data Source
AI summary
A non-transitory computer-readable recording medium stores therein a judgment program that causes a computer to execute a process including acquiring a captured image including a face to which a plurality of markers are attached at a plurality of positions that are associated with a plurality of action units, specifying each of the positions of the plurality of markers included in the captured image, judging an occurrence intensity of a first action unit associated with a first marker from among the plurality of action units based on a judgment criterion of an action unit and a position of the first marker from among the plurality of markers, and outputting the occurrence intensity of the first action unit by associating the occurrence intensity with the captured image.


