Facial Image Preprocessing Selection for Accurate XR Avatar Expressions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing extended reality (XR) technologies face challenges in accurately predicting facial action units for avatars due to variations in lighting conditions and individual user characteristics, leading to inaccuracies in rendering realistic facial expressions.

Innovation Solution

A method is employed to select a combination of preprocessing parameters, such as contrast-limited adaptive histogram equalization (CLAHE) grid size and clip limit, tailored to each user and lighting conditions, to preprocess facial images before applying a machine learning model for predicting facial action units, followed by retargeting these units onto avatars.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If standard preprocessing parameters are used for all users, then device complexity is reduced, but measurement precision of facial action units deteriorates

Engineering Contradiction:
Improvefacial action unit prediction accuracyVSAvoidpreprocessing parameter configuration
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by selecting optimal preprocessing parameters (such as CLAHE grid size and clip limit) specifically tailored to each user's facial characteristics and the current lighting conditions. This customization of parameters improves measurement precision for facial action unit detection while the automated selection process manages the complexity burden.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system implements self-service through automated parameter selection that adapts to each user without requiring manual configuration. The system automatically analyzes user-specific facial features and environmental lighting to select appropriate preprocessing parameters, eliminating the need for users to manually adjust complex settings while maintaining high prediction accuracy.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If preprocessing parameters are customized for each user and lighting condition, then measurement precision improves, but device complexity increases

Engineering Contradiction:
Improvefacial action unit prediction accuracyVSAvoidparameter selection and management
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-establishing multiple sets of preprocessing parameters corresponding to different lighting conditions and user types. During operation, the system quickly selects from these pre-configured parameter sets based on detected conditions, avoiding the need for complex real-time parameter optimization while maintaining high precision.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If lighting conditions vary, then adaptability is improved, but measurement precision deteriorates

Engineering Contradiction:
Improvelighting condition adaptationVSAvoidfacial action unit prediction accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies dynamics by making the preprocessing parameters adaptive to changing lighting conditions. The system dynamically selects appropriate parameter sets based on real-time lighting analysis, allowing it to adapt to various environmental conditions while maintaining consistent measurement precision across different scenarios.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20260073731A1Selecting combination of parameters for preprocessing facial images of wearer of head-mountable display
Publication Date: 2026.03.12 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US20260073731A1 patent drawing
  • US20260073731A1 patent drawing
  • US20260073731A1 patent drawing

AI summary

For each facial image of a wearer of a head-mountable display (HMD), preprocessed facial images corresponding to combinations of preprocessing parameters are generated. A machine learning model is applied to each preprocessed facial image to predict facial action units. The facial action units predicted from each preprocessed facial image are retargeted onto an avatar to render an avatar facial image. Avatar facial landmarks within each avatar facial image and wearer facial landmarks within each facial image are detected. The combination of preprocessing parameters yielding a highest similarity between the avatar facial landmarks and the wearer facial landmarks corresponding to the avatar facial landmarks is selected.