Avatar Facial Expression Synchronization via Vocal and Image Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current avatar facial expression representation technologies in virtual spaces lack the ability to convey natural and sophisticated expressions efficiently, particularly in electronic communication systems, where controlling facial movements and lip expressions is more effective than body movements for conveying user intentions.

Innovation Solution

An apparatus comprising a vocal information processing unit to extract emotional changes and emphasis from voice data, a pronunciation information processing unit to analyze mouth shape changes, and an image information processing unit to track facial movements, all integrated with a facial expression processing unit to represent avatar expressions synchronously with user inputs, while a reliability evaluating unit ensures accurate and natural representation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple processing units (vocal, pronunciation, image) are integrated to represent facial expressions, then the naturalness and sophistication of avatar expressions is improved, but the device complexity increases

Engineering Contradiction:
Improvenaturalness of avatar expressionVSAvoidsystem structure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system divides the facial expression representation task into three independent processing units: vocal information processing unit (extracts emotional changes and emphasis), pronunciation information processing unit (analyzes mouth shape changes), and image information processing unit (tracks facial movements). Each unit processes specific types of input data independently, allowing the system to handle complex expressions through modular, manageable components rather than a monolithic system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The facial expression processing unit serves as a universal component that integrates inputs from all three processing units (vocal, pronunciation, image) and generates comprehensive avatar facial expressions. This multi-functional unit can process various combinations of input data types to produce natural expressions across different communication scenarios, making the system adaptable and versatile.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Loss of information

If vocal, pronunciation, and image information are synchronized to represent facial expressions, then the effectiveness of conveying user intention is improved, but the processing time and computational load increase

Engineering Contradiction:
Improveuser intention conveyance effectivenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary processing of input data in parallel across multiple specialized units: the vocal information processing unit extracts emotional changes and emphasis points, the pronunciation information processing unit analyzes mouth shape parameters, and the image information processing unit tracks facial feature points. By preparing processed outputs from each unit simultaneously rather than sequentially, the system reduces overall processing time while maintaining comprehensive information for accurate intention conveyance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates multiple parallel processing streams that copy and analyze different aspects of the user's communication simultaneously - vocal characteristics, pronunciation patterns, and facial movements - rather than processing one aspect at a time. This parallel copying of processing tasks enables the system to gather comprehensive expression data without sequential delays.

Inventive Principle:
Principle #26Copying

3Measurement precision

If the system processes long-term and short-term parameter changes to estimate emotional changes and emphasis, then the accuracy of emotion detection is improved, but the computational complexity increases

Engineering Contradiction:
Improveemotion detection accuracyVSAvoidprocessing algorithm complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The emotion change estimating portion analyzes parameter changes at two distinct periodic intervals: long-term changes (overall emotional trajectory) and short-term changes (momentary emphasis points). By structuring the analysis as periodic sampling at different time scales rather than continuous analysis, the system achieves accurate emotion detection while managing computational complexity through rhythmically structured processing.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS8396708B2Facial expression representation apparatus
Publication Date: 2013.03.12 SAMSUNG ELECTRONICS CO LTD
  • US8396708B2 patent drawing
  • US8396708B2 patent drawing
  • US8396708B2 patent drawing

AI summary

An avatar facial expression representation technology is provided. The avatar facial expression representation technology estimates changes in emotion and emphasis in a user's voice from vocal information, and changes in mouth shape of the user from pronunciation information of the voice. The avatar facial expression technology tracks a user's facial movements and changes in facial expression from image information and may represent avatar facial expressions based on the result of the these operations. Accordingly, the avatar facial expressions can be obtained which are similar to actual facial expressions of the user.