Avatar Lip-Sync Reconstruction From EEG and Facial Biosignals

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing lip-sync animation technologies require recorded voice data for generating speaking faces, which limits their use for patients who have difficulty speaking or in quiet situations, and they struggle to express detailed emotions and nuances.

Innovation Solution

A multimodal biosignal-based system that collects brain waves and electromyography during speaking imagination to generate avatar lip-sync animation, using a lip-sync reconstruction model to predict mouth shape and facial movements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If recorded voice data is used for lip-sync animation, then the speaking face can be generated, but the system cannot be used for patients who have difficulty speaking or in quiet situations

Engineering Contradiction:
Improveapplicability for patients with speaking difficultiesVSAvoidrequirement for direct speaking
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent replaces the mechanical/physical act of speaking (voice production) with neural signal detection and processing. By using EEG to detect brain waves during speaking imagination and processing these neural signals through AI models, the system generates lip-sync animation without requiring actual vocalization, thus enabling use for patients with speaking difficulties

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an intermediary system consisting of EEG sensors, neural signal processing algorithms, and AI models that translate internal brain activity into external lip-sync animation. This intermediary pathway allows communication of speaking intent without direct voice output, bridging the gap between internal imagination and external expression

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If brain waves during speaking are used for communication, then user intentions can be recognized, but the real-time decoding performance and recognition rate remain low

Engineering Contradiction:
Improverecognition rate of user intentionsVSAvoidreal-time decoding performance
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent employs preliminary action by pre-training AI models (such as autoencoders and classification models) offline with large datasets of brain wave patterns and corresponding speech intents. This pre-processing and pre-learning phase enables the system to achieve high recognition accuracy in real-time applications without requiring complex real-time computation, thus resolving the contradiction between precision and productivity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms brain wave data from raw neural signals into processed feature representations through signal processing techniques and neural network transformations. By changing the parameter representation (from raw EEG to extracted features), the system achieves both high recognition accuracy and efficient real-time processing

Inventive Principle:
Principle #35Parameter changes

3Reliability

If invasive brain wave measurement is used, then communication performance improves, but the system becomes expensive and difficult to use in real life

Engineering Contradiction:
Improvecommunication performanceVSAvoidusability in real life
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent adopts non-invasive, inexpensive EEG sensors and consumer-grade hardware instead of expensive invasive neural implants. By using affordable, easily deployable sensing equipment, the system achieves practical real-world usability while maintaining sufficient communication performance through sophisticated signal processing and AI algorithms

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Data Source

PatentEP4671929A1Device and method for generating avatar lip-sync animation based on multimodal biosignals
Publication Date: 2025.12.31 KOREA UNIV RES & BUSINESS FOUND
  • EP4671929A1 patent drawingFigure 1~2
  • EP4671929A1 patent drawingFigure 3~4
  • EP4671929A1 patent drawingFigure 5~6

AI summary

The present disclosure relates to a device and method for generating avatar lip-sync animation based on multimodal biosignals, The device comprises a multimodal data collection unit configured to collect data including biosignal data including brain waves when a user imagines speaking and image data; a preprocessing unit configured to preprocess the multimodal data; a feature extraction unit configured to extract feature vectors including the user's biosignal feature and facial feature from the preprocessed multimodal data; an avatar generation unit configured to generate an avatar; a lip-sync reconstruction unit configured to predict the mouth shape and facial movement when the user imagines speaking by inputting the extracted feature vectors to a pre-prepared lip-sync reconstruction model; and a lip-sync animation implementation unit for implementing an avatar lip-sync animation by applying the mouth shape and facial movement predicted by the lip-sync reconstruction unit to the avatar generated by the avatar generation unit.