Audio-Driven Face Modification for Synthetic Emotion in Pre-Recorded Video

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods fail to effectively insert synthetic emotions into pre-recorded video streams in a way that appears natural and lifelike, especially when the user's face is decoupled from their real-time appearance, and transitioning between emotions in video clips is difficult and time-consuming.

Innovation Solution

Employing faceswap or facewarp techniques to modify pre-recorded videos with detected real-time emotions, allowing for near-real-time integration of facial expressions and tone, using pre-recorded emotion clips or facial movement rules to synchronize audio and video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If pre-recorded video clips are used to convey emotions, then user privacy is protected by decoupling the user from their real-time appearance, but the video lacks lifelike emotional expression and tone

Engineering Contradiction:
Improveprivacy protectionVSAvoidemotional expression
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The system segments the video generation process into separate components: a pre-recorded video base (for privacy protection) and separately generated emotional facial expressions (for lifelike conveyance). These segments are then composite together, allowing the video to maintain privacy while gaining emotional expression capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system merges two distinct video streams: the original pre-recorded video and a generated emotional expression video. This combination allows the final output to preserve the privacy benefits of pre-recorded video while incorporating the emotional expressiveness of real-time facial analysis.

Inventive Principle:
Principle #5Merging (Combining)

2Loss of information

If real-time camera is used to capture user emotions, then lifelike emotional expression is achieved, but user privacy is compromised by showing real-time appearance

Engineering Contradiction:
Improveemotional expressionVSAvoidprivacy protection
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

Instead of using the user's real-time camera feed, the system creates a synthetic copy of emotional expressions by analyzing audio tone and generating corresponding facial movements on a pre-recorded video avatar. This copying approach preserves privacy while conveying emotion.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If multiple pre-recorded video clips with different emotions are prepared, then emotional variety is achieved, but the transition between clips is difficult and time-consuming to synchronize

Engineering Contradiction:
Improveemotional varietyVSAvoidvideo synchronization
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system dynamically generates emotional expressions in real-time based on audio analysis rather than selecting from pre-recorded clips. This dynamic approach allows seamless transitions between emotions as the audio tone changes, eliminating the synchronization complexity of managing multiple pre-recorded video clips.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes emotional parameters continuously by analyzing audio tone variations and adjusting facial expression parameters accordingly. This parameter-based control enables smooth emotional transitions without the discrete clip-switching complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12361750B2Synthetic emotion in continuously generated voice-to-video system
Publication Date: 2025.07.15 EMC IP HLDG CO LLC
  • US12361750B2 patent drawing
  • US12361750B2 patent drawing
  • US12361750B2 patent drawing

AI summary

One example method includes collecting an audio segment that includes audio data generated by a user, analyzing the audio data to identify an emotion expressed by the user, computing start and end indices of a video segment, selecting video data that shows the emotion expressed by the user, using the video data and the start and end indices of the video segment to modify a face of the user as the face appears in the video segment so as to generate modified face frames, and stitching the modified face frames into the video segment to create a modified video segment with the emotion expressed by the user, and the modified video segment includes the audio data generated by the user.