Video Response Generation with Synchronized Facial Expressions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current virtual assistant systems fail to mimic human-like communication by not synchronizing facial expressions, eye movements, and voice with user interactions, lacking personalization and a genuine human-to-human experience.

Innovation Solution

A method and system that generate video responses by receiving a visual image of a character from the user, creating a frontal face, and producing synchronized audio and video sequences with matching facial expressions, using neural networks for text-to-audio and video conversion, and modulating voice based on the character's gender.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If virtual assistant systems use traditional audio-only or text-based interfaces, then system complexity is reduced, but the human-like communication experience deteriorates

Engineering Contradiction:
Improvehuman-like communication experienceVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent combines multiple components (face generation module, audio sequence generation, video sequence generation with lip-sync, eye movement simulation, and facial expression mapping) into an integrated virtual assistant system. This merging of previously separate functions creates a cohesive human-like communication experience while managing system complexity through modular architecture.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The virtual assistant system performs multiple functions simultaneously: generating frontal faces from images, synthesizing audio responses, creating synchronized video sequences with lip movements, simulating eye movements, and mapping facial expressions. This multi-functionality enables comprehensive human-like interaction without requiring multiple separate systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If the system generates synchronized facial expressions, eye movements, and lip movements, then the naturalness of interaction is improved, but the computational complexity increases

Engineering Contradiction:
Improvenaturalness of interactionVSAvoidcomputational complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-generating frontal faces from input images and pre-mapping facial expressions to video sequences. The audio sequences are generated and synchronized with video frames in advance, allowing the system to deliver natural-looking interactions without real-time computational burden during actual communication.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces intermediary components such as the facial expression mapping module that translates audio sequences into corresponding video sequences with synchronized lip movements. Eye movement simulation acts as an intermediary between the virtual assistant's attention state and visual output, simplifying the overall coordination complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If the system personalizes the virtual assistant's appearance and characteristics, then user engagement is improved, but the data processing requirements increase

Engineering Contradiction:
Improvepersonalization capabilityVSAvoiddata processing requirements
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system applies local quality by allowing users to personalize specific aspects of the virtual assistant (such as selecting a character image for face generation) without requiring complete customization of all system parameters. This selective personalization maintains user engagement while managing data processing requirements.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10740391B2System and method for generation of human like video response for user queries
Publication Date: 2020.08.11 WIPRO LTD
  • US10740391B2 patent drawing
  • US10740391B2 patent drawing
  • US10740391B2 patent drawing

AI summary

Disclosed herein is a method and a video generator for generating video response to user queries. The video generator receives a visual image of a character of interest from the user and generates a frontal face of the visual image. Further, facial expressions of the character of interest are mapped with an audio/video sequence of one or more textual responses for generating a human like video response to the user queries. In an embodiment, the video generator detects gender of the character of interest, and modulates and matches voice of the video response based on the gender of the character of interest. The instant method can synthesize a video with the face of a character of interest to the user, thereby providing a wholesome communication experience to the user.