3D Human Image Mouth Synchronization via Phoneme Sequences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing virtual human image technologies are limited to scripted scenarios and lack dynamic responses, failing to provide natural interactions in applications beyond 3D games and CG movies.

Innovation Solution

A method and apparatus that generate voice response information, phoneme sequences, and mouth movement information to control the mouth and gestures of a three-dimensional human image, allowing for dynamic interactions based on user input, including facial expression identification and decorative image integration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If virtual human image technologies are used in scripted scenarios such as 3D games and CG movies, then high anthropomorphic effects can be achieved, but the system lacks dynamic responses and cannot provide natural interactions in real-world applications

Engineering Contradiction:
Improvedynamic response capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies dynamics by transitioning from static scripted scenarios to dynamic real-time interactions. The system dynamically generates mouth movement information based on phoneme sequences and controls the three-dimensional human image's mouth movements in real-time during voice response playback, enabling adaptive responses to user input rather than fixed scripted behaviors

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements feedback mechanisms by capturing user voice information, generating appropriate responses, and adjusting the three-dimensional human image's mouth movements to match the generated voice response. This closed-loop feedback enables natural interactions where the virtual human responds appropriately to user input rather than following predetermined scripts

Inventive Principle:
Principle #23Feedback

2Manufacturing precision

If mouth movement information is generated based on phoneme sequences to control the three-dimensional human image, then the anthropomorphic effect is enhanced, but the processing complexity and time required increase

Engineering Contradiction:
Improvemouth movement synchronization precisionVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-establishing the correspondence relationship between phoneme sequences and mouth movement information before actual interaction. This pre-computed mapping enables rapid generation of mouth movement controls during real-time voice response playback without requiring complex real-time calculations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces complex mechanical motion control systems with a more efficient information processing approach. Instead of directly controlling physical mouth movements through complex mechanical systems, the system uses phoneme sequence information to generate corresponding mouth movement data, substituting mechanical control with information-based control for better precision and efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11158102B2Method and apparatus for processing information
Publication Date: 2021.10.26 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11158102B2 patent drawing
  • US11158102B2 patent drawing
  • US11158102B2 patent drawing

AI summary

Embodiments of the present disclosure provide a method and apparatus for processing information. A method may include: generating voice response information based on voice information sent by a user; generating a phoneme sequence based on the voice response information; generating mouth movement information based on the phoneme sequence, the mouth movement information being used for controlling a mouth movement of a displayed three-dimensional human image when playing the voice response information; and playing the voice response information, and controlling the mouth movement of the three-dimensional human image based on the mouth movement information.