3D Virtual Portrait Mouth Shape Control via Speech Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for speech and mouth shape synchronization in three-dimensional virtual portraits require manual key frame setting by professionals, which is time-consuming and not real-time, relying heavily on technical skill and lacking automation.

Innovation Solution

A method and apparatus that automatically generate a mouth shape control parameter sequence using speech segments, processed through convolutional neural networks and recurrent neural networks, to control the mouth shape of three-dimensional virtual portraits in real-time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual key frame setting by professional technicians is used, then speech and mouth shape synchronization quality is improved, but production time and labor cost increase significantly

Engineering Contradiction:
Improvemouth shape synchronization qualityVSAvoidproduction time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of key frame setting by technicians with an automated neural network system. The CNN-LSTM model automatically extracts speech features and generates mouth shape control parameters, eliminating the need for manual frame-by-frame adjustment while maintaining synchronization quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service automation where the neural network model independently processes speech input and generates corresponding mouth shape control parameters without human intervention. The automated pipeline includes speech segmentation, feature extraction, and parameter generation all performed by the system itself.

Inventive Principle:
Principle #25Self-service

2Manufacturing precision

If manual key frame setting is used, then control precision over mouth shape is improved, but automation level deteriorates

Engineering Contradiction:
Improvemouth shape control precisionVSAvoidautomation level
Core Design Contradiction:
Manufacturing precisionVSExtent of automation

Solution Approach 1:

The patent substitutes manual control mechanisms with an automated neural network system. The CNN-LSTM architecture automatically learns the mapping between speech features and mouth shape parameters, achieving both high automation and precise control through intelligent algorithms rather than manual adjustment.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system transforms the control approach by changing from manual parameter adjustment to automated parameter generation. The neural network dynamically generates mouth shape control parameters based on speech content, achieving precise control through learned patterns rather than manual specification.

Inventive Principle:
Principle #35Parameter changes

3Extent of automation

If conventional animation engines are used for automatic mouth shape generation, then automation level is improved, but real-time performance deteriorates

Engineering Contradiction:
Improveautomation levelVSAvoidreal-time performance
Core Design Contradiction:
Extent of automationVSSpeed

Solution Approach 1:

The patent segments the speech signal into smaller units and processes them through the neural network model. This segmentation approach, combined with the efficient CNN-LSTM architecture, enables faster processing compared to conventional animation engines while maintaining automation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the computational approach by using a specialized neural network model optimized for real-time processing. The CNN-LSTM architecture with speech feature extraction enables faster generation of mouth shape parameters compared to traditional animation engines, achieving both automation and real-time performance.

Inventive Principle:
Principle #35Parameter changes

4Manufacturing precision

If frame-by-frame manual operation is used, then mouth shape synchronization quality is improved, but labor intensity increases

Engineering Contradiction:
Improvesynchronization qualityVSAvoidlabor intensity
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The system performs self-service by automatically generating mouth shape control parameters from speech input without requiring manual frame-by-frame operation. The neural network model independently completes the synchronization task, eliminating labor-intensive manual adjustments while maintaining high synchronization quality.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual frame-by-frame operation with an automated neural network system. The CNN-LSTM model automatically processes speech and generates corresponding mouth shape parameters, substituting manual labor with intelligent automation while preserving synchronization quality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11308671B2Method and apparatus for controlling mouth shape changes of three-dimensional virtual portrait
Publication Date: 2022.04.19 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11308671B2 patent drawing
  • US11308671B2 patent drawing
  • US11308671B2 patent drawing

AI summary

Embodiments of the present disclosure relate to a method and apparatus for controlling mouth shape changes of a three-dimensional virtual portrait, relating to the field of cloud computing. The method may include: acquiring a to-be-played speech; sliding a preset time window at a preset step length in the to-be-played speech to obtain at least one speech segment; generating, based on the at least one speech segment, a mouth shape control parameter sequence for the to-be-played speech; and controlling, in response to playing the to-be-played speech, a preset mouth shape of the three-dimensional virtual portrait to change based on the mouth shape control parameter sequence.