Multi-Frame Facial Expression Recognition for Jitter-Free Animation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for simulating human facial expressions in virtual characters suffer from a 'jitter phenomenon' due to inconsistent facial expression predictions across consecutive frames, and the lack of integration of image features in expression information determination leads to inaccuracies and discontinuities.

Innovation Solution

A method that involves acquiring a current frame image, a historical image, and a subsequent image, determining target facial image features by fusing these images using various algorithms, and then using a machine learning model to predict expression information, thereby incorporating the features of multiple frames to improve accuracy and continuity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If facial expression prediction is performed on single frame images, then processing speed is fast, but prediction accuracy and continuity are poor leading to jitter phenomenon

Engineering Contradiction:
Improvefacial expression prediction accuracyVSAvoidimage feature integration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges features from multiple frames (current frame, historical frames, and subsequent frames) into a unified feature representation. The multi-frame facial image feature extraction module extracts features from three consecutive frames and integrates them to obtain comprehensive facial expression features, resolving the contradiction by combining information from multiple sources to improve accuracy without excessive complexity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent performs preliminary feature extraction and integration from historical frames and subsequent frames before final expression prediction. By pre-processing and storing facial features from multiple frames in advance, the system prepares comprehensive feature data that improves prediction accuracy while maintaining efficient processing through cached feature representations

Inventive Principle:
Principle #10Preliminary action

2Reliability

If multiple frame images are integrated for expression prediction, then prediction accuracy and continuity improve, but processing time and computational load increase

Engineering Contradiction:
Improvefacial expression prediction consistencyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system extracts and stores facial image features from multiple frames in advance during video capture or preprocessing stages. These pre-extracted features are cached and readily available for rapid integration during expression prediction, reducing real-time computational load while maintaining multi-frame analysis for improved reliability

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces complex real-time multi-frame processing with optimized feature integration mechanisms. By substituting heavy computational operations with efficient feature fusion algorithms and pre-processing techniques, the system achieves reliable multi-frame prediction without proportionally increasing processing time

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20260045118A1Expression information recognition method, apparatus and device, readable storage medium and product
Publication Date: 2026.02.12 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20260045118A1 patent drawing
  • US20260045118A1 patent drawing
  • US20260045118A1 patent drawing

AI summary

The embodiment of the disclosure provides an expression information recognition method, apparatus and device, a readable storage medium and a product. The method includes: acquiring a current frame image, a historical image and a subsequent image of the current frame image, wherein the image comprises a facial image of a target object; determining target facial image feature in the current frame image based on the current frame image, the historical image and the subsequent image; and determining expression information of the target object in the current frame image based on the target facial image feature.