Actor-Replacement System Using Skeletal Pose Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Re-recording a video with a replacement actor or changing the speech of actors from one language to another is time-consuming and labor-intensive, often resulting in undesirable lip synchronization issues.

Innovation Solution

A computing system uses a skeletal detection model to estimate the pose of an original actor in a video, obtains images and speech of a replacement actor, and generates synthetic frames that align the replacement actor's poses and facial expressions with the original actor's, thereby creating a synthetic video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If re-recording a video with a replacement actor is performed, then the actor can be changed, but the process becomes time-consuming and labor-intensive

Engineering Contradiction:
Improveactor replacement capabilityVSAvoidvideo production time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent creates a digital twin or virtual replica of the replacement actor using image generation models and pose estimation techniques. This virtual copy can be seamlessly integrated into the original video without requiring physical re-recording, thus enabling actor replacement while avoiding time-consuming production processes

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of physical re-recording with computational methods including pose estimation, image synthesis, and video generation models. This substitution transforms a labor-intensive physical process into an automated computational workflow, significantly reducing time and effort

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If re-recording dialogue in another language is performed, then speech language can be changed, but lip synchronization issues occur

Engineering Contradiction:
Improvespeech language translation capabilityVSAvoidlip synchronization accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent generates a virtual copy of the actor's face and lips that perfectly synchronizes with the translated speech. By creating a digital replica that can be independently controlled, the system ensures that lip movements match the translated dialogue exactly, eliminating synchronization issues that occur with traditional re-recording methods

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies different processing quality levels to different parts of the video. The region around the actor's mouth and face receives enhanced processing with higher resolution image generation and more precise pose estimation, ensuring that lip synchronization maintains high accuracy while other parts of the video use standard processing

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250191614A1Actor-Replacement System for Videos
Publication Date: 2025.06.12 ROKU INC
  • US20250191614A1 patent drawing
  • US20250191614A1 patent drawing
  • US20250191614A1 patent drawing

AI summary

In one aspect, an example method includes (i) estimating, using a skeletal detection model, a pose of an original actor for each of multiple frames of a video; (ii) obtaining, for each of a plurality of the estimated poses, a respective image of a replacement actor; (iii) obtaining replacement speech in the replacement actor's voice that corresponds to speech of the original actor in the video; (iv) generating, using the estimated poses, the images of the replacement actor, and the replacement speech, synthetic frames corresponding to the multiple frames of the video that depict the replacement actor in place of the original actor, with the synthetic frames including facial expressions for the replacement actor that temporally align with the replacement speech; and (iv) combining the synthetic frames and the replacement speech so as to obtain a synthetic video that replaces the original actor with the replacement actor.