3D Source Object Modeling for Lower-Cost Video Object Replacement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video synthesis technologies face challenges in constructing high-quality animation models due to insufficient visual information of the source object, leading to low-quality synthesized videos and high computational costs.

Innovation Solution

A method involving generating a three-dimensional model from multiple images and animation models from a video, fusing these models to replace the target object with the source object, reducing the need for high-quality visual information and computational resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If video synthesis technology is used to replace target object with source object, then video editing capability is improved, but computational cost increases and quality degrades when visual information of source object is insufficient

Engineering Contradiction:
Improvevideo editing capabilityVSAvoidsynthesized video quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the video synthesis process into distinct components: a three-dimensional model generation module that creates a static 3D representation from multiple images, and an animation model generation module that processes video frames to generate animation models. This segmentation allows each module to specialize in specific tasks, improving overall synthesis quality while managing computational load efficiently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a three-dimensional model as an intermediary representation between the source object images and the final synthesized video. This intermediate 3D model serves as a bridge that captures the geometric structure of the source object, enabling high-quality synthesis even when input images have limited visual information, thereby resolving the quality-degradation problem.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If traditional animation model construction is used, then video synthesis is achieved, but computational resources are excessively consumed

Engineering Contradiction:
Improvevideo synthesis efficiencyVSAvoidcomputational resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent implements a dynamic model construction approach where the three-dimensional model is generated adaptively from multiple images of the source object, and animation models are dynamically created from video frames. This dynamic construction allows the system to process only necessary data at each stage, significantly reducing computational resource consumption compared to traditional methods that process entire video datasets simultaneously.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent performs preliminary actions by pre-generating a three-dimensional model from source object images before the actual video synthesis process. This pre-processing step creates a reusable geometric representation that can be applied across multiple synthesis operations, reducing redundant computational work and improving overall efficiency while lowering resource consumption.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250232510A1Method, device, and computer program product for image processing
Publication Date: 2025.07.17 DELL PROD LP
  • US20250232510A1 patent drawing
  • US20250232510A1 patent drawing
  • US20250232510A1 patent drawing

AI summary

The present disclosure relates to a method, a device, and a computer program product for image processing. The method includes acquiring a plurality of images and a first video, wherein the plurality of images indicate visual information of a source object from a plurality of perspectives, and the first video indicates animation of a target object. The method further includes generating a three-dimensional model for the source object based on the plurality of images, and generating a plurality of animation models for the target object based on the first video. The method further includes fusing the three-dimensional model for the source object and the plurality of animation models for the target object to generate a second video for the source object, wherein in the second video, the target object in the first video is replaced with the source object.