3D Skeleton Retargeting for Editable Volumetric Video

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for generating and utilizing 3D shapes of subjects in volumetric video are cumbersome and costly, requiring manual adjustments and significant effort, especially when errors or partial reshooting is needed.

Innovation Solution

A video processing device that estimates a 3D skeleton of a subject from multi-viewpoint images, applies this skeleton to separate the subject from the background, and generates 3D data using skeleton data for flexible editing and retargeting.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If volumetric video is generated using multi-viewpoint images to create accurate 3D shapes, then manufacturing precision is improved, but device complexity and ease of operation deteriorate due to requiring manual modification of multiple images when errors occur

Engineering Contradiction:
Improve3D shape accuracyVSAvoidease of utilization
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent extracts the essential 3D structural information by generating a skeleton model from multi-viewpoint images, separating the core geometric data from the complete image set. This skeleton can then be applied to multiple views without needing to manually adjust each original image, thus improving ease of operation while preserving 3D accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a reusable skeleton model that can be copied and applied across different viewpoints and scenes. Instead of manually modifying multiple original images when corrections are needed, the skeleton serves as a master template that can be replicated, significantly reducing the operational burden while maintaining consistent 3D shape accuracy.

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If complete volumetric video production is performed with multiple multi-viewpoint cameras, then manufacturing precision is improved, but loss of time and productivity worsen when only partial scenes need reshooting

Engineering Contradiction:
Improve3D shape accuracyVSAvoidreshooting time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the volumetric video production into independent skeleton extraction and application phases. The skeleton can be generated once from multi-viewpoint images and then independently applied to different scenes or time periods, allowing partial scenes to be updated without requiring complete reshooting, thus reducing time loss while maintaining 3D accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary skeleton extraction from multi-viewpoint images in advance, creating a reusable 3D structural template. This preliminary action allows subsequent scene modifications or reshoots to focus only on specific portions, rather than requiring complete reprocessing of all multi-viewpoint data, thereby reducing time loss while preserving manufacturing precision.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If manual modification of multi-viewpoint images is performed to correct 3D shape errors, then manufacturing precision is improved, but productivity deteriorates due to the large workload

Engineering Contradiction:
Improve3D shape accuracyVSAvoidworkload efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent extracts the essential 3D structural information into a skeleton model, separating the core geometric corrections from the complex multi-viewpoint image data. This allows precision improvements to be made by modifying the skeleton once, rather than manually adjusting multiple images, dramatically improving productivity while maintaining 3D shape accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses the skeleton model as a master template that can be copied and applied across multiple viewpoints and scenes. Corrections made to the skeleton are automatically replicated, eliminating the need for repetitive manual modification of each image, thus improving productivity while ensuring consistent manufacturing precision across all views.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250391139A1Video processing device, video processing method, and program
Publication Date: 2025.12.25 SONY GROUP CORP
  • US20250391139A1 patent drawing
  • US20250391139A1 patent drawing
  • US20250391139A1 patent drawing

AI summary

A video processing device according to an aspect of the present disclosure includes: an estimation unit that estimates a 3D skeleton of a subject on the basis of multi-viewpoint images obtained by shooting the subject from a plurality of viewpoints; an application unit that applies a 3D skeleton of the subject estimated by the estimation unit to the subject included in other image different from the multi-viewpoint images and separated from a background in the image; and a generation unit that generates 3D data of a subject to which a 3D skeleton is applied by the application unit.