Skeleton-Based Video Effects and Background Replacement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current mobile devices lack the computational resources to efficiently perform complex filtering operations on live video streams, particularly for skeleton detection and tracking, which is essential for applying filters and recognizing poses and gestures in real-time.

Innovation Solution

The development of a method for generating a multi-view interactive digital media representation (MVIDMR) that includes skeleton detection and background replacement, allowing for the application of filters and effects to live video streams by analyzing spatial relationships between images and using location information to create an immersive and interactive viewing experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If complex filtering operations are applied to live video streams, then visual effects quality is improved, but computational resource consumption increases

Engineering Contradiction:
Improvevisual effects qualityVSAvoidcomputational resource consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The video processing is segmented into keypoint detection, skeleton generation, and filter application stages. Only key skeletal points are processed rather than entire video frames, reducing computational load while maintaining visual effect quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A skeleton representation serves as an intermediary between the original video stream and the applied filters. The skeleton abstracts the video data into essential pose information, enabling efficient filter application with reduced computational resources.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If skeleton detection is performed on live video streams, then pose recognition capability is improved, but processing speed decreases

Engineering Contradiction:
Improvepose recognition capabilityVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The patent extracts only the essential skeletal keypoints from video frames rather than processing entire frames. This extraction approach maintains accurate pose recognition while significantly reducing processing time and computational requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of processing complete video frames, the system performs partial action by detecting only critical skeletal points and generating simplified skeleton representations, achieving sufficient pose recognition speed for real-time applications.

Inventive Principle:
Principle #16Partial or excessive action

3Manufacturing precision

If complex filtering operations are applied to video footage, then visual effect quality is improved, but device capability requirements increase

Engineering Contradiction:
Improvevisual effect qualityVSAvoiddevice capability requirements
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent creates a simplified skeletal copy of the video content that can be processed with basic filtering operations. This copy contains essential pose information needed for visual effects without requiring complex processing of the original high-resolution video footage.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10855936B2Skeleton-based effects and background replacement
Publication Date: 2020.12.01 FUSION INC
  • US10855936B2 patent drawing
  • US10855936B2 patent drawing
  • US10855936B2 patent drawing

AI summary

Various embodiments of the present invention relate generally to systems and methods for analyzing and manipulating images and video. In particular, a multi-view interactive digital media representation (MVIDMR) of a person can be generated from live images of a person captured from a hand-held camera. Using the image data from the live images, a skeleton of the person and a boundary between the person and a background can be determined from different viewing angles and across multiple images. Using the skeleton and the boundary data, effects can be added to the person, such as wings. The effects can change from image to image to account for the different viewing angles of the person captured in each image.