Real-Time Image Compositing Using AI Depth Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Real-time processing of live-action and computer-generated imagery compositing is challenging due to the complexity of accurately matching depth information between the two, requiring efficient methods to determine depth values for accurate integration.

Innovation Solution

The use of auxiliary cameras to obtain stereo depth information, which is correlated with main image capture devices, and processed using pre-processing, disparity detection, feature extraction, and AI techniques like deep neural networks trained with synthetic and live-action data to generate accurate depth maps for real-time compositing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If real-time processing is used to composite CG with live action, then productivity is improved, but measurement precision of depth information deteriorates

Engineering Contradiction:
Improvereal-time processing speedVSAvoiddepth information accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary depth map generation and AI model training before the actual compositing operation. Depth maps are generated in advance using auxiliary cameras and stereo vision algorithms, and AI models are pre-trained with synthetic data to accelerate real-time processing while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary AI-based depth estimation model that bridges the gap between fast but less accurate traditional methods and slow but precise traditional compositing. The AI model serves as a mediator that provides sufficiently accurate depth information at reduced computational cost, enabling real-time performance.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If traditional methods are used to match depth information, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improvedepth matching accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system extracts the complex depth matching task from the main compositing pipeline and handles it separately using dedicated auxiliary cameras and specialized AI processing. This separates the high-precision depth estimation function from the overall compositing system, reducing the complexity burden on the main processing path.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses auxiliary cameras to capture depth information as a separate copy of the visual scene. This duplicated depth channel is processed independently through AI models and then integrated with the main live action footage, simplifying the overall system architecture by dedicating specific hardware to specific functions.

Inventive Principle:
Principle #26Copying

3Measurement precision

If more data processing steps are applied to improve depth map accuracy, then measurement precision is improved, but loss of time increases

Engineering Contradiction:
Improvedepth map accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs computationally intensive preprocessing steps including disparity detection, feature extraction, and AI model training in advance. By completing these time-consuming operations before real-time compositing, the system maintains high depth map accuracy while ensuring real-time performance during actual production.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent dynamically adjusts processing parameters based on scene complexity and required output quality. The AI model can operate at different precision levels, allowing the system to reduce processing time for less critical scenes while maintaining high accuracy for important shots, thus balancing time consumption with depth map quality.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11710247B2System for image compositing including training with synthetic data
Publication Date: 2023.07.25 UNITY TECH SF
  • US11710247B2 patent drawing
  • US11710247B2 patent drawing
  • US11710247B2 patent drawing

AI summary

Embodiments allow live action images from an image capture device to be composited with computer generated images in real-time or near real-time. The two types of images (live action and computer generated) are composited accurately by using a depth map. In an embodiment, the depth map includes a “depth value” for each pixel in the live action image. In an embodiment, steps of one or more of feature extraction, matching, filtering or refinement can be implemented, at least in part, with an artificial intelligence (AI) computing approach using a deep neural network with training. A combination of computer-generated (“synthetic”) and live-action (“recorded”) training data is created and used to train the network so that it can improve the accuracy or usefulness of a depth map so that compositing can be improved.