Real-Time Image Compositing Using AI Depth Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Real-time processing of live-action and computer-generated imagery compositing is challenging due to the complexity of accurately matching depth information between the two, requiring efficient methods to determine depth values for accurate integration.
Innovation Solution
The use of auxiliary cameras to obtain stereo depth information, which is correlated with main image capture devices, and processed using pre-processing, disparity detection, feature extraction, and AI techniques like deep neural networks trained with synthetic and live-action data to generate accurate depth maps for real-time compositing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If real-time processing is used to composite CG with live action, then productivity is improved, but measurement precision of depth information deteriorates
Solution Approach 1:
The system performs preliminary depth map generation and AI model training before the actual compositing operation. Depth maps are generated in advance using auxiliary cameras and stereo vision algorithms, and AI models are pre-trained with synthetic data to accelerate real-time processing while maintaining accuracy.
Solution Approach 2:
The patent introduces an intermediary AI-based depth estimation model that bridges the gap between fast but less accurate traditional methods and slow but precise traditional compositing. The AI model serves as a mediator that provides sufficiently accurate depth information at reduced computational cost, enabling real-time performance.
2Measurement precision
If traditional methods are used to match depth information, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The system extracts the complex depth matching task from the main compositing pipeline and handles it separately using dedicated auxiliary cameras and specialized AI processing. This separates the high-precision depth estimation function from the overall compositing system, reducing the complexity burden on the main processing path.
Solution Approach 2:
The patent uses auxiliary cameras to capture depth information as a separate copy of the visual scene. This duplicated depth channel is processed independently through AI models and then integrated with the main live action footage, simplifying the overall system architecture by dedicating specific hardware to specific functions.
3Measurement precision
If more data processing steps are applied to improve depth map accuracy, then measurement precision is improved, but loss of time increases
Solution Approach 1:
The system performs computationally intensive preprocessing steps including disparity detection, feature extraction, and AI model training in advance. By completing these time-consuming operations before real-time compositing, the system maintains high depth map accuracy while ensuring real-time performance during actual production.
Solution Approach 2:
The patent dynamically adjusts processing parameters based on scene complexity and required output quality. The AI model can operate at different precision levels, allowing the system to reduce processing time for less critical scenes while maintaining high accuracy for important shots, thus balancing time consumption with depth map quality.
Data Source
AI summary
Embodiments allow live action images from an image capture device to be composited with computer generated images in real-time or near real-time. The two types of images (live action and computer generated) are composited accurately by using a depth map. In an embodiment, the depth map includes a “depth value” for each pixel in the live action image. In an embodiment, steps of one or more of feature extraction, matching, filtering or refinement can be implemented, at least in part, with an artificial intelligence (AI) computing approach using a deep neural network with training. A combination of computer-generated (“synthetic”) and live-action (“recorded”) training data is created and used to train the network so that it can improve the accuracy or usefulness of a depth map so that compositing can be improved.


