Novel-View Video Rendering with Scene Flow from Near-Duplicate Photos
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image processing techniques struggle to jointly infer 3D geometry, scene dynamics, and disocclusion from near-duplicate photos with unknown camera poses, resulting in temporally inconsistent and unrealistic animations.
Innovation Solution
A method that represents the scene as feature-based layered depth images augmented with scene flows, using a Transformer-type architecture to model time-varying geometry and appearance, and employs depth-aware bidirectional splatting and rendering to create 3D Moments from near-duplicate photos.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing image processing techniques are used to process near-duplicate photos, then storage space is consumed and locating desirable imagery is time-consuming, but the photos remain static and unenlivened
Solution Approach 1:
The patent transforms static 2D photos into dynamic 3D space-time videos by adding temporal and spatial dimensions. Through novel view synthesis and scene motion interpolation, the system creates animated sequences from single or pairs of photos, enabling cinematic camera motions and realistic scene dynamics without requiring complex multi-camera setups
Solution Approach 2:
The system creates synthetic copies of scenes from near-duplicate photos by inferring 3D geometry and generating novel views. Through techniques like depth estimation, optical flow computation, and neural radiance field synthesis, the patent generates photorealistic video sequences that replicate and animate the original captured moments
2Adaptability or versatility
If 3D geometry and scene dynamics are jointly inferred from near-duplicate photos with unknown camera poses, then cinematic camera motion and scene animation are achieved, but temporal consistency and realism deteriorate
Solution Approach 1:
The patent optimizes multiple parameters simultaneously including camera pose, 3D point positions, scene flow vectors, and rendering weights. Through iterative refinement and loss minimization, the system adjusts these parameters to ensure temporal consistency across generated frames while maintaining photorealistic quality and realistic scene dynamics
Solution Approach 2:
The system employs feedback mechanisms through loss functions that compare synthesized frames against ground truth data during training. The optimization process uses gradient feedback to adjust model parameters, ensuring that generated videos maintain temporal coherence and realistic motion patterns consistent with the input photos
3Productivity
If traditional frame interpolation or view synthesis methods are applied sequentially, then processing simplicity is maintained, but the result is temporally inconsistent and unrealistic animation
Solution Approach 1:
The patent merges novel view synthesis and frame interpolation into a unified joint optimization framework. Rather than applying these methods sequentially, the system simultaneously optimizes both 3D geometry reconstruction and temporal motion interpolation, ensuring temporal consistency and realistic scene dynamics while maintaining processing efficiency through integrated loss functions and shared model components
Data Source
AI summary
The technology introduces 3D Moments, a new computational photography effect. As input a pair of near-duplicate photos is taken (FIG. 1A), i.e., photos of moving subjects from similar viewpoints, which may be very common in people's photo collections. As output, the system produces a video that smoothly interpolates the scene motion from the first photo to the second, while also producing camera motion with parallax that gives a heightened sense of 3D (FIG. 1B). To achieve this effect, the scene is represented as a pair of feature-based layered depth images augmented with scene flow (306). This representation enables motion interpolation along with independent control of the camera viewpoint. The system produces photorealistic space-time videos with motion parallax and scene dynamics (322), while plausibly recovering regions occluded in the original views. Experimentation demonstrating superior performance over baselines on public benchmarks and in-the-wild photos.


