Novel-View Video Rendering with Scene Flow from Near-Duplicate Photos

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image processing techniques struggle to jointly infer 3D geometry, scene dynamics, and disocclusion from near-duplicate photos with unknown camera poses, resulting in temporally inconsistent and unrealistic animations.

Innovation Solution

A method that represents the scene as feature-based layered depth images augmented with scene flows, using a Transformer-type architecture to model time-varying geometry and appearance, and employs depth-aware bidirectional splatting and rendering to create 3D Moments from near-duplicate photos.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing image processing techniques are used to process near-duplicate photos, then storage space is consumed and locating desirable imagery is time-consuming, but the photos remain static and unenlivened

Engineering Contradiction:
Improvephoto animation capabilityVSAvoidprocessing system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent transforms static 2D photos into dynamic 3D space-time videos by adding temporal and spatial dimensions. Through novel view synthesis and scene motion interpolation, the system creates animated sequences from single or pairs of photos, enabling cinematic camera motions and realistic scene dynamics without requiring complex multi-camera setups

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system creates synthetic copies of scenes from near-duplicate photos by inferring 3D geometry and generating novel views. Through techniques like depth estimation, optical flow computation, and neural radiance field synthesis, the patent generates photorealistic video sequences that replicate and animate the original captured moments

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If 3D geometry and scene dynamics are jointly inferred from near-duplicate photos with unknown camera poses, then cinematic camera motion and scene animation are achieved, but temporal consistency and realism deteriorate

Engineering Contradiction:
Improvenovel view synthesis capabilityVSAvoidtemporal consistency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent optimizes multiple parameters simultaneously including camera pose, 3D point positions, scene flow vectors, and rendering weights. Through iterative refinement and loss minimization, the system adjusts these parameters to ensure temporal consistency across generated frames while maintaining photorealistic quality and realistic scene dynamics

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system employs feedback mechanisms through loss functions that compare synthesized frames against ground truth data during training. The optimization process uses gradient feedback to adjust model parameters, ensuring that generated videos maintain temporal coherence and realistic motion patterns consistent with the input photos

Inventive Principle:
Principle #23Feedback

3Productivity

If traditional frame interpolation or view synthesis methods are applied sequentially, then processing simplicity is maintained, but the result is temporally inconsistent and unrealistic animation

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidanimation realism
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent merges novel view synthesis and frame interpolation into a unified joint optimization framework. Rather than applying these methods sequentially, the system simultaneously optimizes both 3D geometry reconstruction and temporal motion interpolation, ensuring temporal consistency and realistic scene dynamics while maintaining processing efficiency through integrated loss functions and shared model components

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250218109A1Rendering Videos with Novel Views from Near-Duplicate Photos
Publication Date: 2025.07.03 GOOGLE LLC
  • US20250218109A1 patent drawing
  • US20250218109A1 patent drawing
  • US20250218109A1 patent drawing

AI summary

The technology introduces 3D Moments, a new computational photography effect. As input a pair of near-duplicate photos is taken (FIG. 1A), i.e., photos of moving subjects from similar viewpoints, which may be very common in people's photo collections. As output, the system produces a video that smoothly interpolates the scene motion from the first photo to the second, while also producing camera motion with parallax that gives a heightened sense of 3D (FIG. 1B). To achieve this effect, the scene is represented as a pair of feature-based layered depth images augmented with scene flow (306). This representation enables motion interpolation along with independent control of the camera viewpoint. The system produces photorealistic space-time videos with motion parallax and scene dynamics (322), while plausibly recovering regions occluded in the original views. Experimentation demonstrating superior performance over baselines on public benchmarks and in-the-wild photos.