Spotlight Training Latent Models Video Communication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for generating photorealistic views of dynamic scenes from novel viewpoints are computationally expensive and memory-intensive, limiting their application to static objects and environments.

Innovation Solution

A computer-implemented system and method for spotlight training of latent models of a scene, which prioritizes training based on view spotlight information to efficiently generate novel views of dynamic scenes using a neural network.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If NeRF is used to generate photorealistic views of dynamic scenes, then image quality and photorealism are improved, but computational cost and memory requirements increase significantly

Engineering Contradiction:
Improveimage qualityVSAvoidcomputational cost
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the scene into static and dynamic components, representing static objects with NeRF while representing dynamic objects with 3D avatars or meshes. This segmentation allows the system to use computationally intensive NeRF only for static portions while using lighter-weight representations for dynamic portions, thereby reducing overall computational cost while maintaining image quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different representation qualities to different parts of the scene based on their dynamics. Static regions receive high-quality NeRF representation, while dynamic regions use simplified avatars or meshes. This local differentiation optimizes the balance between image quality and computational resources by applying intensive processing only where needed.

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If NeRF is used to generate photorealistic views of dynamic scenes, then photorealism is improved, but memory requirements increase significantly

Engineering Contradiction:
ImprovephotorealismVSAvoidmemory requirements
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent segments the scene into static and dynamic components, storing NeRF data only for static objects and using more memory-efficient 3D avatars or meshes for dynamic objects. This segmentation reduces the total memory footprint while preserving photorealism in static regions where it is most visually impactful.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent allocates memory resources locally based on scene content, using high-memory NeRF representations for static regions and lower-memory avatar representations for dynamic regions. This local quality differentiation maintains photorealism where required while significantly reducing overall memory requirements.

Inventive Principle:
Principle #3Local quality

3Productivity

If traditional mesh-based representations are used, then computational efficiency is improved, but ability to represent dynamically changing environments deteriorates

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidrepresentation capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent segments the scene into static and dynamic components, applying mesh-based representations to dynamic objects (which benefit from their efficiency) while using NeRF for static objects (which benefit from their photorealism). This segmentation allows the system to leverage the computational efficiency of meshes for dynamic portions while maintaining the representation capability of NeRF for static portions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a composite representation system that combines multiple techniques (NeRF, 3D avatars, meshes) in a unified framework. Each technique is applied to appropriate scene elements based on their properties, creating a hybrid representation that achieves both computational efficiency and adaptability to dynamically changing environments.

Inventive Principle:
Principle #40Composite materials

4Measurement precision

If NeRF is trained on each frame of dynamic scenes, then temporal accuracy is improved, but computational resources and memory requirements increase significantly

Engineering Contradiction:
Improvetemporal accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent segments the training process by representing dynamic objects with 3D avatars or meshes that can be updated efficiently across frames, while using NeRF for static objects that require less frequent retraining. This segmentation enables temporal accuracy through consistent representation of dynamic elements while significantly reducing the computational resources required compared to training full NeRF on every frame.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250088618A1Spotlight training of latent models used in video communication
Publication Date: 2025.03.13 IKIN INC
  • US20250088618A1 patent drawing
  • US20250088618A1 patent drawing
  • US20250088618A1 patent drawing

AI summary

A computer-implemented method includes receiving training data with training images of a scene and associated camera extrinsics corresponding to three-dimensional (3D) camera locations and camera directions from which the training images are captured. Using the training data, a neural network is trained to represent a latent model of the scene in a latent space where the neural network is configured to synthesize scene images corresponding to novel views of the scene from queried 3D viewpoints and viewing angles. View spotlight information is received. The training of the neural network is prioritized based upon the view spotlight information.