Generative Model Video Generation Via Conditioning Parameters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current rendering engines struggle to provide immersive and dynamic video representations of locations in response to user queries, lacking the ability to accurately depict scenes with specific conditions such as time, weather, and crowd levels.

Innovation Solution

A computer platform utilizing a generative machine-learned model, such as a neural radiance field (NeRF), to generate videos of locations based on user queries. The platform receives queries, generates conditioning parameters, and uses these parameters to create immersive videos that accurately depict scenes with specified conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional rendering engines are used to create 3D scenes, then real-time viewpoint changes can be achieved, but the ability to accurately depict scenes with specific conditions (time, weather, crowd levels) is insufficient

Engineering Contradiction:
Improveability to depict scenes with specific conditionsVSAvoidaccuracy of scene representation
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces conditioning parameters as an intermediary between user queries and the generative model. These parameters (capturing time, weather, crowd levels) mediate the translation of natural language queries into accurate scene representations, enabling the system to depict specific conditions reliably

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the parameters of the generative model by conditioning it on specific parameters extracted from user queries. By adjusting these conditioning parameters (such as time of day, weather conditions, crowd density), the model can accurately represent different scene conditions while maintaining reliability

Inventive Principle:
Principle #35Parameter changes

2Use of energy by moving object

If generative machine-learned models are used to create videos, then computational resources are saved compared to traditional rendering engines, but the complexity of generating accurate scene representations increases

Engineering Contradiction:
Improvecomputational resource efficiencyVSAvoidcomplexity of scene generation process
Core Design Contradiction:
Use of energy by moving objectVSDevice complexity

Solution Approach 1:

The system performs preliminary action by pre-processing user queries to extract conditioning parameters before generating the video. This preliminary extraction and structuring of parameters simplifies the subsequent generation process, reducing the overall complexity while maintaining computational efficiency

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the video generation process into distinct stages: query processing, conditioning parameter extraction, model conditioning, and video generation. This segmentation breaks down the complex task into manageable components, reducing system complexity while preserving resource efficiency

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250190761A1Generation of Video for a Location Via a Generative Machine-Learned Model
Publication Date: 2025.06.12 GOOGLE LLC
  • US20250190761A1 patent drawing
  • US20250190761A1 patent drawing
  • US20250190761A1 patent drawing

AI summary

A computer platform for generating a video includes one or more memories to store instructions and one or more processors to execute the instructions to perform operations, the operations including: receiving a query from a user relating to a location; in response to receiving the query, generating conditioning parameters based at least in part on the query, wherein the conditioning parameters provide values for one or more conditions associated with a scene to be rendered at the location; generating, using a generative machine-learned model, the video, wherein the video depicts the scene at the location and with the values for the one or more conditions; and providing the video for presentation to the user.