Text-to-3D Video Generation Using Parallel Multi-View Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing 3D asset generation methods are slow and inefficient, requiring significant time to generate and edit 3D assets, leading to prolonged user wait times for visualization.

Innovation Solution

A system utilizing two machine learning models to generate and visualize 3D assets quickly by generating multi-view images in parallel, allowing for rapid rotation and editing based on text input, with a predetermined camera offset between image sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If existing 3D asset generation methods are used, then high quality 3D assets can be generated, but the generation time is very long (30 minutes to 2 hours)

Engineering Contradiction:
Improve3D asset qualityVSAvoidgeneration time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the 3D asset generation process into multiple independent view angle generations (front, back, left, right views). Each view angle is generated separately by the machine learning model, allowing parallel processing and significantly reducing total generation time while maintaining quality through focused optimization of each segment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by generating multiple view angle images simultaneously in parallel before assembling them into the final 3D asset. This parallel preliminary generation of all required views eliminates sequential waiting time and accelerates the overall process while preserving quality through comprehensive view coverage.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If existing 3D asset generation methods are used, then detailed 3D assets can be created, but user wait time for visualization is prolonged

Engineering Contradiction:
Improve3D asset detailVSAvoiduser wait time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent divides the visualization task into segmented view angle images (front, back, left, right views) that are generated in parallel. Users can immediately visualize the segmented views without waiting for complete 360-degree rendering, reducing perceived wait time while maintaining detailed quality through focused generation of each view segment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system creates multiple copies of the 3D asset from different view angles simultaneously. Instead of generating one complete slow rendering, it produces multiple view angle copies in parallel that can be immediately displayed to users, reducing wait time while preserving detail through redundant generation approaches.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If existing 3D asset generation methods are used, then accurate 3D models can be produced, but editing and modification time is increased

Engineering Contradiction:
Improve3D model accuracyVSAvoidediting time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the 3D model into independent view angle images that can be individually edited and modified. Users can edit specific view segments without reprocessing the entire model, maintaining accuracy through focused editing while significantly reducing editing time through modular manipulation of segmented views.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12568287B2Generating three-dimensional videos based on text using machine learning models
Publication Date: 2026.03.03 LEMON INC(GB)
  • US12568287B2 patent drawing
  • US12568287B2 patent drawing
  • US12568287B2 patent drawing

AI summary

The present disclosure describes techniques for generating three-dimensional videos based on text using machine learning models. Text and inputting data indicative of a set of multi-view images are input into a machine learning model. Content of the set of multi-view images is associated with the input text. The machine learning model comprises a plurality of sub-models corresponding to a plurality of sets of camera parameters. A plurality of sets of multi-view images is generated based on corresponding camera parameters by the plurality of sub-models. The plurality of sub-models are configured to run in parallel to generate the plurality of sets of multi-view images. A three-dimensional (3D) video is generated based on the plurality of sets of multi-view images. Content of the 3D video is associated with the input text.