Synthetic Video Generation via Text-to-Speech and User Attributes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Content creators face challenges in efficiently generating videos using structured data, as traditional video production methods are time-consuming and costly.

Innovation Solution

A computing system is configured to obtain user attributes, retrieve structured data, generate a textual description, transform it into synthesized speech using a text-to-speech engine, and create a synthetic video for targeted advertisements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional video production methods are used, then video quality and authenticity are maintained, but time consumption and production costs increase significantly

Engineering Contradiction:
Improvevideo generation efficiencyVSAvoidvideo production time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent uses text-to-speech synthesis to generate synthetic speech that copies human speech patterns, and video synthesis models to create synthetic video content that replicates real video appearances. This allows rapid generation of video advertisements without traditional filming, acting, or editing processes, dramatically reducing production time while maintaining visual authenticity.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces mechanical video production processes (filming, physical editing, manual post-production) with automated computational systems. Text-to-speech engines convert written scripts to speech automatically, and video synthesis models generate visual content algorithmically, eliminating the need for physical production equipment and manual labor involved in traditional video creation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If traditional video production methods are used, then authentic video content is created, but production costs increase due to equipment, personnel, and post-processing requirements

Engineering Contradiction:
Improvevideo generation efficiencyVSAvoidvideo production complexity
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

The system copies human speech through text-to-speech synthesis and replicates video content through synthesis models, eliminating the need for expensive production equipment, studio facilities, and professional crews. The synthetic video maintains visual authenticity while removing all cost-associated elements of traditional production.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system enables self-service video production where content can be generated automatically from text inputs without requiring skilled operators for filming, directing, or editing. The automated synthesis processes handle all production aspects independently, reducing dependency on specialized personnel and simplifying the manufacturing process.

Inventive Principle:
Principle #25Self-service

3Productivity

If synthetic video generation is implemented, then production time and costs are reduced, but video authenticity and indistinguishability from real video must be maintained

Engineering Contradiction:
Improvevideo generation efficiencyVSAvoidvideo synthesis quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent employs advanced text-to-speech engines that replicate human speech characteristics including tone, pitch, and intonation, and video synthesis models that generate visuals indistinguishable from real footage. This copying approach ensures synthetic videos maintain high fidelity and authenticity despite being artificially generated.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system adjusts multiple parameters in the synthesis process including audio frequency characteristics, video resolution, frame rates, and temporal consistency to ensure the synthetic output matches real video standards. By optimizing these parameters, the system maintains manufacturing precision while enabling rapid production.

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables the efficient and cost-effective generation of synthetic videos that are indistinguishable from real videos, saving time and resources typically required in traditional video production.

Implementation Method 1

transforming, using a text-to-speech engine, the textual description of the structured data into synthesized speech

Methodology Applied
Scientific EffectText-to-speech synthesis:

Data Source

PatentUS12334115B2Method and system for generating synthetic video advertisements
Publication Date: 2025.06.17 ROKU INC
  • US12334115B2 patent drawing
  • US12334115B2 patent drawing
  • US12334115B2 patent drawing

AI summary

In one aspect, an example method includes (i) obtaining a set of user attributes for a user of a content-presentation device; (ii) based on the set of user attributes, obtaining structured data and determining a textual description of the structured data; (iii) transforming, using a text-to-speech engine, the textual description of the structured data into synthesized speech; and (iv) generating, using the synthesized speech and for display by the content-presentation device, a synthetic video of a targeted advertisement comprising the synthesized speech.