Speech-to-Video System Using Edge Network Functions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies lack efficient and flexible methods for automatically generating high-quality short-form videos from audio inputs, especially in mobile networks, which limits their ability to adapt to diverse user contexts and preferences.

Innovation Solution

A speech-to-video system utilizing speech-processing models to generate text transcripts and audio contexts, combined with generative language models to create text summaries, and video generation models to produce enhanced short-form videos, leveraging edge network functions for quick and efficient processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional video generation methods are used, then video content can be created, but the process is time-consuming and lacks flexibility in adapting to diverse user contexts

Engineering Contradiction:
Improvevideo generation speedVSAvoidadaptation to user contexts
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system segments the video generation process into distinct functional modules: speech processing models for audio analysis, generative language models for text summary creation, and video generation models for visual content synthesis. This modular architecture enables independent optimization of each component, improving overall processing speed while maintaining adaptability through flexible model selection and configuration for different user contexts.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts processing parameters based on user context and preferences, including video duration, resolution, content style, and thematic elements. By allowing parameter changes across different generation scenarios, the system achieves both high productivity through optimized processing settings and high adaptability to diverse user requirements without requiring complete regeneration of content.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If automated video generation is implemented, then efficiency improves, but the complexity of the system increases

Engineering Contradiction:
Improveautomatic video generation efficiencyVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements a universal video generation platform that handles multiple functions through a single integrated system: speech-to-text conversion, text summarization, video synthesis, and parameter optimization. This multi-functional architecture improves productivity by consolidating operations while managing complexity through unified model management and standardized interfaces, allowing the system to adapt to various user contexts without requiring separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Manufacturing precision

If high-quality video output is achieved, then video quality improves, but processing time and computational resources increase

Engineering Contradiction:
Improvevideo output qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary processing of audio inputs using speech processing models to extract key features, generate text transcripts, and identify important segments before initiating full video generation. Generative language models create text summaries in advance that guide the video synthesis process. These preliminary actions reduce the computational burden during final video rendering, maintaining high output quality while significantly reducing overall processing time by pre-processing content and identifying critical elements that require detailed generation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240420404A1Generating enhanced video messages from captured speech
Publication Date: 2024.12.19 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20240420404A1 patent drawing
  • US20240420404A1 patent drawing
  • US20240420404A1 patent drawing

AI summary

This disclosure describes a speech-to-video system that automatically generates enhanced short-form videos from speech. For example, the speech-to-video system utilizes different speech processing models to analyze speech in audio input and determine contextual features. Additionally, the speech-to-video system utilizes various video generation models to create enhanced short-form videos using text summaries, audio contexts, user information, video parameter inputs, and/or other user contexts. In some cases, the speech-to-video system leverages components of a mobile core network to efficiently generate and deliver these features to mobile devices.