Multimedia Data Generation with Segmented Speech and Video

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies for generating video data from user-edited text result in poor-quality videos, lacking emotional depth and user engagement.

Innovation Solution

A multimedia data generating method and apparatus that allow users to input text information, record a reading speech manually, and generate multimedia data comprising the speech and a video image matched with the text, enabling intuitive modification of multimedia segments for enhanced video making efficiency and user experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If text information is directly converted into speech by machine, then video data can be generated automatically, but the quality of the generated videos is poor

Engineering Contradiction:
Improvevideo generation efficiencyVSAvoidvideo quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent segments the video generation process into multiple independent modules: text processing module, speech recording module, video image generation module, and segmentation module. Each module handles specific tasks separately, allowing for quality optimization in each segment while maintaining overall productivity. The speech is divided into multiple speech segments corresponding to text segments, enabling precise control over each portion of the final video.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If manual speech recording is implemented, then emotional depth and user engagement improve, but the complexity of the system increases

Engineering Contradiction:
Improvespeech qualityVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent implements a universal speech acquisition module that can handle both manual speech recording and machine text-to-speech conversion through the same interface. This multi-functional module reduces system complexity by providing a unified entry point for different speech generation methods, while still enabling high-quality manual recording when needed. The module automatically selects the appropriate method based on user input or configuration.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If the multimedia data is segmented into multiple segments, then modification becomes more intuitive, but the processing complexity increases

Engineering Contradiction:
Improvemodification easeVSAvoidprocessing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent divides the multimedia data into multiple segments where each segment contains corresponding text segments, speech segments, and video image segments. This segmentation allows users to modify individual portions of the video independently, improving ease of operation. The system automatically manages the segmentation through predefined rules, reducing the perceived complexity for users while maintaining organizational structure.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements feedback mechanisms where the system automatically adjusts and optimizes the segmentation based on user interactions and modification patterns. The segmentation module receives feedback from user edits and automatically reorganizes segments to maintain coherence, reducing processing complexity while preserving ease of modification. This adaptive feedback loop allows the system to learn from user behavior and optimize its segmentation strategy.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250168464A1Multimedia Data Generating Method, Apparatus, Electronic Device, Medium, and Program Product
Publication Date: 2025.05.22 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250168464A1 patent drawing
  • US20250168464A1 patent drawing
  • US20250168464A1 patent drawing

AI summary

Disclosed are a multimedia data generating method, apparatus, electronic device, medium, and program product, applied to the field of multimedia data processing technologies. The method includes: receiving text information inputted by a user; displaying, in response to a recording trigger operation for the text information, the text information and acquiring a first reading speech of the text information; and generating a first multimedia data based on the text information and the first reading speech and displaying the first multimedia data; the first multimedia data includes the first reading speech and a video image matched with the text information; the first multimedia data includes a plurality of first multimedia segments; the plurality of first multimedia segments correspond to a plurality of text segments included in the text information, respectively. The disclose may improve quality of multimedia data generation.