Audio-Driven Video Content Generation for Personalized Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio generation technologies are limited and complex, leading to poor conversion and lack of personalization in audio content creation.

Innovation Solution

A method and apparatus for content generation that includes presenting an audio edit panel with a control, obtaining text and audio based on a content entity using machine learning models, and adding them to create a video with overlapped text and audio, allowing for personalized and interesting audio playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If current audio generation technologies are used, then audio content can be generated, but the process is limited and complex with poor conversion

Engineering Contradiction:
Improveease of audio generationVSAvoidcomplexity of audio generation process
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The patent combines text generation and audio generation into a single integrated process. The system generates text content and converts it to audio automatically, merging what were previously separate steps into one unified operation that simplifies the overall process.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The audio generation control is designed to handle multiple functions: it can generate text, convert text to audio, and apply audio effects simultaneously. This multi-functional approach eliminates the need for separate tools for each operation, reducing complexity while maintaining ease of use.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If current audio generation technologies are used, then audio content can be generated, but personalization is lacking

Engineering Contradiction:
Improvepersonalization of audio contentVSAvoidefficiency of audio content creation
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system dynamically adjusts audio parameters such as pitch, speed, and volume based on the generated text content and user preferences. This dynamic adaptation allows personalization without requiring manual configuration, maintaining high productivity while enabling customized audio output.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements automatic adjustment of audio parameters including pitch, playback speed, and volume based on the text content and context. These parameter changes enable personalized audio generation without manual intervention, preserving efficiency while enhancing adaptability.

Inventive Principle:
Principle #35Parameter changes

3Extent of automation

If manual audio creation is required, then customization is possible, but user engagement and contribution decrease

Engineering Contradiction:
Improveautomation of audio generationVSAvoidsimplicity of audio editing
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The system performs audio generation, text creation, and parameter adjustment automatically without requiring user intervention. This self-service capability maintains operational simplicity while providing highly automated content creation, resolving the contradiction between automation and ease of use.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250372127A1Content generation
Publication Date: 2025.12.04 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250372127A1 patent drawing
  • US20250372127A1 patent drawing
  • US20250372127A1 patent drawing

AI summary

According to embodiments of the disclosure, a method, an apparatus, a device and a storage medium for content generation are provided. The method includes: in response to an audio edit request, presenting an audio edit panel comprising at least an audio generating control; in response to detecting a trigger on the audio generating control, obtaining a first text and a first audio corresponding to the first text, at least one of the first text or the first audio being determined based on a content entity to be edited; and adding the first text and the first audio into the content entity to obtain a first video, wherein the first text is presented overlapped on the content entity, and the first audio is configured to be at least a part of an audio corresponding to the first video.