Text-to-Audio Generation Using Matching Song Clips

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional text-to-speech technologies only read text in a manner similar to a human voice, lacking the creativity and engagement that music can provide.

Innovation Solution

An audio generation method that uses song clips corresponding to text, where the lyrics of the song clips match the text, allowing for the synthesis of audio that is more engaging and personalized.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional text-to-speech technology is used to convert text into speech, then the text can be converted into audio form, but the audio output lacks creativity and entertainment value

Engineering Contradiction:
Improvecreativity and entertainment valueVSAvoidimplementation complexity
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The system segments the text into multiple subtexts and matches each subtext with corresponding song clips separately. This allows the use of multiple different songs to represent different parts of the text, thereby enhancing creativity and entertainment value while managing complexity through modular processing

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses song clips that serve multiple functions: they represent the semantic meaning of text segments, provide musical entertainment, and maintain lyrical accuracy. This multi-functionality resolves the contradiction by making the audio output both creative/entertaining and meaningful

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If the text is divided into multiple subtexts and matched with different song clips, then the audio generation becomes more flexible and creative, but the processing complexity increases

Engineering Contradiction:
Improveflexibility of audio generationVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The text is divided into multiple subtexts that can be matched with different song clips, enabling flexible and creative audio generation. Each subtext is processed independently, which manages complexity through modularization while maintaining overall flexibility

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts the segmentation and matching process based on the specific text input, allowing flexible adaptation to different content while using standardized processing steps to control complexity

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If song clips are used instead of conventional speech synthesis, then user experience is enhanced with greater creativity, but the system requires access to large music databases

Engineering Contradiction:
Improveuser experience qualityVSAvoidmusic database size
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system extracts only the necessary portions (specific song clips matching specific subtexts) from the music database rather than requiring the entire database to be processed or stored in memory, thus enhancing user experience while managing database resource requirements

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system creates audio representations of text by copying and splicing relevant segments from existing song clips in the database, rather than generating entirely new audio content, which enhances creativity while efficiently utilizing the music database

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12299031B2Audio generation method, related apparatus, and storage medium
Publication Date: 2025.05.13 PETAL CLOUD TECH CO LTD
  • US12299031B2 patent drawing
  • US12299031B2 patent drawing
  • US12299031B2 patent drawing

AI summary

Embodiments of this application provide an audio generation method, a related apparatus, and a storage medium, to provide a better audio generation solution for a user. In embodiments of this application, a text is obtained, a song clip corresponding to the text is obtained through matching, and the song clip is used as audio corresponding to the text. In this way, the text can be expressed in a manner of the song clip.