Text-to-Speech Snippet Fusion for Expressive Audio Output

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional text-to-speech systems lack the ability to incorporate audio snippets from media content items, such as music tracks, to enhance expressiveness and create more interesting or exciting output.

Innovation Solution

A text-to-speech system that includes a forced alignment data store and a combining engine to identify and combine audio snippets from multiple tracks based on forced alignment data, musical style similarities, and quality attributes, creating a more expressive and engaging audio output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional text-to-speech systems are used, then the system structure is simple, but the expressiveness and interest of the audio output are limited

Engineering Contradiction:
Improveexpressiveness of audio outputVSAvoidsystem structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent combines traditional speech synthesis with audio snippet extraction from media content items. The combining engine integrates synthesized speech segments with extracted audio snippets (music, sound effects, voice samples) to create enriched audio output. This merging of multiple audio sources directly addresses the contradiction by adding expressiveness and interest while maintaining a manageable system structure through modular integration.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system segments the audio output into distinct components: synthesized speech portions and extracted audio snippet portions. The combining engine processes these segments separately and then integrates them. This segmentation allows the system to maintain simplicity in individual components while achieving complexity in the overall output, resolving the contradiction between expressiveness and system simplicity.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If audio snippets from multiple tracks are incorporated, then the audio output becomes more interesting, but the processing time and computational resources increase

Engineering Contradiction:
Improveinterest and expressiveness of audio outputVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing media content items to extract and store audio snippets along with their metadata (duration, start time, end time, content description). This pre-extraction and indexing of audio snippets enables faster retrieval and integration during the main processing task, significantly reducing the processing time required when generating interesting audio output.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a database of pre-extracted audio snippets that can be quickly retrieved and reused. Instead of processing entire media files in real-time, the system copies and reuses pre-extracted audio snippet portions, reducing computational resources and processing time while maintaining the interest and expressiveness of the output.

Inventive Principle:
Principle #26Copying

3Measurement precision

If forced alignment data is used to precisely locate audio snippets, then the accuracy of text-to-speech conversion is improved, but the data processing complexity increases

Engineering Contradiction:
Improveaccuracy of audio snippet locationVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system introduces forced alignment data as an intermediary structure that bridges the text input and audio snippet extraction processes. This intermediary data format (containing text segments with corresponding time intervals and audio location information) simplifies the matching process between text and audio, providing precise location accuracy while avoiding complex real-time alignment computations during the main processing task.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12437744B2Text-to-speech from media content item snippets
Publication Date: 2025.10.07 SPOTIFY
  • US12437744B2 patent drawing
  • US12437744B2 patent drawing
  • US12437744B2 patent drawing

AI summary

A text-to-speech engine creates audio output that includes synthesized speech and one or more media content item snippets. The input text is obtained and partitioned into text sets. A track having lyrics that match a part of one of the text sets is identified. The location of the track's audio that contains the lyric is extracted based on forced alignment data. The extracted audio is combined with synthesized speech corresponding to the remainder of the input text to form audio output.