Text-to-Speech Audio Track Multimedia Insertion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis from text methods face challenges in seamlessly integrating multimedia elements into audio tracks, often leading to user dissatisfaction due to interruptions or truncation of content during advertisement insertion.

Innovation Solution

A method and system that identify reading pauses in text to generate markers for precise timing of multimedia content insertion within an audio track, allowing for smooth integration of advertisements without truncating words or concepts, using a text-to-speech conversion system and SSML format.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If advertisement is inserted at the beginning or end of audio track, then advertisement delivery is simple, but user listening experience deteriorates due to forced interruptions

Engineering Contradiction:
Improveease of advertisement insertionVSAvoiduser listening experience
Core Design Contradiction:
Ease of manufactureVSEase of operation

Solution Approach 1:

The system performs preliminary analysis of the audio track structure, identifying pause points and segment boundaries before advertisement insertion. This allows the advertisement to be inserted at optimal locations that minimize disruption to the listening experience while maintaining ease of implementation.

Inventive Principle:
Principle #10Preliminary action

2Extent of automation

If advertisement is inserted after predetermined time interval, then advertisement delivery is automated, but content integrity deteriorates due to word truncation or discourse interruption

Engineering Contradiction:
Improveautomation of advertisement insertionVSAvoidcontent integrity
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The system continuously monitors the audio track playback position and uses feedback from the identified pause points to determine the optimal insertion moment. This feedback mechanism ensures that advertisements are inserted at appropriate boundaries without truncating words or interrupting articulated discourse, maintaining both automation and content integrity.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary analysis to identify all suitable pause points and segment boundaries in the audio track before insertion occurs. This pre-planning ensures that when automated insertion is triggered, it occurs at the next appropriate boundary point, preventing content truncation while maintaining automation.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If advertisement is inserted during audio track reproduction, then user engagement is maintained, but content continuity deteriorates due to interruptions

Engineering Contradiction:
Improveuser engagementVSAvoidcontent continuity
Core Design Contradiction:
Ease of operationVSStability of the object's composition

Solution Approach 1:

The system applies different qualities to different parts of the audio track by identifying specific pause points and segment boundaries. Advertisements are inserted only at locations with appropriate local characteristics (natural pauses, segment boundaries), ensuring that the insertion does not disrupt the overall continuity and stability of the content while maintaining user engagement.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP4239558B1Method and system for inserting multimedia content while playing an audio track generated from a website
Publication Date: 2024.10.16 AUDIOBOOST SRL
  • EP4239558B1 patent drawingFigure 1
  • EP4239558B1 patent drawingFigure 2~3
  • EP4239558B1 patent drawingFigure 4~5

AI summary

Described herein is a method for inserting a multimedia content during the playing of an audio track generated from a text, said method comprising: providing a web page containing a text to be played as audio; identifying a plurality of characters in said text to be played as audio, each character corresponding to a respective reading pause; generating a list of markers, each marker in said list of markers corresponding to a respective character of said plurality of characters; generating a second text, as a function of said text to be played, suitable for being converted into an audio track; generating an audio track as a function of said second text by means of a text-to-speech conversion system, and associating with each marker in said list of markers a respective variable indicating a secondage at which a respective character, and hence a respective reading pause, occurs during the playing of said audio track; selecting a secondage associated with a respective marker; playing said audio track; pausing said audio track at said selected time instant; playing a multimedia content; resuming the playing of said audio track from said selected time instant.