Synchronized Audio Captions Using Timestamp Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in following along with audiobooks due to the lack of synchronized text, as existing solutions require manual reference to physical books, which is cumbersome, and not all audio content has accompanying text.

Innovation Solution

A system and method that provides captions for audio content items, allowing users to opt-in for caption services, generating captions if not available, and synchronizing them with audio output on electronic devices, using timestamps and prioritization for efficient display.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If users acquire a physical book to follow along with the audiobook, then they can reference the text, but it becomes burdensome to simultaneously listen and read

Engineering Contradiction:
Improveability to follow along with audio contentVSAvoidconvenience of using audio content
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent combines audio content and text content into a single integrated system. The electronic device simultaneously outputs audio through speakers and displays synchronized text on the screen, eliminating the need for separate physical books and allowing users to follow along without the burden of coordinating two separate resources.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an electronic device as an intermediary that bridges audio and text content. This device receives audio content, generates or retrieves corresponding text, synchronizes them using timestamps, and presents both to the user in a coordinated manner, resolving the conflict between audio listening and text reading.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If captions are generated for all audio content items, then users can follow along with any audio content, but processing time and computational resources increase

Engineering Contradiction:
Improveavailability of captions for audio contentVSAvoidtime to generate captions
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent implements preliminary action by generating captions for audio content items in advance, before the user requests them. The system processes and stores captions during idle periods or when content is first added to the library, so that when a user wants to follow along, the synchronized text is already available immediately without delays.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system performs self-service by automatically generating and storing captions for audio content items without requiring manual intervention. The caption generation process runs autonomously in the background, managing its own processing queue and resource allocation, which reduces the time impact on user operations.

Inventive Principle:
Principle #25Self-service

3Loss of information

If the electronic device displays the entire caption at once, then users can see all text, but users lose track of the current position in the audio content

Engineering Contradiction:
Improvecompleteness of text informationVSAvoidability to track current position
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent segments the caption text into multiple portions based on timestamp intervals. Instead of displaying the entire caption at once, the system divides the text into manageable segments that correspond to specific time ranges in the audio content, allowing users to see relevant text while maintaining context of their position.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic updating of displayed captions based on the current playback position. As the audio progresses, the system automatically updates the displayed text portion to match the current timestamp, creating a dynamic viewing experience that adapts to the user's position in the content while maintaining complete information availability.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11347379B1Captions for audio content
Publication Date: 2022.05.31 AUDIBLE INC
  • US11347379B1 patent drawing
  • US11347379B1 patent drawing
  • US11347379B1 patent drawing

AI summary

This disclosure describes, in part, techniques for providing captions with audio content. For instance, an electronic device may receive first data representing audio content and second data representing captions that are available for the audio content. The electronic device may then select portions of the captions for display while outputting the audio content. In some instances, the electronic device selects the portions using timestamps represented by the second data. For instance, the electronic device may select a portion of the captions such that the portion of the captions begins at a first pause within the audio content and/or ends at a second pause within the audio content. In some instances, the electronic device may also display graphical elements that indicate the current location within the captions.