Automated Multimedia Transcript Formatting via Phoneme Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for creating transcripts of multimedia content are expensive, time-consuming, and often inaccurate, requiring manual transcription and alignment, which limits their accessibility and searchability.

Innovation Solution

The development of systems and methods that automatically generate formatted, readable transcripts from multimedia content using acoustic and lexical features, extracting phoneme-level transcriptions, inserting punctuation, capitalizing text, and determining paragraph breaks, thereby creating a time-aligned and searchable transcript.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual transcription and alignment is used to create transcripts, then accuracy can be maintained, but the process becomes expensive and time-consuming

Engineering Contradiction:
Improvetranscription accuracyVSAvoidtranscription time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical transcription process with an automated speech-to-text system that uses acoustic and lexical features to generate transcriptions. The system automatically extracts phoneme-level transcriptions from audio data, aligns them with video content, and applies formatting rules without human intervention, thereby eliminating the time-consuming manual process while maintaining accuracy through algorithmic precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces an automated transcription system as an intermediary between the audio-visual content and the final transcript. This intermediary system processes the raw audio data through multiple stages (phoneme extraction, alignment, punctuation insertion, capitalization) to produce accurate transcriptions automatically, resolving the contradiction between speed and accuracy by mediating the transformation process.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If manual transcription is used, then transcript quality can be ensured, but production cost increases significantly

Engineering Contradiction:
Improvetranscript qualityVSAvoidproduction cost
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent replaces expensive manual transcription labor with an automated computational system that processes audio-visual content through algorithmic steps including phoneme extraction, time-alignment, and formatting. This substitution eliminates human labor costs while maintaining transcript quality through consistent application of acoustic and lexical analysis rules.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The transcription system is self-service in that it automatically processes the audio-visual content without requiring human operators. The system extracts features, generates transcriptions, aligns timing, and applies formatting rules autonomously, making the process cost-effective while ensuring reliable output through built-in quality control mechanisms.

Inventive Principle:
Principle #25Self-service

3Productivity

If automated speech recognition is used to generate transcripts, then speed and cost efficiency improve, but accuracy and readability deteriorate

Engineering Contradiction:
Improvetranscription speedVSAvoidtranscription accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the transcription process into distinct stages: phoneme-level transcription extraction, time-alignment with video content, punctuation insertion, and capitalization application. This segmentation allows each stage to be optimized independently, with the phoneme-level extraction ensuring accuracy while the automated formatting stages maintain speed and efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces multiple intermediary processing stages between raw speech recognition output and the final transcript. These intermediaries include phoneme-level transcription extraction, time-alignment processing, and formatting rule application, each serving to improve accuracy while maintaining the speed benefits of automation.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Loss of time

If transcripts are created without formatting, then processing time is reduced, but readability and searchability are compromised

Engineering Contradiction:
Improveprocessing timeVSAvoidreadability
Core Design Contradiction:
Loss of timeVSEase of operation

Solution Approach 1:

The patent segments the formatting process into automatic rule-based operations: punctuation insertion based on phoneme patterns, capitalization applied through algorithmic rules, and paragraph structure determined by timing thresholds. This segmentation enables formatting to be applied automatically without manual intervention, maintaining processing speed while improving readability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The transcription system performs self-service formatting by automatically applying punctuation, capitalization, and paragraph structure rules to the generated transcriptions. This self-service capability ensures that transcripts are immediately readable and searchable without requiring additional manual processing, resolving the contradiction between processing time and readability.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11315546B2Computerized system and method for formatted transcription of multimedia content
Publication Date: 2022.04.26 VERIZON PATENT & LICENSING INC
  • US11315546B2 patent drawing
  • US11315546B2 patent drawing
  • US11315546B2 patent drawing

AI summary

Disclosed are systems and methods for improving interactions with and between computers in content searching, generating, hosting and/or providing systems supported by or configured with personal computing devices, servers and/or platforms. The systems interact to identify and retrieve data within or across platforms, which can be used to improve the quality of data used in processing interactions between or among processors in such systems. The disclosed systems and methods provide systems and methods for automatic creation of a formatted, readable transcript of multimedia content, which is derived, extracted, determined, or otherwise identified from the multimedia content. The formatted, readable transcript can be utilized to increase accuracy and efficiency in search engine optimization, as well as identification of relevant digital content available for communication to a user.