Multimedia Transcript Amalgamation for Closed Captioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for delivering multimedia content face challenges in enabling efficient search and utilization of multimedia content, as consumers struggle to locate specific content without descriptive metadata, and providers lack easy methods to identify objects within content or integrate additional services, leading to inefficient content consumption and usage tracking.

Innovation Solution

A method and system that convert speech in multimedia content to text, allowing for indexing and metadata generation, which includes analyzing multimedia content for closed captioning data, extracting audio for speech-to-text conversion, and selecting amalgamated transcripts to create searchable metadata, enabling better content discovery and integration of additional services.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multiple speech-to-text conversion programs are used to create transcripts, then the accuracy and completeness of the transcript improves, but the processing time and computational resources increase

Engineering Contradiction:
Improvetranscript accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs speech-to-text conversions with multiple programs but selects text from only one or more transcripts rather than amalgamating all results. This partial action approach achieves sufficient transcript accuracy without the full computational overhead of processing and merging all possible transcripts, resolving the contradiction between accuracy and processing time.

Inventive Principle:
Principle #16Partial or excessive action

2Ease of operation

If multimedia content is processed to add descriptive metadata and indexing, then content searchability and user experience improves, but the computational resources and processing requirements increase

Engineering Contradiction:
Improvecontent searchabilityVSAvoidprocessing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system performs speech-to-text conversion and creates indexed transcripts in advance, before users need to search for content. This preliminary action prepares the content with searchable metadata upfront, making future content retrieval efficient and user-friendly without requiring complex real-time processing when users search.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If closed captioning data is detected and used directly, then the process is efficient and quick, but the content may lack accuracy if the closed captioning is poor quality

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidtranscript accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system detects closed captioning data and uses it as an initial transcript, then compares it against transcripts generated by multiple speech-to-text conversion programs. This feedback mechanism allows the system to identify and correct errors in the closed captioning, improving accuracy while maintaining efficiency by starting with the pre-existing captioning data.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9332319B2Amalgamating multimedia transcripts for closed captioning from a plurality of text to speech conversions
Publication Date: 2016.05.03 UNISYS CORP
  • US9332319B2 patent drawing
  • US9332319B2 patent drawing
  • US9332319B2 patent drawing

AI summary

Methods and systems for converting speech to text are disclosed. One method includes analyzing multimedia content to determine the presence of closed captioning data. The method includes, upon detecting closed captioning data, indexing the closed captioning data as associated with the multimedia content. The method also includes, upon failure to detect closed captioning data in the multimedia content, extracting audio data from multimedia content, the audio data including speech data, performing a plurality of speech to text conversions on the speech data to create a plurality of transcripts of the speech data, selecting text from one or more of the plurality of transcripts to form an amalgamated transcript, and indexing the amalgamated transcript as associated with the multimedia content.