Dynamic Grammar Builder for Automated Closed Captioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Human-based closed captioning for live media content is inefficient and prone to errors due to the difficulty in keeping up with fast-paced audio, especially in live broadcasts, and is costly, while existing automated systems struggle with accurately identifying named entities with varying spellings or new terms.

Innovation Solution

An automated closed captioning system that uses a dynamic grammar builder to generate a list of named entities from real-time social network data, rank them based on trending occurrences, and apply this dynamic grammar to improve speech recognition accuracy in creating closed caption text for media content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human typists are employed to create closed captioning text, then the accuracy of closed captioning can be maintained, but the cost and time consumption increase significantly

Engineering Contradiction:
Improveclosed captioning accuracyVSAvoidcaptioning efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables automated closed captioning where the speech recognition component processes audio content independently to generate caption text, eliminating the need for human typists to manually transcribe. The dynamic grammar builder continuously updates the grammar based on real-time content, allowing the system to self-improve accuracy while maintaining high productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The dynamic grammar builder monitors real-time content and trending named entities, then feeds this information back to update the grammar used by the speech recognition component. This continuous feedback loop improves recognition accuracy over time without requiring human intervention, resolving the contradiction between accuracy and efficiency.

Inventive Principle:
Principle #23Feedback

2Productivity

If multiple humans are employed to handle live broadcast captioning, then the productivity increases, but the cost and coordination complexity increase

Engineering Contradiction:
Improvecaptioning throughputVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The automated system performs all captioning tasks independently through the speech recognition component, which processes audio and generates captions without human intervention. The dynamic grammar builder automatically updates the recognition grammar based on real-time content analysis, eliminating the need for multiple human operators and the coordination complexity that would ensue.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system dynamically adjusts the grammar parameters used by speech recognition based on real-time analysis of trending named entities and content. This adaptive parameter change allows a single automated system to handle diverse content types and topics effectively, replacing multiple specialized human typists without the coordination overhead.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If existing automated speech recognition systems are used, then productivity increases, but the accuracy of named entity recognition decreases

Engineering Contradiction:
Improveautomated captioning speedVSAvoidnamed entity recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The grammar used by the speech recognition component is not static but dynamically updated by the dynamic grammar builder in real-time. The system continuously analyzes trending named entities from the content and adjusts the grammar accordingly, allowing the automated system to maintain high named entity recognition accuracy while preserving the productivity benefits of automation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The dynamic grammar builder performs preliminary analysis of content and identifies trending named entities before they appear in the speech stream. This preliminary action allows the speech recognition system to have accurate grammar rules ready in advance, improving named entity recognition accuracy without compromising the speed of automated captioning.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9922095B2Automated closed captioning using temporal data
Publication Date: 2018.03.20 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9922095B2 patent drawing
  • US9922095B2 patent drawing
  • US9922095B2 patent drawing

AI summary

One or more systems and/or techniques are provided for automatic closed captioning for media content. In an example, real-time content, occurring within a threshold timespan of a broadcast of media content (e.g., social network posts occurring during and an hour before a live broadcast of an interview), may be accessed. A list of named entities, occurring within the social network data, may be generated (e.g., Interviewer Jon, Interviewee Kathy, Husband Dave, Son Jack, etc.). A ranked list of named entities may be created based upon trending named entities within the list of named entities (e.g., a named entity may be ranked higher based upon a more frequent occurrence within the social network posts). A dynamic grammar (e.g., library, etc.) may be built based upon the ranked list of named entities. Speech recognition may be performed upon the broadcast of media content utilizing the dynamic grammar to create closed caption text.