On-Screen Caption Text Error Correction via Knowledge Graph Context
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional media systems struggle with accurately converting speech into on-screen caption text due to limitations in dictionary size, human stenographer knowledge, and difficulties with terms of art, newly created buzzwords, foreign names, homophones, and errors that require manual correction.
Innovation Solution
A media guidance application that automatically corrects errors in on-screen caption text by accessing a knowledge graph based on information derived from the media asset, using textual or image recognition, and weighing potential corrections by phonetic similarity and timestamp.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automatic speech recognition is used to convert speech to caption text, then productivity is improved, but manufacturing precision deteriorates due to errors in recognizing terms of art, foreign names, and homophones
Solution Approach 1:
The patent introduces an intermediary correction system that sits between the ASR output and final caption display. This intermediary layer uses multiple correction techniques including pronunciation dictionaries, context analysis, and alternative word matching to fix ASR errors without reducing overall productivity. The system automatically identifies and corrects errors in terms of art, foreign names, and homophones while maintaining high-speed caption generation.
Solution Approach 2:
The patent replaces manual stenographer correction with an automated electronic correction system. Instead of relying on human expertise to manually fix ASR errors, the system uses computational methods including pronunciation pattern matching, contextual analysis, and database lookups to automatically substitute correct terms for erroneous ASR output, thereby maintaining productivity while improving precision.
2Manufacturing precision
If human stenographers are used to correct ASR errors, then manufacturing precision is improved, but productivity deteriorates due to manual intervention requirements
Solution Approach 1:
The patent implements a self-service correction system where the captioning system automatically identifies and corrects its own errors without human intervention. The system uses pronunciation dictionaries, contextual analysis, and alternative word matching algorithms to self-correct ASR errors in real-time, eliminating the need for manual stenographer review while maintaining high accuracy and productivity.
Solution Approach 2:
The patent incorporates feedback mechanisms where the system continuously monitors caption text for errors using multiple validation techniques. When errors are detected, the system automatically retrieves alternative candidates from pronunciation dictionaries and context databases, evaluates them against the original speech context, and implements corrections. This closed-loop feedback system maintains high precision without sacrificing productivity.
3Manufacturing precision
If the dictionary size of ASR systems is increased, then manufacturing precision is improved, but device complexity increases
Solution Approach 1:
The patent segments the vocabulary knowledge into multiple specialized dictionaries organized by domain and language. Instead of one massive dictionary, the system uses separate pronunciation dictionaries for different languages (English, Spanish, French, German), specialized dictionaries for terms of art and jargon, and context-specific vocabularies. This segmentation improves recognition accuracy for specialized terms while keeping each individual dictionary manageable in size and complexity.
Solution Approach 2:
The patent creates a universal correction framework that handles multiple languages, domains, and error types through a single integrated system. The pronunciation dictionary system is designed to be language-agnostic and domain-agnostic, using universal phonetic algorithms and contextual analysis techniques that work across different languages and specialized fields. This multi-functional approach improves precision for diverse terms without proportionally increasing system complexity.
Data Source
AI summary
Systems and methods are described to address shortcomings in conventional systems by correcting an erroneous term in on-screen caption text for a media asset. In some aspects, the systems and methods identify the erroneous term in a text segment of the on-screen caption text, and identify one or more video frames of the media asset corresponding to the text segment. The systems and methods further identify a contextual term related to the erroneous term from the one or more video frames. By accessing a knowledge graph, the systems and methods identify a candidate correction based on the contextual term and a portion of the text segment. Lastly, the systems and methods replaces the erroneous term with the candidate correction.


