Machine-Learned OCR for Real-Time Sports Highlight Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing television systems lack the ability to automatically and efficiently extract and interpret embedded information cards in real-time from sporting event broadcasts to generate metadata for highlights, limiting interactive and enhanced television applications.

Innovation Solution

A machine-learned character classification model is trained using a multidimensional vector space and principal component analysis to recognize and interpret embedded information cards in real-time, generating metadata that is associated with video highlights.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If traditional television systems are used to display sporting events, then the broadcast content is delivered to viewers, but the system cannot automatically extract and interpret embedded information cards in real-time

Engineering Contradiction:
Improveautomatic extraction and interpretation of embedded information cardsVSAvoidreal-time processing capability
Core Design Contradiction:
Extent of automationVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-training machine learning models with extensive training data containing embedded information card patterns, formats, and content types. This pre-training enables the system to rapidly process and interpret information cards in real-time during actual broadcasts without requiring complex runtime analysis, thus achieving both automation and real-time performance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional mechanical or rule-based text extraction systems with machine learning-based optical character recognition (OCR) and natural language processing (NLP) models. These intelligent systems can automatically recognize, extract, and interpret embedded information cards in various formats and languages, enabling automated metadata generation without manual intervention while maintaining real-time processing speeds.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If no automated information extraction system is implemented, then the television system remains simple, but metadata for highlights cannot be generated automatically

Engineering Contradiction:
Improvemetadata extraction from embedded information cardsVSAvoidmachine learning model integration
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system introduces an intermediary layer consisting of machine learning models and processing algorithms that bridge the gap between raw broadcast content and structured metadata. These intermediaries automatically extract information from embedded cards, interpret the content, and generate standardized metadata that can be used for highlights and enhanced television applications, thus preserving information without requiring direct complex integration into the core television system.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system creates copies of embedded information card content by extracting text and data from the visual cards displayed during broadcasts. These copied information elements are then processed, validated, and transformed into structured metadata formats, enabling automated highlight generation and enhanced TV applications without modifying the original broadcast signal or requiring direct access to the source data.

Inventive Principle:
Principle #26Copying

3Productivity

If real-time processing of embedded information cards is implemented, then interactive television applications are enhanced, but the processing complexity and computational resources increase

Engineering Contradiction:
Improvereal-time metadata generation for highlightsVSAvoidprocessing system architecture
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex processing task into distinct modular components: optical character recognition (OCR) for text extraction, natural language processing (NLP) for content interpretation, metadata generation for structured output, and quality assurance for validation. Each module can be independently optimized, trained, and processed in parallel, enabling real-time performance while managing computational complexity through distributed processing and specialized hardware acceleration.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12387493B2Machine learning for recognizing and interpreting embedded information card content
Publication Date: 2025.08.12 STATS LLC
  • US12387493B2 patent drawing
  • US12387493B2 patent drawing
  • US12387493B2 patent drawing

AI summary

Metadata for highlights of a video stream is extracted from card images embedded in the video stream. The highlights may be segments of a video stream, such as a broadcast of a sporting event, that are of particular interest to one or more users. Card images embedded in video frames of the video stream are identified and processed to extract text. The text characters may be recognized by applying a machine-learned model trained with a set of characters extracted from card images embedded in sports television programming contents. The training set of character vectors may be pre-processed to maximize metric distance between the training set members. The text may be interpreted to obtain the metadata. The metadata may be stored in association with the portion of the video stream. The metadata may provide information regarding the highlights, and may be presented concurrently with playback of the highlights.