Media Stream Annotator for Automated Ground Truth Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing systems for annotating media streams, such as live sports events, are inefficient and error-prone, requiring significant human effort and time, and existing pattern recognition systems operate with less than absolute accuracy due to noise, necessitating the development of Ground Truth Metadata validated by human annotators.

Innovation Solution

A Media Stream Annotator system that automates the generation of Ground Truth Metadata using a Human-Computer Interface, adjusts Pattern Recognition System input parameters, and merges metadata from third parties to improve recognition accuracy, enabling efficient annotation and real-time operation on mobile devices with low bit-rate internet connections.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human annotators manually create Ground Truth Metadata, then accuracy is high, but time consumption and effort are significant

Engineering Contradiction:
ImproveGround Truth Metadata accuracyVSAvoidTime required for annotation
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automated self-service annotation where the Pattern Recognition System processes media streams and generates Ground Truth Metadata automatically, eliminating the need for manual human annotation while maintaining high accuracy through multiple recognition systems working in parallel

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual annotation process (human annotators using keyboards and mice) with automated computer vision and speech recognition systems that process media streams electronically, dramatically reducing time consumption while maintaining accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Extent of automation

If Pattern Recognition Systems process media streams, then automation is achieved, but accuracy is less than absolute due to noise

Engineering Contradiction:
ImproveAnnotation automation levelVSAvoidRecognition accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system merges multiple Pattern Recognition Systems (computer vision, speech recognition, and third-party data sources) to process media streams simultaneously, combining their results to achieve higher overall accuracy than any single system could provide alone

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces an intermediary validation layer where Human Annotators review and validate automated recognition results, acting as a mediator between the automated systems and the final Ground Truth Metadata to correct errors and maintain high accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If PRS input parameters are adjusted to maximize accuracy, then recognition performance improves, but system complexity increases

Engineering Contradiction:
ImproveRecognition accuracyVSAvoidParameter adjustment complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system implements feedback mechanisms where Human Annotators validate recognition results and provide corrections, which are then fed back into the Pattern Recognition Systems to automatically adjust and optimize their parameters, simplifying the complexity through automated learning rather than manual tuning

Inventive Principle:
Principle #23Feedback

4Speed

If live media is processed in real-time, then timeliness is improved, but computational resources and processing time increase

Engineering Contradiction:
ImproveReal-time processing speedVSAvoidComputational resource consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system segments the media stream processing into multiple independent parallel tracks, with different Pattern Recognition Systems processing different aspects of the media simultaneously (e.g., object detection, speech recognition, event detection), allowing real-time processing through parallel computation rather than sequential processing

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10553252B2Annotating media content for automatic content understanding
Publication Date: 2020.02.04 LIVECLIPS LLC
  • US10553252B2 patent drawing
  • US10553252B2 patent drawing
  • US10553252B2 patent drawing

AI summary

A system for annotating frames in a media stream 114 includes a pattern recognition system (PRS) 108 to generate PRS output metadata for a frame; an archive 106 for storing ground truth metadata (GTM); a device to merge the GTM and PRS output metadata and thereby generate proposed annotation data (PAD) 110; and a user interface 109 for use by the human annotator HA 118. The user interface 104 includes an editor 111 and an input device 107 used by the HA 118 to approve GTM for the frame. An optimization system 105 receives the approved GTM and metadata output by the PRS 108, and adjusts input parameters for the PRS to minimize a distance metric corresponding to a difference between the GTM and PRS output metadata.