On-Screen Text Detection and Alert Generation in Video Streams

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video content systems lack the ability to effectively detect and utilize on-screen text and objects within video data for generating alerts and supplemental information, limiting user interaction and experience.

Innovation Solution

A system that analyzes video data to detect on-screen text and objects using optical character recognition (OCR) and metadata, allowing for real-time monitoring and generation of alerts and supplemental information based on user-defined regions and key items, with the ability to perform actions such as volume adjustments and superimposing text over other video content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If video data is analyzed to detect on-screen text and objects using OCR and metadata, then user interaction and experience are enhanced, but system complexity increases

Engineering Contradiction:
Improveuser interaction capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments video data analysis into distinct functional modules: OCR processing for text detection, object detection algorithms for identifying visual elements, metadata extraction, and alert generation components. Each module handles specific aspects of analysis independently, making the overall complex system manageable and maintainable while enabling enhanced user interaction through multiple detection capabilities.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary components such as the alert engine that mediates between the complex detection processes and the user interface. This intermediary layer processes detector output data, manages alert generation, and coordinates between various detection modules, thereby hiding system complexity from users while delivering enhanced interaction capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If real-time monitoring of video data is implemented to generate alerts, then responsiveness to content is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveresponse speedVSAvoidcomputational resource consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system implements partial monitoring by allowing users to specify particular regions of interest within video frames and select specific types of content (text, objects, or both) to monitor. This selective approach reduces the overall computational burden compared to analyzing every pixel and element in the entire video stream, enabling real-time alert generation with reduced resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent employs preliminary processing techniques including pre-defined regions of interest, pre-configured detection parameters, and cached metadata information that reduce the computational load during real-time monitoring. By preparing detection criteria and parameters in advance, the system achieves faster real-time response with lower processing requirements.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If multiple detection algorithms are used to identify on-screen text and objects, then detection accuracy is improved, but processing complexity increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple detection algorithms (OCR for text, object detection for visual elements) into a unified processing framework that shares common infrastructure such as image preprocessing, feature extraction, and result integration modules. This merging approach maintains high detection accuracy through multiple specialized algorithms while reducing overall processing complexity by eliminating redundant operations and sharing computational resources across different detection tasks.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP3044725B1Generating alerts based upon detector outputs
Publication Date: 2021.04.07 ARRIS ENTERPRISES LLC
  • EP3044725B1 patent drawingFigure 1A
  • EP3044725B1 patent drawingFigure 1B
  • EP3044725B1 patent drawingFigure 2A

AI summary

Systems and methods for generating alerts and enhanced viewing experience features using on-screen data are disclosed. Textual data corresponding to on-screen text is determined from the visual content of video data. The textual data is associated with corresponding regions and frames of the video data in which the corresponding on-screen text was detected. Users can select regions in the frames of the visual content to monitor for a particular triggering item (e.g., a triggering word, name, or phrase). During play back of the video data, the textual data associated with the selected regions in the frames can be monitored for the triggering item. When the triggering item is detected in the textual data, an alert can be generated. Alternatively, the textual data for the selected region can be extracted to compile supplemental information that can be rendered over the playback of the video data or over other video data.