Caption Event Fingerprinting for Audio-Visual Content Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in monitoring and ensuring the integrity and timing consistency of audio-visual content captions across different formats and delivery systems, particularly due to their non-periodic nature and varied formats such as text-based and image-based captions, which complicates error detection and comparison.

Innovation Solution

The method involves generating caption event fingerprints based on word lengths, regardless of character identity, and using fingerprint processors to compare these fingerprints across different locations in a content delivery chain, allowing for the identification of matching caption events and measurement of errors like missing events, timing errors, and discrepancies, without the need for Optical Character Recognition (OCR).

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If caption monitoring is performed using traditional text-based methods, then caption integrity can be verified, but the system becomes complex and cannot handle various caption formats (text-based and image-based captions)

Engineering Contradiction:
Improvecaption format compatibilityVSAvoidmonitoring system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent replaces traditional text-based caption verification methods with audio fingerprinting technology. Instead of analyzing caption text directly, the system generates audio fingerprints from caption audio signals and compares these fingerprints for verification. This substitution enables the system to handle both text-based and image-based captions uniformly through audio signal processing, thereby improving format compatibility while maintaining manageable system complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If OCR is used to process image-based captions, then text extraction is possible, but processing overhead increases significantly

Engineering Contradiction:
Improvecaption information retrievalVSAvoidprocessing efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent extracts only the essential audio fingerprint features from caption audio signals rather than performing full OCR processing on image-based captions. By taking out only the necessary acoustic characteristics for verification, the system retrieves caption information efficiently without the heavy processing overhead associated with complete optical character recognition, thus improving productivity while maintaining information retrieval capability.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If detailed caption text analysis is performed, then caption accuracy can be verified, but timing synchronization becomes more difficult to maintain

Engineering Contradiction:
Improvecaption accuracy verificationVSAvoidtiming synchronization
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary generation of audio fingerprints from caption audio signals before the actual verification process. By preparing the fingerprint data in advance, the system enables rapid comparison and verification of caption accuracy without requiring time-consuming detailed text analysis during playback, thus maintaining precise timing synchronization while still achieving accurate caption verification through fingerprint matching.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10091543B2Monitoring audio-visual content with captions
Publication Date: 2018.10.02 GRASS VALLEY LTD
  • US10091543B2 patent drawing
  • US10091543B2 patent drawing
  • US10091543B2 patent drawing

AI summary

To monitor audio-visual content which includes captions, caption fingerprints are derived from a length of each word in the caption, without regard to the identity of the character or characters forming the word. Audio-visual content is searched to identify a caption event having a matching fingerprint and missing captions; caption timing errors and caption discrepancies are measured.