Second-Screen Metadata Synchronization via Automatic Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users of two-screen audiovisual experiences face difficulties in obtaining and synchronizing supplemental data about audiovisual programs, such as actor information, locations, and music, during playback, as existing methods require manual internet searches or lack efficient mechanisms.

Innovation Solution

Automatic generation of metadata through facial recognition, audio recognition, and object recognition during pre-processing of audiovisual content, which is then synchronized with playback time and displayed on a second-screen device, using a networked computer system that includes a control computer, content delivery network, and mobile computing device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If manual internet searches are used to obtain supplemental data, then information can be obtained, but the process is time-consuming and inefficient

Engineering Contradiction:
Improvesupplemental dataVSAvoidtime for manual searches
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary actions by automatically generating metadata about audiovisual content in advance of playback. Facial recognition, audio recognition, and object recognition are executed during pre-processing to create comprehensive metadata that is stored and ready for rapid retrieval during playback, eliminating the need for manual searches

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system enables self-service by automatically generating and retrieving metadata without user intervention. The metadata generation process autonomously analyzes audiovisual content using recognition technologies, and the system automatically retrieves and displays relevant metadata during playback, freeing users from manual information gathering

Inventive Principle:
Principle #25Self-service

2Reliability

If metadata is generated and stored for synchronization, then synchronized display is achieved, but system complexity increases

Engineering Contradiction:
Improvesynchronization accuracyVSAvoidsystem architecture
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system introduces a metadata store as an intermediary component that bridges the audiovisual playback system and the second-screen device. This centralized metadata repository receives pre-generated metadata and delivers it to the mobile device for synchronization with playback timing, managing complexity through modular architecture

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system replaces manual synchronization mechanisms with automated electronic synchronization based on timing metadata. Instead of manual coordination between devices, the system uses electronic signals and timing data to automatically synchronize metadata display with audiovisual playback events

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If automatic recognition technologies are used during pre-processing, then metadata generation is efficient, but computational resources are consumed

Engineering Contradiction:
Improvemetadata generation speedVSAvoidcomputational energy
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system performs computationally intensive recognition operations in advance during pre-processing rather than during real-time playback. Facial recognition, audio recognition, and object recognition are executed before the audiovisual content is played back, allowing resources to be consumed when not under time pressure and enabling rapid metadata retrieval during viewing

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates metadata copies of the audiovisual content information rather than processing the original content during playback. Recognition technologies generate descriptive metadata that replicates key information about faces, audio, and objects, allowing the system to work with lightweight data structures instead of heavy media files during synchronization and display

Inventive Principle:
Principle #26Copying

Data Source

PatentEP3138296B1Displaying data associated with a program based on automatic recognition
Publication Date: 2020.06.17 NETFLIX INC
  • EP3138296B1 patent drawingFigure 1
  • EP3138296B1 patent drawingFigure 2
  • EP3138296B1 patent drawingFigure 3A

AI summary

In one approach, a controller computer performs a pre-processing phase involves applying automatic facial recognition, audio recognition, and/or object recognition to frames or static images of a media item to identify actors, music, locations, vehicles, and props or other items that are depicted in the program. The recognized data is used as the basis of queries to one or more data sources to obtain descriptive metadata about people, items, and places that have been recognized in the program. The resulting metadata is stored in a database in association with time point values indicating when the recognized things appeared in the particular program. Thereafter, when an end user plays the same program using a first- screen device, the stored metadata is downloaded to a second-screen device of the end user. When playback reaches the same time point values on the first-screen device, one or more windows, panels or other displays are formed on the second- screen device to display the metadata associated with those time point values.