Video Entity Annotation Using Automatic Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current tools for generating interactive video content are limited, requiring manual indication of clickable areas that may not align spatially or temporally with the object of interest, limiting user engagement and interaction.

Innovation Solution

A system that automatically identifies entities within video content using techniques like face recognition, audio recognition, and text recognition, linking these entities to supplemental content stored in a database, allowing users to interact with and access additional information during playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual indication of clickable areas is used, then interactive content can be created, but the clickable areas may not align spatially or temporally with the object of interest

Engineering Contradiction:
Improveease of creating interactive contentVSAvoidspatial and temporal alignment accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent replaces manual mechanical indication methods with automatic computer vision-based entity identification. The system uses face recognition, audio recognition, and text recognition algorithms to automatically detect and identify entities in video content, eliminating the need for manual clickable area placement while achieving precise spatial and temporal alignment with objects of interest.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If automatic entity identification is implemented, then precise alignment with objects of interest is achieved, but system complexity increases

Engineering Contradiction:
Improveentity identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the complex entity identification task into separate specialized modules: face recognition module, audio recognition module, and text recognition module. Each module handles a specific type of entity identification independently, and their results are integrated to provide comprehensive entity detection. This segmentation reduces overall system complexity by making each component more manageable and specialized.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If multiple recognition techniques are used, then entity identification accuracy improves, but processing time increases

Engineering Contradiction:
Improveentity identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements a selective recognition approach where the system applies recognition techniques based on the specific context and content type. Not all recognition algorithms are applied to every video segment - instead, the system selects appropriate recognition methods based on what entities are present or expected, reducing unnecessary processing while maintaining high identification accuracy when needed.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10070170B2Content annotation tool
Publication Date: 2018.09.04 GOOGLE LLC
  • US10070170B2 patent drawing
  • US10070170B2 patent drawing
  • US10070170B2 patent drawing

AI summary

A content annotation tool is disclosed. In a configuration, a portion of a movie may be obtained from a database. Entities, such as an actor, background music, text, etc. may be automatically identified in the movie. A user, such as a content producer, may associate and/or provide supplemental content for an identified entity to the database. A selection of one or more automatically identified entities may be received. A database entry may be generated that links the identified entity with the supplemental content. The selected automatically identified one or more entities and//or supplemental content associated therewith may be presented to an end user.