Speech Recognition Video Annotation System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Home movies lack metadata for efficient cataloging, searching, and retrieval, making it cumbersome for users to access specific segments, unlike commercially produced videos which have standardized metadata.

Innovation Solution

A method and apparatus that uses speech recognition technology to annotate video content with metadata, allowing users to verbally describe video segments, which are then converted to text and stored for easy retrieval, using a voice-based metadata generation module integrated with video storage devices like DVRs and set-top boxes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech recognition technology is used to generate metadata, then the ease of operation is improved, but the device complexity increases

Engineering Contradiction:
Improveease of annotating video contentVSAvoidcomplexity of speech recognition system
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces a speech recognition system as an intermediary component that mediates between the user's voice input and the metadata generation process. This intermediary automatically converts spoken annotations into structured metadata, resolving the contradiction by providing ease of operation through voice-based input while managing the complexity through automated processing rather than manual metadata creation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If manual metadata creation is used, then the device complexity is reduced, but the productivity decreases

Engineering Contradiction:
Improvespeed of cataloging video contentVSAvoidcomplexity of metadata generation system
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system implements self-service by enabling users to create metadata through natural speech annotations during video playback. The speech recognition technology automatically processes these annotations and generates structured metadata without requiring manual intervention or complex configuration, thereby improving productivity while keeping the user interface simple and intuitive.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If detailed metadata is stored for each video segment, then the ease of searching and retrieving is improved, but the loss of storage space increases

Engineering Contradiction:
Improveease of searching and retrieving video segmentsVSAvoidstorage space for metadata
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential and relevant information from speech annotations to create compact metadata representations. By selectively extracting key descriptors from detailed speech transcripts and storing only the most important metadata elements, the system improves searchability and retrieval ease while minimizing the storage space required for metadata.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10482168B2Method and apparatus for annotating video content with metadata generated using speech recognition technology
Publication Date: 2019.11.19 GOOGLE TECHNOLOGY HOLDINGS LLC
  • US10482168B2 patent drawing
  • US10482168B2 patent drawing
  • US10482168B2 patent drawing

AI summary

A method and apparatus is provided for annotating video content with metadata generated using speech recognition technology. The method begins by rendering video content on a display device. A segment of speech is received from a user such that the speech segment annotates a portion of the video content currently being rendered. The speech segment is converted to a text-segment and the text-segment is associated with the rendered portion of the video content. The text segment is stored in a selectively retrievable manner so that it is associated with the rendered portion of the video content.