Audio-Video Annotation System with Real-Time Transcription and Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for learning more about an audio-video sequence, such as searching for information about identified individuals or objects, are inefficient and interrupt the viewing experience, often leading to misidentification and requiring significant effort to open new browser windows for queries.

Innovation Solution

A facility that automatically annotates audio-video sequences by performing voice transcription and image recognition during playback, displaying relevant information such as speaker names and object identifications near the video frame, and caching these annotations for real-time or near-real-time access, allowing seamless integration with web searches and reducing hardware resource usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If conventional search methods are used to learn about audio-video sequences, then information can be obtained, but the viewing experience is interrupted and significant effort is required

Engineering Contradiction:
Improveinformation accessVSAvoidviewing experience
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system performs preliminary actions by automatically generating annotations, transcribing speech, and identifying objects/people during video playback before the user needs the information. This eliminates the need for users to manually search for information during viewing, thus maintaining the viewing experience while ensuring information is readily available when needed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary annotation system that mediates between the video content and the user's information needs. Instead of users directly searching for information (which interrupts viewing), the intermediary system automatically processes video content and presents relevant information through annotations, thus resolving the contradiction between information access and viewing continuity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If manual search queries are performed to identify individuals or objects, then information can be obtained, but misidentification occurs and significant effort is required

Engineering Contradiction:
Improveinformation accuracyVSAvoidsearch effort
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system enables self-service by automatically performing speech transcription, object identification, and information retrieval without user intervention. The annotation system processes video content autonomously, eliminating manual search efforts and reducing misidentification errors that occur when users manually query information.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual search process with an automated computational system. Instead of users manually typing queries and interpreting results (prone to error), the system uses automatic speech recognition, image recognition, and information retrieval algorithms to accurately identify and provide information about video content.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of operation

If annotations are generated in real-time during playback, then viewing experience is maintained, but processing speed and latency increase

Engineering Contradiction:
Improveviewing continuityVSAvoidprocessing speed
Core Design Contradiction:
Ease of operationVSSpeed

Solution Approach 1:

The system performs preliminary processing of video content during playback, generating annotations and transcriptions in advance of when users need the information. This preliminary action allows the system to maintain viewing continuity while managing processing loads efficiently, reducing latency by preparing information before it is requested.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10743085B2Automatic annotation of audio-video sequences
Publication Date: 2020.08.11 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10743085B2 patent drawing
  • US10743085B2 patent drawing
  • US10743085B2 patent drawing

AI summary

In some examples, a facility augments an audio-video sequence playback display with respect to a current playback position of the audio-video sequence within a time index range of the sequence. For a first portion of the time index range of the sequence containing the current playback position (“CPP”), the facility performs automatic voice transcription against the audio component to obtain speech text for at least one speaker. For a second portion of the time index range of the sequence containing the CPP, the facility performs automatic image recognition against the video component to obtain identifying information identifying at least one person, object, or location. Simultaneously with the sequence playback display and proximate to the sequence playback display, the facility displays one or more annotations each based upon (a) at least a portion of the obtained speech text, (b) at least a portion of the obtained identifying information, or (c) both.