Video Augmentation System with Contextual Overlays

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video telepresence technologies lack the ability to provide enhanced context and insights into the content being displayed, requiring users to manually search for relevant information, which is time-consuming and inefficient.

Innovation Solution

A system and method that augment video content by processing speech input, gaze direction, level of understanding, and environmental conditions to overlay additional contextual information, allowing users to pause and examine specific portions of the video while maintaining the rest in the background, enhancing the viewing experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If video telepresence technologies are used to enable real-time communication, then individuals can communicate using audio and video, but the technologies provide relatively little insights into the content being displayed

Engineering Contradiction:
Improvecontextual informationVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary system that processes video content and generates augmented reality overlays. This intermediary layer adds contextual information (such as object identification, transcription, and metadata) without requiring direct modification of the core video telepresence system, thus resolving the contradiction between information completeness and system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces manual information gathering (mechanical user action) with automated computer vision and natural language processing systems. The system automatically analyzes video content, identifies objects, generates transcriptions, and creates contextual overlays, eliminating the need for users to manually search for or annotate content.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of time

If users manually search for relevant information in video content, then they can find specific details, but the process is time-consuming and inefficient

Engineering Contradiction:
Improveinformation retrieval timeVSAvoidinformation access efficiency
Core Design Contradiction:
Loss of timeVSProductivity

Solution Approach 1:

The patent performs preliminary analysis of video content by pre-processing footage to identify objects, generate transcriptions, and create metadata tags before user viewing. This preliminary action allows the system to instantly retrieve and display relevant information when users interact with the content, eliminating manual searching and significantly reducing information retrieval time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms where the system continuously monitors user interactions with video content and dynamically adjusts the display of augmented information. Based on user gaze, selection, and engagement patterns, the system provides relevant contextual information in real-time, creating an efficient adaptive information delivery system.

Inventive Principle:
Principle #23Feedback

3Loss of information

If augmentation content is displayed in an overlaid manner on video, then contextual information is enhanced, but the display complexity increases

Engineering Contradiction:
Improvecontextual contextVSAvoiddisplay complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies local quality by displaying augmentation content only in specific regions of the screen where it is most relevant, rather than uniformly across the entire display. The system strategically positions overlays near the corresponding video content and adjusts their prominence based on local context, maintaining information enhancement while managing display complexity.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3465620B1Shared experience with contextual augmentation
Publication Date: 2023.08.16 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3465620B1 patent drawingFigure 1
  • EP3465620B1 patent drawingFigure 2
  • EP3465620B1 patent drawingFigure 3

AI summary

Techniques for augmenting video content to enhance context of the video content are described herein. In some instances, a video may be captured at a first location and transmitted to a second location, where the video is output in real-time. A context surrounding a user that is capturing the video and/or a user that is viewing the video may be used to augment the video with additional content. For example, the techniques may process speech or other input associated with either user, a gaze associated with either user, a previous conversation for either user, an area of interest identified by either user, a level of understanding of either user, an environmental condition, and so on. Based on the processing, the techniques may determine augmentation content. The augmentation content may be displayed with the video in an overlaid manner to enhance the experience of the user viewing the video.