Activity-Triggered Video Augmentation for Context-Aware Viewing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video presentation systems fail to adequately provide supplemental content corresponding to the video content and elements within the video content, limiting user engagement and interaction.

Innovation Solution

A system that uses a first device, such as an HMD or tablet, to detect user activity directed at a portion of a video content viewing area on a second device, like a TV, and provides augmentations by identifying the element of interest within the video content, allowing for enhanced interaction and information delivery.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If video content is presented on a display device, then visual information is delivered to users, but supplemental content corresponding to video elements is not adequately provided

Engineering Contradiction:
Improvesupplemental contentVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary system comprising a display device, a computing device, and an augmented reality device that mediates between the video content source and the user. The computing device receives video content, identifies elements within it, generates supplemental content, and coordinates with the AR device to present this supplemental information to the user, thereby resolving the information deficiency without requiring direct modification of the video source

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system segments the video content into distinct elements (such as objects, characters, or scenes) and processes each element separately to generate corresponding supplemental content. This segmentation allows the system to provide targeted supplemental information for specific video elements rather than treating the entire video as a single unit, improving information delivery while maintaining manageable system complexity

Inventive Principle:
Principle #1Segmentation

2Ease of operation

If user activity detection is implemented to provide context-aware augmentations, then user engagement is enhanced, but device complexity and processing requirements increase

Engineering Contradiction:
Improveuser engagementVSAvoiddetection system complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent merges the functionality of multiple devices into a coordinated system where the display device, computing device, and augmented reality device work together. The computing device handles the complex task of detecting user activity (such as gaze direction or device proximity) and processing video content, while the display and AR devices focus on presenting the appropriate content. This distribution of functions enhances user engagement while preventing any single device from becoming overly complex

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system implements feedback loops where user activity detected by sensors (such as eye tracking or device positioning) continuously informs the generation and presentation of supplemental content. The system monitors user responses to the presented augmentations and adjusts subsequent content delivery accordingly, creating an adaptive experience that enhances engagement while the automated feedback mechanisms reduce the need for complex manual control systems

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12573154B1Augmented video based on user activity
Publication Date: 2026.03.10 APPLE INC
  • US12573154B1 patent drawing
  • US12573154B1 patent drawing
  • US12573154B1 patent drawing

AI summary

Various implementations that use user activity determined via a first device (e.g., an HMD, tablet, etc.) to provide augmentations corresponding to video content (e.g., a TV show, movie, etc.) that is being presented on a second device (e.g., a TV, monitor, etc.). The first device (e.g., HMD) detects that a user activity (e.g., a user's gaze, gesture, interest) is directed to a portion of a surface corresponding to a video content viewing area provided by the second device (e.g., the TV, monitor, etc.). An element of the video content currently being displayed at that portion of the surface by the second device is identified and used to provide an augmentation that is displayed by the first or second device.