First-Person POV Device Attention Mapping for Video Feed Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automated systems for determining regions-of-interest in video streams, such as those used in video conferencing or event recording, rely on unreliable indicators like sound volume or motion, which fail to accurately identify the focus of attention, especially in situations where speakers are not actively moving.

Innovation Solution

The method involves using first-person point-of-view devices to capture images and video, which are then correlated with reference camera data to determine the region-of-interest by mapping user attention data to a global reference frame, allowing for the identification of salient areas or individuals in a scene.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If sound volume is used as a basis for determining the best video feed, then automated video editing can be performed without human operators, but sound volume may be a poor indicator when sound signals are amplified by sound amplification systems

Engineering Contradiction:
Improveautomated video editingVSAvoidregion-of-interest identification
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent introduces first-person point-of-view device data as an intermediary indicator to identify regions-of-interest. Instead of directly using sound volume (which can be misleading due to amplification), the system uses POV device orientation and captured imagery as a mediator to more accurately determine what the audience is actually looking at, thereby resolving the contradiction between automation and measurement precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Extent of automation

If the amount of motion in video streams is used as an indicator of the region-of-interest, then automated detection can be performed, but the amount of motion may not be reliable for certain situations such as when the speaker moves too little

Engineering Contradiction:
Improveautomated detectionVSAvoidregion-of-interest detection
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent uses first-person point-of-view device data as an intermediary to detect regions-of-interest. Rather than relying on motion amplitude (which fails when speakers are stationary), the system uses POV device orientation and field-of-view data as a mediator to identify what subjects the audience is observing, even when those subjects are not moving significantly.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Area of stationary object

If multiple cameras are deployed to capture video streams from different angles, then comprehensive coverage of the event is achieved, but the complexity of selecting the best video feed increases

Engineering Contradiction:
Improveevent coverageVSAvoidvideo feed selection
Core Design Contradiction:
Area of stationary objectVSDevice complexity

Solution Approach 1:

The patent implements a feedback mechanism where data from first-person point-of-view devices is used to automatically determine which camera angle best captures the region-of-interest. The POV device orientation and captured imagery provide real-time feedback about audience attention, which automatically controls the selection among multiple camera feeds, thereby managing the complexity of multi-camera systems through intelligent automation.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10721439B1Systems and methods for directing content generation using a first-person point-of-view device
Publication Date: 2020.07.21 GOOGLE LLC
  • US10721439B1 patent drawing
  • US10721439B1 patent drawing
  • US10721439B1 patent drawing

AI summary

A method for personalizing a content item using captured footage is disclosed. The method includes receiving a first video feed from a first camera, wherein the first camera is designated as a source camera for capturing an event during a first time duration. The method also includes receiving data from a second camera, and determining, based on the received data from the second camera, that an action was performed using the second camera, the action being indicative of a region of interest (ROI) of the user of the second camera occurring within a second time duration. The method further includes designating the second camera as the source camera for capturing the event during the second time duration.