First-Person POV Device Attention Mapping for Video Feed Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current automated systems for determining regions-of-interest in video streams, such as those used in video conferencing or event recording, rely on unreliable indicators like sound volume or motion, which fail to accurately identify the focus of attention, especially in situations where speakers are not actively moving.
Innovation Solution
The method involves using first-person point-of-view devices to capture images and video, which are then correlated with reference camera data to determine the region-of-interest by mapping user attention data to a global reference frame, allowing for the identification of salient areas or individuals in a scene.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If sound volume is used as a basis for determining the best video feed, then automated video editing can be performed without human operators, but sound volume may be a poor indicator when sound signals are amplified by sound amplification systems
Solution Approach 1:
The patent introduces first-person point-of-view device data as an intermediary indicator to identify regions-of-interest. Instead of directly using sound volume (which can be misleading due to amplification), the system uses POV device orientation and captured imagery as a mediator to more accurately determine what the audience is actually looking at, thereby resolving the contradiction between automation and measurement precision.
2Extent of automation
If the amount of motion in video streams is used as an indicator of the region-of-interest, then automated detection can be performed, but the amount of motion may not be reliable for certain situations such as when the speaker moves too little
Solution Approach 1:
The patent uses first-person point-of-view device data as an intermediary to detect regions-of-interest. Rather than relying on motion amplitude (which fails when speakers are stationary), the system uses POV device orientation and field-of-view data as a mediator to identify what subjects the audience is observing, even when those subjects are not moving significantly.
3Area of stationary object
If multiple cameras are deployed to capture video streams from different angles, then comprehensive coverage of the event is achieved, but the complexity of selecting the best video feed increases
Solution Approach 1:
The patent implements a feedback mechanism where data from first-person point-of-view devices is used to automatically determine which camera angle best captures the region-of-interest. The POV device orientation and captured imagery provide real-time feedback about audience attention, which automatically controls the selection among multiple camera feeds, thereby managing the complexity of multi-camera systems through intelligent automation.
Data Source
AI summary
A method for personalizing a content item using captured footage is disclosed. The method includes receiving a first video feed from a first camera, wherein the first camera is designated as a source camera for capturing an event during a first time duration. The method also includes receiving data from a second camera, and determining, based on the received data from the second camera, that an action was performed using the second camera, the action being indicative of a region of interest (ROI) of the user of the second camera occurring within a second time duration. The method further includes designating the second camera as the source camera for capturing the event during the second time duration.


