Heat Ranking Media Objects via Spatial-Temporal Feature Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current digital media advertising techniques for determining viewer attention in videos are time-consuming, costly, and inaccurate due to reliance on manual analysis of eye-tracking data and static feature analysis, leading to unreliable and inefficient content placement.
Innovation Solution
A system comprising a heat map creation engine, semantic segmentation engine, and heat ranking engine that extracts spatial-temporal features from video frames to automatically create heat maps and rank media objects based on viewer interest, enabling efficient and accurate content placement without manual intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If wearable eye-tracking devices are used to track viewers' eye positions and movements, then viewer attention data can be collected, but the process becomes time-consuming and cost-intensive
Solution Approach 1:
The patent uses computer vision algorithms to capture and analyze visual information from video frames, creating a digital copy of viewer attention patterns without requiring physical eye-tracking devices. The system processes video frames to generate heat maps that replicate the functionality of eye-tracking data collection, thereby eliminating the need for expensive and time-consuming wearable devices while maintaining measurement precision
Solution Approach 2:
The patent replaces the mechanical wearable eye-tracking device system with a software-based computer vision system that processes video frames algorithmically. By substituting the physical tracking mechanism with computational image processing, the system achieves the same measurement objectives without the associated time and cost overhead of device deployment and manual data collection
2Ease of operation
If heat maps are manually examined by marketers to identify hotspots, then content placement decisions can be made, but the process becomes time-intensive and error-prone
Solution Approach 1:
The patent implements an automated system where the computer vision algorithm independently analyzes video frames, generates heat maps, identifies hotspots, and provides content placement recommendations without requiring manual marketer intervention. The system serves itself by performing all analytical tasks algorithmically, eliminating time-intensive manual examination while reducing human error through consistent automated processing
Solution Approach 2:
The patent transforms the heat map analysis process from a manual visual inspection parameter to an automated computational parameter. By changing the analysis method from human visual processing to algorithmic image processing, the system maintains the ability to identify hotspots and make content placement decisions while dramatically reducing time consumption and eliminating manual errors
3Productivity
If static features are used in algorithmic-based heat map techniques, then heat maps can be generated, but the accuracy of viewer attention data is limited
Solution Approach 1:
The patent transitions from static feature analysis to dynamic feature analysis by processing multiple video frames sequentially and analyzing temporal changes in visual content. The system extracts features that evolve over time, such as object motion, scene transitions, and temporal attention patterns, thereby maintaining heat map generation efficiency while significantly improving the accuracy of viewer attention data through dynamic rather than static analysis
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A heat map creation engine receives a series of image frames from amongst a plurality of image frames of a media, where the series of image frames includes a target frame and at least one neighboring frame. The heat map creation engine also extracts spatial-temporal features of the image frames, rescales the spatial-temporal features to obtain heat distribution over the target frame, and creates a heat map for the target frame based on the heat distribution. A semantic segmentation engine segments the image frames into multiple media objects based on pre-defined classes, and selects one or more media objects from amongst the media objects based on pre-defined conditions. A heat ranking engine ranks the selected one or more media objects in the media based on the heat scores and the created heat map.