Deduplicating Viewership Data via Hashing and Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges in accurately collecting and reporting viewership data from diverse devices due to noisy data from smart TVs and STBs, with issues like devices being left on inadvertently and duplicate data from different sources, leading to unreliable reporting for advertisers and content providers.
Innovation Solution
A method and system that determine similarity scores between sets of viewership data from different devices using deterministic functions and hashing techniques to identify candidate pairs and eliminate redundant data, ensuring accurate and reliable viewership reporting by distinguishing between actual viewing and device status.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If viewership data is collected from multiple devices (smart TVs and STBs), then data coverage and completeness are improved, but data redundancy and double-counting increase
Solution Approach 1:
The patent segments the viewership data collection process by device type (smart TV vs. STB) and applies device-specific processing rules. Smart TV data is processed through ACR-based content identification while STB data uses different validation criteria, allowing each data source to be handled appropriately while maintaining overall data integrity
Solution Approach 2:
The patent introduces an intermediary processing layer that receives data from multiple devices, applies deterministic functions and hashing techniques, and produces deduplicated viewership records. This intermediary system acts as a mediator between raw device data and final reporting, eliminating duplicates while preserving valid viewing events
2Adaptability or versatility
If ACR technology is used to identify content on smart TVs, then content recognition capability is improved, but reliability decreases when content is not pre-scheduled
Solution Approach 1:
The patent applies preliminary action by pre-scheduling known content in a database before ACR processing. When ACR detects content playback, it first checks against the pre-built schedule of expected content. This preliminary preparation enables reliable verification of ACR results and distinguishes between actual viewing and device idle states
3Ease of operation
If devices are left on indefinitely, then availability for viewing is improved, but data noise and false viewing reports increase
Solution Approach 1:
The patent implements feedback mechanisms that continuously monitor device state and viewership patterns. When a device remains in a viewing state beyond expected durations or shows inconsistent patterns, the system adjusts its interpretation of subsequent data, using feedback from device behavior patterns to filter out false positives from indefinitely-left-on devices
4Measurement precision
If deterministic functions and hashing are applied to all device pairs, then deduplication accuracy is improved, but computational burden increases
Solution Approach 1:
The patent applies partial action by performing deterministic functions and hashing only on candidate device pairs that meet preliminary matching criteria, rather than all possible pairs. This selective application maintains high deduplication accuracy for relevant cases while significantly reducing unnecessary computational overhead from comparing unrelated devices
Data Source
AI summary
A system identifies redundancies among a plethora of viewership data of each of several markets. Some embodiments may obtain, from each of a plurality of different devices, a different set of viewership data, each set comprising different subsets that respectively relate to different entities; each of the subsets may indicate several time-based views of content over a period of time. These or other embodiments may determine a set of values by performing a set of deterministic functions using the time-based views of each of the subsets; compare the values, which relate to same entities and to same time intervals, of the determined sets of each distinct pair of the devices, and identify, from among the device pairs, candidate pairs. One exemplary output of this approach may be identifiers of the devices of a first pair of devices that is identified from among the candidate pairs.


