Cross-Device User Identification via Hybrid Data Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for identifying devices associated with a particular user are limited in scalability and accuracy, especially when dealing with large data sets of millions of devices and users, leading to ineffective tracking of user interactions across multiple devices.
Innovation Solution
The system combines deterministic and probabilistic data to generate a hybrid cross-device data structure, using user identifiers and common usage patterns to link devices, allowing for accurate identification of user devices in large-scale analytics environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If deterministic methods are used to identify devices associated with a user, then accuracy in identifying user devices is improved, but scalability to large data sets deteriorates
Solution Approach 1:
The patent segments the device identification process into two distinct phases: a deterministic phase that accurately identifies devices with available login data, and a probabilistic phase that scales to process remaining devices using usage patterns. This segmentation allows each phase to optimize for its specific strength while avoiding the weaknesses of the other approach.
Solution Approach 2:
The patent merges deterministic identification results with probabilistic identification results into a unified device-to-user mapping. By combining these two approaches, the system achieves both the accuracy of deterministic methods and the scalability of probabilistic methods, resolving the contradiction between precision and productivity.
2Measurement precision
If deterministic data collection is used, then accuracy in linking devices to users is improved, but data availability deteriorates due to limited deterministic identifiers
Solution Approach 1:
The patent introduces probabilistic usage pattern data as an intermediary between devices that lack deterministic identifiers and user profiles. This intermediary layer allows the system to bridge the gap where deterministic data is unavailable, maintaining data availability while preserving accuracy through the hybrid approach.
Solution Approach 2:
The patent changes the identification parameters from relying solely on deterministic login identifiers to incorporating probabilistic usage pattern parameters. This parameter change expands the available data pool significantly, allowing identification of devices that would otherwise remain unlinked to users.
3Productivity
If probabilistic data alone is used for device identification, then scalability to large data sets is improved, but accuracy in identifying user devices deteriorates
Solution Approach 1:
The patent performs preliminary deterministic identification on all devices before applying probabilistic methods. This preliminary action ensures that devices with available deterministic data are accurately identified first, establishing a foundation of high-confidence mappings that prevent probabilistic errors from propagating.
Solution Approach 2:
The patent implements feedback mechanisms where deterministic identification results are used to validate and refine probabilistic identifications. The system continuously compares probabilistic results against deterministic ground truth where available, adjusting probabilistic models to maintain high accuracy while preserving scalability.
Data Source
AI summary
Systems and methods are disclosed for clustering multiple devices that are associated with particular users by utilizing both probabilistic and deterministic data derived from analytics information on the users. An analytics computing system generates at least one deterministic device cluster that groups a first set of devices associated with a first user. The first set of devices share deterministic user identifiers specific to the first user. The analytics computing system also identifies a probabilistic link between a device in the first set of devices and additional devices. The probabilistic link indicates common usage patterns between two devices. Based on the probabilistic link, the analytics computing system generates a data structure that includes the deterministic device cluster and the additional devices.


