Probabilistic Entity Linking Across Disjoint Identifier Spaces
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges in attributing information about entities across disparate operating environments, such as internet and mobile devices, due to disjointed identifier spaces, which prevents seamless content serving and decision-making.
Innovation Solution
Building models based on features associated with a source panel of entities operating in multiple identifier spaces to predict the likelihood of an entity's presence in another space, using a server system with a model building module and scoring module to compute a score indicating the entity's presence across different spaces.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If information is captured for entities in one operating environment, then content consumption habits can be analyzed, but the information cannot be attributed to the same entity in a different operating environment
Solution Approach 1:
The patent introduces a probabilistic linkage mechanism as an intermediary between disparate identifier spaces. Instead of requiring direct mapping between identifier spaces (which would be complex and invasive), the system uses behavioral patterns and contextual signals to create probabilistic associations that enable information attribution across environments without complex direct mapping infrastructure
Solution Approach 2:
The patent replaces the mechanical approach of direct identifier mapping with a statistical/probabilistic model. Instead of using deterministic mapping mechanisms that require complex infrastructure, the system substitutes a machine learning-based probabilistic framework that infers entity presence across environments based on behavioral patterns, thereby reducing system complexity
2Reliability
If direct mapping between identifier spaces is implemented, then information can be attributed across environments, but privacy concerns and system complexity increase
Solution Approach 1:
The probabilistic linkage model serves as an intermediary layer that enables reliable information attribution without requiring direct exposure or complex mapping between identifier spaces. This intermediary approach maintains reliability by using multiple contextual signals while preserving privacy and reducing complexity
Solution Approach 2:
The patent uses temporary, context-specific probabilistic linkages rather than permanent direct mappings. These probabilistic associations are created only when needed for specific content serving decisions and are discarded afterward, avoiding the need for persistent complex mapping infrastructure while maintaining attribution reliability
3Adaptability or versatility
If probabilistic models are used to infer entity presence, then information can be utilized across disparate environments, but computational resources increase
Solution Approach 1:
The system applies probabilistic modeling selectively rather than universally. It uses probabilistic inference only for entities where cross-environment attribution is beneficial and feasible, rather than applying complex models to all entities, thereby reducing overall computational energy consumption while maintaining adaptability for content serving
Solution Approach 2:
The patent dynamically adjusts the complexity and precision of probabilistic models based on contextual parameters such as data availability, entity type, and content serving requirements. By changing model parameters adaptively rather than using fixed high-complexity models, the system achieves cross-environment versatility with optimized energy consumption
Data Source
AI summary
Embodiments of the invention build models to predict the likelihood of entities that operate in a given identifier space also operating in a disjoined identifier space based on a source panel of entities that operate in one or both of the identifier spaces. In operation, a model building engine builds a model based on features associated with the source panel and features associated with standard populations in the given identifier space. The model is used to determine whether the target entity is more similar to those entities in the source panel that operate only in the given identifier space or those entities in the source panel that operate in both identifier spaces.


