Cross-entity Heterogeneous Data Categorization via Seed Audiences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current approaches for categorizing heterogeneous data fail to capture global user interests and are expensive to scale, as they are often entity-specific and unable to incorporate travel-related data or competitor entity interactions in seed audience definitions.
Innovation Solution
The method generates category-specific seed audiences using unsupervised embedding representations of heterogeneous user events, such as search queries and purchases, which can co-occur with conversion events across multiple clients, allowing for cross-entity audience expansion and leveraging existing audience expansion approaches like Predictive Audiences/Predictive Segments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If entity-specific seed definitions are created from users who navigate to a webpage operated by the entity, then the categorization is specific to the entity, but global user interests such as travel indications from competitor entity interactions are not captured
Solution Approach 1:
The patent merges entity-specific user event data with global user event data from multiple entities into a unified seed audience definition. This combines the benefits of entity-specific precision with global user interest capture, resolving the contradiction between specific categorization and comprehensive information coverage.
Solution Approach 2:
The patent creates a universal seed audience definition that serves multiple entities simultaneously while capturing global user interests. This multi-functional approach allows the same seed definition to be used across different entities, preventing information loss about global user behaviors while maintaining entity-specific applicability.
2Measurement precision
If current entity-specific approaches are used for categorizing heterogeneous data, then results are specific to each entity, but the approach is expensive to scale for heterogeneous events
Solution Approach 1:
The patent creates a universal seed audience definition that can be applied across multiple entities and heterogeneous events simultaneously. This single universal definition replaces the need for separate entity-specific definitions, enabling efficient scaling while maintaining categorization accuracy through the unified approach.
Solution Approach 2:
The patent segments the categorization process into a universal seed definition phase and an entity-specific application phase. This segmentation allows the expensive categorization work to be done once universally, then efficiently applied to multiple entities, improving scaling productivity while preserving accuracy.
3Productivity
If automated audience expansion approaches are used, then audience definitions can be generated efficiently, but capturing latent user interests across multiple clients requires sophisticated processing
Solution Approach 1:
The patent introduces an intermediary processing layer that aggregates user events from multiple clients and entities before generating seed audience definitions. This intermediary layer simplifies the complexity by providing a unified view of user interests, enabling automated efficient audience generation while handling the sophisticated processing requirements centrally.
Data Source
AI summary
The techniques described herein relate to constructing and using seed audiences. In an embodiment, a method includes loading, by a processing device, a user event sequence, the user event sequence including a plurality of user events and a plurality of corresponding conversions; generating, by the processing device, a plurality of conversion neighborhoods based on the user event sequence, a given conversion neighborhood in the plurality of conversion neighborhood including at least one conversion rule and a set of user events from the plurality of user events; annotating, by the processing device, each conversion neighborhood in the plurality of conversion neighborhoods with categorical labels; and generating, by the processing device, seed audiences for each conversion neighborhood, a given seed audience including a ranked list of user events for each conversion rule associated with the conversion neighborhood.


