Secure Data Linkage Orchestration Across Disparate Repositories
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generating accurate computerized predictions is challenging when different data sets are stored in disparate data storage by disparate parties, each having their own data security and data privacy obligations, compounded by high uncertainty and complexity in interrelationships, especially when dealing with probabilistic linkages and a large number of potential paths and spurious relationships.
Innovation Solution
A data processing orchestrator device or service securely interoperates with data sets at various points in time, combining data sets representing user intents and outcomes to establish probabilistic mappings, using machine learning and reinforcement learning mechanisms like multi-armed bandits to refine model representations and improve relevance of offers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data sets are stored in disparate data storage by disparate parties with data security and privacy obligations, then data security and privacy are maintained, but generating accurate computerized predictions becomes challenging
Solution Approach 1:
The patent introduces a trusted execution environment (TEE) as an intermediary component that securely stores and processes data from multiple disparate sources. The TEE acts as a mediator that enables data linkage and predictive analytics while maintaining security boundaries, allowing parties to collaborate without directly sharing sensitive data outside the secure environment.
Solution Approach 2:
The patent combines multiple disparate data sets from different sources within a unified trusted execution environment. By merging data storage and processing capabilities in a secure centralized location, the system enables holistic predictive analytics while maintaining the security and privacy requirements of individual data owners.
2Measurement precision
If multiple data sets are combined for holistic predictive analytics, then prediction accuracy improves, but uncertainty and complexity of interrelationships increase
Solution Approach 1:
The patent segments the complex predictive analytics process into distinct modules: data ingestion, data linkage establishment, model training, and prediction generation. Each module handles specific aspects of the data processing pipeline, making the overall system more manageable and interpretable despite handling complex multi-dataset relationships.
Solution Approach 2:
The patent implements feedback mechanisms where prediction outcomes are tracked and used to refine data linkages and model parameters. This iterative feedback loop helps reduce uncertainty by continuously improving the model based on actual outcomes, making the complex system more predictable over time.
3Adaptability or versatility
If probabilistic linkages are established between data sets, then comprehensive predictions are achieved, but the number of potential paths and spurious relationships increases
Solution Approach 1:
The patent applies parameter changes by adjusting the strength and confidence thresholds of data linkages based on statistical analysis. Linkages with low confidence or high spurious relationship potential are filtered or down-weighted, while strong probabilistic relationships are emphasized, simplifying the overall relationship graph while maintaining comprehensiveness.
4Measurement precision
If reinforcement learning mechanisms are used to distinguish offers, then offer relevance improves, but computational resources and training time increase
Solution Approach 1:
The patent performs preliminary actions by pre-training reinforcement learning models offline using historical data before deployment. This allows the system to learn offer distinction patterns in advance, reducing the computational burden and training time required during live operation while maintaining high offer relevance.
Data Source
AI summary
Systems and methods for establishing data linkages are described in various embodiments. A system architecture is described which provides a data processing orchestrator device or service which securely interoperates with data sets at various points in time associated with a set of interactions a user may have with computer systems. The data sets are obtained from different data repositories, and are combined together for analysis such that a first data set representing intents (e.g., web search/browse history) can be combined together with a second data set representing outcomes (e.g., purchase transaction history, web site shopping carts).


