Predictive Layers for Siloed Data Privacy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in utilizing siloed data without sharing it, particularly due to regulatory restrictions and data privacy concerns, which limits the ability of systems to enhance predictive accuracy by combining data from multiple sources.
Innovation Solution
The development of systems and methods that generate and utilize predictive layers and base models to allow participating systems to benefit from each other's data without exchanging the actual data, using a common-data layer and model-configuration layer to determine associations and fit models for shared data types.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If data is shared between systems to improve predictive accuracy, then predictive accuracy is improved, but data privacy and regulatory compliance deteriorate
Solution Approach 1:
The patent introduces trained machine learning models as intermediaries between data sources and systems needing predictive capabilities. These models are trained on siloed data locally and then deployed to other systems, enabling predictive accuracy improvement without direct data sharing. The models act as mediators that transfer knowledge rather than raw data, thus maintaining privacy compliance while achieving the desired predictive performance.
Solution Approach 2:
The patent creates copies of trained models that can be deployed across multiple systems. Instead of sharing the actual data, the system generates model copies that encapsulate the learned patterns and relationships. These model copies can be distributed and executed locally, providing predictive capabilities without exposing the underlying sensitive data, thereby resolving the contradiction between accuracy improvement and privacy protection.
2Reliability
If data is siloed to maintain privacy compliance, then data privacy is protected, but predictive accuracy and data sample size deteriorate
Solution Approach 1:
Trained models serve as intermediaries that bridge the gap between siloed data and predictive needs. Systems can deploy these model intermediaries to leverage patterns learned from other systems' data without directly accessing or sharing the actual data. This maintains privacy compliance while still achieving improved predictive accuracy through the knowledge embedded in the model intermediaries.
Solution Approach 2:
The patent segments the data utilization process into two distinct phases: local data training and model deployment. The training phase occurs locally within each system's data silo, maintaining privacy compliance. The deployment phase uses the trained models externally, achieving predictive accuracy improvement. This segmentation allows the system to benefit from both data isolation and model sharing.
3Quantity of substance
If data is aggregated from multiple sources to increase sample size, then predictive accuracy is improved, but system complexity and data transfer requirements deteriorate
Solution Approach 1:
The patent creates model copies that encapsulate the aggregated knowledge from multiple data sources. Instead of physically aggregating large volumes of data across systems, the system generates compact model copies that contain the essential patterns and relationships. These model copies can be deployed and executed with minimal infrastructure complexity, avoiding the need for complex data aggregation pipelines while still achieving the benefits of large sample sizes.
Solution Approach 2:
The patent extracts the essential predictive knowledge from large datasets and consolidates it into trained models. This extraction process removes the bulk of the raw data while retaining the critical patterns and relationships. The resulting models are compact and can be deployed without requiring the original large datasets to be stored or transferred, thereby reducing system complexity while maintaining predictive accuracy.
Data Source
AI summary
Systems and methods for models utilizing siloed data are disclosed. For example, data stored with and/or available to one or more systems may be siloed such that it may not be aggregated and/or shared with other systems. The presently-disclosed systems and methods generate and utilize predictive layers and models to allow each system to predict outcomes using its own data and then models are shared between systems to allow each associated system to gain the benefits of the data of other systems without aggregating such data or otherwise sharing the data.


