Synthetic Data Protocols for Privacy-Preserving Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data generation methods for user behavior analysis and predictive modeling face challenges in balancing data privacy with resource efficiency, as they require direct utilization of customer data, which poses privacy concerns and is resource-intensive, straining storage and computational resources without addressing privacy concerns or optimizing efficiency.
Innovation Solution
A system and method for altered data generation using electronic arrangement protocols that modify real-world data in real-time to obfuscate sensitive information, generating synthetic data through electronic arrangement protocols on a distributed ledger, providing data similar to real-world data without compromising privacy, and enabling transient use to reduce resource consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If direct customer data is used for user behavior analysis and predictive modeling, then data accuracy and model performance are improved, but data privacy is compromised and storage resources are consumed
Solution Approach 1:
The patent creates synthetic copies of customer data that replicate the statistical properties and patterns of real data without containing actual personal information. These synthetic data copies are generated through electronic arrangement protocols that mimic real-world data relationships, enabling accurate predictive modeling while eliminating privacy risks associated with using real customer data
Solution Approach 2:
The patent generates transient synthetic data that can be created and discarded as needed for specific analytical tasks. Instead of storing persistent copies of real customer data, the system creates temporary synthetic data representations that fulfill analytical requirements and are then removed, reducing both privacy risks and long-term storage resource consumption
2Reliability
If test data is generated for user behavior analysis, then comprehensive data coverage is achieved, but storage space and computational resources are significantly consumed
Solution Approach 1:
The patent implements a dynamic data generation system where synthetic test data is created on-demand based on specific analytical requirements rather than pre-generating and storing all possible test scenarios. The electronic arrangement protocols dynamically generate data with appropriate characteristics for each specific use case, reducing overall storage requirements while maintaining comprehensive data coverage for analysis
Solution Approach 2:
The patent changes the fundamental parameter of data representation by using synthetic data with modified characteristics that preserve statistical properties without replicating actual customer information. This parameter change allows the system to achieve comprehensive data coverage for reliable analysis while using significantly less storage space compared to storing actual customer data or exhaustive test datasets
Data Source
AI summary
Systems, computer program products, and methods are described herein for altered data generation and transient use via electronic arrangement protocols. The present disclosure includes retrieving data of a first database, aggregating the data with additional data to form an aggregated data object, tagging the aggregated data object with at least one summary tag, searching an electronic arrangement protocol repository to identify at least one electronic arrangement protocol related to the at least one summary tag, wherein the at least one electronic arrangement protocol may include a first electronic arrangement protocol rule, for a first element of the at least one summary tag, comprising instructions to alter identifiable information, fabricating synthetic data based on the aggregated data object using the at least one electronic arrangement protocol, generating and appending the non-fungible token to a distributed ledger.


