Web Data Emulator System for Privacy-Preserving Synthetic Data Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users are hesitant to share information about their device interactions due to privacy concerns, which limits content providers' ability to understand user behavior and market appropriate products, services, and content effectively.
Innovation Solution
A web data emulator system that processes user interaction data to generate anonymous, synthetic data preserving statistical properties, allowing content providers to analyze user behavior without compromising user privacy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If user interaction data is collected and shared with content providers, then content providers can understand user behavior and market products effectively, but user privacy is compromised
Solution Approach 1:
The patent creates synthetic copies of user interaction data that replicate the statistical properties and patterns of real user behavior without containing actual personal information. The synthetic data generation system produces artificial clickstream data, browsing patterns, and interaction metrics that mirror real user behavior distributions, enabling content providers to analyze user behavior while preserving privacy through substitution of real data with realistic synthetic alternatives
Solution Approach 2:
The synthetic data generation system acts as an intermediary layer between raw user interaction data and content provider analysis. Instead of directly sharing sensitive user data, the system processes real data through empirical estimation and simulation to produce anonymized synthetic data that serves as a safe mediator, allowing content providers to gain behavioral insights without direct exposure to personal information
2Object-affected harmful factors
If anonymous synthetic data is generated from web data, then user privacy is preserved, but data utility for analysis may be reduced
Solution Approach 1:
The patent transforms the parameters and structure of raw user interaction data through empirical estimation processes, converting detailed individual-level data into aggregated statistical distributions and probability models. This parameter transformation maintains the essential behavioral patterns and statistical properties needed for analysis while removing identifying information, achieving privacy protection without sacrificing analytical utility
Solution Approach 2:
The patent replaces direct mechanical use of raw user data with a computational simulation system that uses empirical estimations and probability distributions to generate synthetic data. Instead of directly processing and sharing actual user interaction records, the system substitutes a mathematical modeling approach that preserves behavioral insights while eliminating privacy risks associated with direct data sharing
Data Source
AI summary
A device receives web data, associated with user devices, that is generated based on interactions of the user devices with a network and one or more content provider devices. The device removes erroneous or objectionable web data from the web data to generate a subset of the web data, and categorizes the subset of the web data by assigning categories to the subset of the web data. The device performs an empirical estimation of the categorized subset of the web data to generate empirical estimations. The device performs a simulation of the empirical estimations to generate synthetic data that corresponds to the web data and removes private information relating to the user devices and users of the user devices, and stores the synthetic data in a storage device.


