Web Data Emulator System for Privacy-Preserving Synthetic Data Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users are hesitant to share information about their device interactions due to privacy concerns, which limits content providers' ability to understand user behavior and market appropriate products, services, and content effectively.

Innovation Solution

A web data emulator system that processes user interaction data to generate anonymous, synthetic data preserving statistical properties, allowing content providers to analyze user behavior without compromising user privacy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If user interaction data is collected and shared with content providers, then content providers can understand user behavior and market products effectively, but user privacy is compromised

Engineering Contradiction:
Improveuser behavior informationVSAvoidprivacy concern
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic copies of user interaction data that replicate the statistical properties and patterns of real user behavior without containing actual personal information. The synthetic data generation system produces artificial clickstream data, browsing patterns, and interaction metrics that mirror real user behavior distributions, enabling content providers to analyze user behavior while preserving privacy through substitution of real data with realistic synthetic alternatives

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The synthetic data generation system acts as an intermediary layer between raw user interaction data and content provider analysis. Instead of directly sharing sensitive user data, the system processes real data through empirical estimation and simulation to produce anonymized synthetic data that serves as a safe mediator, allowing content providers to gain behavioral insights without direct exposure to personal information

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If anonymous synthetic data is generated from web data, then user privacy is preserved, but data utility for analysis may be reduced

Engineering Contradiction:
Improveprivacy protectionVSAvoidbehavioral analysis accuracy
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The patent transforms the parameters and structure of raw user interaction data through empirical estimation processes, converting detailed individual-level data into aggregated statistical distributions and probability models. This parameter transformation maintains the essential behavioral patterns and statistical properties needed for analysis while removing identifying information, achieving privacy protection without sacrificing analytical utility

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces direct mechanical use of raw user data with a computational simulation system that uses empirical estimations and probability distributions to generate synthetic data. Instead of directly processing and sharing actual user interaction records, the system substitutes a mathematical modeling approach that preserves behavioral insights while eliminating privacy risks associated with direct data sharing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9866454B2Generating anonymous data from web data
Publication Date: 2018.01.09 VERIZON PATENT & LICENSING INC
  • US9866454B2 patent drawing
  • US9866454B2 patent drawing
  • US9866454B2 patent drawing

AI summary

A device receives web data, associated with user devices, that is generated based on interactions of the user devices with a network and one or more content provider devices. The device removes erroneous or objectionable web data from the web data to generate a subset of the web data, and categorizes the subset of the web data by assigning categories to the subset of the web data. The device performs an empirical estimation of the categorized subset of the web data to generate empirical estimations. The device performs a simulation of the empirical estimations to generate synthetic data that corresponds to the web data and removes private information relating to the user devices and users of the user devices, and stores the synthetic data in a storage device.