Configuration-Driven Dataset Generation for Private Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for training content-serving machine learning models lack efficiency and privacy, relying on individual user data collection and failing to provide accurate and secure training datasets.
Innovation Solution
A configuration-driven pipeline that aggregates anonymized contextual data from multiple client profiles, ensuring data privacy and accuracy by extracting and correlating labels and features, and validating the dataset to meet predefined criteria before training the model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If individual user data collection is used for training models, then model training can be performed, but data privacy is compromised and processing burden increases
Solution Approach 1:
The patent merges data from multiple client profiles into aggregated datasets before model training. By combining individual user data points into collective statistical patterns, the system maintains training effectiveness while eliminating personally identifiable information, thus resolving the contradiction between training accuracy and data privacy.
Solution Approach 2:
The patent introduces an intermediary aggregation layer that processes individual user data before it reaches the model training pipeline. This intermediary step transforms raw individual data into anonymized aggregated statistics, serving as a mediator that protects privacy while preserving the essential patterns needed for accurate model training.
2Quantity of substance
If individual user data is collected and processed, then training data can be obtained, but processing burden and resource consumption increase
Solution Approach 1:
The patent performs preliminary aggregation of user data into anonymized datasets before model training begins. By pre-processing and consolidating data from multiple clients in advance, the system reduces the processing burden during actual model training, improving overall processing efficiency while maintaining adequate training data availability.
3Measurement precision
If comprehensive user data is collected, then model accuracy can be improved, but data security and privacy compliance become problematic
Solution Approach 1:
The patent extracts only the essential statistical patterns and aggregated features from individual user data, leaving behind personally identifiable information. By taking out only the necessary aggregated metrics needed for model training, the system achieves adequate prediction accuracy while ensuring data security and privacy compliance.
Data Source
AI summary
The present disclosure provides methods, systems, and media for a computing device that provide a configuration-driven pipeline enabling coalescing numerous data sources to curate customized datasets that can be used in training and/or inference operations relating to content machine learning models deployed for identifying digital components to provide to client devices. In one aspect, the methods include receiving a request for generation of training data for training a contextual model used to identify digital components; identifying, using configuration files and the received request, a key and corresponding value type for extraction from the data; extracting, from the data corresponding to the plurality of client profiles, data for the identified key and value type; aggregating the data for the identified key and value type, to obtain an aggregated dataset; determining that the aggregated dataset satisfies a set of validation criteria; and in response, providing the aggregated dataset as training data.


