Configuration-Driven Dataset Generation for Private Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for training content-serving machine learning models lack efficiency and privacy, relying on individual user data collection and failing to provide accurate and secure training datasets.

Innovation Solution

A configuration-driven pipeline that aggregates anonymized contextual data from multiple client profiles, ensuring data privacy and accuracy by extracting and correlating labels and features, and validating the dataset to meet predefined criteria before training the model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If individual user data collection is used for training models, then model training can be performed, but data privacy is compromised and processing burden increases

Engineering Contradiction:
Improvemodel training accuracyVSAvoiddata privacy risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent merges data from multiple client profiles into aggregated datasets before model training. By combining individual user data points into collective statistical patterns, the system maintains training effectiveness while eliminating personally identifiable information, thus resolving the contradiction between training accuracy and data privacy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an intermediary aggregation layer that processes individual user data before it reaches the model training pipeline. This intermediary step transforms raw individual data into anonymized aggregated statistics, serving as a mediator that protects privacy while preserving the essential patterns needed for accurate model training.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If individual user data is collected and processed, then training data can be obtained, but processing burden and resource consumption increase

Engineering Contradiction:
Improvetraining data availabilityVSAvoidprocessing efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent performs preliminary aggregation of user data into anonymized datasets before model training begins. By pre-processing and consolidating data from multiple clients in advance, the system reduces the processing burden during actual model training, improving overall processing efficiency while maintaining adequate training data availability.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If comprehensive user data is collected, then model accuracy can be improved, but data security and privacy compliance become problematic

Engineering Contradiction:
Improvemodel prediction accuracyVSAvoiddata security and compliance
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent extracts only the essential statistical patterns and aggregated features from individual user data, leaving behind personally identifiable information. By taking out only the necessary aggregated metrics needed for model training, the system achieves adequate prediction accuracy while ensuring data security and privacy compliance.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250265359A1Configuration based dataset generation for content serving systems
Publication Date: 2025.08.21 GOOGLE LLC
  • US20250265359A1 patent drawing
  • US20250265359A1 patent drawing
  • US20250265359A1 patent drawing

AI summary

The present disclosure provides methods, systems, and media for a computing device that provide a configuration-driven pipeline enabling coalescing numerous data sources to curate customized datasets that can be used in training and/or inference operations relating to content machine learning models deployed for identifying digital components to provide to client devices. In one aspect, the methods include receiving a request for generation of training data for training a contextual model used to identify digital components; identifying, using configuration files and the received request, a key and corresponding value type for extraction from the data; extracting, from the data corresponding to the plurality of client profiles, data for the identified key and value type; aggregating the data for the identified key and value type, to obtain an aggregated dataset; determining that the aggregated dataset satisfies a set of validation criteria; and in response, providing the aggregated dataset as training data.