Synthetic Data Generation for Privacy-Compliant Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data protection and privacy regulations restrict the use of real data for training artificial intelligence and machine learning models, making it challenging to create effective training data while ensuring compliance.

Innovation Solution

A computer-implemented process that accesses static and dynamic system data to derive relevant data, using distance-based and importance-based metrics, generates synthetic data that complies with privacy regulations, and is used to train machine learning models for classification purposes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real data is used for training machine learning models, then model training effectiveness is improved, but data protection and privacy compliance deteriorates

Engineering Contradiction:
Improvemodel training effectivenessVSAvoiddata protection compliance
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic data that copies the statistical properties, distributions, and relationships of real data without containing actual sensitive information. This allows machine learning models to be trained on data that mimics real-world patterns while ensuring privacy compliance, as the synthetic data contains no personally identifiable information

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces synthetic data as an intermediary between real data and machine learning models. This intermediary preserves the essential characteristics needed for model training while eliminating privacy risks, allowing the model to learn from data that is statistically representative but legally safe to use

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If synthetic data is generated to comply with privacy regulations, then data protection compliance is improved, but model training effectiveness may deteriorate

Engineering Contradiction:
Improvedata protection complianceVSAvoidmodel training effectiveness
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent carefully adjusts parameters of the synthetic data generation process to preserve critical statistical properties such as data distributions, correlations, and relationships. By maintaining these parameters while removing sensitive information, the synthetic data remains effective for model training while ensuring privacy compliance

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If manual data selection and exploration is performed, then data relevance accuracy is improved, but productivity deteriorates

Engineering Contradiction:
Improvedata relevance accuracyVSAvoiddata processing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements automated systems that self-select and explore data without requiring manual intervention. The system automatically identifies relevant data, explores its properties, and generates synthetic data, eliminating the time-consuming manual processes while maintaining accuracy through algorithmic precision

Inventive Principle:
Principle #25Self-service

4Manufacturing precision

If comprehensive data exploration is performed to understand data structures and correlations, then synthetic data quality is improved, but time consumption increases

Engineering Contradiction:
Improvesynthetic data qualityVSAvoiddata exploration time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs comprehensive data exploration and analysis as a preliminary step before synthetic data generation. By understanding data structures, distributions, and correlations upfront, the system can efficiently generate high-quality synthetic data without needing to revisit exploration phases, reducing overall time consumption while maintaining quality

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11630881B2Data insight automation
Publication Date: 2023.04.18 SAP SE
  • US11630881B2 patent drawing
  • US11630881B2 patent drawing
  • US11630881B2 patent drawing

AI summary

Static and dynamic process data of a system are accessed. Thereafter, using this accessed process data, a subset of such data forming relevant data for a particular context is derived. The data is then explored using a computer-implemented process or processes to automatically get insight into information about structures, distributions and correlations of the relevant data. Rules can be generated based on the exploring of relevant data that describe data dependencies within the relevant data. These generated rules can later be used to generate synthetic data. Such synthetic data, in turn, can be used to for a variety of purposes including the training of a machine learning model while, at the same time, complying with applicable privacy and data protection laws and regulations.