ML Model Training via DMZ Data Views

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Software vendors face challenges in developing customer-specific machine learning (ML) models due to the reluctance of customers to share production data, leading to suboptimal performance when trained on synthetic or anonymized data.

Innovation Solution

A data access control and workload management framework that allows software vendors to train ML models within a demilitarized zone (DMZ) of each customer using transformed production data, ensuring secure and controlled access while maintaining data privacy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If customers share production data with software vendors for ML model training, then ML model performance and accuracy improve, but data privacy and security risks increase

Engineering Contradiction:
ImproveML model accuracyVSAvoiddata privacy risk
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent introduces a DMZ (demilitarized zone) environment as an intermediary between the customer's production data and the vendor's ML training processes. The DMZ acts as a secure buffer that allows data access for training while maintaining isolation and control, thus improving model accuracy without compromising data privacy. This is implemented through a data science pool with controlled views that enable vendor access to transformed production data in a secure manner.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates transformed copies of production data through the data science pool, where views provide processed representations of the original data. These copies maintain the statistical properties needed for accurate ML training while removing direct access to sensitive production data. The transformation process creates data that can be used for training without exposing the actual production data, thus improving model performance while protecting privacy.

Inventive Principle:
Principle #26Copying

2Object-affected harmful factors

If vendors use synthetic or anonymized data for ML model training, then data privacy is protected, but ML model performance deteriorates

Engineering Contradiction:
Improvedata privacy protectionVSAvoidML model performance
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The patent transforms production data through various parameter changes in the data science pool, creating views that modify data representation while preserving statistical properties. This transformation allows the data to maintain its predictive power for ML training while altering its form to protect privacy. The parameter changes enable the data to serve dual purposes: maintaining model training effectiveness while providing privacy protection.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The transformed data views in the data science pool serve as an intermediary between raw production data and ML training processes. This intermediary layer preserves the statistical characteristics needed for accurate training while removing direct exposure to sensitive data, thus maintaining model performance while protecting privacy.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If vendors access customer production data directly for ML training, then model customization improves, but system complexity and security management increase

Engineering Contradiction:
Improvemodel customizationVSAvoidaccess control complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the data access architecture into distinct components: production data systems, data science pools with multiple views, DMZ environments, and training systems. This segmentation allows different levels of access and control for different purposes. The data science pool is divided into multiple views that can be selectively accessed, enabling customized ML training while simplifying security management through modular access control.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a new dimensional layer (the data science pool with views) between the traditional customer data system and vendor training system. This additional dimension provides multiple access paths and control points, enabling customized model training while managing complexity through structured view definitions rather than direct access management.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11551141B2Data access control and workload management framework for development of machine learning (ML) models
Publication Date: 2023.01.10 SAP SE
  • US11551141B2 patent drawing
  • US11551141B2 patent drawing
  • US11551141B2 patent drawing

AI summary

Methods, systems, and computer-readable storage media for providing a software system to each customer in a set of customers, each customer being associated with a customer system in a set of customer systems, the software system including a set of views in a data science pool, each of the views in the set of views providing a data set based on production data of respective customers; for each customer system: accessing at least one data set within the customer system through a released view provided in a DMZ within the customer system and corresponding to a respective view in the set of views, and triggering training of a ML model in the DMZ to provide and results; and selectively publishing the ML model for consumption by each of the customers in the set of customers based on a set of results comprising the results from each customer system.