Trusted Third-Party ML Model Training With Restricted Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning technologies face challenges in training models across disparate data features from different entities, particularly when these features are restricted and inaccessible to each other, hindering collaboration and data sharing while ensuring privacy and security.

Innovation Solution

A system and method for training a machine learning model on behalf of a first entity using restricted data features from a second entity, where the model is trained by an independent modeling system that accesses and combines metadata and encrypted data features, ensuring secure and limited access, facilitating collaboration without exposing sensitive data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If machine learning models are trained using data features from multiple entities, then model performance and diversity are improved, but data privacy and security are compromised

Engineering Contradiction:
Improvemodel performanceVSAvoiddata privacy risk
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent introduces a trusted third-party modeling system as an intermediary that trains machine learning models using data features from multiple entities without allowing direct access to the raw data. The modeling system acts as a mediator that processes encrypted or secured data features, generates trained models, and returns them to requesting entities while maintaining data isolation and privacy boundaries between entities.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates copies of data features in the form of trained machine learning models that capture the essential patterns and relationships from the original data without exposing the actual data itself. Entities receive model copies that replicate the predictive capabilities derived from their data features while the original sensitive data remains protected and inaccessible to other entities.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If restricted data features are made accessible for model training, then collaboration between entities is improved, but data security and access control are weakened

Engineering Contradiction:
Improvecollaboration capabilityVSAvoiddata security
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The trusted third-party modeling system serves as an intermediary that enables collaboration between entities with restricted data features. It manages access control, authentication, and authorization mechanisms while processing data features from multiple entities. The intermediary ensures that entities can collaborate on model training without directly sharing or exposing their restricted data to each other.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the data access and model training process into distinct isolated environments. Each entity's data features are processed in separate secured contexts within the modeling system, with fine-grained access controls that prevent unauthorized cross-entity data access while still allowing the model to learn from aggregated patterns across all entities' data features.

Inventive Principle:
Principle #1Segmentation

3Ease of operation

If direct access to data features is provided for model training, then training flexibility and customization are improved, but data exposure and security risks increase

Engineering Contradiction:
Improvetraining flexibilityVSAvoiddata exposure
Core Design Contradiction:
Ease of operationVSObject-generated harmful factors

Solution Approach 1:

The modeling system acts as an intermediary that provides training flexibility through configurable model architectures, hyperparameters, and training protocols while maintaining data security. Entities can specify their training requirements and preferences without directly accessing or exposing their data features to the requesting party. The intermediary translates these requirements into secure training operations on isolated data copies.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system creates secure copies of data features that are used for model training instead of providing direct access to the original data. These copies contain the necessary information for training but are isolated from the original data sources and cannot be used to reconstruct or access the sensitive underlying data, thus enabling training flexibility without data exposure.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20240095579A1Restricted reuse of machine learning model data features
Publication Date: 2024.03.21 AT&T INTELLECTUAL PROPERTY I L P
  • US20240095579A1 patent drawing
  • US20240095579A1 patent drawing
  • US20240095579A1 patent drawing

AI summary

A processing system including at least one processor may obtain a request from a first entity to train a machine learning model, access at least one data feature of at least a second entity, and train the machine learning model on behalf of the first entity in accordance with the at least one data feature of the at least the second entity to generate a trained machine learning model, where the at least one data feature of the at least the second entity is a restricted data feature that is inaccessible to the first entity. The processing system may then provide the trained machine learning model to the first entity.