Deidentified Content Hashing for Privacy-Compliant Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing medical data privacy laws and contractual limitations restrict the long-term storage and usage of user data for training machine learning models, necessitating a secure method to deidentify personal information while allowing model training.

Innovation Solution

A two-step hashing process using a first hashing algorithm to generate deidentified content and a second hashing algorithm to create surrogates, enabling secure sharing and training of machine learning models without revealing personal information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If personal information is stored and used for training machine learning models, then model training accuracy is improved, but data privacy compliance deteriorates

Engineering Contradiction:
Improvemodel training accuracyVSAvoiddata privacy compliance
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent introduces hash values as an intermediary representation between personal information and machine learning models. The hashing function transforms personal information into hash values that preserve statistical properties for model training while preventing direct access to original data, thus serving as a mediator that satisfies both training accuracy and privacy compliance requirements

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a copied representation of personal information through hashing. The hash values are mathematical transformations that replicate the statistical characteristics needed for model training without being the original personal information, allowing models to learn from data patterns while the original sensitive information remains protected

Inventive Principle:
Principle #26Copying

2Reliability

If personal information is deidentified using hashing, then data privacy compliance is improved, but model training effectiveness deteriorates

Engineering Contradiction:
Improvedata privacy complianceVSAvoidmodel training effectiveness
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent changes the parameter representation of personal information by transforming it into hash values through a mathematical function. This parameter transformation maintains the statistical properties necessary for model training while ensuring privacy compliance, as the hash values retain distribution characteristics without exposing original data

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If user data is stored long-term for model training, then model accuracy is improved, but data security risks worsen

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata security risks
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent uses hash values as an intermediary storage form instead of storing original personal information. This intermediary representation allows long-term storage for model training purposes while minimizing security risks, as the hashed data cannot be easily reversed to obtain original sensitive information even if storage systems are compromised

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12393728B2System and method for generating deidentified content
Publication Date: 2025.08.19 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12393728B2 patent drawing
  • US12393728B2 patent drawing
  • US12393728B2 patent drawing

AI summary

A method, computer program product, and computing system for processing raw content to identify personal information; replacing the personal information within the raw content with a first mathematical representation of the personal information generated using a first hashing algorithm, thus defining deidentified content; and defining a selected surrogate for the first mathematical representation, wherein the selected surrogate has a second mathematical representation generated using a second hashing algorithm that is equivalent to the first mathematical representation of the personal information that was generated using the first hashing algorithm.