Modifying Public Datasets for Statistical Resemblance to Target Profiles

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models struggle to produce accurate results when trained on publicly available data that does not statistically resemble the target data, especially when the target data contains protected information, as it may lack relevant terminology or characteristics present in the target dataset.

Innovation Solution

The system modifies a corpus of publicly available data to statistically resemble the target dataset by characterizing both datasets with profiles and iteratively adjusting the publicly available data to match the target profile, allowing for accurate training and validation without accessing the protected data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If publicly available data is used for machine learning model training, then data accessibility is improved, but model accuracy deteriorates due to statistical differences from target data

Engineering Contradiction:
Improvedata accessibilityVSAvoidmodel accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent creates a modified copy of publicly available data that statistically resembles the target data distribution. Instead of directly using the original public data, the system generates a transformed version that copies the essential statistical characteristics of the protected target data while maintaining accessibility, thereby resolving the contradiction between data accessibility and model accuracy

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies parameter transformations to the publicly available data to alter its statistical distribution. By changing parameters such as data sampling rates, feature transformations, and distribution matching, the system modifies the public data to better match the target data's statistical properties, thus improving model accuracy while keeping the data accessible

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If target data containing protected information is accessed for model training, then model accuracy is improved, but data security deteriorates

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata security risk
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent introduces a data modification system as an intermediary between the protected target data and the machine learning model. This intermediary transforms the target data's statistical characteristics into a modified public dataset, allowing the model to learn from data that resembles the target without directly accessing sensitive information, thus maintaining both accuracy and security

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a statistical copy of the target data distribution using publicly available data. Instead of copying the actual sensitive data, the system copies only the statistical properties and patterns, enabling accurate model training while preserving the security of the original protected information

Inventive Principle:
Principle #26Copying

3Ease of operation

If publicly available data is used without modification, then data accessibility is maintained, but statistical resemblance to target data deteriorates

Engineering Contradiction:
Improvedata accessibilityVSAvoidstatistical resemblance
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent systematically changes the parameters of publicly available data to match the statistical distribution of the target data. This includes adjusting data sampling, applying transformations, and modifying feature distributions to achieve statistical resemblance, thereby maintaining accessibility while improving precision

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements a feedback mechanism where the statistical properties of the modified data are continuously evaluated against the target data distribution. Based on this feedback, the system iteratively adjusts the modification parameters to achieve better statistical resemblance, ensuring both accessibility and precision are optimized

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11630835B2Modifications of user datasets to support statistical resemblance
Publication Date: 2023.04.18 SALESFORCE INC
  • US11630835B2 patent drawing
  • US11630835B2 patent drawing
  • US11630835B2 patent drawing

AI summary

A method, apparatus, and system for modifying user datasets to support statistical resemblance is described. To support modifying user datasets to support statistical resemblance, an application may generate a first profile from a first corpus that includes first user data, generate a set of modified profiles from a second profile from a second corpus including second user data, wherein the first profile and the set of modified profiles includes respective sets of first and second attributes corresponding to one or both of text or metadata associated with the respective first and second user data, determine a mathematical distance between the first profile and each modified profile of the set of modified profiles based at least in part on a comparison between the first set of attributes and the second set of attributes, and finally, select a modified profile having a smallest determined mathematical distance.