Synthetic Data Generator for ML Model Validation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Candidate recommendation systems face challenges in validating the performance of machine learning models trained on non-industry relevant data, as they lack industry-specific testing data to accurately verify the models' performance before deployment in the intended industry.

Innovation Solution

A system and method that generate synthetic test data based on computed parameters for applicants and job requisitions, using a different data source than the training data, to validate the performance of machine learning ranking models by populating work experience and job requisition parameters with text data from various sources, and applying this synthetic data to the model to generate a ranking list for evaluation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a machine learning model is trained on non-industry relevant data, then the model can be developed and tested initially, but the model cannot be accurately validated for industry-specific performance

Engineering Contradiction:
ImproveModel development easeVSAvoidModel validation reliability
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent creates synthetic test data that copies the structural and statistical properties of industry-specific data without requiring actual industry data during model development. This synthetic data serves as a faithful replica that enables validation while maintaining the ability to develop models on general data first.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary generation of synthetic industry-specific test data before actual model deployment. This allows the validation framework to be prepared in advance with appropriate test cases that reflect industry conditions, enabling thorough validation before real-world deployment.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If industry-specific test data is used for validation, then accurate performance verification is possible, but data availability and accessibility are limited

Engineering Contradiction:
ImprovePerformance measurement accuracyVSAvoidData source accessibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent transforms the availability constraint by changing the parameter of data source from requiring actual industry data to generating synthetic data with controlled parameters that match industry data characteristics. This allows validation without dependency on scarce industry-specific datasets.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces synthetic data as an intermediary between the model validation process and actual industry data. This intermediary enables accurate validation by mediating the interaction between models and industry conditions without requiring direct access to proprietary industry datasets.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Stability of the object's composition

If synthetic test data is generated from the same data source as training data, then data consistency is maintained, but model validation accuracy is reduced due to data leakage

Engineering Contradiction:
ImproveData consistencyVSAvoidValidation accuracy
Core Design Contradiction:
Stability of the object's compositionVSMeasurement precision

Solution Approach 1:

The patent segments the data usage into distinct functions: training data from one source and synthetic test data generated from different sources or through different processes. This segmentation prevents data leakage while maintaining the structural consistency needed for valid validation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates test data as a synthetic copy that replicates the statistical properties and patterns of industry data without being identical to the training data. This copying approach maintains data consistency in terms of distribution and structure while ensuring independence from the training set for unbiased validation.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11556870B2System and method for validating a candidate recommendation model
Publication Date: 2023.01.17 ORACLE INT CORP
  • US11556870B2 patent drawing
  • US11556870B2 patent drawing
  • US11556870B2 patent drawing

AI summary

In some examples a first parameter for respective applicants or candidates can be computed based on respective text data from a text dataset that can include a plurality of different types of text data. The first parameter can be populated with a given portion of text of the respective text data. A second parameter for a job requisition can be computed based on the respective text data used to compute the first parameter for a given applicant or candidate. The second parameter can be populated with a different portion of text of the respective text data used to compute the first parameter. Synthetic test data can be generated based on the computed parameters to test a machine learning (ML) ranking model that has been trained on training data that is from a different data source than the text dataset to validate a performance of the ML ranking model.