Synthetic Data Generator for ML Model Validation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Candidate recommendation systems face challenges in validating the performance of machine learning models trained on non-industry relevant data, as they lack industry-specific testing data to accurately verify the models' performance before deployment in the intended industry.
Innovation Solution
A system and method that generate synthetic test data based on computed parameters for applicants and job requisitions, using a different data source than the training data, to validate the performance of machine learning ranking models by populating work experience and job requisition parameters with text data from various sources, and applying this synthetic data to the model to generate a ranking list for evaluation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a machine learning model is trained on non-industry relevant data, then the model can be developed and tested initially, but the model cannot be accurately validated for industry-specific performance
Solution Approach 1:
The patent creates synthetic test data that copies the structural and statistical properties of industry-specific data without requiring actual industry data during model development. This synthetic data serves as a faithful replica that enables validation while maintaining the ability to develop models on general data first.
Solution Approach 2:
The patent performs preliminary generation of synthetic industry-specific test data before actual model deployment. This allows the validation framework to be prepared in advance with appropriate test cases that reflect industry conditions, enabling thorough validation before real-world deployment.
2Measurement precision
If industry-specific test data is used for validation, then accurate performance verification is possible, but data availability and accessibility are limited
Solution Approach 1:
The patent transforms the availability constraint by changing the parameter of data source from requiring actual industry data to generating synthetic data with controlled parameters that match industry data characteristics. This allows validation without dependency on scarce industry-specific datasets.
Solution Approach 2:
The patent introduces synthetic data as an intermediary between the model validation process and actual industry data. This intermediary enables accurate validation by mediating the interaction between models and industry conditions without requiring direct access to proprietary industry datasets.
3Stability of the object's composition
If synthetic test data is generated from the same data source as training data, then data consistency is maintained, but model validation accuracy is reduced due to data leakage
Solution Approach 1:
The patent segments the data usage into distinct functions: training data from one source and synthetic test data generated from different sources or through different processes. This segmentation prevents data leakage while maintaining the structural consistency needed for valid validation.
Solution Approach 2:
The patent creates test data as a synthetic copy that replicates the statistical properties and patterns of industry data without being identical to the training data. This copying approach maintains data consistency in terms of distribution and structure while ensuring independence from the training set for unbiased validation.
Data Source
AI summary
In some examples a first parameter for respective applicants or candidates can be computed based on respective text data from a text dataset that can include a plurality of different types of text data. The first parameter can be populated with a given portion of text of the respective text data. A second parameter for a job requisition can be computed based on the respective text data used to compute the first parameter for a given applicant or candidate. The second parameter can be populated with a different portion of text of the respective text data used to compute the first parameter. Synthetic test data can be generated based on the computed parameters to test a machine learning (ML) ranking model that has been trained on training data that is from a different data source than the text dataset to validate a performance of the ML ranking model.


