Production Dataset Preparation Using Metadata-Guided Merging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data processing systems face challenges in accurately emulating production environments during API testing and training machine learning models due to the use of raw data that may lose realistic trends, leading to undetected errors and inefficiencies.
Innovation Solution
A method involving the processing of raw datasets to generate index metadata, associating these metadata with the datasets, and merging qualified datasets based on user-defined inputs to create a production dataset, utilizing natural language processing and machine learning models like word2vec and N-gram to enhance dataset quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If raw datasets are used directly for API testing and machine learning model training, then data processing speed is improved, but data quality and realism are worsened
Solution Approach 1:
The patent applies preliminary action by generating index metadata and augmenting datasets before they are used for API testing and machine learning model training. The system processes raw datasets in advance to create augmented datasets with enhanced quality features, including generated metadata that preserves realistic trends. This pre-processing ensures that when data is eventually used for testing or training, it already possesses the required quality characteristics without needing further processing at the point of use.
2Device complexity
If raw datasets are used without processing, then data processing complexity is reduced, but reliability of testing and training is worsened
Solution Approach 1:
The patent introduces an intermediary processing layer that acts as a mediator between raw datasets and their ultimate use in API testing or machine learning training. The system generates index metadata and creates augmented datasets that serve as an intermediate representation, preserving realistic trends while maintaining structured organization. This intermediary layer ensures reliability by guaranteeing that data quality requirements are met before data reaches the testing or training phase, without requiring complex processing at the point of use.
3Manufacturing precision
If index metadata is generated and datasets are augmented, then data quality is improved, but data processing time is worsened
Solution Approach 1:
The patent applies preliminary action by performing dataset augmentation and index metadata generation in advance, before the data is needed for API testing or machine learning model training. The system processes raw datasets to create augmented datasets with enhanced quality features, including generated metadata that preserves realistic trends. This pre-processing approach shifts the time cost to an earlier stage, allowing faster data retrieval and use when the augmented datasets are already prepared.
Data Source
AI summary
Methods, computer program products, and systems are presented. The methods, computer program products, and systems can include, for example, processing multiple datasets using metadata, wherein the metadata can characterize relationships among datasets. In dependence on such metadata, production datasets can be, e.g., generated, versioned, and/or merged. The resulting production datasets can support subsequent computing uses such as analytics, machine learning, application testing, and/or enterprise processing.


