Production Dataset Preparation Using Metadata-Guided Merging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data processing systems face challenges in accurately emulating production environments during API testing and training machine learning models due to the use of raw data that may lose realistic trends, leading to undetected errors and inefficiencies.

Innovation Solution

A method involving the processing of raw datasets to generate index metadata, associating these metadata with the datasets, and merging qualified datasets based on user-defined inputs to create a production dataset, utilizing natural language processing and machine learning models like word2vec and N-gram to enhance dataset quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If raw datasets are used directly for API testing and machine learning model training, then data processing speed is improved, but data quality and realism are worsened

Engineering Contradiction:
Improvedata processing speedVSAvoiddata quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent applies preliminary action by generating index metadata and augmenting datasets before they are used for API testing and machine learning model training. The system processes raw datasets in advance to create augmented datasets with enhanced quality features, including generated metadata that preserves realistic trends. This pre-processing ensures that when data is eventually used for testing or training, it already possesses the required quality characteristics without needing further processing at the point of use.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If raw datasets are used without processing, then data processing complexity is reduced, but reliability of testing and training is worsened

Engineering Contradiction:
Improvedata processing complexityVSAvoidreliability of testing and training
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces an intermediary processing layer that acts as a mediator between raw datasets and their ultimate use in API testing or machine learning training. The system generates index metadata and creates augmented datasets that serve as an intermediate representation, preserving realistic trends while maintaining structured organization. This intermediary layer ensures reliability by guaranteeing that data quality requirements are met before data reaches the testing or training phase, without requiring complex processing at the point of use.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If index metadata is generated and datasets are augmented, then data quality is improved, but data processing time is worsened

Engineering Contradiction:
Improvedata qualityVSAvoiddata processing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing dataset augmentation and index metadata generation in advance, before the data is needed for API testing or machine learning model training. The system processes raw datasets to create augmented datasets with enhanced quality features, including generated metadata that preserves realistic trends. This pre-processing approach shifts the time cost to an earlier stage, allowing faster data retrieval and use when the augmented datasets are already prepared.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260023783A1Dataset preparation
Publication Date: 2026.01.22 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20260023783A1 patent drawing
  • US20260023783A1 patent drawing
  • US20260023783A1 patent drawing

AI summary

Methods, computer program products, and systems are presented. The methods, computer program products, and systems can include, for example, processing multiple datasets using metadata, wherein the metadata can characterize relationships among datasets. In dependence on such metadata, production datasets can be, e.g., generated, versioned, and/or merged. The resulting production datasets can support subsequent computing uses such as analytics, machine learning, application testing, and/or enterprise processing.