Context-Aware Synthetic Data Generation for ML Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Training machine-learning algorithms for data analysis and content prediction from unstructured data often requires large amounts of data, and data availability can be limited or scarce, especially depending on the context.

Innovation Solution

A computer-implemented method for dynamically modifying content based on user-specified context, which receives input data with annotated parts, extracts content elements, and replaces them with new content from internal or external data sources, such as weather forecasts or social media, to generate synthetic data for machine-learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If machine-learning algorithms are trained using traditional methods with existing data, then model accuracy can be achieved, but data availability is limited or scarce

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata availability
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent creates synthetic copies of real-world data by generating synthetic documents that mirror the structure, format, and semantic relationships of actual documents. These synthetic copies are populated with AI-generated content that maintains realistic patterns while providing unlimited quantity for training machine-learning models, thereby resolving the contradiction between data scarcity and model accuracy requirements

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system employs large language models to autonomously generate synthetic training data without requiring extensive human involvement. The AI model self-services by automatically creating realistic documents, extracting entities, and generating training examples, thereby producing abundant training data independently to overcome data availability limitations while maintaining high model accuracy

Inventive Principle:
Principle #25Self-service

2Reliability

If large amounts of data are collected for training, then model performance improves, but data processing time and computational resources increase

Engineering Contradiction:
Improvemodel performanceVSAvoiddata processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-generating and storing synthetic training data in advance using AI models. This pre-computed synthetic data is ready for immediate use in training machine-learning algorithms, eliminating the need for time-consuming data collection and processing during model development. The preliminary generation of realistic training examples accelerates the overall training process while maintaining high model performance

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

By creating synthetic copies of training data through AI generation, the system provides unlimited training examples without the time constraints of manual data collection. These synthetic copies can be generated on-demand or pre-generated in bulk, significantly reducing data processing time compared to traditional methods while still providing sufficient data volume for high model performance

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11328117B2Automated content modification based on a user-specified context
Publication Date: 2022.05.10 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11328117B2 patent drawing
  • US11328117B2 patent drawing
  • US11328117B2 patent drawing

AI summary

Dynamically changing a content based on a user-defined context includes receiving, by one or more processors, input data from a user, the input data includes at least one document with an annotated part identifying a first content element, the first content element is associated with a first content type. The one or more processors determine a context information associated with the annotated part and extract the annotated part. A first replacement for the first content element is retrieved from a first data source selected based on the content information. The one or more processors replace the first content element in the at least one document with the first replacement.