Inference Model Synthetic Data Generation for Pipeline Reliability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data pipelines face interruptions due to inaccessible data caused by user-selected limitations and regulatory constraints, leading to misalignment of APIs and disruptions in computer-implemented services.

Innovation Solution

A system that utilizes a single inference model trained using self-supervised learning to generate synthetic data for unpopulated fields, ensuring reliability by qualifying predictable fields and using them to supplement data, thereby reducing pipeline failures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If user-selected limitations and regulatory constraints are imposed on data collection, then data privacy and compliance are improved, but data availability and pipeline reliability deteriorate

Engineering Contradiction:
Improvedata availabilityVSAvoiduser-selected limitations and regulatory constraints
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent introduces an inference model as an intermediary component between the data pipeline and inaccessible data. When data is unavailable due to user limitations or regulatory constraints, the inference model generates synthetic data that substitutes for the missing information, allowing the pipeline to continue operating without direct access to the original data. This mediator resolves the contradiction by enabling data availability while respecting privacy and compliance constraints.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates copies of inaccessible data through synthetic data generation. Instead of requiring direct access to the original data that is restricted by user preferences or regulations, the system generates synthetic copies that replicate the necessary data characteristics and structures. These copies allow downstream pipeline operations to proceed while the original restricted data remains inaccessible, thus maintaining both availability and compliance.

Inventive Principle:
Principle #26Copying

2Measurement precision

If multiple inference models are trained for different user limitations, then data supplementation accuracy is improved, but device complexity and computational resources worsen

Engineering Contradiction:
Improvedata supplementation accuracyVSAvoidnumber of inference models
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent trains a single inference model to handle multiple types of user-selected limitations and regulatory constraints. Rather than creating separate specialized models for each limitation type, the universal model is trained to recognize and generate appropriate synthetic data replacements for various restriction scenarios. This multi-functional approach maintains high supplementation accuracy while significantly reducing the number of models required and the associated computational complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges the functionality of multiple specialized inference models into a single unified model. By combining the learning objectives and training data from various limitation types into one comprehensive model, the system achieves the data supplementation capabilities of multiple models while using only one computational resource. This merging strategy reduces device complexity while preserving the accuracy needed for different user limitation scenarios.

Inventive Principle:
Principle #5Merging (Combining)

3Duration of action of stationary object

If synthetic data is generated to supplement unpopulated fields, then data pipeline continuity is improved, but data quality and reliability worsen

Engineering Contradiction:
Improvepipeline continuityVSAvoiddata quality
Core Design Contradiction:
Duration of action of stationary objectVSReliability

Solution Approach 1:

The patent implements feedback mechanisms to continuously monitor and improve the quality of synthetic data generated by the inference model. The system evaluates the accuracy and usefulness of generated synthetic data against actual data patterns and user requirements, using this feedback to refine the model's performance. This feedback loop ensures that while pipeline continuity is maintained through synthetic data generation, the data quality remains reliable and progressively improves over time.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250005394A1Generation of supplemented data for use in a data pipeline
Publication Date: 2025.01.02 DELL PROD LP
  • US20250005394A1 patent drawing
  • US20250005394A1 patent drawing
  • US20250005394A1 patent drawing

AI summary

Methods and systems for managing operation of a data pipeline are disclosed. To manage operation of a data pipeline when a portion of data is inaccessible may require generating a synthetic portion of data to generalize the inaccessible portion of data. Prior to the generation of the synthetic portion of data, it may be determined whether the type of information associated with the inaccessible portion of the data may be reliably predicted using an inference model and the available portion of the data. If the inaccessible portion of the data may be reliably predicted, the inference model may utilize the available portion of the data to predict the inaccessible portion of the data to obtain supplemented data. The supplemented data may then be used in the data pipeline.