Data Conversion Operator for Machine Learning Pipelines

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data pipeline technologies are inefficient in processing data from external relational databases within machine learning pipelines, prone to data leakage during cross-validation, and leave decision-making to users, increasing the risk of errors and security vulnerabilities.

Innovation Solution

The implementation of a computer system and method that analyzes and converts data using relational algebra algorithms to map data routes, predict outcomes, rank them based on positive match percentages, and dynamically transmit the data into a machine learning pipeline, enhancing processing efficiency and security by reducing sensitive data processing outside the pipeline.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If data is processed outside the machine learning pipeline using traditional data wrangling steps, then data preparation can be performed extensively, but data leakage occurs during cross-validation and security vulnerabilities increase

Engineering Contradiction:
Improvedata preparation capabilityVSAvoiddata leakage prevention
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent applies preliminary action by performing data conversion to the target domain before data enters the machine learning pipeline. The conversion operator transforms source domain data into target domain data upfront, ensuring that all subsequent processing occurs within the secure pipeline environment. This prevents data leakage during cross-validation while maintaining comprehensive data preparation capabilities.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces a conversion operator as an intermediary component that bridges the source domain and target domain. This operator handles the data transformation within the pipeline architecture, acting as a mediator that ensures data security while enabling extensive data preparation. The intermediary approach allows data wrangling to occur within the secure pipeline boundary rather than outside it.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If users manually make decisions about data processing, then flexibility is maintained, but user errors increase and security risks rise

Engineering Contradiction:
Improvedata processing flexibilityVSAvoiderror reduction
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements self-service by enabling the machine learning pipeline to automatically perform data conversion and processing without requiring manual user intervention. The conversion operator autonomously transforms data between domains based on predefined configurations, eliminating human error while maintaining flexibility through programmable conversion rules. The system serves itself by handling data preparation tasks internally within the pipeline.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies parameter changes by transforming data from the source domain to the target domain through systematic parameter transformations. The conversion operator modifies data parameters and structures according to predefined conversion rules, providing flexibility through configurable transformation parameters while ensuring reliability through automated, error-free execution. This allows adaptable data processing without manual intervention.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If multiple data conversion methods are evaluated, then optimal conversion can be selected, but processing time increases

Engineering Contradiction:
Improvedata conversion accuracyVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-defining multiple conversion operators and their associated conversion rules before data processing begins. During runtime, the system automatically selects and applies the appropriate pre-configured conversion operator based on the data characteristics. This eliminates the need for time-consuming evaluation of multiple conversion methods during actual processing, as the optimal conversion path has already been determined in advance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms that allow the system to automatically evaluate and select the most appropriate conversion operator based on data characteristics and conversion success metrics. The feedback loop enables the system to learn from conversion outcomes and optimize the selection process, maintaining high conversion accuracy while minimizing processing time through intelligent, data-driven operator selection.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11720586B2Automatic conversion of data within data pipeline
Publication Date: 2023.08.08 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11720586B2 patent drawing
  • US11720586B2 patent drawing
  • US11720586B2 patent drawing

AI summary

Embodiments of the present invention provide a computer system a computer program product, and a method that comprises analyzing identified data for a determined conversion of the identified data, wherein the identified data is input data stored on an external database; automatically converting the analyzed data to a uniform domain by mapping a data route within the analyzed data, predicting a plurality of outcomes based on an application of a plurality of scenarios associated with the mapped data route, ranking the predicted outcomes based on a positive match percentage for the analyzed data, and converting the analyzed data associated with at least one ranked outcome using a relational algebra algorithm; and dynamically transmitting the converted, analyzed data into at least one section of a machine learning data pipeline.