Automated Data Type Mapping via Statistical Profiling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computing environments face challenges in efficiently integrating and mapping complex data structures across different software applications and execution environments, requiring substantial manual effort and expertise.

Innovation Solution

A system leveraging machine learning (DataFlow Machine Learning, DFML) for automated mapping of complex data structures, driven by metadata, schema, and statistical profiling, to facilitate data integration and flow management across various data sources and targets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual mapping of data structures is performed, then mapping precision can be maintained, but development time and labor cost increase significantly

Engineering Contradiction:
Improvemapping precisionVSAvoiddevelopment time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-service by automatically inferring data type mappings through statistical profiling and pattern recognition. The machine learning model analyzes source data structures and target schema definitions to autonomously generate mappings without requiring manual intervention from domain experts, thus resolving the contradiction between precision and time consumption.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual mapping process with an automated machine learning-based inference system. The system uses statistical profiling, pattern decomposition, and ML models to substitute human expertise in curating data integrations, thereby reducing development time while maintaining mapping quality through algorithmic precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of time

If automated mapping is implemented, then development time is reduced, but mapping accuracy may deteriorate

Engineering Contradiction:
Improvedevelopment timeVSAvoidmapping accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The system incorporates feedback mechanisms where the machine learning model continuously refines its mappings based on statistical profiling of the data. The model analyzes patterns in the source data and compares them against target schema constraints, adjusting mappings iteratively to maintain accuracy while achieving automation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent employs parameter changes by transforming the mapping problem into a statistical inference task. The system changes parameters such as probability distributions, statistical moments, and pattern weights to optimize mapping accuracy. The ML model adjusts these parameters based on the complexity of data structures and the quality of available metadata.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If complex data structures are mapped manually, then mapping reliability can be ensured, but system complexity increases

Engineering Contradiction:
Improvemapping reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the complex data mapping process into distinct manageable components: statistical profiling, pattern decomposition, machine learning inference, and validation. This segmentation allows the system to handle complex data structures by breaking them down into simpler analytical steps, maintaining reliability while reducing overall system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary machine learning model that mediates between the source data structures and target schema. This intermediary layer simplifies the mapping process by translating complex source formats into standardized internal representations, thereby ensuring reliable mappings without requiring the entire system to handle full complexity directly.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Measurement precision

If manual domain model expert curation is used, then mapping quality can be maintained, but productivity decreases

Engineering Contradiction:
Improvemapping qualityVSAvoidapplication development productivity
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables self-service by allowing it to autonomously perform the curation work that previously required domain model experts. The machine learning model independently analyzes data structures, infers mappings, and generates integrations without human intervention, thereby maintaining mapping quality while dramatically improving productivity by eliminating the bottleneck of expert availability.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250156160A1System and method for automated mapping of data types for use with dataflow environments
Publication Date: 2025.05.15 ORACLE INT CORP
  • US20250156160A1 patent drawing
  • US20250156160A1 patent drawing
  • US20250156160A1 patent drawing

AI summary

In accordance with various embodiments, described herein is a system (Data Artificial Intelligence system, Data AI system), for use with a data integration or other computing environment, that leverages machine learning (ML, DataFlow Machine Learning, DFML), for use in managing a flow of data (dataflow, DF), and building complex dataflow software applications (dataflow applications, pipelines). In accordance with an embodiment, the system can provide support for auto-mapping of complex data structures, datasets or entities, between one or more sources or targets of data, referred to herein in some embodiments as HUBs. The auto-mapping can be driven by a metadata, schema, and statistical profiling of a dataset; and used to map a source dataset or entity associated with an input HUB, to a target dataset or entity or vice versa, to produce an output data prepared in a format or organization (projection) for use with one or more output HUBs.