Automated Data Type Mapping via Statistical Profiling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing environments face challenges in efficiently integrating and mapping complex data structures across different software applications and execution environments, requiring substantial manual effort and expertise.
Innovation Solution
A system leveraging machine learning (DataFlow Machine Learning, DFML) for automated mapping of complex data structures, driven by metadata, schema, and statistical profiling, to facilitate data integration and flow management across various data sources and targets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual mapping of data structures is performed, then mapping precision can be maintained, but development time and labor cost increase significantly
Solution Approach 1:
The system performs self-service by automatically inferring data type mappings through statistical profiling and pattern recognition. The machine learning model analyzes source data structures and target schema definitions to autonomously generate mappings without requiring manual intervention from domain experts, thus resolving the contradiction between precision and time consumption.
Solution Approach 2:
The patent replaces the mechanical manual mapping process with an automated machine learning-based inference system. The system uses statistical profiling, pattern decomposition, and ML models to substitute human expertise in curating data integrations, thereby reducing development time while maintaining mapping quality through algorithmic precision.
2Loss of time
If automated mapping is implemented, then development time is reduced, but mapping accuracy may deteriorate
Solution Approach 1:
The system incorporates feedback mechanisms where the machine learning model continuously refines its mappings based on statistical profiling of the data. The model analyzes patterns in the source data and compares them against target schema constraints, adjusting mappings iteratively to maintain accuracy while achieving automation.
Solution Approach 2:
The patent employs parameter changes by transforming the mapping problem into a statistical inference task. The system changes parameters such as probability distributions, statistical moments, and pattern weights to optimize mapping accuracy. The ML model adjusts these parameters based on the complexity of data structures and the quality of available metadata.
3Reliability
If complex data structures are mapped manually, then mapping reliability can be ensured, but system complexity increases
Solution Approach 1:
The system segments the complex data mapping process into distinct manageable components: statistical profiling, pattern decomposition, machine learning inference, and validation. This segmentation allows the system to handle complex data structures by breaking them down into simpler analytical steps, maintaining reliability while reducing overall system complexity.
Solution Approach 2:
The patent introduces an intermediary machine learning model that mediates between the source data structures and target schema. This intermediary layer simplifies the mapping process by translating complex source formats into standardized internal representations, thereby ensuring reliable mappings without requiring the entire system to handle full complexity directly.
4Measurement precision
If manual domain model expert curation is used, then mapping quality can be maintained, but productivity decreases
Solution Approach 1:
The system enables self-service by allowing it to autonomously perform the curation work that previously required domain model experts. The machine learning model independently analyzes data structures, infers mappings, and generates integrations without human intervention, thereby maintaining mapping quality while dramatically improving productivity by eliminating the bottleneck of expert availability.
Data Source
AI summary
In accordance with various embodiments, described herein is a system (Data Artificial Intelligence system, Data AI system), for use with a data integration or other computing environment, that leverages machine learning (ML, DataFlow Machine Learning, DFML), for use in managing a flow of data (dataflow, DF), and building complex dataflow software applications (dataflow applications, pipelines). In accordance with an embodiment, the system can provide support for auto-mapping of complex data structures, datasets or entities, between one or more sources or targets of data, referred to herein in some embodiments as HUBs. The auto-mapping can be driven by a metadata, schema, and statistical profiling of a dataset; and used to map a source dataset or entity associated with an input HUB, to a target dataset or entity or vice versa, to produce an output data prepared in a format or organization (projection) for use with one or more output HUBs.


