Automated Data Source Mapping via Binding Conditions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data integration methods, particularly in the discovery phase, are manual and time-consuming, lacking automation in identifying relationships and mappings between disparate data sources across software applications, databases, files, reports, and systems.
Innovation Solution
An automated method and apparatus that analyze metadata and data using rules, techniques, and statistics to deduce relationships between systems, converting schemas into normalized relational models, discovering binding conditions, correlations, and transformation functions to establish mappings between data sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual methods are used to identify relationships and mappings between data sources, then the integration can be performed with simple tools, but the process becomes time-consuming and labor-intensive
Solution Approach 1:
The system performs self-service by automatically analyzing metadata and data from multiple data sources to discover relationships and mappings without human intervention. The automated discovery engine examines data patterns, correlations, and semantic relationships to generate integration mappings autonomously, eliminating the need for manual analysis and significantly reducing the discovery phase duration.
Solution Approach 2:
The patent replaces manual mechanical processes with automated computational systems. Instead of human analysts manually examining data sources and creating mappings, an automated engine uses algorithms to analyze metadata, detect patterns, and generate integration rules, substituting human labor with machine-based intelligent processing.
2Productivity
If automated methods are used to discover relationships and mappings, then productivity increases, but the system complexity increases
Solution Approach 1:
The system introduces metadata as an intermediary layer between raw data sources and integration mappings. By analyzing metadata (data about data) rather than raw data directly, the system simplifies the discovery process. Metadata provides structured information about data schemas, relationships, and semantics, making automated analysis more manageable and reducing system complexity.
Solution Approach 2:
The automated discovery process is segmented into distinct phases: metadata analysis, pattern detection, correlation analysis, and mapping generation. Each phase handles specific tasks independently, allowing the complex overall process to be managed through modular, manageable components rather than a monolithic complex system.
3Measurement precision
If comprehensive data analysis is performed to discover semantics and relationships, then mapping accuracy improves, but the computational resources required increase
Solution Approach 1:
The system performs preliminary analysis by examining metadata and data samples before conducting full-scale discovery. This preliminary action identifies potential relationships and filters out irrelevant data, allowing the comprehensive analysis to focus only on promising candidates. This staged approach maintains high mapping accuracy while reducing overall computational resource consumption.
Data Source
AI summary
In one aspect, semantics and relationships and mappings are identified between a first and a second data source. Data between the first and second data source is compared. A binding condition is discovered between portions of data in the first and the second data source based upon the comparison, wherein the binding condition identifies data within the first and second data sources that map to each other. The binding condition is used to discover correlations between portions of data in the first and the second data source, wherein the correlations identify data in the first data source that correspond to values in the second data source. The binding condition and the correlations are used to discover a transformation function between portions of data in the first and the second data source, wherein the transformation function generates data in the second data source data in the first data source.


