Canonical Data Model for Schema Integration Conflict Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The integration of diverse business data interfaces, schemas, and data models across different applications hinders application integration, leading to significant costs in enterprise IT budgets due to heterogeneity and the complexity of mapping between different standards and domains.
Innovation Solution
The implementation of a computer-implemented method to generate a unified data model (UDM) by merging and resolving conflicts in hierarchical schemas, resulting in a canonical data model (CDM) that consolidates schema correspondences and reduces heterogeneity, allowing for cross-domain and cross-standard integration through relevance rating, context logic, transitive mappings, and iterative improvement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If diverse business data interfaces, schemas and data models are integrated across applications, then application integration capability is improved, but integration cost and complexity increase due to heterogeneity
Solution Approach 1:
The patent introduces a canonical data model (CDM) as an intermediary layer between diverse source schemas and target schemas. The CDM acts as a mediator that consolidates multiple heterogeneous schemas into a unified structure, enabling integration without requiring direct mapping between each pair of source and target schemas. This intermediary approach reduces the overall integration complexity by breaking down the mapping problem into manageable steps through the standardized CDM layer.
Solution Approach 2:
The canonical data model serves as a universal data structure that can represent multiple different business domains and schemas simultaneously. By designing the CDM to be multi-functional and domain-agnostic, the system can integrate various heterogeneous schemas (e.g., customer data, product data, order data) into a single unified model, thereby improving integration capability while reducing the need for domain-specific integration solutions.
2Ease of operation
If schema correspondence knowledge is consolidated into a unified data model, then data model homogeneity is improved, but the process requires conflict resolution and iterative refinement
Solution Approach 1:
The patent performs preliminary actions by automatically detecting and identifying schema correspondences between source schemas before final integration. The system pre-processes the schema mappings, identifies potential conflicts in advance, and prepares resolution strategies beforehand. This preliminary action reduces the time required during the actual integration process, as conflicts are already identified and prepared for resolution rather than discovered and resolved in real-time during integration.
Solution Approach 2:
The canonical data model maintenance process incorporates feedback mechanisms that continuously monitor and evaluate the consistency of the CDM against incoming schema correspondences. When conflicts are detected or when new schema mappings are introduced, the system provides feedback to adjust and refine the CDM accordingly. This feedback loop enables iterative refinement of the data model, improving homogeneity over time while managing the complexity of conflict resolution through systematic evaluation and adjustment.
Data Source
AI summary
The present disclosure describes methods, systems, and computer program products for providing and maintaining an evolving canonical data model (CDM) which consolidates current knowledge of the correspondences of existing schemas. One computer-implemented method includes receiving the plurality of source hierarchical schemas, each source hierarchical schema being stored as a computer-readable document in computer-readable memory, processing, using a computer, the source hierarchical schemas to generate a merged graph, the merged graph comprising a plurality of merged nodes, each merged node being provided based on one or more nodes from at least two of the source hierarchical schemas, and determining, using the computer, that the merged graph includes one or more conflicts and, in response, resolving each conflict of the one or more conflicts to generate a computed-transitive-edge-free, conflict-free merged graph as a unified data model (UDM), wherein resolving comprises splitting one or more merged nodes into respective sub-sets of merged nodes.


