Metadata-Based Schema Mapping With ML Field Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data mapping processes in integration flows are time-consuming and complex, often requiring months of development due to the disparity and hierarchy of source and target schemas, and typically rely on metadata without leveraging multi-modal metadata for efficient schema matching.
Innovation Solution
A method and system that utilizes a machine learning model trained on schema metadata to generate fixed-size vector representations for each field, calculating field-to-field similarity and confidence scores to suggest data mappings without requiring actual data access or relying on previous user mappings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional data mapping processes are used to handle schema disparity and hierarchy, then mapping accuracy can be maintained, but development time increases to months and process complexity increases
Solution Approach 1:
The patent replaces manual mechanical data mapping processes with an automated machine learning system. The ML model automatically generates mapping suggestions by analyzing source and target schema metadata, eliminating the need for manual field-by-field mapping that traditionally took months of development time while maintaining mapping accuracy through intelligent algorithmic analysis.
Solution Approach 2:
The system performs preliminary analysis of schema metadata before actual data mapping is required. By pre-processing and understanding the structural characteristics, data types, and relationships in both source and target schemas upfront, the system prepares mapping suggestions in advance, significantly reducing the time needed during actual implementation while ensuring accuracy through thorough preliminary analysis.
2Measurement precision
If manual data mapping processes are used, then mapping quality can be maintained, but device complexity and operational complexity increase
Solution Approach 1:
The patent replaces complex manual mapping processes with an automated machine learning system that handles schema analysis, field matching, and mapping generation automatically. This substitution reduces operational complexity by eliminating manual intervention while maintaining mapping quality through the intelligent capabilities of the ML model in understanding and matching schema structures.
Solution Approach 2:
The system performs self-service by automatically analyzing schema metadata, generating mapping suggestions, and identifying corresponding fields without requiring manual configuration or complex user intervention. The ML model independently processes the mapping task, reducing both device complexity and operational complexity while maintaining high mapping quality through autonomous intelligent analysis.
3Ease of operation
If traditional mapping approaches relying on metadata are used, then implementation is straightforward, but mapping efficiency and accuracy decrease
Solution Approach 1:
The patent enhances traditional metadata-based mapping by replacing simple rule-based approaches with a machine learning system that automatically analyzes and interprets metadata. This substitution maintains implementation ease as the system still operates on metadata but dramatically improves mapping efficiency and accuracy through intelligent pattern recognition and automated field matching capabilities.
Solution Approach 2:
The system changes the parameters of metadata analysis by applying machine learning techniques that go beyond traditional string matching and simple attribute comparison. The ML model transforms how metadata is processed, extracting deeper semantic meanings and relationships, thereby improving mapping efficiency and accuracy while keeping the implementation approach straightforward by still operating on existing metadata without requiring additional data collection.
Data Source
AI summary
Methods, computer program products, and/or systems are provided that perform the following operations: obtaining source schema metadata, wherein the source schema metadata is associated with fields of a source schema; obtaining target schema metadata with a target schema, wherein the target schema metadata is associated with fields of a target schema; determining, for each field of the source schema and each field of the target schema, a representation for each field based, at least in part, on the source schema metadata or the target schema metadata associated with each field; and providing the representation for each field of the source schema and each field of the target schema for use in generating data mappings between the source schema and the target schema.


