Ontology Mapping System for Heterogeneous Data Integration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The integration of data from semantically heterogeneous sources is hindered by differences in vocabulary and terminology, requiring significant human effort and time, as computers struggle to understand the context and meaning of terms across different data sources.

Innovation Solution

A method that calculates semantic distance measures between data elements using a linked top ontology data structure, allowing for automated mapping between different representations of concepts or data structures by comparing names with vocabulary terms and determining commonality between top ontology nodes, supplemented by syntactic similarity evaluation and preprocessing techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated mapping methods are used to integrate data from semantically heterogeneous sources, then productivity is improved, but the accuracy and reliability of mappings deteriorate due to semantic heterogeneity and context loss

Engineering Contradiction:
Improvedata integration speedVSAvoidmapping accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces an intermediary system comprising a source ontology, target ontology, and mapping ontology that mediates between heterogeneous data sources. The mapping ontology acts as a bridge containing mapping rules that connect source concepts to target concepts, enabling automated integration while maintaining accuracy by preserving semantic relationships through the intermediary mapping layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the mapping problem from direct source-to-target matching into a multi-parameter transformation process involving source ontology parameters, target ontology parameters, and mapping ontology parameters. By changing the approach from binary matching to multi-parameter transformation with configurable mapping rules, the system achieves both automation and reliability.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If manual mapping methods are used to ensure accurate mappings between heterogeneous data sources, then mapping accuracy is improved, but the time and human effort required increase significantly

Engineering Contradiction:
Improvemapping accuracyVSAvoidmapping generation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements self-service automation where the mapping system automatically generates mappings by retrieving source ontology, target ontology, and mapping ontology from storage, executing mapping rules, and producing mapping results without human intervention. This eliminates manual mapping time while maintaining accuracy through the structured intermediary mapping ontology and predefined mapping rules.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs preliminary actions by pre-defining mapping rules in the mapping ontology that encode semantic relationships between source and target concepts. These pre-established mapping rules are stored and reused automatically, eliminating the need for repeated manual mapping efforts while ensuring consistent accurate mappings across different data integration tasks.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If comprehensive ontology mappings are created to handle semantic heterogeneity, then adaptability is improved, but the complexity of the mapping system increases

Engineering Contradiction:
Improvehandling semantic heterogeneityVSAvoidmapping system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the complex mapping problem into three distinct ontology components: source ontology for source data structure, target ontology for target data structure, and mapping ontology for mapping rules. This segmentation divides the complexity into manageable, independent modules that can be developed, stored, and executed separately, reducing overall system complexity while maintaining comprehensive adaptability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The mapping ontology serves as a universal intermediary that can handle multiple different source ontologies and target ontologies through a single standardized mapping rule structure. This multi-functional mapping ontology reduces complexity by providing a universal solution that adapts to various semantic heterogeneity scenarios without requiring separate custom mapping systems for each case.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9773053B2Method and apparatus for processing electronic data
Publication Date: 2017.09.26 BRITISH TELECOM PLC
  • US9773053B2 patent drawing
  • US9773053B2 patent drawing
  • US9773053B2 patent drawing

AI summary

A system (100) for generating a computer readable data file representative of a mapping between a first representation of a set of concepts or of a data structure (e.g. a database schema) and a second representation of a set of concepts or of a data structure (e.g. an ontology), each representation comprising a plurality of complex representational elements (e.g. tables in a database schema and concepts in an ontology) each of which may itself include a number of associated subordinate representational elements (e.g. columns/fields of a table in a database schema and attributes of a concept in an ontology). The system (100) includes a semantic similarity calculation module (134) for calculating a semantic similarity measure between a subordinate element of the first representation and each of the subordinate elements in the second representation and a mapping generation module (137) for generating a mapping between the subordinate element of the first representation and one of the subordinate elements of the second representation selected in dependence upon the calculated semantic similarity measures between the subordinate elements.