Graph Mining for Automated Schema Concept Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data integration methods are inefficient in handling diverse data formats and large-scale datasets, particularly in scenarios where data structures are complex and not well-defined, leading to manual and error-prone processes for schema mapping and transformation.

Innovation Solution

A machine-implemented method that generates data concepts by converting schema representations into graph forms, mining frequent closed subgraphs, and using a relevancy metric to filter and store schema representations for automated schema matching and transformation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If conventional database integrator methods are used to determine mappings between data formats, then translation can be performed, but the process becomes manual and error-prone when dealing with complex and not well-defined data structures

Engineering Contradiction:
Improveautomation of schema mappingVSAvoidaccuracy of schema mapping
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent replaces manual mechanical schema mapping processes with automated graph mining algorithms. The system automatically derives graph representations from schemas, mines frequent closed subgraphs, and generates concept hierarchies without human intervention, thereby increasing automation while maintaining reliability through systematic algorithmic approaches.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the representation parameters of schema data by converting traditional schema formats into graph representations. This parameter transformation enables the application of graph mining techniques to automatically discover concepts and relationships, resolving the contradiction between automation and reliability.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If manual schema mapping is performed for diverse data formats, then translation accuracy can be maintained, but productivity decreases due to the time-consuming nature of the process

Engineering Contradiction:
Improvespeed of data integrationVSAvoidtime for schema mapping
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by automatically deriving graph representations from schemas and pre-mining frequent closed subgraphs before actual data integration tasks. This preliminary processing creates reusable concept hierarchies that accelerate subsequent data integration operations, significantly improving productivity while reducing time loss.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates copies of schema information in graph representation format and uses these copies for automated concept extraction. This copying approach enables rapid processing of multiple schemas without repeatedly analyzing the original complex data structures, thereby increasing productivity.

Inventive Principle:
Principle #26Copying

3Extent of automation

If automated graph mining is performed on large sets of schemas, then concept extraction can be automated, but the complexity of processing and analyzing the data increases

Engineering Contradiction:
Improveautomation of concept extractionVSAvoidcomplexity of processing system
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent segments the complex schema processing task into distinct phases: deriving graph representations, mining frequent closed subgraphs, and generating concept hierarchies. This segmentation reduces the complexity of each individual processing step while maintaining high levels of automation across the entire concept extraction pipeline.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8719299B2Systems and methods for extraction of concepts for reuse-based schema matching
Publication Date: 2014.05.06 SAP SE
  • US8719299B2 patent drawing
  • US8719299B2 patent drawing
  • US8719299B2 patent drawing

AI summary

In one embodiment, an approach to automated recurring concept extraction, from a plurality of input data models (schemas) is presented. The approach converts input data models to graphs, with typed elements. The graphs are mined for closed subgraphs that have a defined minimum support. The identified subgraphs can be filtered with a relevance metric. These subgraphs are converted to schemas or an appropriate representation, and stored for reuse in a repository. The repository can be used to automate further transformation or mapping of schemas presented to a system that uses the repository. In one example, the repository is used in a schema covering process to perform schema transformation.