Graph Mining for Automated Schema Concept Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data integration methods are inefficient in handling diverse data formats and large-scale datasets, particularly in scenarios where data structures are complex and not well-defined, leading to manual and error-prone processes for schema mapping and transformation.
Innovation Solution
A machine-implemented method that generates data concepts by converting schema representations into graph forms, mining frequent closed subgraphs, and using a relevancy metric to filter and store schema representations for automated schema matching and transformation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If conventional database integrator methods are used to determine mappings between data formats, then translation can be performed, but the process becomes manual and error-prone when dealing with complex and not well-defined data structures
Solution Approach 1:
The patent replaces manual mechanical schema mapping processes with automated graph mining algorithms. The system automatically derives graph representations from schemas, mines frequent closed subgraphs, and generates concept hierarchies without human intervention, thereby increasing automation while maintaining reliability through systematic algorithmic approaches.
Solution Approach 2:
The patent changes the representation parameters of schema data by converting traditional schema formats into graph representations. This parameter transformation enables the application of graph mining techniques to automatically discover concepts and relationships, resolving the contradiction between automation and reliability.
2Productivity
If manual schema mapping is performed for diverse data formats, then translation accuracy can be maintained, but productivity decreases due to the time-consuming nature of the process
Solution Approach 1:
The patent performs preliminary actions by automatically deriving graph representations from schemas and pre-mining frequent closed subgraphs before actual data integration tasks. This preliminary processing creates reusable concept hierarchies that accelerate subsequent data integration operations, significantly improving productivity while reducing time loss.
Solution Approach 2:
The patent creates copies of schema information in graph representation format and uses these copies for automated concept extraction. This copying approach enables rapid processing of multiple schemas without repeatedly analyzing the original complex data structures, thereby increasing productivity.
3Extent of automation
If automated graph mining is performed on large sets of schemas, then concept extraction can be automated, but the complexity of processing and analyzing the data increases
Solution Approach 1:
The patent segments the complex schema processing task into distinct phases: deriving graph representations, mining frequent closed subgraphs, and generating concept hierarchies. This segmentation reduces the complexity of each individual processing step while maintaining high levels of automation across the entire concept extraction pipeline.
Data Source
AI summary
In one embodiment, an approach to automated recurring concept extraction, from a plurality of input data models (schemas) is presented. The approach converts input data models to graphs, with typed elements. The graphs are mined for closed subgraphs that have a defined minimum support. The identified subgraphs can be filtered with a relevance metric. These subgraphs are converted to schemas or an appropriate representation, and stored for reuse in a repository. The repository can be used to automate further transformation or mapping of schemas presented to a system that uses the repository. In one example, the repository is used in a schema covering process to perform schema transformation.


