Cross-System Table Similarity Detection via Semantic and Cell Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data systems often operate as isolated 'data islands' due to limited external connectivity, making it difficult to efficiently detect and utilize similarities among tables across different data systems, which hinders the integration and analysis of data.
Innovation Solution
A method is developed to generate a semantic dataset from column names across multiple data systems, computing three measures of similarity (X, Y, and Z) based on text and data analysis, and combining these measures to improve data system integration and analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data systems operate as isolated data islands with limited external connectivity, then each data system maintains its own data integrity and security, but the ability to detect and utilize similarities among tables across different data systems deteriorates
Solution Approach 1:
The patent introduces an intermediary similarity detection mechanism that bridges isolated data systems. This mechanism computes multiple similarity measures (X, Y, Z) between tables in different data systems without requiring direct connectivity or data movement, thus maintaining data integrity while enabling cross-system similarity detection through a mediating analytical layer
Solution Approach 2:
The patent replaces the mechanical requirement of direct system connectivity with a computational similarity measurement approach. Instead of physically connecting data systems to detect similarities, the invention uses algorithmic computation of similarity measures based on table schemas and data characteristics, substituting physical connectivity with intellectual computation
2Difficulty of detecting and measuring
If multiple data systems are integrated to enable similarity detection, then the ability to utilize similarities among tables improves, but the complexity of the system increases
Solution Approach 1:
The patent segments the similarity detection process into three distinct measurable components (X, Y, and Z measures), each evaluating different aspects of table similarity. This segmentation allows the complex problem of cross-system table similarity to be broken down into manageable, independent computational tasks that can be executed separately and then integrated
Solution Approach 2:
The patent transforms the abstract concept of table similarity into concrete, measurable parameters (three distinct similarity measures). By changing the parameter representation from qualitative assessment to quantitative measurement, the system enables automated similarity detection without requiring complex qualitative judgment mechanisms
3Measurement precision
If comprehensive similarity measures are computed across multiple data systems, then the accuracy of data integration improves, but the computational time and resources increase
Solution Approach 1:
The patent computes three different similarity measures (X, Y, and Z) between tables, which may seem excessive at first glance. However, this partial redundancy serves a purpose: by computing multiple measures rather than a single comprehensive measure, the system achieves more accurate and nuanced similarity assessment while keeping each individual computation relatively simple and fast
Data Source
AI summary
A method, computer program product, and computer system for detection and utilization of similarities among tables in multiple data systems that include a first data system and a second data system. A semantic dataset is generated. A first measure (X) of similarity between semantic data in columns of the first and second data systems is computed using the semantic dataset. A first, different measure (Y) of similarity between semantic data in columns of the first and second data systems is computed using the semantic dataset. A third measure (Z) similarity between columns of the first and second data systems is computed based on data in cells within the columns. A weighted combination (U) of X, Y, and Z between the columns of tables in the first and second data systems is computed. X, Y, Z, U or combinations thereof are used to improve a computer system.


