Automated Data Relationship Identification for Query Execution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data processing systems face challenges in efficiently analyzing and querying large, complex data sets due to unknown or lost relationships between diverse data sources, requiring laborious manual analysis and expert knowledge, and making data accessible to ordinary users.
Innovation Solution
A method for identifying relationships between data collections by computing relationship metrics such as distinctness and overlap between data sets, using candidate keys to evaluate potential relationships and generate relationship indicators for data query execution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual analysis is used to determine data structure and create queries, then relationship accuracy can be maintained, but productivity decreases and requires expert knowledge
Solution Approach 1:
The system enables self-service by automatically analyzing data collections and identifying relationships without requiring expert intervention. The automated relationship identification system processes data collections, computes metrics, and generates query recommendations independently, allowing ordinary users to access complex data sets without manual analysis or expert knowledge.
Solution Approach 2:
The patent replaces the mechanical manual analysis process with an automated computational system. Instead of experts manually examining data relationships, the system uses automated metric computation (distinctness, overlap, correlation) to identify relationships, substituting human cognitive work with algorithmic processing that maintains accuracy while dramatically improving productivity.
2Ease of operation
If data is extracted from original sources, then data accessibility improves, but relationship information is lost
Solution Approach 1:
The system performs preliminary action by automatically identifying and preserving relationships between data collections during the data extraction process. Before data is moved or transformed from original sources, the system analyzes the data collections, computes relationship metrics, and stores relationship information, ensuring that accessibility improvements do not result in relationship information loss.
Solution Approach 2:
The system implements feedback by continuously analyzing data collections and using computed relationship metrics to refine query generation. The relationship identification process provides feedback about data structures and connections, allowing the system to adapt and improve query recommendations while maintaining accessibility across diverse data sources.
3Productivity
If automated relationship identification is implemented, then productivity increases, but device complexity increases
Solution Approach 1:
The system applies segmentation by dividing the relationship identification process into distinct computational stages: data collection loading, candidate relationship generation, metric computation (distinctness, overlap, correlation), relationship scoring, and query recommendation. This segmented approach manages complexity by handling each aspect separately while maintaining high productivity through automated processing of each segment.
Solution Approach 2:
The patent implements universality by creating a multi-functional system that can handle diverse data collections from various sources, compute multiple relationship metrics simultaneously, identify different types of relationships, and generate various query formats. This universal approach increases productivity across different data scenarios while managing complexity through a unified framework.
Data Source
AI summary
A method and software tool for identifying relationships between columns of one or more data tables are disclosed. In the disclosed method, a relationship indicator is computed for each of a plurality of column pairs, each column pair comprising respective first and second columns selected from the one or more data tables. The relationship indicator comprises a measure of a relationship (e.g. indicating a strength or likelihood of a relationship) between data of the first column and data of the second column. Relationships between columns of the data tables are then identified in dependence on the computed relationship indicators. The identified relationships may be used to create and execute data queries.


