Automated Data Relationship Identification for Query Execution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data processing systems face challenges in efficiently analyzing and querying large, complex data sets due to unknown or lost relationships between diverse data sources, requiring laborious manual analysis and expert knowledge, and making data accessible to ordinary users.

Innovation Solution

A method for identifying relationships between data collections by computing relationship metrics such as distinctness and overlap between data sets, using candidate keys to evaluate potential relationships and generate relationship indicators for data query execution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual analysis is used to determine data structure and create queries, then relationship accuracy can be maintained, but productivity decreases and requires expert knowledge

Engineering Contradiction:
Improverelationship identification accuracyVSAvoiddata query creation efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables self-service by automatically analyzing data collections and identifying relationships without requiring expert intervention. The automated relationship identification system processes data collections, computes metrics, and generates query recommendations independently, allowing ordinary users to access complex data sets without manual analysis or expert knowledge.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual analysis process with an automated computational system. Instead of experts manually examining data relationships, the system uses automated metric computation (distinctness, overlap, correlation) to identify relationships, substituting human cognitive work with algorithmic processing that maintains accuracy while dramatically improving productivity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If data is extracted from original sources, then data accessibility improves, but relationship information is lost

Engineering Contradiction:
Improvedata accessibilityVSAvoidrelationship information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The system performs preliminary action by automatically identifying and preserving relationships between data collections during the data extraction process. Before data is moved or transformed from original sources, the system analyzes the data collections, computes relationship metrics, and stores relationship information, ensuring that accessibility improvements do not result in relationship information loss.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback by continuously analyzing data collections and using computed relationship metrics to refine query generation. The relationship identification process provides feedback about data structures and connections, allowing the system to adapt and improve query recommendations while maintaining accessibility across diverse data sources.

Inventive Principle:
Principle #23Feedback

3Productivity

If automated relationship identification is implemented, then productivity increases, but device complexity increases

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system applies segmentation by dividing the relationship identification process into distinct computational stages: data collection loading, candidate relationship generation, metric computation (distinctness, overlap, correlation), relationship scoring, and query recommendation. This segmented approach manages complexity by handling each aspect separately while maintaining high productivity through automated processing of each segment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements universality by creating a multi-functional system that can handle diverse data collections from various sources, compute multiple relationship metrics simultaneously, identify different types of relationships, and generate various query formats. This universal approach increases productivity across different data scenarios while managing complexity through a unified framework.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11360950B2System for analysing data relationships to support data query execution
Publication Date: 2022.06.14 HITACHI VANTARA LLC
  • US11360950B2 patent drawing
  • US11360950B2 patent drawing
  • US11360950B2 patent drawing

AI summary

A method and software tool for identifying relationships between columns of one or more data tables are disclosed. In the disclosed method, a relationship indicator is computed for each of a plurality of column pairs, each column pair comprising respective first and second columns selected from the one or more data tables. The relationship indicator comprises a measure of a relationship (e.g. indicating a strength or likelihood of a relationship) between data of the first column and data of the second column. Relationships between columns of the data tables are then identified in dependence on the computed relationship indicators. The identified relationships may be used to create and execute data queries.