Automated Correlated Column Identification in Database Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Identifying correlated columns in multiple database tables for efficient merging is a tedious process, requiring comparison of each column across tables, which is inefficient as the number of columns increases.

Innovation Solution

A method and system for identifying correlated columns by determining correlation attributes, such as metadata, statistical parameters, ontological properties, and measurement units, and comparing these attributes to determine similarities and potential merging of columns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual comparison of each column across database tables is performed to identify correlated columns, then accurate identification of mergeable columns can be achieved, but the process becomes extremely tedious and time-consuming as the number of columns increases

Engineering Contradiction:
Improveaccuracy of column correlation identificationVSAvoidtime required for column comparison
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical comparison with automated computational analysis using natural language processing and machine learning algorithms. The system automatically extracts semantic meaning from column names, descriptions, and data values to identify correlations without human intervention, thereby maintaining accuracy while dramatically reducing time consumption.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces intermediate processing layers including natural language processing modules, semantic analysis components, and machine learning models that act as intermediaries between raw column data and correlation identification. These intermediaries transform unstructured column metadata into structured semantic representations that can be efficiently compared and matched.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If the number of columns in database tables increases, then more comprehensive data can be stored, but the complexity of identifying correlated columns increases significantly

Engineering Contradiction:
Improveamount of data storedVSAvoidcomplexity of column comparison process
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments the column comparison process into independent analytical modules: natural language processing of column names, semantic analysis of column descriptions, statistical analysis of data values, and correlation scoring. This segmentation allows the system to handle large numbers of columns by processing them through specialized sub-routines rather than requiring comprehensive pairwise comparison of all columns.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the comparison task from comparing raw column data directly to comparing extracted semantic parameters and features. By converting column names, descriptions, and values into standardized semantic parameters, the system reduces the dimensionality of the comparison problem and makes it scalable to tables with many columns.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automated methods are used to identify correlated columns, then the process efficiency is improved, but the precision of identification may be reduced compared to manual expert analysis

Engineering Contradiction:
Improvespeed of column correlation identificationVSAvoidaccuracy of column correlation identification
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where the system's correlation identification results are continuously refined based on user interactions, validation data, and performance metrics. The machine learning models are trained on feedback from manual expert corrections and validation cases, allowing the automated system to improve its precision over time while maintaining high productivity.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent employs a composite approach combining multiple analytical techniques: natural language processing, semantic analysis, statistical methods, and machine learning. By integrating these different methods into a unified correlation identification system, the patent achieves both high speed automation and high precision through the complementary strengths of each component.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS7725498B2Techniques for identifying mergeable data
Publication Date: 2010.05.25 X CORP
  • US7725498B2 patent drawing
  • US7725498B2 patent drawing
  • US7725498B2 patent drawing

AI summary

A system, method and article of manufacture for identifying mergeable data in a data processing system and, more particularly, for identifying correlated columns from one or more database tables. One embodiment comprises determining correlation attributes for a first column and a second column from one or more database tables. The correlation attributes describe for each column at least one of the column and content of the column. The correlation attributes from the first and second column are compared and similarities between the first and second column are identified on the basis of the comparison. Then, on the basis of the identified similarities, it is determined whether the first and second columns are correlated. Only if the columns are determined to be correlated, the first and second columns are merged.