Datalake Master Table Identification via Feature Grouping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In large datalakes, managing and processing hundreds of thousands of computer readable tables is challenging due to the difficulty in identifying and tracking master tables, which are essential for analysis and generating additional tables, as existing methods rely on manual notes and naming conventions prone to human error.

Innovation Solution

A computer-implemented method and system that automatically groups computer readable tables based on their features, generates neighborhoods of similar tables, and identifies master tables within these neighborhoods, using processors and memory to execute instructions for obtaining, grouping, and identifying tables, thereby reducing manual effort and error.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual notes and naming conventions are used to identify master tables, then the process is simple to implement, but it is prone to human error and time-consuming

Engineering Contradiction:
Improveaccuracy of master table identificationVSAvoidtime required for manual identification
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables automatic self-identification of master tables through computational algorithms that analyze table features, relationships, and metadata. The automated process eliminates the need for manual intervention by having the system serve itself in identifying master tables through feature extraction and pattern recognition.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes (human review of notes and naming conventions) with automated computational mechanisms. The system uses processors to execute algorithms that automatically analyze table characteristics, substitute human judgment with machine-based feature extraction and comparison methods.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Extent of automation

If automated feature-based grouping is implemented, then manual effort is reduced, but system complexity increases

Engineering Contradiction:
Improveautomation of table processingVSAvoidcomplexity of processing system
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The system segments the complex task of master table identification into distinct manageable stages: feature extraction, feature comparison, neighborhood generation, and master table identification. Each stage is handled by specialized computational modules that work sequentially, reducing overall system complexity through decomposition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate representations such as feature vectors, similarity matrices, and neighborhood graphs that mediate between raw table data and final master table identification. These intermediaries simplify the processing by breaking down complex relationships into manageable computational steps.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If all tables are processed individually, then processing accuracy is maintained, but processing time and resources increase

Engineering Contradiction:
Improveprocessing speedVSAvoidnumber of tables processed
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The system merges similar tables into neighborhoods based on feature similarity, allowing batch processing of related tables together. By grouping tables with common characteristics, the system can process multiple tables simultaneously using shared computational resources, improving productivity while maintaining accuracy through unified analysis.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent dynamically adjusts processing parameters such as similarity thresholds, neighborhood sizes, and feature weights based on the characteristics of the data being processed. This adaptive approach optimizes processing efficiency by tuning parameters to match the specific dataset, reducing computational resources while maintaining high accuracy.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11556566B1Processing of computer readable tables in a datalake
Publication Date: 2023.01.17 INTUIT INC
  • US11556566B1 patent drawing
  • US11556566B1 patent drawing
  • US11556566B1 patent drawing

AI summary

Systems and methods for identifying one or more master tables of a datalake are described. A system may obtain a plurality of computer readable tables of a datalake (with each computer readable table including one or more features). The system may also group the plurality of computer readable tables into a plurality of groups based on a number of features of each computer readable table of the plurality of computer readable tables. The system may further generate, for each of one or more groups of the plurality of groups, one or more neighborhoods based on a similarity of features between computer readable tables of the group. The system may also identify, for each neighborhood, one or more master tables from the one or more computer readable tables of the group. The system may further provide an indication of one or more master tables identified in the datalake.