Meta-Database System for Locating Variables Across Heterogeneous Sources

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in identifying key variables and associated data relevant for modeling within extensive storage architectures, as important data is often dispersed across distinct and separate databases, making it challenging to ascertain relevant data for accurate inference and model training, leading to increased processing time and power consumption.

Innovation Solution

A meta-database system is created that summarizes data from multiple sources using AI algorithms to categorize, group, and profile data, allowing a graphical user interface to search for important variables, their locations, and associations, thereby reducing processing time and power requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If data from multiple distinct databases is used to improve model accuracy, then prediction accuracy is improved, but processing time increases

Engineering Contradiction:
Improveprediction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary scanning and profiling of multiple databases to identify key variables and their locations before the actual modeling process. This advance preparation creates a roadmap that guides subsequent data access, eliminating the need to search through entire databases during model training and inference, thus resolving the contradiction between using comprehensive data and maintaining processing speed

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention introduces an intermediary layer (the scanning and profiling system) between the user and the distributed databases. This intermediary identifies and locates relevant data across multiple databases, acting as a guide that enables accurate modeling without requiring users to manually navigate extensive storage architectures, thereby reducing processing time while maintaining prediction accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If data from multiple distinct databases is used to improve model accuracy, then prediction accuracy is improved, but computing power consumption increases

Engineering Contradiction:
Improveprediction accuracyVSAvoidcomputing power consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system extracts only the essential information needed for modeling - specifically, the locations and identities of key variables - from the distributed databases. Rather than processing or transferring large volumes of raw data, the profiling algorithm extracts metadata about where relevant data is stored, significantly reducing computing power requirements while still enabling accurate models to be built from the identified data sources

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

By performing data identification and location profiling in advance, the system prepares a comprehensive guide that reduces the computational burden during actual model training. The preliminary scanning phase identifies all relevant variables and their locations, allowing subsequent modeling operations to directly access needed data without extensive searching or processing, thus lowering overall computing power consumption

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If extensive storage architectures are searched to identify relevant data, then data completeness is improved, but ease of operation deteriorates

Engineering Contradiction:
Improvedata completenessVSAvoidease of locating data
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The invention introduces a profiling algorithm as an intermediary that automatically searches and maps data across extensive storage architectures. This intermediary handles the complex task of locating relevant variables in distributed databases, presenting users with a simplified view of where data is located without requiring them to navigate the underlying storage complexity, thus maintaining data completeness while dramatically improving ease of operation

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs self-service by automatically scanning, profiling, and organizing information about data locations across multiple databases. The profiling algorithm independently identifies key variables, determines their locations, and creates a structured representation of the data landscape, eliminating the need for users to manually search through extensive storage architectures while ensuring comprehensive data coverage

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11822564B1Graphical user interface enabling interactive visualizations using a meta-database constructed from autonomously scanned disparate and heterogeneous sources
Publication Date: 2023.11.21 TRUIST BANK
  • US11822564B1 patent drawing
  • US11822564B1 patent drawing
  • US11822564B1 patent drawing

AI summary

A system for interfacing with a meta-database representing data from a plurality of source databases. A key variable repository module operably couples the databases and includes an AI program with a scanner algorithm and a profiler algorithm. The scanner algorithm receives the source data from a source interface, compresses the data, and synchronizes the data with the meta-data using a meta-database interface. The profiler algorithm receives the meta-data from the meta-database interface, generates granular data types for the meta-data, determines variables indicative of the meta-data, generates variable probability distributions, produces variable associations, and modifies the meta-database to include the probability distributions and associations using the meta-data interface. A graphical user interface allows for generating a representation of a location of data associated with a user identified variable with the source databases, the probability distribution for the user identified variable, or any associations for the user identified variable.