Meta-Database System for Locating Variables Across Heterogeneous Sources
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulties in identifying key variables and associated data relevant for modeling within extensive storage architectures, as important data is often dispersed across distinct and separate databases, making it challenging to ascertain relevant data for accurate inference and model training, leading to increased processing time and power consumption.
Innovation Solution
A meta-database system is created that summarizes data from multiple sources using AI algorithms to categorize, group, and profile data, allowing a graphical user interface to search for important variables, their locations, and associations, thereby reducing processing time and power requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If data from multiple distinct databases is used to improve model accuracy, then prediction accuracy is improved, but processing time increases
Solution Approach 1:
The system performs preliminary scanning and profiling of multiple databases to identify key variables and their locations before the actual modeling process. This advance preparation creates a roadmap that guides subsequent data access, eliminating the need to search through entire databases during model training and inference, thus resolving the contradiction between using comprehensive data and maintaining processing speed
Solution Approach 2:
The invention introduces an intermediary layer (the scanning and profiling system) between the user and the distributed databases. This intermediary identifies and locates relevant data across multiple databases, acting as a guide that enables accurate modeling without requiring users to manually navigate extensive storage architectures, thereby reducing processing time while maintaining prediction accuracy
2Measurement precision
If data from multiple distinct databases is used to improve model accuracy, then prediction accuracy is improved, but computing power consumption increases
Solution Approach 1:
The system extracts only the essential information needed for modeling - specifically, the locations and identities of key variables - from the distributed databases. Rather than processing or transferring large volumes of raw data, the profiling algorithm extracts metadata about where relevant data is stored, significantly reducing computing power requirements while still enabling accurate models to be built from the identified data sources
Solution Approach 2:
By performing data identification and location profiling in advance, the system prepares a comprehensive guide that reduces the computational burden during actual model training. The preliminary scanning phase identifies all relevant variables and their locations, allowing subsequent modeling operations to directly access needed data without extensive searching or processing, thus lowering overall computing power consumption
3Loss of information
If extensive storage architectures are searched to identify relevant data, then data completeness is improved, but ease of operation deteriorates
Solution Approach 1:
The invention introduces a profiling algorithm as an intermediary that automatically searches and maps data across extensive storage architectures. This intermediary handles the complex task of locating relevant variables in distributed databases, presenting users with a simplified view of where data is located without requiring them to navigate the underlying storage complexity, thus maintaining data completeness while dramatically improving ease of operation
Solution Approach 2:
The system performs self-service by automatically scanning, profiling, and organizing information about data locations across multiple databases. The profiling algorithm independently identifies key variables, determines their locations, and creates a structured representation of the data landscape, eliminating the need for users to manually search through extensive storage architectures while ensuring comprehensive data coverage
Data Source
AI summary
A system for interfacing with a meta-database representing data from a plurality of source databases. A key variable repository module operably couples the databases and includes an AI program with a scanner algorithm and a profiler algorithm. The scanner algorithm receives the source data from a source interface, compresses the data, and synchronizes the data with the meta-data using a meta-database interface. The profiler algorithm receives the meta-data from the meta-database interface, generates granular data types for the meta-data, determines variables indicative of the meta-data, generates variable probability distributions, produces variable associations, and modifies the meta-database to include the probability distributions and associations using the meta-data interface. A graphical user interface allows for generating a representation of a location of data associated with a user identified variable with the source databases, the probability distribution for the user identified variable, or any associations for the user identified variable.


