Enterprise Data Retrieval Using Centralized Dataset Knowledge Bases
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In enterprise environments, data generated by different groups is often stored in isolated databases, making it difficult for users to discover and access relevant datasets across geographic locations, leading to inefficiencies and resource consumption.
Innovation Solution
A network-based data storage and access system that stores knowledge base data about datasets, including links to their locations, allowing users to search and access them centrally, with features like question-and-answer forums, use case feedback, and machine learning for recommendations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If data is stored in distributed databases across different geographic locations, then data accessibility and local storage efficiency are improved, but data discovery difficulty and access coordination complexity increase
Solution Approach 1:
The patent introduces a centralized metadata database as an intermediary that stores information about datasets distributed across multiple geographic locations. This metadata database acts as a mediator between users and the distributed data storage system, allowing users to discover and access data without directly managing the complexity of distributed storage coordination.
Solution Approach 2:
The system segments data management into two independent parts: (1) distributed data storage at multiple geographic locations for accessibility and performance, and (2) centralized metadata management for discovery and coordination. This segmentation allows each part to optimize for its specific function without compromising the other.
2Ease of operation
If all generated datasets are stored in a central location, then data discovery and access are simplified, but storage costs and resource consumption increase
Solution Approach 1:
The system separates metadata storage (centralized) from actual data storage (distributed). Only lightweight metadata descriptions, not the full datasets, are stored centrally in the metadata database. This allows simplified data discovery through central indexing while avoiding the resource burden of duplicating or centralizing actual data storage.
Solution Approach 2:
Instead of storing complete dataset copies in a central location, the system creates and stores only metadata copies (descriptions, indices, and references) centrally. Users access actual data from distributed storage locations using the metadata as guidance, significantly reducing central storage requirements while maintaining discovery capabilities.
3Speed
If users directly access distributed databases without a centralized system, then access speed to local data is improved, but the ability to discover and locate relevant datasets across groups deteriorates
Solution Approach 1:
The centralized metadata database serves as an intermediary information layer that enhances rather than hinders data access. It provides discovery information (what data exists, where it is located, what it contains) without interfering with the actual data access path. Users query the metadata intermediary to locate data, then access the actual data directly from distributed storage, maintaining access speed while adding discovery capability.
Solution Approach 2:
The system performs preliminary action by pre-indexing and storing metadata information about all distributed datasets in the centralized metadata database. This preliminary organization of information allows users to quickly discover and locate relevant datasets before accessing them, without affecting the speed of actual data retrieval from distributed storage locations.
Data Source
AI summary
A computer system is provided and is programmed to: (1) store, in a first database, knowledge base data sets for each of a plurality of datasets, wherein each knowledge base data set includes at least data relating to the use of the associated dataset and a link to a separate database storing the corresponding dataset; (2) receive, at the first database from a user computer device, a request for knowledge base data for a first dataset; (3) instruct the user computer device to display the knowledge base data for the first dataset; (4) receive a request for access to the first dataset, wherein the request for access includes one or more database operations to be performed on the first dataset; (5) access the first dataset; and/or (6) execute the one or more database operations on the first dataset to provide results to the user computer device.


