Uniform Data Access API for Multi-Source Cost Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems fail to provide a uniform interface for accessing data across varied data stores in big data ecosystems, leading to increased client application development time, unknown costs for clients, and instability in data lake storage due to improper query patterns.
Innovation Solution
A uniform data access web API platform that predicts costs, identifies optimal data sources, and streamlines data access across multiple data sources, enabling transparent cost communication and optimization, while ensuring stability and security.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional systems access data from multiple data sources without a uniform interface, then data can be retrieved from various stores, but client application development time increases and operation becomes complex
Solution Approach 1:
The patent implements a universal data access API layer that provides a standardized interface for accessing multiple types of data stores (relational databases, NoSQL databases, data lakes, cloud storage). This uniform interface allows client applications to interact with diverse data sources through a single consistent API, eliminating the need for separate access logic for each data store type and significantly reducing development time.
Solution Approach 2:
The patent introduces an intermediary data access layer positioned between client applications and multiple data sources. This intermediary layer handles the complexity of accessing different data store types, performing data transformation, format conversion, and protocol adaptation. Client applications interact only with the standardized intermediary interface, while the intermediary manages the diversity of underlying data sources.
2Ease of operation
If conventional systems allow clients to access data without cost prediction, then data access is simple, but clients incur unexpected high costs for large-scale computations and data movement
Solution Approach 1:
The patent implements cost prediction and estimation functionality that calculates and presents expected costs to clients before they execute data access operations. The system analyzes the requested operation, estimates computational resources required (CPU, memory, disk), data movement volume (network bandwidth), and storage implications, then presents this cost information to the client in advance. This allows clients to make informed decisions about their data access patterns and optimize for cost while maintaining operational simplicity.
3Productivity
If conventional systems process queries without query pattern analysis, then processing is straightforward, but data lake stability deteriorates due to improper storage and query pattern mapping
Solution Approach 1:
The patent implements dynamic query pattern analysis and adaptation that monitors and learns from query patterns to optimize data access strategies. The system dynamically adjusts query execution plans, data retrieval strategies, and resource allocation based on observed patterns, ensuring stable and efficient operation of the data lake while maintaining high query processing speeds.
4Adaptability or versatility
If conventional systems lack dataset footprint curation, then data storage is flexible, but resource efficiency decreases and costs increase
Solution Approach 1:
The patent implements automated dataset footprint curation that dynamically adjusts data storage parameters, retention policies, and format optimizations based on access patterns, data importance, and resource constraints. The system transforms raw data into optimized storage representations, applies compression where appropriate, and manages data lifecycle transitions between different storage tiers, thereby improving resource efficiency while preserving storage flexibility for diverse data types.
Data Source
AI summary
The present disclosure provides a system (110) and a method for accessing datasets across data sources. The system (110) receives, from a user device (104), a query to fetch at least one dataset from at least one data source of a data lake (150). The system (110) sends a list including a plurality of parameters to the user device (104) and receives datasets of interest and a query pattern defined based on a user requirement from the user device (104), predicts an estimated cost for the datasets of interest, and identifies an optimal data source corresponding to the datasets of interest. The system (110) sends the estimated cost to the user device (104), pre-processes the optimal data source, and provides access to the user device (104) to fetch the datasets of interest from the optimal data source based on a positive response being received from the user device (104).


