Discovery Engine Automating Data Exploration via API Calls
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual data discovery methods using discovery engines are resource-intensive, inefficient, and fail to scale with increasing data volumes, often leading to unnecessary API calls and resource wastage, as users must manually navigate and prioritize data sources without systematic guidance.
Innovation Solution
A method utilizing a discovery engine with APIs to automate data exploration, where initial metadata is determined, a discovery space is rendered, and API calls are issued to generate further metadata, with strategies selected from a pool to optimize data processing, ensuring efficient resource allocation and prioritization of important data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual data discovery methods are used, then users can explore data sources, but resource consumption increases and efficiency decreases
Solution Approach 1:
The discovery engine automatically performs data exploration tasks without requiring manual user intervention. The system self-manages the discovery process by autonomously navigating data sources, executing API calls, and generating metadata, thereby eliminating resource waste associated with manual operations while maintaining ease of use through automated service delivery.
Solution Approach 2:
The system implements feedback mechanisms where discovery results are continuously evaluated against stopping criteria. The engine receives feedback from each API call and metadata generation step, using this information to dynamically adjust the discovery process and determine when to terminate exploration, thereby optimizing resource usage while ensuring thorough data exploration.
2Measurement precision
If comprehensive data exploration is performed, then data discovery quality improves, but resource consumption and time increase
Solution Approach 1:
The discovery engine performs partial exploration by selectively executing API calls based on evaluated priorities and stopping criteria. Rather than exhaustively exploring all possible data aspects, the system performs sufficient exploration to meet quality thresholds while avoiding unnecessary time consumption through intelligent termination decisions based on feedback from intermediate results.
Solution Approach 2:
The system performs preliminary evaluation of discovery space content and initial metadata before committing to extensive exploration paths. By assessing data sources and potential discovery outcomes in advance, the engine prioritizes high-value exploration paths and avoids time-consuming exploration of low-yield areas, thereby maintaining discovery quality while reducing overall time investment.
3Productivity
If systematic guidance is implemented, then resource allocation optimizes, but system complexity increases
Solution Approach 1:
The discovery system is segmented into distinct functional modules: metadata determination, discovery space rendering, API call generation, result evaluation, and stopping criterion assessment. Each module handles a specific aspect of the discovery process independently, allowing systematic resource allocation through modular control while managing complexity through clear separation of concerns and defined interfaces between segments.
Data Source
AI summary
The present disclosure relates to a method for accessing data of one or more data sources using a discovery engine. The method comprises: determining a discovery space content from initial metadata of a data source indicated in a data exploration request. The discovery space content may be rendered. The rendered content may be used for determining a set of one or more tasks for generating further metadata from at least part of the data of the data source, wherein the set of tasks comprises a combination of API calls. The API calls may be issued to the discovery engine. Discovery results of the issued API calls may be received. A data discovery status may be devalued using the discovery results. The discovery space content may be augmented using the further metadata and the data discovery status. The augmented discovery space content may be rendered for receiving further API calls.


