Autonomous Data Source Discovery via API Microservices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data source discovery methods are manual and inefficient, requiring interviews and network audits to identify and access data sources across multiple applications and locations, with periodic updates often involving redundant work.
Innovation Solution
An autonomous data source discovery system that uses a discovery job processor, streaming processor, API microservices manager, and queue manager to identify and associate data sources with their custodians, leveraging APIs and webhooks for real-time updates and efficient data management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual interviews and network audits are used to identify data sources, then data source identification can be performed, but the process is time-consuming and inefficient
Solution Approach 1:
The system enables autonomous data source discovery where the discovery engine automatically identifies data sources, applications, and custodians without requiring manual interviews or audits. The system self-updates the data source catalog by autonomously accessing applications and extracting metadata, eliminating the need for repeated manual discovery efforts.
Solution Approach 2:
The patent replaces manual mechanical processes (interviews, audits) with an automated software-based discovery engine that uses APIs and webhooks to programmatically identify and catalog data sources across the organization's IT infrastructure.
2Reliability
If periodic updates are performed by re-doing previous work, then data source catalogs are updated, but redundant work is performed
Solution Approach 1:
The discovery engine operates continuously or on scheduled intervals to monitor and detect changes in applications and data sources. When changes are detected through webhook notifications or periodic checks, the system performs incremental updates to the data source catalog, maintaining currency without re-doing the entire discovery process.
Solution Approach 2:
The system implements feedback mechanisms where applications send webhook notifications to the discovery engine when data sources are created, modified, or deleted. This feedback loop enables the catalog to be automatically updated in real-time or near-real-time, ensuring reliability without redundant manual work.
3Adaptability or versatility
If each application requires its own interface to pull data sources, then application-specific data can be accessed, but device complexity increases
Solution Approach 1:
The discovery engine implements a universal interface that can interact with multiple different applications through standardized API connections. Rather than requiring separate custom interfaces for each application, the system uses a unified architecture that adapts to different data sources through configurable connection parameters and standardized data models.
Solution Approach 2:
The discovery engine acts as an intermediary layer between the data source catalog and various applications. It manages all application-specific interface complexities internally while presenting a unified, simplified interface to users and other systems, thereby reducing perceived complexity while maintaining adaptability.
4Ease of manufacture
If manual processes are used for data source discovery, then implementation is straightforward, but ease of operation decreases
Solution Approach 1:
The system performs autonomous data source discovery without requiring manual intervention. The discovery engine automatically connects to applications, extracts data source information, and populates the catalog, making the operation extremely easy while maintaining straightforward implementation through standardized protocols.
Data Source
AI summary
A computer-implemented method of discovering data sources includes receiving a request at a computing device through a user interface identifying applications and any additional data source types associated with the applications, and parameters used to access the applications, automatically authenticating the computing device to applications that require authentication, using the parameters, making calls through a programming interface for each application requesting identification of data sources, receiving a list identified data sources through the programming interface, providing unique identifiers for each of the identified data sources, providing an access identifier that identifies users that have access to the data sources, and storing the identified data sources, unique identifiers, and access identifiers as a data source catalog.


