Application Asset Discovery for Automated Data Lineage Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data governance tools lack the capability to provide a precise, automated understanding of data lineage and asset relationships, leading to inaccurate inventory and manual tracking that is time-consuming and error-prone, especially in API-driven environments.
Innovation Solution
Implementing a language-agnostic data discovery module that utilizes machine learning algorithms to scan application source code, identify technical assets, harvest metadata, and create a knowledge map for fine-grained data understanding, including data stores, APIs, and services, with automated data quality checks and policy enforcement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual tracking methods are used for data inventory and lineage, then implementation simplicity is maintained, but accuracy and reliability of data tracking deteriorate
Solution Approach 1:
The patent replaces manual tracking mechanisms with automated code scanning and machine learning algorithms. The system automatically scans application source code, identifies data assets and relationships, and builds lineage graphs without human intervention, thereby improving accuracy while managing complexity through automation.
Solution Approach 2:
The system enables self-service by having applications automatically generate and update their own data lineage information through code scanning. The metadata extraction and relationship mapping occur autonomously, allowing the system to self-maintain accurate data inventory without requiring manual tracking efforts.
2Productivity
If manual tracking of data assets is performed, then resource consumption is low, but time required for data governance increases
Solution Approach 1:
The system performs preliminary action by scanning and analyzing application source code during the development or deployment phase, extracting metadata and establishing data relationships before production. This proactive approach creates the data lineage information in advance, eliminating the need for time-consuming manual tracking later.
Solution Approach 2:
The automated scanning and metadata extraction processes run continuously or periodically, maintaining up-to-date data lineage information as applications evolve. This continuous automation eliminates idle time and ensures productivity gains are sustained as the system adapts to changing data assets.
3Measurement precision
If fine-grained analysis of technical assets is implemented, then measurement precision improves, but complexity of detection and measurement increases
Solution Approach 1:
The patent segments the complex task of data lineage tracking into distinct components: code scanning, metadata extraction, relationship identification, and lineage graph construction. Each component handles a specific aspect of the analysis, making the overall fine-grained measurement manageable through modular processing of individual data assets and their relationships.
4Productivity
If automated code scanning and metadata harvesting are implemented, then productivity increases, but device complexity increases
Solution Approach 1:
The data discovery module is designed as a universal, language-agnostic system that can scan and analyze code across multiple programming languages and frameworks. This multi-functionality consolidates what would otherwise require separate tools for each language, managing device complexity while maintaining high productivity through a single automated platform.
Data Source
AI summary
Various methods, apparatuses/systems, and media for implementing a data discovery module are disclosed. A repository includes one or more memories that stores application code for each application among a plurality of applications. A processor is operatively connected to the repository via a communication network. The processor scans the application source code for each application among the plurality of applications; identifies, in response to scanning, all technical assets and their relationships within each application; harvests technical metadata from the technical assets and their relationships to identify what information is used, stored, created, and moved by the application; implements machine learning algorithms to automatically assign descriptive and administrative metadata at a field level; loads the assigned descriptive and administrative metadata into an enterprise data catalog; and creates, in response to loading, a knowledge map, thereby providing a fine-grain level understanding of data within the technical assets.


