Machine Learning Data Lineage Search for Self-Service Transformation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data lineage request processes require significant user effort and time, often taking days to obtain results, and lack self-service capabilities for users to efficiently trace and manipulate data lineage information.
Innovation Solution
A self-service data lineage search apparatus using machine learning (ML) that enables users to quickly identify and transform data lineage information through a graphical user interface, utilizing a five-part key system with wildcard capabilities to match and update data identifiers, reducing processing requirements and improving efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If traditional drill down method is used for data lineage requests, then data lineage information can be obtained, but the process takes a day or longer and requires significant user effort
Solution Approach 1:
The patent replaces the traditional mechanical drill-down process with machine learning algorithms that automatically infer data lineage relationships. The ML model processes data identifiers and metadata to predict lineage connections, eliminating the need for manual step-by-step drilling through data flow paths, thereby reducing search time from days to minutes.
Solution Approach 2:
The system enables users to perform data lineage searches independently through a user-friendly interface without requiring technical expertise or special access permissions. The ML-powered platform automatically processes queries, handles data identifier matching, and presents results without requiring user involvement in the complex underlying data lineage tracking mechanisms.
2Ease of operation
If technical interface is used for data lineage requests, then data lineage information can be accessed, but the interface is challenging for users to self-serve
Solution Approach 1:
The patent introduces machine learning algorithms as an intermediary layer between users and the complex data lineage system. The ML model handles the complexity of data identifier matching, metadata parsing, and lineage relationship inference, presenting simplified results to users through intuitive interfaces without exposing users to the underlying system complexity.
Solution Approach 2:
The system provides self-service capabilities through automated query processing and intelligent data identifier matching. Users can perform lineage searches without technical expertise as the ML-powered platform automatically handles data identifier resolution, wildcard matching, and result presentation, eliminating the need for complex technical interfaces.
3Adaptability or versatility
If wildcard search capability is added to data lineage search, then search flexibility is improved, but processing requirements and bandwidth increase
Solution Approach 1:
The patent implements partial action by processing only the portions of data identifiers that are relevant to wildcard matching patterns. Instead of processing entire data lineage paths, the ML model focuses computational resources on matching the wildcard characters and inferring lineage relationships only for the affected segments, reducing overall processing requirements while maintaining search flexibility.
Solution Approach 2:
The system changes the parameter of data identifier matching from exact matching to wildcard pattern matching handled by ML algorithms. This allows flexible searches while the ML model optimizes processing by learning which wildcard patterns are most common and pre-computing lineage relationships for frequently queried data elements, thereby managing bandwidth and processing requirements.
Data Source
AI summary
A machine learning (“ML”) apparatus for self-service search and transformation of data lineage information is provided. The apparatus may include machine readable memory configured to store technical data element identifiers (“TDEIs”) and a computer configured to receive a query for data lineage information corresponding to a first TDEI. The apparatus may include a processor configured to identify a level of commonality between the first TDEI and a second TDEI and determine whether the first TDEI and the second TDEI share a threshold level of commonality. Following a determination that the first TDEI and the second TDEI share a threshold level of commonality, the processor may identify any mismatches between the first TDEI and the second TDEI and determine whether a threshold number of mismatches exists between the first TDEI and the second TDEI. The processor may then transform the data lineage using ML.


