Machine Learning Data Lineage Search for Self-Service Transformation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data lineage request processes require significant user effort and time, often taking days to obtain results, and lack self-service capabilities for users to efficiently trace and manipulate data lineage information.

Innovation Solution

A self-service data lineage search apparatus using machine learning (ML) that enables users to quickly identify and transform data lineage information through a graphical user interface, utilizing a five-part key system with wildcard capabilities to match and update data identifiers, reducing processing requirements and improving efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If traditional drill down method is used for data lineage requests, then data lineage information can be obtained, but the process takes a day or longer and requires significant user effort

Engineering Contradiction:
Improvetime to obtain data lineage resultsVSAvoiddata lineage search efficiency
Core Design Contradiction:
Loss of timeVSProductivity

Solution Approach 1:

The patent replaces the traditional mechanical drill-down process with machine learning algorithms that automatically infer data lineage relationships. The ML model processes data identifiers and metadata to predict lineage connections, eliminating the need for manual step-by-step drilling through data flow paths, thereby reducing search time from days to minutes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables users to perform data lineage searches independently through a user-friendly interface without requiring technical expertise or special access permissions. The ML-powered platform automatically processes queries, handles data identifier matching, and presents results without requiring user involvement in the complex underlying data lineage tracking mechanisms.

Inventive Principle:
Principle #25Self-service

2Ease of operation

If technical interface is used for data lineage requests, then data lineage information can be accessed, but the interface is challenging for users to self-serve

Engineering Contradiction:
Improveuser friendliness of data lineage searchVSAvoidcomplexity of data lineage request process
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces machine learning algorithms as an intermediary layer between users and the complex data lineage system. The ML model handles the complexity of data identifier matching, metadata parsing, and lineage relationship inference, presenting simplified results to users through intuitive interfaces without exposing users to the underlying system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system provides self-service capabilities through automated query processing and intelligent data identifier matching. Users can perform lineage searches without technical expertise as the ML-powered platform automatically handles data identifier resolution, wildcard matching, and result presentation, eliminating the need for complex technical interfaces.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If wildcard search capability is added to data lineage search, then search flexibility is improved, but processing requirements and bandwidth increase

Engineering Contradiction:
Improvesearch flexibility for data lineage queriesVSAvoidprocessing requirements and bandwidth
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent implements partial action by processing only the portions of data identifiers that are relevant to wildcard matching patterns. Instead of processing entire data lineage paths, the ML model focuses computational resources on matching the wildcard characters and inferring lineage relationships only for the affected segments, reducing overall processing requirements while maintaining search flexibility.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system changes the parameter of data identifier matching from exact matching to wildcard pattern matching handled by ML algorithms. This allows flexible searches while the ML model optimizes processing by learning which wildcard patterns are most common and pre-computing lineage relationships for frequently queried data elements, thereby managing bandwidth and processing requirements.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12436961B2Machine learning apparatus for data lineage transformation
Publication Date: 2025.10.07 BANK OF AMERICA CORP
  • US12436961B2 patent drawing
  • US12436961B2 patent drawing
  • US12436961B2 patent drawing

AI summary

A machine learning (“ML”) apparatus for self-service search and transformation of data lineage information is provided. The apparatus may include machine readable memory configured to store technical data element identifiers (“TDEIs”) and a computer configured to receive a query for data lineage information corresponding to a first TDEI. The apparatus may include a processor configured to identify a level of commonality between the first TDEI and a second TDEI and determine whether the first TDEI and the second TDEI share a threshold level of commonality. Following a determination that the first TDEI and the second TDEI share a threshold level of commonality, the processor may identify any mismatches between the first TDEI and the second TDEI and determine whether a threshold number of mismatches exists between the first TDEI and the second TDEI. The processor may then transform the data lineage using ML.