Data Lineage Identification and Change Impact Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data lineage analysis systems in distributed computing environments face challenges in understanding end-to-end data lineage information due to incomplete metadata, multiple data formats, and lack of connection between application-level terminology and technical data sources, limiting their ability to assess data object change impact effectively.
Innovation Solution
The system leverages unstructured data and technical attributes to identify data lineage across multiple data sources using advanced artificial intelligence classification algorithms, determining direct and indirect relationships, and predicting change impact scores by analyzing metadata and incident tickets, enabling seamless integration with new modules and platforms without configuration changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If advanced AI classification algorithms are used to identify data lineage and predict change impact, then measurement precision and reliability of data object change impact assessment are improved, but device complexity and computational resources required increase
Solution Approach 1:
The patent introduces an intermediary AI classification model that acts as a mediator between raw metadata/incident ticket data and change impact assessment results. This model processes unstructured data through trained algorithms to produce standardized impact scores, resolving the contradiction by adding computational complexity only where needed for precision while keeping the overall system architecture manageable through clear data flow stages
Solution Approach 2:
The system performs preliminary actions by pre-training classification models on historical metadata and incident ticket data before actual change impact assessment. This advance preparation creates reusable knowledge structures that reduce real-time computational complexity while maintaining high measurement precision during actual change impact predictions
2Loss of information
If unstructured data and incident tickets are processed to identify indirect relationships between data sources, then completeness of data lineage information is improved, but loss of time for data processing increases
Solution Approach 1:
The system performs preliminary processing of unstructured incident ticket data and metadata to pre-identify relationships between data sources before actual lineage analysis is needed. By training classification models in advance on historical data, the system creates pre-computed knowledge structures that enable rapid retrieval and analysis, reducing real-time processing time while maintaining complete lineage information
Solution Approach 2:
The patent creates simplified copies or representations of complex unstructured data through structured metadata extraction and classification. Instead of processing entire incident tickets and raw metadata each time, the system creates condensed feature vectors and relationship representations that capture essential lineage information, reducing processing time while preserving data completeness
3Adaptability or versatility
If multiple data formats and metadata schemas are reconciled to create unified data lineage representation, then adaptability of the system to different data sources is improved, but device complexity increases
Solution Approach 1:
The patent implements a universal metadata classification framework that can handle multiple data formats and schemas through a single AI classification model. This model is trained to recognize and process diverse metadata structures, incident ticket formats, and data source schemas, providing universal adaptability without requiring separate processing logic for each format, thus managing complexity through unified multi-functional processing
Data Source
AI summary
Methods and systems are described for data lineage identification and change impact prediction. Servers capture metadata that defines data objects associated with data sources. The servers determine direct relationships between data sources based upon the captured metadata. The servers identify indirect relationships between the data sources. The servers generate a data lineage across the data sources for the data objects. The servers extract unstructured text from database incident tickets and match the unstructured text to the metadata. The servers generate a multidimensional vector for the data objects based upon the data lineage and the unstructured text. The servers train a classification model using the vectors to predict a change impact score for each data object. The servers receive a request to change a data object. The servers determine a change impact score for the data object. When the score is below a threshold, the servers execute the change.


