Automated Data Labeling via Lineage Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data asset labeling methods, especially in large enterprises with complex naming rules, struggle to automatically label technical data assets effectively, as auto-discovery tools often fail to work with vendor-defined naming conventions, leading to challenges in data curation.
Innovation Solution
A computer-implemented method that utilizes data lineage exploration to identify technical data assets and map them to corresponding business items in UI reports, employing similarity analysis and knowledge graphs to automatically assign relevant labels, even in cases with complex naming rules.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated data discovery tools are used to label technical data assets, then labeling efficiency is improved, but the tools fail to work effectively with vendor-defined complex naming rules
Solution Approach 1:
The patent introduces data lineage information as an intermediary element that connects technical data assets to business concepts. Instead of directly matching technical asset names with business terms (which fails due to complex naming rules), the system uses lineage data showing how technical assets are transformed into business items in UI reports. This intermediary approach enables automated labeling to work effectively even with vendor-defined naming conventions by focusing on the functional relationship rather than name similarity.
2Reliability
If manual labeling methods are used to ensure accurate mapping of technical data assets to business items, then labeling accuracy is improved, but the process becomes time-consuming and inefficient
Solution Approach 1:
The patent applies preliminary action by pre-establishing the data lineage information that documents the transformation relationships from technical data assets through ETL processes to business items in UI reports. This preliminary structuring of relationship data enables subsequent automated matching operations to achieve both high accuracy and efficiency, as the complex mapping logic has already been captured in the lineage information before the labeling process begins.
Solution Approach 2:
The system uses feedback by analyzing the lineage data to identify corresponding business items and using their names to determine relevant labels through similarity analysis. The confidence threshold mechanism provides feedback control, where labels are automatically assigned only when the similarity match exceeds the threshold, ensuring accuracy while maintaining automation. This feedback loop enables the system to achieve manual-level accuracy with automated efficiency.
3Extent of automation
If similarity analysis is used to match business item names with labels, then automated label assignment is improved, but false matches may occur without confidence threshold validation
Solution Approach 1:
The patent implements feedback control through the confidence threshold mechanism. The similarity analysis between business item names and labels produces a confidence score, and the system automatically compares this score against a predefined threshold. Only when the confidence score exceeds the threshold is the label automatically assigned. This feedback loop maintains high automation while ensuring measurement precision, as incorrect matches with low confidence scores are automatically rejected rather than falsely assigned.
Data Source
AI summary
Systems and methods for computer-automated labeling of data are disclosed. In embodiments, a method includes: identifying technical data assets in lineage data and corresponding business items in User Interface (UI) data of a user, wherein the lineage data includes a data source for the UI data; mapping the technical data assets to the corresponding business items; determining relevant labels to assign to the technical data assets from a label repository based on a similarity analysis of names of the corresponding business items and labels in the label repository; determining that one or more of the relevant labels meet a confidence threshold based on the similarity analysis; and automatically assigning the one or more of the relevant labels to associated ones of the technical data assets based on the determining that the one or more of the relevant labels meet the confidence threshold.


