Automated Data Labeling via Lineage Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data asset labeling methods, especially in large enterprises with complex naming rules, struggle to automatically label technical data assets effectively, as auto-discovery tools often fail to work with vendor-defined naming conventions, leading to challenges in data curation.

Innovation Solution

A computer-implemented method that utilizes data lineage exploration to identify technical data assets and map them to corresponding business items in UI reports, employing similarity analysis and knowledge graphs to automatically assign relevant labels, even in cases with complex naming rules.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated data discovery tools are used to label technical data assets, then labeling efficiency is improved, but the tools fail to work effectively with vendor-defined complex naming rules

Engineering Contradiction:
Improvelabeling efficiencyVSAvoidlabeling accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces data lineage information as an intermediary element that connects technical data assets to business concepts. Instead of directly matching technical asset names with business terms (which fails due to complex naming rules), the system uses lineage data showing how technical assets are transformed into business items in UI reports. This intermediary approach enables automated labeling to work effectively even with vendor-defined naming conventions by focusing on the functional relationship rather than name similarity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If manual labeling methods are used to ensure accurate mapping of technical data assets to business items, then labeling accuracy is improved, but the process becomes time-consuming and inefficient

Engineering Contradiction:
Improvelabeling accuracyVSAvoidlabeling efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies preliminary action by pre-establishing the data lineage information that documents the transformation relationships from technical data assets through ETL processes to business items in UI reports. This preliminary structuring of relationship data enables subsequent automated matching operations to achieve both high accuracy and efficiency, as the complex mapping logic has already been captured in the lineage information before the labeling process begins.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback by analyzing the lineage data to identify corresponding business items and using their names to determine relevant labels through similarity analysis. The confidence threshold mechanism provides feedback control, where labels are automatically assigned only when the similarity match exceeds the threshold, ensuring accuracy while maintaining automation. This feedback loop enables the system to achieve manual-level accuracy with automated efficiency.

Inventive Principle:
Principle #23Feedback

3Extent of automation

If similarity analysis is used to match business item names with labels, then automated label assignment is improved, but false matches may occur without confidence threshold validation

Engineering Contradiction:
Improveautomated label assignmentVSAvoidlabel matching precision
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent implements feedback control through the confidence threshold mechanism. The similarity analysis between business item names and labels produces a confidence score, and the system automatically compares this score against a predefined threshold. Only when the confidence score exceeds the threshold is the label automatically assigned. This feedback loop maintains high automation while ensuring measurement precision, as incorrect matches with low confidence scores are automatically rejected rather than falsely assigned.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11269844B2Automated data labeling
Publication Date: 2022.03.08 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11269844B2 patent drawing
  • US11269844B2 patent drawing
  • US11269844B2 patent drawing

AI summary

Systems and methods for computer-automated labeling of data are disclosed. In embodiments, a method includes: identifying technical data assets in lineage data and corresponding business items in User Interface (UI) data of a user, wherein the lineage data includes a data source for the UI data; mapping the technical data assets to the corresponding business items; determining relevant labels to assign to the technical data assets from a label repository based on a similarity analysis of names of the corresponding business items and labels in the label repository; determining that one or more of the relevant labels meet a confidence threshold based on the similarity analysis; and automatically assigning the one or more of the relevant labels to associated ones of the technical data assets based on the determining that the one or more of the relevant labels meet the confidence threshold.