NLP Dataset Field Labeling for Ambiguous Names

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data processing systems struggle to efficiently and accurately assign labels to dataset fields, particularly when field names are ambiguous, use inconsistent naming conventions, and require metadata-driven processing without manual intervention.

Innovation Solution

A method and system that utilize natural language processing (NLP) to analyze field names and data content, generating or selecting field labels from a glossary, and optionally merging scores to ensure accurate label assignment, even when conventional methods fail.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual label assignment is used to ensure accuracy, then label precision is improved, but processing time and labor cost increase significantly

Engineering Contradiction:
Improvelabel assignment accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service through automated label assignment using NLP and machine learning models that process field names and data samples independently, achieving both accuracy and efficiency without manual intervention for routine cases

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an intermediary automated labeling system that sits between raw data and final labeled output, using trained models to translate field characteristics into appropriate labels, reducing both manual labor and errors

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If comprehensive data analysis is performed to improve label accuracy, then measurement precision is improved, but computational resources and processing time increase

Engineering Contradiction:
Improvelabel assignment accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system applies partial action by analyzing only necessary portions of data (field names and selective data samples) rather than complete datasets, achieving sufficient accuracy while reducing computational overhead through targeted analysis

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent implements preliminary action through pre-trained machine learning models and pre-established label glossaries that are prepared in advance, enabling rapid inference during actual label assignment without performing exhaustive analysis from scratch

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated labeling systems are implemented to reduce manual intervention, then productivity is improved, but measurement precision deteriorates due to ambiguity in field names

Engineering Contradiction:
Improvelabel assignment efficiencyVSAvoidlabel assignment accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system segments the labeling process into distinct stages: field name analysis, data sample analysis, candidate label generation, and final selection. This segmentation allows each component to specialize and improve overall accuracy while maintaining automation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent incorporates feedback mechanisms where the system evaluates candidate labels against multiple criteria (field name matching, data consistency, glossary constraints) and iteratively refines selections, with options for human feedback to improve future automated decisions

Inventive Principle:
Principle #23Feedback

4Reliability

If multiple analysis methods are combined to handle ambiguous field names, then reliability is improved, but device complexity increases

Engineering Contradiction:
Improvelabel assignment reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple analysis approaches (NLP-based field name analysis, statistical data analysis, pattern matching) into a unified automated labeling framework, achieving improved reliability through complementary methods while managing complexity through integrated architecture

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250217388A1Techniques for assigning labels to dataset fields
Publication Date: 2025.07.03 AB INITIO TECHNOLOGY LLC
  • US20250217388A1 patent drawing
  • US20250217388A1 patent drawing
  • US20250217388A1 patent drawing

AI summary

Techniques for processing a dataset comprising data stored in fields to identify field labels. The field labels describe data stored in the dataset fields. The techniques determine whether any field labels in a field label glossary match a field. If none of the field labels in the field label glossary match the field, the techniques generate a new field label using the name of the field. The generated field label may be assigned to the field.