Automated Textual Data Labeling via Unsupervised Topic Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual labeling of textual data, such as media content, is expensive, time-consuming, and prone to consistency issues due to human biases, and existing labels are often ambiguous and generic, making it difficult to accurately classify and categorize content.

Innovation Solution

An automated method using unsupervised models and Natural Language Processing (NLP) techniques to generate detailed labels and descriptors for textual data, including preprocessing, topic modeling, and deep learning models to create enhanced genres and microgenres, reducing the need for manual labeling and improving label accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling is performed using in-house experts or crowd sourcing, then labeling can be completed with human understanding and context, but it becomes expensive and time consuming

Engineering Contradiction:
Improvelabeling accuracyVSAvoidlabeling time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing textual data through tokenization, stop word removal, and feature extraction before labeling. This preparation work is done automatically to reduce the subsequent manual labeling time while maintaining accuracy requirements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary automated labeling system that uses machine learning models and natural language processing techniques to bridge between raw textual data and final labels. This intermediary process handles the time-consuming aspects while human experts focus on quality assurance and complex cases.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual labeling is performed by human curators, then contextual understanding can be achieved, but consistency issues arise due to human biases

Engineering Contradiction:
Improvelabeling accuracyVSAvoidlabeling consistency
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system implements self-service labeling capabilities where the automated model labels data consistently according to learned patterns from training data. This reduces human bias and improves consistency while maintaining accuracy through the model's ability to understand contextual nuances.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent incorporates feedback mechanisms where labeled data is continuously evaluated and used to retrain and improve the labeling model. This feedback loop ensures consistency by applying the same learned criteria across all labeling tasks while allowing the system to adapt and improve over time.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If existing generic labels and keywords are used for media content, then broad categorization is achieved, but unambiguous identification becomes difficult

Engineering Contradiction:
Improveclassification flexibilityVSAvoidcontent identification accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies local quality by generating specific, context-aware labels for different portions and aspects of the textual data rather than applying generic labels uniformly. This allows the system to maintain broad categorization flexibility while providing precise, unambiguous identification through localized, detailed labeling.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system segments the labeling process into multiple levels and dimensions, creating hierarchical labels that range from broad categories to specific descriptors. This segmentation allows the system to maintain versatility for broad classification while achieving precise identification through detailed, segmented labeling of different content aspects.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20220277351A1Method and apparatus for labeling data
Publication Date: 2022.09.01 AT&T INTELLECTUAL PROPERTY I L P
  • US20220277351A1 patent drawing
  • US20220277351A1 patent drawing
  • US20220277351A1 patent drawing

AI summary

Aspects of the subject disclosure may include, for example, determining classes from a corpus based on topic modeling, data clustering and unsupervised learning. Labels are determined for each of the classes and trained models are generated for each of the classes by assignment of a plurality of textual documents to labels based on a highest number of matches. A raw textual document can be tokenized and stop words removed. A corresponding one of the trained models can be selected according to a class that is applicable to subject matter of the raw textual document. The processed document can be assigned to a target label based on a highest number of matches of words. Other embodiments are disclosed.