Automated Textual Data Labeling via Unsupervised Topic Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual labeling of textual data, such as media content, is expensive, time-consuming, and prone to consistency issues due to human biases, and existing labels are often ambiguous and generic, making it difficult to accurately classify and categorize content.
Innovation Solution
An automated method using unsupervised models and Natural Language Processing (NLP) techniques to generate detailed labels and descriptors for textual data, including preprocessing, topic modeling, and deep learning models to create enhanced genres and microgenres, reducing the need for manual labeling and improving label accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual labeling is performed using in-house experts or crowd sourcing, then labeling can be completed with human understanding and context, but it becomes expensive and time consuming
Solution Approach 1:
The system performs preliminary actions by pre-processing textual data through tokenization, stop word removal, and feature extraction before labeling. This preparation work is done automatically to reduce the subsequent manual labeling time while maintaining accuracy requirements.
Solution Approach 2:
The patent introduces an intermediary automated labeling system that uses machine learning models and natural language processing techniques to bridge between raw textual data and final labels. This intermediary process handles the time-consuming aspects while human experts focus on quality assurance and complex cases.
2Measurement precision
If manual labeling is performed by human curators, then contextual understanding can be achieved, but consistency issues arise due to human biases
Solution Approach 1:
The system implements self-service labeling capabilities where the automated model labels data consistently according to learned patterns from training data. This reduces human bias and improves consistency while maintaining accuracy through the model's ability to understand contextual nuances.
Solution Approach 2:
The patent incorporates feedback mechanisms where labeled data is continuously evaluated and used to retrain and improve the labeling model. This feedback loop ensures consistency by applying the same learned criteria across all labeling tasks while allowing the system to adapt and improve over time.
3Adaptability or versatility
If existing generic labels and keywords are used for media content, then broad categorization is achieved, but unambiguous identification becomes difficult
Solution Approach 1:
The patent applies local quality by generating specific, context-aware labels for different portions and aspects of the textual data rather than applying generic labels uniformly. This allows the system to maintain broad categorization flexibility while providing precise, unambiguous identification through localized, detailed labeling.
Solution Approach 2:
The system segments the labeling process into multiple levels and dimensions, creating hierarchical labels that range from broad categories to specific descriptors. This segmentation allows the system to maintain versatility for broad classification while achieving precise identification through detailed, segmented labeling of different content aspects.
Data Source
AI summary
Aspects of the subject disclosure may include, for example, determining classes from a corpus based on topic modeling, data clustering and unsupervised learning. Labels are determined for each of the classes and trained models are generated for each of the classes by assignment of a plurality of textual documents to labels based on a highest number of matches. A raw textual document can be tokenized and stop words removed. A corresponding one of the trained models can be selected according to a class that is applicable to subject matter of the raw textual document. The processed document can be assigned to a target label based on a highest number of matches of words. Other embodiments are disclosed.


