Weakly Supervised Classifier Creation for Feedback Text Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The development of machine learning models for data processing is hindered by the labor-intensive process of labeling data for supervised learning, with advanced models requiring even more training data and existing weakly supervised models being insufficient for many applications.
Innovation Solution
A system and method for creating a large weakly supervised training set using a keyword-based classifier to automatically generate a training dataset, which is then used to construct deep learning classifiers, allowing for quicker and easier model development, and facilitating the construction of human-validated validation sets for both keyword-based and machine-learning-based topics in databases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual labeling is used to create training data for supervised machine learning models, then the quality and accuracy of the model can be improved, but the time and labor required increases significantly
Solution Approach 1:
The system performs preliminary automated labeling using keyword-based classifiers and weakly supervised models to generate initial training data before final manual validation. This preliminary action creates a ready-to-use training dataset that requires minimal manual intervention, significantly reducing the time and labor needed for complete manual labeling while maintaining model accuracy through subsequent validation steps.
2Measurement precision
If advanced neural network models are used to improve classification performance, then the model capability increases, but the amount of training data required increases even more
Solution Approach 1:
The system introduces weakly supervised models and keyword-based classifiers as intermediary components that bridge the gap between limited manually labeled data and the requirements of advanced neural network models. These intermediaries generate additional pseudo-labeled training data from unlabeled corpora, enabling advanced models to achieve high classification performance without requiring proportionally large amounts of manually labeled training data.
3Ease of manufacture
If keyword-based classification is used to simplify the modeling process, then the ease of creation improves, but the ability to understand topic trends and sentiment decreases
Solution Approach 1:
The system merges keyword-based classification with weakly supervised machine learning models and neural network-based sentiment analysis to create a hybrid classification system. This combination maintains the simplicity and ease of creation associated with keyword-based methods while incorporating the advanced topic understanding and sentiment analysis capabilities of machine learning models, thereby preserving both ease of manufacture and information quality.
Data Source
AI summary
A system and method for creating a machine learning (ML) classifier for a database uses a weakly-supervised training data set created automatically from database items on the basis of a human-created keyword set. The automatically created training data set is used to construct one or more deep learning classifier checkpoints, which can then be compared with one another and with a classifier based on the original keyword set in order to select a classifier for use by other users viewing the database.


