Weakly Supervised Classifier Creation for Feedback Text Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The development of machine learning models for data processing is hindered by the labor-intensive process of labeling data for supervised learning, with advanced models requiring even more training data and existing weakly supervised models being insufficient for many applications.

Innovation Solution

A system and method for creating a large weakly supervised training set using a keyword-based classifier to automatically generate a training dataset, which is then used to construct deep learning classifiers, allowing for quicker and easier model development, and facilitating the construction of human-validated validation sets for both keyword-based and machine-learning-based topics in databases.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling is used to create training data for supervised machine learning models, then the quality and accuracy of the model can be improved, but the time and labor required increases significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoidlabeling time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary automated labeling using keyword-based classifiers and weakly supervised models to generate initial training data before final manual validation. This preliminary action creates a ready-to-use training dataset that requires minimal manual intervention, significantly reducing the time and labor needed for complete manual labeling while maintaining model accuracy through subsequent validation steps.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If advanced neural network models are used to improve classification performance, then the model capability increases, but the amount of training data required increases even more

Engineering Contradiction:
Improveclassification performanceVSAvoidtraining data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system introduces weakly supervised models and keyword-based classifiers as intermediary components that bridge the gap between limited manually labeled data and the requirements of advanced neural network models. These intermediaries generate additional pseudo-labeled training data from unlabeled corpora, enabling advanced models to achieve high classification performance without requiring proportionally large amounts of manually labeled training data.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of manufacture

If keyword-based classification is used to simplify the modeling process, then the ease of creation improves, but the ability to understand topic trends and sentiment decreases

Engineering Contradiction:
Improvemodel creation easeVSAvoidtopic understanding
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The system merges keyword-based classification with weakly supervised machine learning models and neural network-based sentiment analysis to create a hybrid classification system. This combination maintains the simplicity and ease of creation associated with keyword-based methods while incorporating the advanced topic understanding and sentiment analysis capabilities of machine learning models, thereby preserving both ease of manufacture and information quality.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12086543B2Rule-based machine learning classifier creation and tracking platform for feedback text analysis
Publication Date: 2024.09.10 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12086543B2 patent drawing
  • US12086543B2 patent drawing
  • US12086543B2 patent drawing

AI summary

A system and method for creating a machine learning (ML) classifier for a database uses a weakly-supervised training data set created automatically from database items on the basis of a human-created keyword set. The automatically created training data set is used to construct one or more deep learning classifier checkpoints, which can then be compared with one another and with a classifier based on the original keyword set in order to select a classifier for use by other users viewing the database.