Chatbot-Based Data Labeling for Machine Learning Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The complexity of cloud-based systems necessitates efficient monitoring and maintenance, which is hindered by the time-consuming and costly process of manually labeling data for machine learning model training, especially in high-volume, fast-paced environments.

Innovation Solution

Implementing a chatbot-based system that analyzes data points, generates alert tickets, communicates with users, and trains machine learning models using user feedback, enabling real-time data labeling without requiring specialized knowledge.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual data labeling is used for machine learning model training, then data accuracy can be ensured, but the process becomes time-consuming and expensive

Engineering Contradiction:
Improvedata labeling accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automated data labeling through chatbot interactions, where the machine learning model itself performs the labeling task by analyzing data points and generating labels based on user feedback, eliminating the need for manual human labeling while maintaining accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system incorporates feedback loops where user responses to chatbot queries are used to continuously improve and refine the automated labeling process, allowing the model to learn from interactions and enhance labeling accuracy over time without additional manual intervention

Inventive Principle:
Principle #23Feedback

2Measurement precision

If specialized knowledge is required for data labeling, then labeling quality improves, but the complexity of the system increases

Engineering Contradiction:
Improvelabeling qualityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The chatbot serves as an intermediary between users and the machine learning model, translating user feedback into structured labeling data. This intermediary layer simplifies the interface, allowing users without specialized knowledge to contribute to high-quality data labeling through natural conversation

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system automatically processes and structures user feedback from chatbot interactions into training data, eliminating the need for specialized manual intervention. The automated pipeline handles complexity internally while presenting a simple interface to end users

Inventive Principle:
Principle #25Self-service

3Speed

If real-time data labeling is implemented, then system responsiveness improves, but the computational resources required increase

Engineering Contradiction:
Improveresponse timeVSAvoidcomputational resource consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary analysis of data points through chatbot interactions before full model training is required. By pre-processing and pre-labeling data in real-time through automated chatbot queries, the system prepares training data ahead of time, reducing the computational burden during actual training while maintaining fast response times

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10891560B2Supervised learning system training using chatbot interaction
Publication Date: 2021.01.12 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10891560B2 patent drawing
  • US10891560B2 patent drawing
  • US10891560B2 patent drawing

AI summary

An apparatus comprises a memory and a processor coupled to the memory. The processor is configured to receive input from a cloud service data source, wherein the input comprises at least one data point, analyze the data point via a machine learning model to determine characteristics indicated by the data point, determine whether the characteristics indicated by the data point meet an alert threshold that indicates a problem in a network, generate an alert ticket when the characteristics indicated by the data point meet the alert threshold, wherein the alert ticket indicates the problem in the network, communicate with a user based on contents of the alert ticket, receive feedback from the user relating to the alert ticket, and train the machine learning model according to the feedback received from the user.