Context Category Dataset Generation via Human-Machine Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating datasets for natural language processing, such as crowdsourcing and machine prediction, face challenges in accuracy and efficiency, particularly in human-machine collaboration for context category classification.

Innovation Solution

A context category dataset generation apparatus and method using a user interface that provides a hashtag list and word embedding vectors to predict context categories, allowing user feedback for updating the dataset, thereby improving prediction accuracy through human-machine collaboration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If crowdsourcing method is used to generate dataset, then productivity is improved, but manufacturing precision deteriorates

Engineering Contradiction:
Improvedata generation speedVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system implements a feedback mechanism where machine prediction results are presented to users for review and correction. Users provide feedback by confirming or correcting the predicted context categories, which is then used to update and improve the model's future predictions. This feedback loop enables the system to maintain high productivity while improving manufacturing precision over time through human-in-the-loop learning.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary classification using machine prediction before human review. By pre-processing the data with AI models to generate initial context category predictions, the system reduces the workload on human annotators and accelerates the overall data generation process while maintaining quality through subsequent human verification.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If machine prediction method is used to generate dataset, then productivity is improved, but manufacturing precision deteriorates

Engineering Contradiction:
Improvedata generation efficiencyVSAvoidcontext category accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system introduces an intermediary human review step between machine prediction and final dataset generation. The user interface presents machine prediction results to users, who act as intermediaries to validate or correct the predictions. This intermediary layer bridges the gap between automated efficiency and human accuracy, allowing the system to leverage both machine speed and human judgment.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback loops where user corrections to machine predictions are captured and used to retrain and improve the prediction model. This continuous learning process allows the system to maintain high productivity while progressively improving manufacturing precision through accumulated human feedback.

Inventive Principle:
Principle #23Feedback

3Manufacturing precision

If human-machine collaboration method is used, then manufacturing precision is improved, but device complexity increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidsystem structure complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system segments the data generation process into distinct modules: machine prediction component, user interface component, and feedback processing component. Each module handles a specific aspect of the workflow, making the overall complex system more manageable and maintainable while enabling human-machine collaboration to improve manufacturing precision.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system incorporates self-service elements where users can review and correct their own predictions, and the system automatically processes feedback to improve future predictions. This reduces the need for complex manual intervention processes while maintaining high manufacturing precision through user involvement.

Inventive Principle:
Principle #25Self-service

4Manufacturing precision

If human review process is added to machine prediction, then manufacturing precision is improved, but loss of time increases

Engineering Contradiction:
Improvedataset accuracyVSAvoiddata generation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system implements partial human review by presenting only predictions that require verification to users, while accepting high-confidence machine predictions without human intervention. This selective approach maintains manufacturing precision for critical cases while minimizing time loss by avoiding unnecessary human review for obvious predictions.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs preliminary filtering and confidence assessment of machine predictions before presenting them to users. By pre-processing predictions to identify only those requiring human review, the system reduces the time users need to spend while maintaining manufacturing precision through targeted human verification of uncertain cases.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11580101B2Method and apparatus for generating context category dataset
Publication Date: 2023.02.14 KOREA ADVANCED INST OF SCI & TECH
  • US11580101B2 patent drawing
  • US11580101B2 patent drawing
  • US11580101B2 patent drawing

AI summary

The present disclosure provides an apparatus for and method of generating a context category dataset. According to some embodiments, the present disclosure provides a context category dataset generating apparatus and method which predict a context category to which a user-inputted hashtag belongs, receive from the user the user's context category to which the hashtag belongs, and generate and update the context category dataset.