Training Corpus Refinement via Autonomous Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for creating and refining training corpora for text classification are manual, time-consuming, error-prone, and lack quality assurance, leading to poor classifier accuracy due to issues like inter-class overlap and intra-class noise, and they rely heavily on manual intervention and lack self-learning capabilities.

Innovation Solution

An autonomous regenerative feedback mechanism using a Corpus Advisor with a diagnosis machine learning model for overlap and noise treatment, and a self-learning AI control system that refines and augments the training corpus based on user feedback, incorporating a reinforcement learning model to validate and integrate new intelligence in a controlled manner.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual methods are used to create and refine training corpora, then the process allows for human judgment and correction, but the process becomes time-consuming, error-prone, and lacks self-learning capabilities

Engineering Contradiction:
Improveclassifier accuracyVSAvoidtime for corpus refinement
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-diagnosis of training corpus quality issues and self-refinement by automatically identifying and removing overlapping and noisy samples, eliminating the need for manual intervention while maintaining high classifier accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements a feedback mechanism where classification results are analyzed to identify quality issues in the training corpus, which then triggers automated refinement processes to improve future classification accuracy

Inventive Principle:
Principle #23Feedback

2Reliability

If the training corpus includes all available samples, then the corpus is comprehensive, but it contains overlapping and noisy samples that reduce classification accuracy

Engineering Contradiction:
Improveclassification accuracyVSAvoidcorpus refinement process
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system extracts and removes harmful elements (overlapping samples and noisy samples) from the training corpus through automated diagnosis and filtering, retaining only high-quality samples that improve classification accuracy

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary diagnosis and refinement of the training corpus before classification tasks, identifying and removing quality issues in advance to prevent them from affecting classification accuracy

Inventive Principle:
Principle #10Preliminary action

3Extent of automation

If manual refinement of training corpora is performed, then quality control can be applied, but the process lacks automation and scalability

Engineering Contradiction:
Improvecorpus refinement automationVSAvoidquality assurance
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system replaces manual mechanical refinement processes with automated computational algorithms that diagnose corpus quality, identify overlapping and noisy samples, and perform refinement without human intervention, achieving both automation and precision

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Adaptability or versatility

If the training corpus is frequently updated with new data, then the system adapts to new information, but the quality and consistency of the corpus may deteriorate

Engineering Contradiction:
Improvesystem adaptabilityVSAvoidcorpus consistency
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The system continuously monitors the training corpus for quality degradation through automated diagnosis, detecting overlapping and noisy samples that arise from new data additions, and triggers refinement processes to restore corpus consistency while preserving adaptability

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs periodic quality checks and refinement cycles on the training corpus, maintaining consistent quality standards through regular automated diagnosis and cleaning operations

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS11488055B2Training corpus refinement and incremental updating
Publication Date: 2022.11.01 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11488055B2 patent drawing
  • US11488055B2 patent drawing
  • US11488055B2 patent drawing

AI summary

Training corpus refinement and incremental updating includes obtaining a training corpus having training samples, refining the training corpus to produce a refined training corpus of data, by applying to the training corpus overlap and noise reduction treatments, maintaining an incremental intelligence database based on filtered user feedback and having candidate feedback training samples to augment the refined training corpus, controlling integration of the candidate feedback training samples with the refined training corpus, and augmenting the refined training corpus with at least some of the candidate feedback training samples to produce an augmented training corpus.