Self-service Classification System Using Iterative ML Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current document classification methods require extensive training sets and are labor-intensive, time-consuming, and prone to errors, making them inaccessible to non-technical users who need to classify large volumes of unstructured documents efficiently.

Innovation Solution

A self-service classification system that employs a combination of machine learning and user-interaction techniques, using a smaller initial data set and a novel workflow to generate a customized classification model, allowing users to tune and improve the model through triage rules, feature selection, and user feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual classification is used to identify specific information within large amounts of unstructured documents, then classification accuracy can be maintained, but labor intensity and time consumption increase significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service classification by automatically generating classification models that users can apply to their documents without manual intervention. The model generation process is automated, requiring users only to provide initial topics and optional seed documents, after which the system performs iterative training and model generation independently.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual classification process with an automated machine learning system. Instead of professionals manually reading and classifying documents, the system uses trained classification models to automatically categorize documents, substituting human cognitive effort with computational processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If Machine Learning models are trained to perform classification, then productivity increases, but the requirement for Machine Learning expertise and extensive training data creates barriers for non-technical users

Engineering Contradiction:
Improveclassification throughputVSAvoiduser accessibility
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system provides self-service capabilities by automatically generating classification models without requiring users to have Machine Learning expertise. Users simply input their topics and optional seed documents, and the system handles model training, evaluation, and generation autonomously, making advanced classification accessible to non-technical users.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary actions by automatically conducting iterative training runs, evaluating model performance, and generating classification models before users need them. The system proactively prepares models based on user inputs, eliminating the need for users to understand or perform complex training procedures.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If extensive training sets are used to generate classification models, then model quality improves, but data preparation becomes labor intensive and time consuming

Engineering Contradiction:
Improvemodel qualityVSAvoiddata preparation effort
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The system applies partial action by using only the minimum necessary training data - users provide just topics and optional seed documents rather than extensive labeled datasets. The system then performs multiple iterative training runs with this partial data, achieving good model quality without requiring complete or excessive training sets.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs preliminary iterative training runs automatically before generating the final classification model. These preliminary actions allow the system to evaluate different model configurations and select the best performing model, achieving high quality without requiring users to manually prepare extensive training data.

Inventive Principle:
Principle #10Preliminary action

4Device complexity

If ad hoc rule based solutions are used for classification, then implementation is simpler, but the solutions are inadequate and very hard to maintain

Engineering Contradiction:
Improvesystem simplicityVSAvoidclassification effectiveness
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent replaces simple but inadequate rule-based systems with automated machine learning models. Instead of using basic keyword matching or manual rules that are easy to implement but ineffective, the system employs trained classification models that provide both simplicity of use and high reliability in classification effectiveness.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10990897B2Self-service classification system
Publication Date: 2021.04.27 REFINITIV US ORGANIZATION LLC
  • US10990897B2 patent drawing
  • US10990897B2 patent drawing
  • US10990897B2 patent drawing

AI summary

Systems, technologies and techniques for generating a customized classification model are disclosed. The system and technologies, such as THOMSON REUTERS SELF-SERVICE CLASSIFICATION™, employ part machine learning and part an user interactive approach to generate a customized classification model. The system combines a novel approach for text classification using a smaller initial set of data to initiate training, with a unique workflow and user interaction for customization.