Automated Document Categorization via Question Answering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual content categorization in cybersecurity is labor-intensive and relies heavily on human experts, lacking an effective automated solution to reduce operational costs and skill shortages.

Innovation Solution

An automated document categorization method using Natural Language Processing (NLP) and Question and Answer (Q&A) technology, where a bank of questions is composed based on domain expertise and answered by language models to generate features for documents, mimicking human analyst interactions and reducing manual workload.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual classification by subject matter experts is used, then classification accuracy is improved, but productivity deteriorates due to labor-intensive work

Engineering Contradiction:
Improveclassification accuracyVSAvoidcategorization efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent creates a corpus of categorized documents by copying and structuring expert classification knowledge into a standardized format. This copied knowledge base enables automated systems to replicate expert-level categorization without requiring continuous human expert intervention, thereby maintaining accuracy while improving productivity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces an intermediary question-answering system that mediates between the document to be categorized and the classification categories. This intermediary automatically generates questions and answers based on the document content, translating unstructured text into structured categorical information without direct human involvement in each categorization task

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated question answering is used, then productivity is improved, but measurement precision may deteriorate compared to expert manual classification

Engineering Contradiction:
Improvecategorization throughputVSAvoidcategorization accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary action by pre-processing documents to generate questions and answers before final categorization. The question-answering mechanism proactively extracts key information and formulates it in a structured format, preparing the data in advance for accurate automated classification without requiring expert review of every document

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the mechanical system of human expert analysis with an automated question-answering language model. This substitution uses natural language processing to perform categorization tasks that previously required human cognitive effort, achieving both high productivity and maintained precision through AI-driven document understanding

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20240303259A1Imitating analyst's content categorization with automatic question answering
Publication Date: 2024.09.12 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20240303259A1 patent drawing
  • US20240303259A1 patent drawing
  • US20240303259A1 patent drawing

AI summary

A document categorization method, system, and computer program product that includes forming a corpus of categorized documents by relying on a manual classification of a subject matter expert, composing a bank of questions, and answering each question automatically using a question answering language model with respect to each document and generating a set of features for each document.