Automatic Rule Generation for Document Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document classification and extraction systems are complex, time-consuming, and difficult to scale due to the need for manual rule creation and updates, especially with the constant evolution of document types and industries, leading to inefficiencies in handling large and diverse volumes of data.

Innovation Solution

A method and system that uses a combination of natural language processing, machine learning, and deep learning algorithms to automatically generate and update classification, extraction, and validation rules, enabling predictive and adaptive document processing by identifying similar document types and applying relevant rules based on attributes and user feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual rule creation and updates are used for document classification, then initial system setup is possible, but the system becomes complex and difficult to maintain as document types evolve

Engineering Contradiction:
Improveclassification accuracyVSAvoidrule management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system automatically generates, updates, and maintains classification rules without manual intervention. The rule generation module continuously learns from new document types and automatically adapts the classification model, eliminating the need for manual rule creation and updates while maintaining high classification accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The classification rules are made dynamic and adaptive rather than static. The system continuously evolves the rules based on new document types and patterns, allowing the rule set to automatically adjust and improve over time without increasing management complexity

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If manual rule updates are performed to handle evolving document types, then adaptability improves, but time consumption and processing delays increase

Engineering Contradiction:
Improveadaptability to new document typesVSAvoidtime for rule updates
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The rule generation operates continuously and automatically in the background. The system continuously monitors new document types and continuously updates classification rules without interruption, eliminating time losses associated with manual rule updates while maintaining high adaptability

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system implements feedback loops where classification results and new document types are continuously fed back to the rule generation module. This automatic feedback mechanism enables the system to adapt to new document types in real-time without manual intervention or time delays

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If comprehensive manual classification rules are created for all document types, then classification coverage is improved, but the initial setup time and resource requirements increase

Engineering Contradiction:
Improvedocument type coverageVSAvoidinitial setup time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary automatic rule generation based on available sample documents without requiring comprehensive manual rule creation. It prepares initial classification rules automatically and continuously improves them, reducing initial setup time while achieving comprehensive document type coverage through iterative learning

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The rule generation module serves multiple functions: it generates initial rules, updates existing rules, handles new document types, and validates classification results. This multi-functional automatic system replaces multiple manual processes, achieving comprehensive coverage without proportional increases in setup time

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Productivity

If automated classification is implemented, then processing speed improves, but accuracy may decrease without proper validation rules

Engineering Contradiction:
Improveclassification speedVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system implements multiple feedback mechanisms including validation rules that check classification results, confidence score thresholds that trigger re-evaluation, and continuous learning from corrected classifications. This feedback loop maintains high accuracy while preserving the speed benefits of automated classification

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The validation and accuracy checking processes are automated rather than manual. The system self-validates classification results using generated validation rules and automatically corrects or re-processes low-confidence classifications, maintaining accuracy without sacrificing processing speed

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11810381B2Automatic rule prediction and generation for document classification and validation
Publication Date: 2023.11.07 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11810381B2 patent drawing
  • US11810381B2 patent drawing
  • US11810381B2 patent drawing

AI summary

A method is provided. The method may include, in response to electronically receiving a document, automatically classifying the document and different parts of the document, by electronically identifying a document type associated with the document and electronically tagging data associated with the different parts of the document based on classification rules. The method may further include automatically extracting the tagged data associated with the automatically classified document based on data extraction rules. The method may further include detecting first feedback associated with the classification rules and second feedback associated with the data extraction rules. The method may further include automatically generating and updating validation rules based on the identified document type, the detected first feedback, and the detected second feedback to validate the automatically classified document and the automatically tagged and extracted data.