Document Classification Class Vectors for False Positive Disambiguation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional automated document classification methods using natural language processing struggle with low accuracy, particularly when documents are similar to each other but belong to distinct classes, and require significant effort and time to train new neural networks for new document classes.

Innovation Solution

A method and system that adjusts class definitions by analyzing misclassified documents, moving class centers in semantic space based on misclassified documents, and iteratively refining the classification system to improve accuracy without retraining neural networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If conventional neural network approaches are used to classify documents, then automation is achieved, but classification accuracy deteriorates when documents are similar but belong to distinct classes

Engineering Contradiction:
Improveautomation of document classificationVSAvoidclassification accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent implements feedback by analyzing misclassified documents and using them to adjust class definitions. The system identifies documents that were misclassified, extracts features from these documents, and uses this information to refine the class definitions and vectors, thereby improving classification accuracy for similar documents in future iterations.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes parameters by adjusting class definitions and class vectors based on misclassified documents. Specifically, it modifies the class vectors in semantic space by incorporating features from misclassified documents, which changes the parameter representation of each class and improves the system's ability to distinguish between similar document types.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If neural networks are trained to recognize each document class, then classification capability is improved, but training time and effort increase significantly when new classes are added

Engineering Contradiction:
Improveclassification capabilityVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements universality by creating a single classification system that can handle multiple document classes without requiring separate neural networks for each class. The system uses a unified approach with class vectors and semantic space representation that can accommodate any number of document classes, making the system adaptable and eliminating the need for extensive retraining when new classes are introduced.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies dynamics by making the class definitions and class vectors adaptable and modifiable. The system can dynamically adjust class definitions based on misclassified documents and can easily incorporate new document classes by creating new class vectors, without requiring static, pre-trained neural networks for each specific class.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If class definitions are adjusted based on misclassified documents, then classification accuracy is improved, but system complexity increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements self-service by enabling the classification system to automatically identify its own weaknesses through misclassified documents and self-correct by adjusting its class definitions. The system autonomously analyzes errors, extracts relevant features from misclassified documents, and modifies its own parameters without requiring external intervention or complex manual tuning processes.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12572746B2Deep learning systems and methods to disambiguate false positives in natural language processing analytics
Publication Date: 2026.03.10 NUIX
  • US12572746B2 patent drawing
  • US12572746B2 patent drawing
  • US12572746B2 patent drawing

AI summary

Embodiments improve a document classification system by adjusting the class vector of each class based on embedded vectors of documents known to have been misclassified by the system. Some embodiments use a set of misclassified documents to adjust the class vector of the class into which a document has been misclassified. Some embodiments use a set of misclassified documents to adjust the class vector of the pre-assigned class of each misclassified document. Embodiments of adjusting the class vector may be described as machine learning.