Probabilistic Document Classification via Confidence Level Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document classification approaches face challenges with high false positive and false negative rates, and optimization difficulties due to their accumulative and weight-based systems, which hinder effective identification of sensitive information in documents.

Innovation Solution

A probabilistic document classification method that analyzes patterns and evidences within documents to construct classification rules, evaluates these rules against acceptance criteria, and adjusts confidence levels, allowing for iterative rule optimization based on precision and recall thresholds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If weight based systems are used for document classification, then the system can accumulate classification evidence, but the system cannot be easily optimized over time and produces high false positive and false negative rates

Engineering Contradiction:
Improveclassification accuracyVSAvoidoptimization difficulty
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent transforms the classification system from using fixed weights to using probabilistic parameters that can be dynamically adjusted. The classification rules now incorporate confidence levels and probability thresholds that can be optimized over time through iterative testing and feedback, allowing the system to adapt to changing document patterns while maintaining mathematical rigor through Bayesian probability updates.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system transitions from static weight-based classification to dynamic probabilistic classification. The confidence levels and probability values are continuously updated based on new evidence and testing results, allowing the classification rules to evolve and improve over time. This dynamic approach enables easy optimization through iterative testing while maintaining system reliability.

Inventive Principle:
Principle #15Dynamics

2Productivity

If traditional classification approaches are used, then the system can process documents, but false positive and false negative rates remain high

Engineering Contradiction:
Improvedocument processing capabilityVSAvoidclassification precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements a feedback mechanism where classification results are continuously evaluated against expected outcomes through testing. The system uses precision and recall metrics to measure performance and automatically adjusts confidence levels and probability thresholds based on testing feedback. This closed-loop feedback system reduces false positives and negatives while maintaining efficient document processing throughput.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces the mechanical weight-based accumulation system with a probabilistic mathematical model. Instead of simply adding weights when patterns are found, the system uses Bayesian probability to calculate confidence levels, providing a more precise measurement of classification certainty. This substitution enables better discrimination between true positives and false positives, improving measurement precision while maintaining processing efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9495639B2Determining document classification probabilistically through classification rule analysis
Publication Date: 2016.11.15 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9495639B2 patent drawing
  • US9495639B2 patent drawing
  • US9495639B2 patent drawing

AI summary

A classification application identifies patterns and evidences within representative documents. The application constructs a classification rule according to an entity and an affinity determined from the patterns and evidences. The application processes the representative documents with the classification rule to evaluate whether the rules meet acceptance requirements. Subsequent to a successful evaluation, the application identifies confidence levels for patterns and evidences within other documents.