Lexicon-Based Classifier Models for Tunable Compliance Error Rates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional classifier models in compliance enforcement are prone to deficiencies such as high false positive and false negative error rates, lack of transparency, and inability to tailor performance to specific compliance tasks, due to their black box nature and reliance on substantial labeled training data.

Innovation Solution

The development of lexicon-based classifier models that allow users to tune the tradeoff between false positive and false negative error rates via a balance parameter, using segmented lexicons and probabilistic classification to classify content as belonging to a positive or negative class, enabling more effective compliance enforcement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional classifier models are used for compliance enforcement, then automated monitoring can be performed, but high false positive and false negative error rates occur

Engineering Contradiction:
Improveautomated monitoring capabilityVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments the training data into multiple subsets with different error-rate characteristics, and segments the classifier into multiple base classifiers that can be independently tuned. This allows the system to divide the classification task into multiple specialized components, each optimized for different operational requirements, thereby maintaining high productivity while improving reliability through selective application of appropriate classifiers.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a balance parameter that controls the tradeoff between false positive rate and false negative rate. By adjusting this parameter, the system can dynamically change the classification threshold and error-rate characteristics to match different compliance enforcement scenarios, thereby improving reliability without sacrificing automated monitoring capability.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If conventional classifier models are used, then classification can be performed, but the models lack transparency and are black box in nature

Engineering Contradiction:
Improveclassification speedVSAvoidmodel interpretability
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent introduces lexicon-based classifiers as intermediary components that bridge the gap between black box conventional classifiers and interpretable rule-based systems. These lexicon-based classifiers use predefined term lists and scoring rules that are inherently interpretable, while still providing automated classification functionality. The system can explain decisions by referencing specific lexicon terms and scoring criteria, thereby maintaining productivity while improving ease of operation through enhanced transparency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If conventional classifier models are used, then automated compliance monitoring can be implemented, but the models cannot be tailored to specific compliance tasks

Engineering Contradiction:
Improvemonitoring efficiencyVSAvoidtask-specific customization
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates a dynamic classifier selection mechanism where the system can adaptively choose or combine different base classifiers based on the specific compliance task requirements. The balance parameter and lexicon configurations can be dynamically adjusted for different compliance scenarios, allowing the system to maintain high monitoring efficiency while being highly adaptable to specific task requirements through configurable parameters and modular architecture.

Inventive Principle:
Principle #15Dynamics

4Measurement precision

If conventional classifier models are used, then classification can be performed, but substantial labeled training data is required

Engineering Contradiction:
Improveclassification performanceVSAvoidtraining data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent employs lexicon-based classifiers that use precompiled term lists and scoring rules prepared in advance, rather than requiring extensive labeled training data to be processed during deployment. The lexicons and scoring mechanisms are created beforehand and can be applied directly to new content, thereby achieving good classification performance without requiring substantial volumes of labeled training data.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20230214707A1Enhanced lexicon-based classifier models with tunable error-rate tradeoffs
Publication Date: 2023.07.06 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20230214707A1 patent drawing
  • US20230214707A1 patent drawing
  • US20230214707A1 patent drawing

AI summary

The disclosure is directed to systems, methods, and computer storage media, for, among other things, generating, training, and tuning lexicon-based classifier models. The models may be employed in various compliance enforcement applications and/or tasks. The tradeoff between the model's false positive error rate (FPR) and the model's false negative rate (FNR) may be “tuned” via a balance parameter supplied by the user. The classifier model may classify content (e.g., text records) as either belonging to a “positive” class or a “negative” class. The positive class may be associated with non-compliance, while the negative class may be associated with compliance (or vice-versa). In some embodiments, the classifier model may be a probabilistic probability model that provides a probability (or degree of belief) that the content is associated with the positive and/or negative class.