Automated Data Categorization via Unity and Reliability Metrics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data categorization methods are inefficient and prone to errors, particularly when handling complex transaction data, requiring significant time and processing resources, and often necessitate manual intervention to correct incorrect categorizations, making it impractical for large volumes of data.

Innovation Solution

An automated method that generates reliability metrics by determining categorical distributions and unity metrics for transaction rules, allowing for the assessment of rule effectiveness and providing feedback for improvement, thereby enhancing the efficiency and accuracy of data categorization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional categorization processes are used, then processing can be performed, but the processes are inefficient and require significant time and processing resources

Engineering Contradiction:
Improvecategorization efficiencyVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-defining categorization rules with text strings and categories before processing transactions. Rules are established in advance with specific text string patterns that map to categories, enabling rapid automated categorization without time-consuming manual analysis during processing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual mechanical categorization processes with automated computer-based systems. The mechanical process of human reviewers manually categorizing transactions is substituted with an automated system that applies predefined rules and text string matching algorithms, significantly reducing processing time and resource requirements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If conventional categorization processes are used, then processing can be performed, but they are prone to errors and require manual intervention to correct

Engineering Contradiction:
Improvecategorization accuracyVSAvoidmanual intervention requirement
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system incorporates feedback mechanisms where transactions are first processed through automated rule-based categorization, then reviewed by humans only when necessary. The feedback loop allows automated systems to learn from corrections and refine their categorization rules, improving accuracy over time while reducing manual intervention for routine transactions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The categorization system performs self-service by automatically applying predefined rules to categorize transactions without requiring continuous human intervention. The system self-corrects by using feedback from manual corrections to refine automated rules, enabling it to handle categorization tasks independently for the majority of transactions.

Inventive Principle:
Principle #25Self-service

3Reliability

If manual intervention is required to correct erroneous categorizations, then accuracy can be improved, but it becomes completely impractical to process large volume of data

Engineering Contradiction:
Improvecategorization accuracyVSAvoidprocessing volume capacity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system segments the categorization process into two distinct parts: automated processing for the majority of transactions and manual review only for exceptional cases. This segmentation allows high-volume data to be processed automatically while maintaining accuracy through targeted human review of only the necessary transactions, making large-volume processing practical.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies partial action by using automated categorization for all transactions and reserving manual intervention only for the partial subset that requires correction. This approach avoids the excessive action of manually reviewing every transaction, instead applying manual review only where necessary to maintain accuracy at scale.

Inventive Principle:
Principle #16Partial or excessive action

4Speed

If automated categorization is implemented, then processing speed increases, but rule complexity and difficulty of detection increase

Engineering Contradiction:
Improveprocessing speedVSAvoidrule complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The system applies local quality by creating specific, targeted text string patterns for different transaction types rather than using complex global rules. Each rule is locally optimized with specific text string combinations that match particular transaction characteristics, making the overall system simpler and easier to manage while maintaining high processing speed through automated application.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12039267B2Automated categorization of data by generating unity and reliability metrics
Publication Date: 2024.07.16 INTUIT INC
  • US12039267B2 patent drawing
  • US12039267B2 patent drawing
  • US12039267B2 patent drawing

AI summary

Certain aspects of the present disclosure provide techniques for generating a metric, include receiving a rule defining one or more text strings; determining a set of transactions based on a user attribute; determining a first subset of transactions; determining a second subset of transactions; generating a first categorical distribution based on each transaction of the first subset of transactions being associated with a transaction description containing at least one text string of the one or more text strings; calculating a first unity metric based on the first categorical distribution; generating a second categorical distribution based on each transaction of the second subset of transactions being associated with a transaction description that does not contain a text string of the one or more text strings; calculating a second unity metric based on the second categorical distribution; determining a reliability metric for the rule; and providing the reliability metric.