Automated Data Categorization via Unity and Reliability Metrics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data categorization methods are inefficient and prone to errors, particularly when handling complex transaction data, requiring significant time and processing resources, and often necessitate manual intervention to correct incorrect categorizations, making it impractical for large volumes of data.
Innovation Solution
An automated method that generates reliability metrics by determining categorical distributions and unity metrics for transaction rules, allowing for the assessment of rule effectiveness and providing feedback for improvement, thereby enhancing the efficiency and accuracy of data categorization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional categorization processes are used, then processing can be performed, but the processes are inefficient and require significant time and processing resources
Solution Approach 1:
The system performs preliminary actions by pre-defining categorization rules with text strings and categories before processing transactions. Rules are established in advance with specific text string patterns that map to categories, enabling rapid automated categorization without time-consuming manual analysis during processing.
Solution Approach 2:
The patent replaces manual mechanical categorization processes with automated computer-based systems. The mechanical process of human reviewers manually categorizing transactions is substituted with an automated system that applies predefined rules and text string matching algorithms, significantly reducing processing time and resource requirements.
2Reliability
If conventional categorization processes are used, then processing can be performed, but they are prone to errors and require manual intervention to correct
Solution Approach 1:
The system incorporates feedback mechanisms where transactions are first processed through automated rule-based categorization, then reviewed by humans only when necessary. The feedback loop allows automated systems to learn from corrections and refine their categorization rules, improving accuracy over time while reducing manual intervention for routine transactions.
Solution Approach 2:
The categorization system performs self-service by automatically applying predefined rules to categorize transactions without requiring continuous human intervention. The system self-corrects by using feedback from manual corrections to refine automated rules, enabling it to handle categorization tasks independently for the majority of transactions.
3Reliability
If manual intervention is required to correct erroneous categorizations, then accuracy can be improved, but it becomes completely impractical to process large volume of data
Solution Approach 1:
The system segments the categorization process into two distinct parts: automated processing for the majority of transactions and manual review only for exceptional cases. This segmentation allows high-volume data to be processed automatically while maintaining accuracy through targeted human review of only the necessary transactions, making large-volume processing practical.
Solution Approach 2:
The system applies partial action by using automated categorization for all transactions and reserving manual intervention only for the partial subset that requires correction. This approach avoids the excessive action of manually reviewing every transaction, instead applying manual review only where necessary to maintain accuracy at scale.
4Speed
If automated categorization is implemented, then processing speed increases, but rule complexity and difficulty of detection increase
Solution Approach 1:
The system applies local quality by creating specific, targeted text string patterns for different transaction types rather than using complex global rules. Each rule is locally optimized with specific text string combinations that match particular transaction characteristics, making the overall system simpler and easier to manage while maintaining high processing speed through automated application.
Data Source
AI summary
Certain aspects of the present disclosure provide techniques for generating a metric, include receiving a rule defining one or more text strings; determining a set of transactions based on a user attribute; determining a first subset of transactions; determining a second subset of transactions; generating a first categorical distribution based on each transaction of the first subset of transactions being associated with a transaction description containing at least one text string of the one or more text strings; calculating a first unity metric based on the first categorical distribution; generating a second categorical distribution based on each transaction of the second subset of transactions being associated with a transaction description that does not contain a text string of the one or more text strings; calculating a second unity metric based on the second categorical distribution; determining a reliability metric for the rule; and providing the reliability metric.


