ML Spend Classification Multi-Classifier Pipeline
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Accurate spend classification is challenging due to inconsistent and poor categorization of transactions, exacerbated by evolving organizational spend patterns and user fatigue, leading to flawed spend analysis and misrepresentation of spending data.
Innovation Solution
A machine learning-based system using unsupervised algorithms and a multi-classifier pipeline, including index, Wikipedia, Bayes' rule, and embeddings classifiers, to automatically assign spend categories, reducing the need for labeled data and improving classification accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual spend classification is performed by users, then categorization can be applied to transactions, but user fatigue and inconsistency lead to poor classification quality
Solution Approach 1:
The system enables self-service classification by automatically categorizing spend transactions using machine learning algorithms. The classifier processes transactions autonomously without requiring manual user intervention, eliminating user fatigue while maintaining high classification accuracy through continuous learning from labeled data
Solution Approach 2:
The patent replaces the mechanical manual classification process with an automated machine learning system. The mechanical action of users manually categorizing transactions is substituted with an electronic classification system that uses trained models to automatically assign categories, improving both ease of operation and classification precision
2Device complexity
If traditional classification methods are used, then processing is simpler, but classification accuracy deteriorates due to inconsistent categorization
Solution Approach 1:
The classification system is segmented into multiple specialized classifiers (index-based classifier, encyclopedia-based classifier, Bayes' rule-based classifier, and embeddings classifier). Each classifier handles specific aspects of transaction categorization, and their results are combined to achieve high overall accuracy while managing system complexity through modular design
Solution Approach 2:
The patent employs a composite classification approach by integrating multiple different classification algorithms and data sources (transaction indexes, encyclopedia knowledge bases, statistical models, and embedding representations). This composite system leverages the strengths of each individual classifier to achieve superior classification accuracy that exceeds what any single method could provide
3Quantity of substance
If comprehensive spend tracking is implemented across all transactions, then complete spend data is captured, but classification consistency deteriorates due to patchwork catalogs and compliance issues
Solution Approach 1:
The machine learning classifier is designed as a universal system that can handle diverse transaction types across different catalogs and data sources. It processes various spend categories (goods, services, expenses) through a unified classification framework, maintaining consistent categorization standards across the entire organization's spend data regardless of source or type
4Adaptability or versatility
If evolving spend categories are continuously updated, then categorization remains current, but reclassification of historic data becomes necessary
Solution Approach 1:
The system performs preliminary classification of transactions at the time they occur, assigning categories based on the taxonomy available at that moment. When the category taxonomy evolves, the system can efficiently reclassify only the affected transactions using the updated models, rather than requiring complete reclassification of all historic data, thus reducing the time loss associated with taxonomy updates
Data Source
AI summary
Embodiments classify a product to a product category. Embodiments receive a textual description of the product and create an index of product categories. Embodiments use a plurality of classifiers to classify the textual description to one of the product categories, the classifiers including an index based classifier, an encyclopedia based classifier, a Bayes' rule based classifier, and an embeddings classifier.


