AI Spend Categorization via Clustering and Human-in-the-Loop
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Zero-based budgeting (ZBB) is a time-consuming process due to inaccurate and inefficient classification of transactional data, requiring significant human effort and being specific to individual clients and industries, making it difficult to achieve accurate cost categorization across different sectors.
Innovation Solution
The method involves processing and consolidating spend data to generate a cleaned data set, clustering logs based on similarity, and using a hierarchical category structure with machine learning techniques and Human-in-the-Loop validation to categorize spend data, allowing for accurate and efficient categorization without prior knowledge of client practices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual categorization of transactional data is performed, then accuracy in cost classification can be achieved, but significant human effort and time are required
Solution Approach 1:
The patent replaces manual mechanical categorization with an automated machine learning system that uses natural language processing and classification algorithms to categorize transactional data, eliminating the need for human reviewers while maintaining or improving accuracy
Solution Approach 2:
The system enables self-service automated categorization where the machine learning model independently processes and categorizes transactional data without requiring human intervention, allowing the system to serve itself in the categorization task
2Productivity
If machine learning techniques are used to automate categorization, then processing speed increases, but model accuracy is insufficient across different industry sectors
Solution Approach 1:
The patent creates a universal machine learning model that can be applied across multiple industry sectors and client types, making the system multi-functional and transferable beyond single-client specific applications while maintaining high accuracy through continuous learning and adaptation
Solution Approach 2:
The system dynamically adjusts model parameters and thresholds based on the specific characteristics of different industry sectors and client data, allowing the same base model to adapt to varying accuracy requirements across different domains
3Reliability
If conventional manual approaches are used for data classification, then in-depth knowledge of client practices can be incorporated, but the process requires hundreds of individuals and is highly specific to each client
Solution Approach 1:
The patent replaces the complex human resource system with an automated machine learning platform that can process and understand client-specific practices without requiring hundreds of individual reviewers, simplifying the operational complexity while maintaining reliability
Data Source
AI summary
A largely automated method of categorizing spend data is provided that does not require a prior in-depth knowledge of an organization's transactional data. Natural language processing is applied to text data from transactional data to generate a consolidated cleaned data set (CDS) containing information for categorization. Logs for transactions are clustered based on similarity, forming the minimal data set (MDS). An automated algorithm selects a subset of high-value clusters that are categorized by requesting users to manually categorize one or more representative logs from each cluster of the subset. A model is then trained using the subset of manually categorized clusters and used to predict spend categories for the remaining logs with high accuracy. The AI engine automatically analyzes the predictions based on client context and either auto-tunes the machine learning model or identifies a new subset of clusters to be manually categorized. This loop may continue until 95%-100% of the spend is categorized.


