ERP Transaction Categorization Using Clustering and Rule-Based Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The complexity of coded ERP data formats across different companies necessitates significant human expertise for interpretation, leading to challenges in automating transaction processing, delayed scalability, and increased error rates in transaction classification, which affects financial reporting and decision-making.
Innovation Solution
A machine learning-based computing system and method that utilizes clustering models like DBSCAN and KNN to identify transaction categories from ERP data, generating features and applying pre-configured rules to automate the classification process, reducing human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual decoding of ERP data is used, then human expertise can interpret complex coded formats, but it results in time-consuming processes and high error rates in transaction classification
Solution Approach 1:
The patent replaces the manual mechanical decoding process with an automated machine learning system. The ML model trained on historical ERP data automatically classifies transactions by learning patterns in coded formats, eliminating the need for human experts to manually interpret each transaction code while maintaining or improving classification accuracy.
Solution Approach 2:
The system enables self-service automation where the ML model independently processes and classifies transactions without human intervention. The model learns from historical data and autonomously applies classification rules to new transactions, reducing dependency on human expertise and significantly decreasing processing time.
2Extent of automation
If standardized interpretation of ERP codes is implemented, then automation of processes can be achieved, but it requires overcoming the diversity of codes across different companies
Solution Approach 1:
The patent creates a universal ML-based classification system that can handle multiple diverse ERP code formats from different companies. The model is trained on varied historical data representing different code structures, enabling it to adaptively classify transactions across diverse formats without requiring company-specific customization, thus achieving both automation and versatility.
Solution Approach 2:
The system transforms unstructured or semi-structured ERP code data into standardized numerical features that the ML model can process. By converting diverse code formats into a common feature space through data preprocessing and feature engineering, the system enables automated classification while accommodating parameter variations across different ERP systems.
3Reliability
If human experts manually classify transactions, then complex coded formats can be interpreted accurately, but it leads to scalability issues and delays in processing
Solution Approach 1:
The patent substitutes human expert interpretation with an automated ML system that maintains reliability through rigorous training on historical data. The model learns reliable classification patterns from expert-labeled historical transactions and applies them consistently to new data, achieving both high reliability and scalable throughput without human intervention.
Solution Approach 2:
The system performs preliminary action by pre-training the ML model on extensive historical ERP data before deployment. This preliminary training phase allows the model to learn reliable classification patterns in advance, enabling it to quickly and accurately process new transactions at high throughput without requiring real-time human expert involvement.
Data Source
AI summary
A machine learning based (ML-based) computing method for identifying transaction categories from enterprise resource planning (ERP) data, is disclosed. The machine learning based computing method includes steps of: receiving inputs from electronic devices associated with first users; retrieving first data associated with first transactions from databases, based on the inputs received from the electronic devices associated with the first users; generating features associated with the transaction categories based on the first data associated with the first transactions; identifying first transaction categories based on at least one of: the first data associated with the first transactions and the features associated with the transaction categories, by clustering based machine learning models; identifying second transaction categories based on pre-configured rules; and providing an output of at least one of: the first transaction categories and the second transaction categories, to the first users on user interfaces associated with the electronic devices.


