Supplier-Driven Commodity Classification for Ambiguous E-Procurement Records
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing large-scale database systems struggle to accurately correlate transactions when commodities are identified using different descriptions or supplier names in different documents, leading to inefficiencies in e-procurement systems, despite the underlying commodities being the same.
Innovation Solution
A tiered, supplier-driven approach using machine learning models and heuristic constraints to classify commodities, supplemented by manual curation for high-spend suppliers, reduces the need for extensive training data and improves categorization accuracy by focusing on supplier categories.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If traditional machine learning models are used to classify commodities based on invoice descriptions and supplier names, then categorization can be automated, but accuracy deteriorates when different descriptions or supplier names refer to the same commodity
Solution Approach 1:
The patent combines multiple data sources (invoice descriptions, supplier names, and community intelligence from multiple enterprises) into a unified classification approach. By merging these diverse inputs and using consensus-based categorization across multiple enterprises, the system achieves more accurate commodity classification than traditional single-source machine learning models.
Solution Approach 2:
The patent introduces community intelligence as an intermediary layer between raw invoice data and final classification. This intermediary aggregates categorization information from multiple enterprises and suppliers, mediating the classification process to resolve ambiguities that single-enterprise systems cannot handle.
2Measurement precision
If extensive training datasets are used to improve machine learning model accuracy, then categorization precision improves, but data preparation complexity and time increase significantly
Solution Approach 1:
The system enables self-service categorization by leveraging existing community intelligence and supplier-provided information. Rather than requiring extensive manual training data preparation, the system automatically aggregates and utilizes categorization data that already exists across the enterprise community, reducing the burden of data preparation while maintaining high accuracy.
Solution Approach 2:
The patent performs preliminary categorization actions by multiple independent enterprises and suppliers before final classification is needed. Each enterprise and supplier independently categorizes commodities they supply, creating a pre-prepared knowledge base that eliminates the need for extensive centralized training data preparation when the actual classification is performed.
3Ease of operation
If invoice description is given equal weight with supplier name in classification, then processing is simplified, but prediction accuracy deteriorates because description tends to dominate incorrectly
Solution Approach 1:
The patent applies local quality by giving different weights to different input sources based on their reliability and context. Supplier names receive appropriate weighting that reflects their importance in identifying the source of commodities, while invoice descriptions are weighted according to their specific context. This localized weighting approach prevents any single source from dominating and improves overall prediction accuracy.
Data Source
AI summary
Commodity category values can be determined automatically for suppliers in an e-procurement system using a computer-implemented process that is supplier-focused and uses successive heuristics, supplemented with machine learning models that predict category and subcategory values based on supplier names and invoice descriptions. Embodiments can support community intelligence applications to enable buyer computers to query and obtain lists of suppliers corresponding to categories and to generate graphs or charts that aggregate historic invoice data based on canonical category values that have been determined for suppliers.


