Spend Data Enrichment via Web Mining for Procurement Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data classification systems for procurement spend data face challenges due to high volume and limited information in transaction records, leading to inaccurate classifications and difficulties in categorizing spend data without complete vendor information.
Innovation Solution
A method and system utilizing artificial intelligence and web mining tools to cleanse and enrich spend data, applying classification models with confidence scoring, and leveraging historical data or web information to classify spend data, even when vendor information is incomplete.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional data classification methods are used on spend data, then the process is simple, but the classification accuracy is low due to limited information in transaction records
Solution Approach 1:
The system performs preliminary data enrichment by web mining vendor information before classification. It queries external sources to obtain vendor descriptions, addresses, and other contextual data, then integrates this enriched information into the transaction records. This preliminary action ensures that when classification occurs, the data contains sufficient information for accurate categorization, directly addressing the limitation of sparse transaction data.
Solution Approach 2:
The patent introduces an intermediary web mining layer between the raw spend data and the classification algorithm. This intermediary component acts as a bridge that transforms incomplete transaction records into enriched datasets by extracting vendor information from external web sources. The intermediary processing enables accurate classification without requiring changes to the core classification algorithm, thus managing complexity while improving accuracy.
2Reliability
If vendor information is removed or made redundant due to incompleteness, then data consistency is maintained, but classification accuracy deteriorates
Solution Approach 1:
The system implements self-service by automatically detecting missing vendor information and autonomously querying external web sources to retrieve it. Rather than requiring manual data entry or deletion, the system independently enriches its own dataset by mining vendor profiles, descriptions, and contextual information from public sources. This self-service approach maintains data consistency through automated validation while ensuring classification accuracy through comprehensive information gathering.
Solution Approach 2:
Before classification occurs, the system performs preliminary data completion by web mining vendor information. It identifies records with missing vendor data, queries external sources to retrieve vendor profiles and descriptions, and integrates this enriched information into the dataset. This preliminary enrichment ensures that all records have sufficient information for accurate classification while maintaining consistency through structured data integration.
3Measurement precision
If manual data enrichment is performed to obtain unknown vendor information, then classification accuracy improves, but processing time increases
Solution Approach 1:
The patent replaces manual mechanical data gathering with automated web mining technology. Instead of人工 querying vendor information, the system uses automated bots and web crawlers to extract vendor profiles, descriptions, and contextual data from external sources. This substitution of automated digital processes for manual mechanical operations dramatically reduces processing time while maintaining comprehensive data enrichment, enabling accurate classification at scale.
Solution Approach 2:
The web mining component serves as an automated intermediary that efficiently bridges the gap between raw spend data and vendor information. It implements intelligent query generation, parallel processing of multiple vendor lookups, and automated data integration, replacing time-consuming manual research with streamlined automated processes that maintain accuracy while reducing processing time.
4Loss of information
If all spend data is processed through web mining to obtain unknown terms, then data enrichment is comprehensive, but computational resources are excessively consumed
Solution Approach 1:
The system applies partial web mining by selectively enriching only those records that contain unknown or missing vendor information. Rather than performing comprehensive web mining on all spend data, it identifies records needing enrichment through preliminary analysis and applies web mining only to those cases. This partial action approach ensures information completeness for records that need it while avoiding unnecessary computational resource consumption on already-complete records.
Solution Approach 2:
The patent implements local quality by applying different processing strategies to different subsets of data based on their specific needs. Records with complete vendor information undergo standard classification, while records with missing information receive targeted web mining enrichment. This localized approach ensures that computational resources are concentrated where they are most needed, achieving information completeness for critical records while minimizing overall resource consumption.
Data Source
AI summary
The present invention discloses a method, a system and a computer program product for spend data classification using selected taxonomies. The invention provides refresh classification tool and implementation classification tools for spend data classification. The invention further provides a web mining tool for determining unknown terms in spend data to obtain an enriched spend data for classification.


