Spend Data Enrichment via Web Mining for Procurement Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data classification systems for procurement spend data face challenges due to high volume and limited information in transaction records, leading to inaccurate classifications and difficulties in categorizing spend data without complete vendor information.

Innovation Solution

A method and system utilizing artificial intelligence and web mining tools to cleanse and enrich spend data, applying classification models with confidence scoring, and leveraging historical data or web information to classify spend data, even when vendor information is incomplete.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional data classification methods are used on spend data, then the process is simple, but the classification accuracy is low due to limited information in transaction records

Engineering Contradiction:
Improveclassification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary data enrichment by web mining vendor information before classification. It queries external sources to obtain vendor descriptions, addresses, and other contextual data, then integrates this enriched information into the transaction records. This preliminary action ensures that when classification occurs, the data contains sufficient information for accurate categorization, directly addressing the limitation of sparse transaction data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary web mining layer between the raw spend data and the classification algorithm. This intermediary component acts as a bridge that transforms incomplete transaction records into enriched datasets by extracting vendor information from external web sources. The intermediary processing enables accurate classification without requiring changes to the core classification algorithm, thus managing complexity while improving accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If vendor information is removed or made redundant due to incompleteness, then data consistency is maintained, but classification accuracy deteriorates

Engineering Contradiction:
Improvedata consistencyVSAvoidclassification accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The system implements self-service by automatically detecting missing vendor information and autonomously querying external web sources to retrieve it. Rather than requiring manual data entry or deletion, the system independently enriches its own dataset by mining vendor profiles, descriptions, and contextual information from public sources. This self-service approach maintains data consistency through automated validation while ensuring classification accuracy through comprehensive information gathering.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Before classification occurs, the system performs preliminary data completion by web mining vendor information. It identifies records with missing vendor data, queries external sources to retrieve vendor profiles and descriptions, and integrates this enriched information into the dataset. This preliminary enrichment ensures that all records have sufficient information for accurate classification while maintaining consistency through structured data integration.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If manual data enrichment is performed to obtain unknown vendor information, then classification accuracy improves, but processing time increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical data gathering with automated web mining technology. Instead of人工 querying vendor information, the system uses automated bots and web crawlers to extract vendor profiles, descriptions, and contextual data from external sources. This substitution of automated digital processes for manual mechanical operations dramatically reduces processing time while maintaining comprehensive data enrichment, enabling accurate classification at scale.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The web mining component serves as an automated intermediary that efficiently bridges the gap between raw spend data and vendor information. It implements intelligent query generation, parallel processing of multiple vendor lookups, and automated data integration, replacing time-consuming manual research with streamlined automated processes that maintain accuracy while reducing processing time.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Loss of information

If all spend data is processed through web mining to obtain unknown terms, then data enrichment is comprehensive, but computational resources are excessively consumed

Engineering Contradiction:
Improveinformation completenessVSAvoidcomputational resource consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The system applies partial web mining by selectively enriching only those records that contain unknown or missing vendor information. Rather than performing comprehensive web mining on all spend data, it identifies records needing enrichment through preliminary analysis and applies web mining only to those cases. This partial action approach ensures information completeness for records that need it while avoiding unnecessary computational resource consumption on already-complete records.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent implements local quality by applying different processing strategies to different subsets of data based on their specific needs. Records with complete vendor information undergo standard classification, while records with missing information receive targeted web mining enrichment. This localized approach ensures that computational resources are concentrated where they are most needed, achieving information completeness for critical records while minimizing overall resource consumption.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10438133B2Spend data enrichment and classification
Publication Date: 2019.10.08 GLOBAL EPROCURE
  • US10438133B2 patent drawing
  • US10438133B2 patent drawing
  • US10438133B2 patent drawing

AI summary

The present invention discloses a method, a system and a computer program product for spend data classification using selected taxonomies. The invention provides refresh classification tool and implementation classification tools for spend data classification. The invention further provides a web mining tool for determining unknown terms in spend data to obtain an enriched spend data for classification.