AI Spend Data Classification with Blockchain Support
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data classification systems face challenges in accurately processing and classifying large volumes of spend data due to issues like missing vendor names, inconsistent data entry, redundant transactions, and the difficulty in handling non-text data formats, especially in blockchain networks, leading to inaccurate results and inefficiencies in enterprise and supply chain management applications.
Innovation Solution
A method and system that utilize AI engines and machine learning algorithms for data cleansing, enrichment, and classification, involving stratified sampling, transfer learning, and dynamic processing logic to generate reference data and confidence scores, while also addressing the complexities of blockchain networks through AI-based processing nodes and dynamic data models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional data classification methods are used on large volumes of spend data, then processing speed is maintained, but classification accuracy deteriorates due to missing vendor names, inconsistent data entry, and redundant transactions
Solution Approach 1:
The patent segments the large volume of spend data into smaller, manageable batches for processing. The system divides incoming data streams into discrete transaction groups, processes each batch through the classification engine independently, and aggregates results. This segmentation enables accurate processing of each data point while maintaining overall processing throughput by parallelizing batch operations.
Solution Approach 2:
The patent applies preliminary data cleansing and enrichment actions before classification. The system performs vendor name standardization, removes redundant transactions, fills missing data fields, and validates data consistency prior to the classification process. This preliminary processing ensures that the classification engine receives high-quality data, improving accuracy without requiring re-processing of erroneous records.
2Reliability
If manual data verification is performed to ensure accuracy, then classification reliability improves, but processing time increases significantly
Solution Approach 1:
The patent implements self-service verification mechanisms where the classification system automatically validates its own outputs. The engine incorporates confidence scoring that automatically verifies classification results against multiple data sources and cross-references vendor information, transaction patterns, and spend categories. Records with low confidence scores are automatically flagged for review, while high-confidence records are processed through, eliminating the need for manual verification of all transactions.
3Measurement precision
If comprehensive data validation is applied to all data sources, then data quality improves, but system complexity increases due to multiple data sources and formats
Solution Approach 1:
The patent implements a universal data validation framework that handles multiple data sources and formats through a single standardized interface. The system employs a unified validation engine that can process purchase orders, invoices, payment transactions, and blockchain data using the same validation rules and protocols. This multi-functional approach ensures consistent data quality across all sources without requiring separate validation systems for each data type.
Solution Approach 2:
The patent dynamically adjusts validation parameters based on data source characteristics and transaction types. The system modifies validation thresholds, required data fields, and verification depth according to the specific data being processed. For example, blockchain transactions may require different validation parameters compared to traditional ERP data, allowing the system to maintain high data quality while adapting to varying complexity requirements of different sources.
4Measurement precision
If AI-based annotation is performed on all data subsets, then reference data accuracy improves, but computational resources and time consumption increase
Solution Approach 1:
The patent applies AI-based annotation selectively to only those data subsets that require it. The system identifies data records with ambiguous characteristics, missing vendor information, or unusual transaction patterns that would benefit from AI annotation. High-confidence, routine transactions are processed through standard classification rules without AI intervention, while only problematic or novel records undergo computationally intensive AI annotation, optimizing resource utilization while maintaining reference data accuracy where needed.
Data Source
AI summary
The present invention discloses a method, a system and a computer program product for data processing and classification. The invention provides warm start and cold start classification tools for classification of data obtained from known or unknown entities. The system and method are also configured to be employed over blockchain based networks.


