Synthetic Data Training for Duplicate Invoice Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Accounts payable teams face difficulties in handling duplicate invoices due to OCR errors, late payments, and vendor practices, leading to potential financial losses and high manual review costs, as existing solutions struggle to efficiently detect and eliminate duplicates among large volumes of invoices.
Innovation Solution
A machine learning system is trained using synthetic data to generate candidate invoice pairs and predict duplicate invoices, reducing the need for labeled data and enabling automated detection of duplicates through a gradient boosting model and blocking strategy to efficiently identify potential duplicates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review methods are used to detect duplicate invoices, then detection accuracy can be maintained, but the cost and time required increases significantly
Solution Approach 1:
The patent replaces manual mechanical review processes with an automated machine learning system that uses synthetic data training to detect duplicate invoices. The system substitutes human analysts with algorithms that can process invoices automatically, maintaining high detection accuracy while eliminating the time cost of manual review.
Solution Approach 2:
The system enables self-service duplicate detection by training machine learning models to autonomously identify duplicate invoices without requiring human intervention for each review. The synthetic data generation and model training create a self-sufficient system that continuously improves its detection capabilities while operating independently.
2Productivity
If traditional duplicate detection methods are used, then implementation is simple, but the system cannot handle large volumes of invoices efficiently
Solution Approach 1:
The patent applies preliminary action by generating synthetic training data and training machine learning models before actual duplicate detection is needed. This preparatory phase creates a ready-to-use detection system that can immediately handle large invoice volumes without requiring complex real-time processing logic during operation.
Solution Approach 2:
The system changes parameters by transforming the detection approach from rule-based threshold comparisons to machine learning probability predictions. This parameter transformation enables the system to handle large volumes efficiently by leveraging the pattern recognition capabilities of trained models rather than complex real-time computational logic.
3Measurement precision
If synthetic training data is generated and machine learning models are trained, then duplicate detection accuracy improves, but data preparation and model training time increases
Solution Approach 1:
The patent performs preliminary action by generating synthetic training data and training machine learning models in advance before production use. This upfront investment in data preparation and model training creates a optimized detection system that delivers high accuracy predictions without requiring extensive training time during actual operation.
Data Source
AI summary
Embodiments detect duplicate invoices, each invoice including a plurality of fields. Embodiments generate synthetic training data using a plurality of training invoices and generating one or more modified fields for each of the plurality of training invoices. Embodiments train a machine learning model using the synthetic training data and generate a plurality of candidate invoice pairs. Embodiments input the plurality of candidate invoice pairs to the trained machine learning model and generate, by the trained machine learning model, a prediction of whether each of the candidate invoices pairs is a duplicate invoice pair.


