Data Processing Model for Document Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data processing systems are inefficient and prone to errors when handling large volumes of business data, requiring significant manual effort and time to identify anomalies, leading to sub-optimal auditing processes.
Innovation Solution
A data processing system utilizing a processor configured to obtain labelled training data, process it using machine learning algorithms to generate a data processing model, and compare extracted key information from test documents with reference information to identify inconsistencies, reducing manual effort and training time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual evaluation of data elements is performed, then accuracy in identifying anomalies can be maintained through human judgment, but the processing time and effort increase significantly
Solution Approach 1:
The system performs preliminary actions by pre-processing training data and training the machine learning model in advance. The model is trained on labelled training data to learn patterns of anomalies, so that during actual operation, it can quickly process test documents without requiring manual evaluation, thus reducing processing time while maintaining accuracy through pre-learned knowledge
Solution Approach 2:
The patent replaces the mechanical system of manual human evaluation with an automated machine learning-based system. The processor automatically processes test documents, extracts key information, and identifies anomalies by comparing with reference information, eliminating the need for manual human judgment while maintaining consistent accuracy across large volumes of data
2Productivity
If automated systems are used to process data, then processing speed and efficiency are improved, but significant manual effort is still required for preprocessing training data
Solution Approach 1:
The system implements self-service by automatically performing the preprocessing of training data. The processor reads training documents, extracts key information automatically, and prepares the labelled training data without requiring manual intervention. This eliminates the manual preprocessing step while maintaining high processing speed and efficiency
3Loss of time
If only a limited number of data elements are evaluated manually, then processing time is reduced, but the analysis becomes prone to errors and inaccuracies
Solution Approach 1:
The patent replaces manual sampling and evaluation with an automated machine learning system that processes all test documents systematically. The processor evaluates every test document by extracting key information and comparing it with reference information stored in the database, eliminating sampling errors and ensuring comprehensive coverage without increasing processing time
4Measurement precision
If comprehensive manual auditing of all financial data is performed, then accuracy and completeness of audit evidence are improved, but the time and resources required become prohibitive
Solution Approach 1:
The patent replaces comprehensive manual auditing with an automated system that processes all financial documents using the trained machine learning model. The processor extracts key information from each document and compares it with reference information in the database, providing comprehensive audit coverage with minimal human resources while maintaining high accuracy through systematic automated evaluation
Data Source
AI summary
Disclosed is a data processing system comprising a database arrangement and at least one processor coupled in communication with the database arrangement. The at least one processor is configured to obtain, from the database arrangement, labelled training data; process 5 the labelled training data using at least one machine learning algorithm to generate a data processing model; obtain, from the database arrangement or a user device, at least one test document; process the at least one test document using the data processing model to extract key information from the at least one test document; and compare the 10 extracted key information with reference key information to identify whether or not there exist inconsistencies between the extracted key information and the reference key information.


