Data Normalization via Fuzzy Comparison for Financial Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The complexity of modern financial allocation models in businesses makes it difficult to accurately determine total cost of ownership and generate reliable reporting, especially for larger enterprises, due to the sheer number of items and entities that need to be modeled, leading to challenges in data integrity and automated data entry processes.
Innovation Solution
The implementation of a system that normalizes ingested data sets based on fuzzy comparisons to known data sets, using an ingestion engine that applies ingestion rules to transform raw data into model records, provides a confidence score, and allows for interactive modification of model records, ensuring data accuracy and integrity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated data entry processes are used to populate financial models, then productivity is improved, but data integrity deteriorates
Solution Approach 1:
The system performs self-correction by automatically comparing ingested data against known datasets and applying normalization rules without human intervention. The ingestion engine autonomously identifies and corrects data quality issues, maintaining both high productivity and data integrity through automated self-service mechanisms.
Solution Approach 2:
The system implements feedback loops where ingested data is continuously validated against known datasets and normalization rules. Confidence scores are calculated and used to determine whether manual review is needed, creating a feedback mechanism that maintains data integrity while preserving automated processing efficiency.
2Measurement precision
If the number of tracked activities and elements increases to improve budgeting accuracy, then measurement precision is improved, but device complexity worsens
Solution Approach 1:
The system segments the complex data normalization task into distinct components: ingestion rules, known datasets, confidence score calculation, and manual review triggers. This segmentation allows the financial model to handle numerous tracked elements without proportionally increasing overall system complexity.
Solution Approach 2:
The system introduces an intermediary layer (the ingestion engine with normalization rules) between raw data input and the financial model. This intermediary automatically standardizes data before it enters the model, enabling accurate tracking of numerous elements without requiring proportional increases in model complexity.
3Reliability
If fuzzy comparison normalization is applied to ingested data, then data integrity is improved, but processing time worsens
Solution Approach 1:
The system applies partial normalization by calculating confidence scores and only flagging records that fall below a threshold for manual review. Most records are processed automatically without full manual verification, achieving adequate data integrity while minimizing processing time loss.
Solution Approach 2:
The system changes the parameter of data validation from binary (pass/fail) to a continuous confidence score. This allows flexible threshold adjustment where high-confidence records are processed quickly automatically, while only low-confidence records require additional processing time for manual review.
Data Source
AI summary
Embodiments are directed towards normalizing ingested data sets based on fuzzy comparisons to known data sets. Raw data sets that each include raw records may be provided to an ingestion engine. Ingestion rules and known data sets may be provided based on the raw records. The ingestion engine may be employed to iteratively execute the ingestion rules. A comparison of the raw records to the known data sets may be performed. Contents of the raw records may be transformed into model record values and stored in model records. A score value that indicates a confidence level that the model records are correct may be provided. An association of the one or more ingestion rules used to transform the raw record contents into the model record values for each of the one or more model records may be added to a data model.


