ML Ensemble Data Digitization for Capital Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Heterogeneous computing systems face challenges in integrating a computing system with a centralized processing infrastructure due to large volumes of data and excessive data transformations, read/write database calls, or generating erroneous computing actions.
Innovation Solution
The use of custom integrated machine learning ensembles to digitize data by identifying, extracting, and mapping data sets from one type to another, thereby transforming data into a format compatible with electronic transaction systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional data formatting methods are used to transform heterogeneous data sets, then data can be converted to a standardized type, but computational resources are consumed excessively and error rates increase
Solution Approach 1:
The patent replaces traditional mechanical data transformation methods (manual formatting, rule-based conversion) with machine learning-based automated processing. The system uses trained ML models to automatically identify, extract, and map data from heterogeneous sources to standardized formats, eliminating the need for extensive manual intervention and reducing computational overhead while improving accuracy.
Solution Approach 2:
The patent implements self-service through automated data transformation pipelines where the system autonomously processes heterogeneous data sets without requiring manual intervention. The machine learning models automatically adapt to different data formats, perform extraction and mapping operations, and generate standardized outputs, enabling the system to serve itself in processing diverse data sources efficiently.
2Adaptability or versatility
If multiple data sets from various sources are formatted manually, then they can be integrated into a unified type, but the process is tedious and error-prone
Solution Approach 1:
The patent implements a universal data transformation system that can handle multiple data formats, sources, and types through a single integrated platform. The machine learning models are trained to recognize and process various data structures (JSON, XML, CSV, databases, APIs) and automatically adapt to different schemas, enabling one system to serve multiple data integration purposes without requiring separate processing pipelines for each format.
Solution Approach 2:
The patent introduces machine learning models as intermediary components between heterogeneous data sources and the target standardized format. These ML models act as intelligent mediators that automatically translate between different data structures, schemas, and formats, eliminating the need for manual mapping rules and reducing errors in data transformation while maintaining high adaptability to various source formats.
3Productivity
If extensive data transformations and database calls are performed to integrate heterogeneous systems, then data can be processed, but onboarding time increases significantly
Solution Approach 1:
The patent applies preliminary action by pre-training machine learning models on diverse data formats and schemas before deployment. The models are prepared in advance with knowledge of various data structures, extraction patterns, and mapping rules, enabling them to quickly process new heterogeneous data sources without requiring extensive runtime transformations or multiple database calls, thereby reducing onboarding time while maintaining high productivity.
Data Source
AI summary
The present disclosure relates generally to the digitization of documents and more particularly, to a system, method and computer program which integrates multiple trained machine learning ensembles to identify, extract, and map a data set. The method, for example, includes receiving a data set from sources; identifying ensembles, each ensemble comprising machine learning models and each ensemble to determine an outcome; identifying a type for the data set based on a vendor type and the data set; executing a section detection module to identify sections of the data set and classify the sections; executing a page classification module; generating associations between the sections and the classifications; transforming, based on the association, the sections, the classifications, and the type, the data set into a second file type; and presenting the transformed data set for integration into a capital management system.


