ML Ensemble Data Digitization for Capital Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Heterogeneous computing systems face challenges in integrating a computing system with a centralized processing infrastructure due to large volumes of data and excessive data transformations, read/write database calls, or generating erroneous computing actions.

Innovation Solution

The use of custom integrated machine learning ensembles to digitize data by identifying, extracting, and mapping data sets from one type to another, thereby transforming data into a format compatible with electronic transaction systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional data formatting methods are used to transform heterogeneous data sets, then data can be converted to a standardized type, but computational resources are consumed excessively and error rates increase

Engineering Contradiction:
Improvedata transformation accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent replaces traditional mechanical data transformation methods (manual formatting, rule-based conversion) with machine learning-based automated processing. The system uses trained ML models to automatically identify, extract, and map data from heterogeneous sources to standardized formats, eliminating the need for extensive manual intervention and reducing computational overhead while improving accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent implements self-service through automated data transformation pipelines where the system autonomously processes heterogeneous data sets without requiring manual intervention. The machine learning models automatically adapt to different data formats, perform extraction and mapping operations, and generate standardized outputs, enabling the system to serve itself in processing diverse data sources efficiently.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If multiple data sets from various sources are formatted manually, then they can be integrated into a unified type, but the process is tedious and error-prone

Engineering Contradiction:
Improvedata format compatibilityVSAvoiddata formatting complexity
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent implements a universal data transformation system that can handle multiple data formats, sources, and types through a single integrated platform. The machine learning models are trained to recognize and process various data structures (JSON, XML, CSV, databases, APIs) and automatically adapt to different schemas, enabling one system to serve multiple data integration purposes without requiring separate processing pipelines for each format.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces machine learning models as intermediary components between heterogeneous data sources and the target standardized format. These ML models act as intelligent mediators that automatically translate between different data structures, schemas, and formats, eliminating the need for manual mapping rules and reducing errors in data transformation while maintaining high adaptability to various source formats.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If extensive data transformations and database calls are performed to integrate heterogeneous systems, then data can be processed, but onboarding time increases significantly

Engineering Contradiction:
Improvedata integration efficiencyVSAvoidsystem onboarding time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-training machine learning models on diverse data formats and schemas before deployment. The models are prepared in advance with knowledge of various data structures, extraction patterns, and mapping rules, enabling them to quickly process new heterogeneous data sources without requiring extensive runtime transformations or multiple database calls, thereby reducing onboarding time while maintaining high productivity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250029417A1Data digitization via custom integrated machine learning ensembles
Publication Date: 2025.01.23 ADP INC
  • US20250029417A1 patent drawing
  • US20250029417A1 patent drawing
  • US20250029417A1 patent drawing

AI summary

The present disclosure relates generally to the digitization of documents and more particularly, to a system, method and computer program which integrates multiple trained machine learning ensembles to identify, extract, and map a data set. The method, for example, includes receiving a data set from sources; identifying ensembles, each ensemble comprising machine learning models and each ensemble to determine an outcome; identifying a type for the data set based on a vendor type and the data set; executing a section detection module to identify sections of the data set and classify the sections; executing a page classification module; generating associations between the sections and the classifications; transforming, based on the association, the sections, the classifications, and the type, the data set into a second file type; and presenting the transformed data set for integration into a capital management system.