Modular Training Data Storage for Low-Latency Model Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing model generation systems face inefficiencies in processing heterogeneous data from multiple sources, requiring manual input and imputation, leading to increased data storage and transmission requirements, and lack scalability and effective data visualization.

Innovation Solution

A model generation platform that intelligently stores and processes data from various sources and formats based on performance attributes, using a graphical user interface for data selection and visualization, enabling efficient storage in suitable memory types like RAM for parallel processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If data is stored in traditional storage systems, then data capacity is sufficient, but data accessibility and latency are poor

Engineering Contradiction:
Improvedata accessibilityVSAvoiddata storage capacity
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The patent segments data storage into multiple tiers based on performance attributes. Frequently accessed training data is stored in high-speed memory (RAM) while less frequently accessed data is stored in lower-cost storage systems. This segmentation allows the system to optimize for both speed and capacity by placing only the necessary subset of data in high-performance storage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by assigning different storage locations to different portions of data based on their specific access patterns and performance requirements. Training data that requires fast access is stored in RAM with high accessibility, while archival data is stored in slower but larger-capacity storage systems. Each data portion receives storage quality matched to its specific needs.

Inventive Principle:
Principle #3Local quality

2Productivity

If manual data input and imputation is performed, then data processing accuracy can be maintained, but processing time and labor costs increase

Engineering Contradiction:
Improvedata processing efficiencyVSAvoiddata processing accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent implements self-service by enabling the system to automatically select, retrieve, and prepare training data from stored representations without requiring manual input or imputation. The data pipeline automatically queries the storage system for required training data based on model generation requests, eliminating manual data processing steps while maintaining accuracy through systematic automated retrieval.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If data is stored in a centralized manner, then data management is simplified, but scalability and multi-user access are limited

Engineering Contradiction:
Improvesystem scalabilityVSAvoiddata management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal data storage and retrieval system that serves multiple functions: storing training data for different machine learning models, supporting multiple users simultaneously, providing data for various data pipeline configurations, and enabling both fast access and long-term retention. This multi-functional system achieves scalability without proportionally increasing management complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Manufacturing precision

If large amounts of training data are stored, then model generation accuracy can be improved, but storage costs and data transmission requirements increase

Engineering Contradiction:
Improvemodel generation accuracyVSAvoiddata storage volume
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by pre-storing representations of training data in an optimized format and location before they are needed for model generation. When a model generation request arrives, the system can quickly retrieve the pre-prepared training data representations from storage without requiring extensive real-time processing or transmission, thus providing accurate model generation with reduced storage and transmission overhead.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12608405B2Model generation based on modular storage of training data
Publication Date: 2026.04.21 CITIBANK N A
  • US12608405B2 patent drawing
  • US12608405B2 patent drawing
  • US12608405B2 patent drawing

AI summary

The model generation platform enables generation of machine learning models based on modular storage of user-specified training data. The platform can obtain a dataset from a different heterogenous source (e.g., a structured database, an unstructured database, a semi-structured file system, manual upload of a comma-separated value file, a spreadsheet, and/or big data) through an associated application programming interface and store this data in a first storage medium. The model generation platform can obtain an indication of a portion of the dataset from a user via a user interface and determine a second storage medium for this portion of the dataset based on an associated estimated performance metric. In response to a request for generation of a machine learning model, the model generation platform can generate a machine learning model using training data comprising a subset of the portion of the dataset.