Training And Testing Data Generation Using SQL for Simpler ML Pipelines
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning libraries and tools are overly complex, requiring significant expertise and time to build and deploy machine learning models, making them inaccessible to many software developers and increasing the cost and time required for training and deployment.
Innovation Solution
A computer-implemented method and system that automates the generation of training and testing datasets using Structured Query Language (SQL) code, allowing for the creation of ingestible files for machine-learning pipelines, which can handle multiple data sources, perform custom splits, and generate datasets on-the-fly, reducing the need for human intervention and computational costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing machine learning libraries and tools are used, then powerful components spanning the entire machine learning workflow are provided, but the system becomes overly complex and requires high level of infrastructure sophistication and engineering resources
Solution Approach 1:
The patent segments the complex machine learning workflow into distinct, manageable components: data extraction module, data preprocessing module, model training module, and model deployment module. Each module handles a specific aspect of the workflow, allowing users to access powerful ML capabilities without being overwhelmed by the complexity of the entire system. The segmentation enables modular configuration and reduces the learning curve for users.
Solution Approach 2:
The patent introduces an intermediary layer (the system architecture described in claims 1-20) that mediates between the user and the complex machine learning libraries. This intermediary automatically handles data wrangling, pipeline configuration, and modeling decisions, translating high-level user requirements into detailed technical implementations. The intermediary shields users from complexity while still providing access to powerful ML components.
2Adaptability or versatility
If existing machine learning libraries and tools are used, then comprehensive machine learning components are available, but significant time is required to build and deploy models
Solution Approach 1:
The patent implements preliminary action by pre-configuring data extraction patterns, preprocessing pipelines, and model training templates. The system includes pre-built connectors for common data sources, pre-defined data transformation workflows, and pre-configured model training parameters. This preliminary preparation allows users to quickly deploy ML models without spending weeks on routine setup tasks, while still providing comprehensive ML components for complex scenarios.
Solution Approach 2:
The patent enables rapid model development through parameter changes by allowing users to adjust high-level configuration parameters (such as data source connections, model types, and training parameters) without modifying the underlying complex pipeline structure. The system automatically adapts the detailed implementation based on these parameter changes, enabling fast iteration and deployment while maintaining access to comprehensive ML components.
3Adaptability or versatility
If comprehensive data processing and model training is performed manually, then customized machine learning models can be created, but significant engineering resources and expertise are required
Solution Approach 1:
The patent implements self-service capabilities where the system automatically performs data wrangling, pipeline configuration, and modeling decisions based on user inputs. The system includes automated data quality assessment, automatic feature engineering suggestions, and self-configuring training pipelines. This allows users with limited expertise to create customized ML models by simply specifying their requirements, while the system handles the complex implementation details automatically.
Solution Approach 2:
The patent creates a universal platform that handles multiple ML tasks (data extraction, preprocessing, training, deployment) through a single unified interface. The system provides multi-functional components that can adapt to different data types, model types, and deployment scenarios without requiring users to master separate tools for each task. This universality maintains full customization capability while dramatically improving ease of operation for users with varying expertise levels.
Data Source
AI summary
Provided are computing systems, methods, and platforms for generating training and testing data for machine-learning models. The operations can include receiving signal extraction information that has instructions to query a data store. Additionally, the operations can include accessing, using Structured Query Language (SQL) code generated based on the signal extraction information, raw data from the data store. Moreover, the operations can include processing the raw data using signal configuration information to generate a plurality of signals. The signal configuration information can have instructions on how to generate the plurality of signals from the raw data. Furthermore, the operations can include joining, using SQL code, the plurality of signals with a first label source to generate training data and testing data. Subsequently, the operations can include processing the training data and the testing data to generate the input data. The input data being an ingestible for a machine-learning pipeline.


