Machine-Learning Service Configuration GUI for Diverse Data Preparation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine-learning (ML) model development processes are hindered by the significant time spent identifying, acquiring, and formatting data from diverse sources, leading to inefficiencies in data pipeline management.

Innovation Solution

A data pipeline tool provides a graphical user interface (GUI) for designing and configuring ML workflows, allowing users to drag-and-drop templates, connect data sources and models, and automate data preparation, security, and metadata management, facilitating efficient model development and deployment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If programmers manually identify, acquire, and format data from diverse sources, then data can be processed by ML models, but significant time is spent on data preparation tasks

Engineering Contradiction:
ImproveML model development speedVSAvoidTime spent on data identification, acquisition, and formatting
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system segments the complex data preparation process into distinct modular components including data source identification modules, data acquisition modules, data formatting modules, and ML model training modules. Each module handles a specific aspect of the workflow, allowing parallel processing and reducing overall preparation time while maintaining data quality standards.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by pre-identifying and pre-formatting data from multiple sources before ML model training begins. Data is collected, cleaned, and structured in advance using automated pipelines, so that when model training starts, the data is already ready for immediate processing, eliminating the time-consuming manual preparation step.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If multiple data sources with different formats, languages, and configurations are accessed, then comprehensive data can be gathered, but complexity of data management increases

Engineering Contradiction:
ImproveAbility to access diverse data sourcesVSAvoidComplexity of data pipeline configuration
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements a universal data pipeline framework that can handle multiple data sources with different formats, languages, and configurations through a single standardized interface. The framework includes adaptive translators and format converters that automatically adjust to various data sources, allowing comprehensive data gathering without increasing operational complexity for users.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces intermediary components including data translators, format converters, and standardized interface layers that mediate between diverse data sources and the ML model training process. These intermediaries handle the complexity of format conversion and data normalization, shielding users from the underlying complexity while maintaining versatility in data source access.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated data pipelines are implemented, then data preparation time is reduced, but initial setup and configuration complexity increases

Engineering Contradiction:
ImproveData preparation efficiencyVSAvoidEase of pipeline setup and configuration
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

The automated data pipeline system implements self-service capabilities where the pipeline automatically configures itself based on detected data sources and requirements. The system includes automatic data source discovery, self-configuration of data formats, and automated testing capabilities that reduce the need for manual setup while maintaining high data preparation efficiency. The pipeline adapts to new data sources autonomously without requiring complex reconfiguration.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12367422B2GUI for configuring machine-learning services
Publication Date: 2025.07.22 STATE FARM MUTAL AUTOMOBILE INSURANCE COMPANY
  • US12367422B2 patent drawing
  • US12367422B2 patent drawing
  • US12367422B2 patent drawing

AI summary

A data pipeline tool provides a machine-learning design interface that a user can utilize (e.g., via an electronic device such as a personal computer, tablet, or smart phone) to design or configure data pipelines or workflows defining the manner in which ML models are developed, trained, tested, validated, or deployed. Once deployed, a designed ML model may generate predictive results based on input data fed to the ML model. The tool may present the predictive results via a GUI, and may enable a user to mark-up or otherwise interact with those predictive results. The tool may enable the user to share the results (which may include a mark-up or annotation provided by a user).