Machine Learning Workflow Clustering via 3D Spatial Embedding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning workflow management systems struggle to accurately cluster and manage complex machine learning workflows due to cognitive overload and limitations in handling time-dependent and operation-dependent flows.

Innovation Solution

The system employs a method to cluster machine learning workflows by embedding the graph in a coordinate system, allowing for comparison based on properties such as operation type, adjacent operators, data processing, and infrastructure, thereby capturing essential workflow characteristics beyond mere graph properties.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If machine learning workflows are managed using traditional graph-based approaches, then the system structure is simple, but the cognitive overload increases and time-dependent relationships cannot be captured

Engineering Contradiction:
Improveworkflow clustering accuracyVSAvoidworkflow representation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent transforms the workflow representation from a traditional 2D graph structure to a 3D spatial embedding where nodes are positioned in three-dimensional space. This dimensional transformation enables the system to capture temporal relationships (through vertical positioning), operational dependencies (through horizontal positioning), and data flow connections (through spatial proximity), thereby resolving the contradiction between representation simplicity and cognitive load while maintaining clustering accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If workflows are clustered based on graph properties alone, then the clustering process is simple, but time-dependent and operation-dependent relationships are lost

Engineering Contradiction:
Improveworkflow clustering precisionVSAvoidclustering process complexity
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent changes the clustering parameters from traditional graph-based metrics (degree, connectivity) to spatial embedding coordinates that encode temporal and operational information. By transforming workflow nodes into 3D space where position reflects timing and operation dependencies, the system achieves precise clustering that captures time-dependent relationships without requiring complex multi-parameter analysis, thus improving measurement precision while maintaining ease of processing.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If detailed workflow properties are captured for accurate clustering, then clustering accuracy improves, but cognitive overload increases

Engineering Contradiction:
Improveworkflow matching accuracyVSAvoiduser cognitive load
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system performs self-service by automatically generating 3D spatial embeddings of workflows based on their inherent temporal and operational structures. The embedding process automatically extracts and encodes time-dependent and operation-dependent relationships without requiring manual annotation or complex user input. This automated self-embedding enables accurate workflow matching and clustering while keeping the user interface simple and cognitive load minimal, as the system handles the complex representation internally.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250190874A1System and method of suggesting machine learning workflows through machine learning
Publication Date: 2025.06.12 GEIGEL ARTURO
  • US20250190874A1 patent drawing
  • US20250190874A1 patent drawing
  • US20250190874A1 patent drawing

AI summary

A system and method of processing a machine learning flows by decomposing the flows on an x-y grid and extracting relevant information about their utilization on a particular category of machine learning workflow. This information is utilized to extract N-gram sequences that can be used as training for a machine learning algorithm that will suggest to the user which operator to put in a new machine learning workflow.