ML Application Orchestration Service for Large-Scale Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The complexity of machine learning algorithms and the difficulty in generating, updating, and deploying machine learning models make it time and resource consumptive, and these models are often specific to particular use cases and environments, requiring complete regeneration with any changes, which is different from traditional software engineering practices.

Innovation Solution

A framework for building, orchestrating, and deploying complex machine learning inference applications through a ML application orchestration service that enables users to define models, perform data transformations, and coordinate workflow logic, automatically scaling computing resources based on traffic patterns and optimizing performance metrics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If machine learning models are built using traditional academic research approaches, then model accuracy and specialization are improved, but development time and resource consumption increase significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoiddevelopment time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the machine learning development process into distinct modular components: data preprocessing modules, model training modules, evaluation modules, and deployment modules. Each module can be independently developed, tested, and reused across different projects, significantly reducing development time while maintaining model accuracy through systematic organization of complex workflows

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary action through pre-built template workflows, pre-trained model components, and standardized data preprocessing pipelines that are prepared in advance. These pre-configured elements can be directly applied to new projects, eliminating the need to build everything from scratch and drastically reducing development time without compromising model performance

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If machine learning models are highly specialized for particular use cases, then model performance for specific tasks is improved, but adaptability to environment changes deteriorates

Engineering Contradiction:
Improvemodel performanceVSAvoidenvironment adaptability
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates universal model architectures and workflow templates that can serve multiple use cases. The modular design allows the same base model structure to be adapted for different tasks by swapping specific processing modules or adjusting parameters, enabling one model framework to handle diverse applications while maintaining high performance for each specific use case

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements dynamic adaptability through configurable model parameters and flexible workflow orchestration that can adjust to changing environmental conditions. The system allows runtime configuration changes and model retraining triggers based on performance monitoring, enabling specialized models to adapt to new conditions without complete regeneration

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If machine learning model deployment follows academic research practices, then research flexibility is improved, but deployment efficiency and reliability deteriorate

Engineering Contradiction:
Improveresearch flexibilityVSAvoiddeployment efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent introduces an intermediary layer between research and deployment: a standardized workflow orchestration system that translates flexible research prototypes into reliable production deployments. This intermediary layer handles the conversion of experimental models into deployed services, managing concerns like scalability, monitoring, and resource management while preserving research flexibility through configurable parameters and modular components

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If complete model regeneration is performed for any environment change, then model relevance to current conditions is improved, but resource consumption and time loss increase

Engineering Contradiction:
Improvemodel relevanceVSAvoidregeneration time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements partial regeneration through incremental model updates and selective retraining of only the affected model components rather than complete regeneration. When environmental conditions change, the system identifies and retrains only the specific modules or parameters that need adjustment, maintaining model relevance while significantly reducing the time and resources required compared to full regeneration

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11373119B1Framework for building, orchestrating and deploying large-scale machine learning applications
Publication Date: 2022.06.28 AMAZON TECH INC
  • US11373119B1 patent drawing
  • US11373119B1 patent drawing
  • US11373119B1 patent drawing

AI summary

Techniques for a framework for building, orchestrating, and deploying complex, large-scale Machine Learning (ML) or deep learning (DL) inference applications is described. A ML application orchestration service is disclosed that enables the construction, orchestration, and deployment of complex ML inference applications in a provider network. The disclosed service provides customers with the ability to define machine learning (ML) models and define transformation operations on data before and/or after being provided to the ML models to construct a complex ML inference application. The service provides a framework for the orchestration (co-ordination) of the workflow logic (e.g., of the request and/or response flows) involved in building and deploying a complex ML inference application in the provider network.