AI Workload Partitioning for Edge Resource Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional approaches to managing AI workloads on edge clusters with heterogeneous compute capacity are inefficient, leading to suboptimal performance, resource utilization, and increased latency due to static workload distribution and lack of appreciation for unique AI workload characteristics.

Innovation Solution

An AI workload partitioning system that analyzes input streams and AI models to characterize compute and memory resources, dynamically partitions the workload, and distributes it across heterogeneous devices, optimizing execution on specialized hardware units to enhance performance and resource utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If static workload distribution is used on heterogeneous edge devices, then device complexity is reduced and ease of operation is improved, but productivity deteriorates and loss of time increases due to inefficient execution and longer latency

Engineering Contradiction:
Improveworkload management simplicityVSAvoidAI workload execution efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent implements dynamic workload partitioning and distribution that adapts to changing conditions. The system continuously monitors workload characteristics, device status, and performance metrics to dynamically adjust the partitioning strategy and task allocation across heterogeneous edge devices, transforming the static system into a dynamic one that optimizes productivity while maintaining operational simplicity through automated adaptation

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes key parameters including workload partitioning granularity, device selection criteria, and resource allocation ratios based on real-time analysis of AI workload characteristics and edge device capabilities. By dynamically adjusting these parameters, the system resolves the contradiction between operational simplicity and execution efficiency

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If uniform workload distribution approach is applied across all AI workloads, then device complexity is reduced and ease of manufacture is improved, but productivity deteriorates due to inability to meet performance and resource utilization standards for diverse AI workloads

Engineering Contradiction:
Improvesystem implementation simplicityVSAvoidAI workload performance
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent applies local quality by analyzing unique characteristics of each AI workload (such as compute requirements, memory needs, and data characteristics) and applying specialized partitioning and distribution strategies tailored to each workload type. This allows the system to optimize performance for diverse AI workloads while maintaining a unified framework that preserves ease of implementation through standardized analysis and adaptation mechanisms

Inventive Principle:
Principle #3Local quality

3Device complexity

If static partitioning of AI workloads is used, then device complexity is reduced and ease of operation is improved, but loss of time increases due to longer latency in workload execution

Engineering Contradiction:
Improveworkload partitioning complexityVSAvoidworkload execution latency
Core Design Contradiction:
Device complexityVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of AI workload characteristics and pre-computes optimal partitioning strategies before actual execution. By analyzing workload patterns, device capabilities, and performance requirements in advance, the system prepares optimized task distributions that minimize execution latency while avoiding the complexity of real-time decision-making during runtime

Inventive Principle:
Principle #10Preliminary action

4Ease of operation

If conventional workload management approaches are used, then ease of operation is maintained, but resource utilization deteriorates leading to inefficient execution and increased power consumption

Engineering Contradiction:
Improveworkload management simplicityVSAvoidpower consumption
Core Design Contradiction:
Ease of operationVSLoss of energy

Solution Approach 1:

The patent implements feedback mechanisms that continuously monitor resource utilization metrics, execution performance, and power consumption across edge devices. This feedback is used to dynamically adjust workload partitioning and distribution decisions, optimizing energy efficiency by directing workloads to devices with optimal resource availability and performance characteristics, thereby reducing overall power consumption while maintaining operational simplicity through automated control

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12106154B2Serverless computing architecture for artificial intelligence workloads on edge for dynamic reconfiguration of workloads and enhanced resource utilization
Publication Date: 2024.10.01 INTEL CORP
  • US12106154B2 patent drawing
  • US12106154B2 patent drawing
  • US12106154B2 patent drawing

AI summary

Systems, apparatuses and methods include technology that analyzes an input stream and an artificial intelligence (AI) model graph to generate a workload characterization. The workload characterization characterizes one or more of compute resources or memory resources, and the one or more of the compute resources or the memory resources is associated with execution of the AI model graph based on the input stream. The technology partitions the AI model graph into subgraphs based on the workload characterization. The technology selects a plurality of hardware devices to execute the subgraphs.