Managed AI Infrastructure Service for Distributed Workloads
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current general-purpose cloud-based infrastructure services are inadequate for handling the exponential growth of artificial intelligence (AI) workloads, such as Deep Learning Training (DLT) jobs, due to their workload-agnostic design, leading to inefficiencies and limitations in scalability and resource utilization.
Innovation Solution
A fully managed, globally distributed, multi-tenant AI infrastructure service that integrates diverse infrastructure resources via native support interfaces, allowing for efficient scheduling and execution of AI workloads, including training and inferencing tasks, using a hierarchical scheduling system that optimizes resource allocation and utilization across a distributed pool of hardware accelerators, networking, and storage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If general-purpose cloud-based infrastructure services are used to handle AI workloads, then existing infrastructure can be leveraged with minimal additional investment, but resource utilization efficiency deteriorates and scalability is limited due to workload-agnostic design
Solution Approach 1:
The patent creates a purpose-built AI infrastructure service that specializes in handling AI workloads (training and inferencing) while maintaining universal compatibility across different hardware accelerators, networking configurations, and storage systems. This specialized universal design resolves the contradiction by optimizing resource utilization for AI-specific patterns while maintaining broad adaptability to various AI workload types and scales.
Solution Approach 2:
The system dynamically adjusts resource allocation parameters based on workload characteristics, changing allocation granularity, scheduling priorities, and resource grouping strategies to match AI-specific requirements. This parameter optimization enables high resource utilization efficiency while maintaining adaptability to different AI workload patterns and scales.
2Productivity
If general-purpose cloud infrastructure is incrementally extended to support AI, then existing investments are protected, but performance and scalability deteriorate due to fundamental differences in AI workload requirements
Solution Approach 1:
The patent segments the AI infrastructure into distinct functional layers (compute, storage, networking) with specialized optimization for each layer while maintaining standardized interfaces. This segmentation enables high AI workload performance through purpose-built optimizations without requiring complete system redesign, thus managing complexity through modular architecture.
Solution Approach 2:
The system introduces an intermediary abstraction layer between the physical infrastructure and AI workloads that handles complexity management, resource allocation, and workload scheduling. This intermediary enables high performance by translating AI-specific requirements into optimized resource allocation while shielding users from underlying infrastructure complexity.
3Quantity of substance
If distributed infrastructure resources are integrated into a unified cloud platform, then resource pool scalability is improved, but management complexity increases due to heterogeneous resource coordination
Solution Approach 1:
The patent implements a universal resource abstraction layer that provides standardized interfaces and management mechanisms for heterogeneous distributed resources. This universal approach enables the resource pool to scale while managing complexity through consistent allocation, scheduling, and monitoring mechanisms across diverse hardware accelerators, networking, and storage systems.
Solution Approach 2:
The system creates homogeneous management interfaces and allocation mechanisms for heterogeneous physical resources, standardizing how resources are provisioned, monitored, and allocated regardless of their underlying diversity. This homogenization of management approaches enables scalable resource pooling while controlling integration complexity through unified policies and procedures.
4Productivity
If multi-tenant AI workload execution is implemented on shared infrastructure, then resource utilization efficiency is improved, but security and isolation challenges worsen due to workload separation requirements
Solution Approach 1:
The patent segments workloads into isolated execution environments with enforced separation boundaries, allocating specific resource subsets to each workload while maintaining physical or virtual isolation mechanisms. This segmentation enables high resource utilization through shared infrastructure while ensuring security and reliability through enforced workload separation and access control.
Data Source
AI summary
The disclosure herein describes managing artificial intelligence (AI) workloads in a cloud infrastructure platform. A set of distributed infrastructure resources are integrated into the cloud infrastructure platform via native support interfaces. AI workloads are received from a plurality of tenants, wherein the AI workloads include training workloads and inferencing workloads and resource subsets of the set of distributed infrastructure resources are assigned to the received AI workloads. The received AI workloads are scheduled for execution on the assigned resource subsets and based on the scheduling of the AI workloads, they are executed on the assigned resource subsets. The described cloud infrastructure platform provides efficient, secure execution of AI workloads for many different tenants and enables the flexible use of a wide variety of both third-party and first-party infrastructure resources.


