Managed AI Infrastructure Service for Distributed Workloads

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current general-purpose cloud-based infrastructure services are inadequate for handling the exponential growth of artificial intelligence (AI) workloads, such as Deep Learning Training (DLT) jobs, due to their workload-agnostic design, leading to inefficiencies and limitations in scalability and resource utilization.

Innovation Solution

A fully managed, globally distributed, multi-tenant AI infrastructure service that integrates diverse infrastructure resources via native support interfaces, allowing for efficient scheduling and execution of AI workloads, including training and inferencing tasks, using a hierarchical scheduling system that optimizes resource allocation and utilization across a distributed pool of hardware accelerators, networking, and storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If general-purpose cloud-based infrastructure services are used to handle AI workloads, then existing infrastructure can be leveraged with minimal additional investment, but resource utilization efficiency deteriorates and scalability is limited due to workload-agnostic design

Engineering Contradiction:
Improveadaptability to AI workloadsVSAvoidresource utilization efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent creates a purpose-built AI infrastructure service that specializes in handling AI workloads (training and inferencing) while maintaining universal compatibility across different hardware accelerators, networking configurations, and storage systems. This specialized universal design resolves the contradiction by optimizing resource utilization for AI-specific patterns while maintaining broad adaptability to various AI workload types and scales.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically adjusts resource allocation parameters based on workload characteristics, changing allocation granularity, scheduling priorities, and resource grouping strategies to match AI-specific requirements. This parameter optimization enables high resource utilization efficiency while maintaining adaptability to different AI workload patterns and scales.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If general-purpose cloud infrastructure is incrementally extended to support AI, then existing investments are protected, but performance and scalability deteriorate due to fundamental differences in AI workload requirements

Engineering Contradiction:
ImproveAI workload performanceVSAvoidinfrastructure complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the AI infrastructure into distinct functional layers (compute, storage, networking) with specialized optimization for each layer while maintaining standardized interfaces. This segmentation enables high AI workload performance through purpose-built optimizations without requiring complete system redesign, thus managing complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary abstraction layer between the physical infrastructure and AI workloads that handles complexity management, resource allocation, and workload scheduling. This intermediary enables high performance by translating AI-specific requirements into optimized resource allocation while shielding users from underlying infrastructure complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If distributed infrastructure resources are integrated into a unified cloud platform, then resource pool scalability is improved, but management complexity increases due to heterogeneous resource coordination

Engineering Contradiction:
Improveresource pool sizeVSAvoidresource integration complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent implements a universal resource abstraction layer that provides standardized interfaces and management mechanisms for heterogeneous distributed resources. This universal approach enables the resource pool to scale while managing complexity through consistent allocation, scheduling, and monitoring mechanisms across diverse hardware accelerators, networking, and storage systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system creates homogeneous management interfaces and allocation mechanisms for heterogeneous physical resources, standardizing how resources are provisioned, monitored, and allocated regardless of their underlying diversity. This homogenization of management approaches enables scalable resource pooling while controlling integration complexity through unified policies and procedures.

Inventive Principle:
Principle #33Homogeneity

4Productivity

If multi-tenant AI workload execution is implemented on shared infrastructure, then resource utilization efficiency is improved, but security and isolation challenges worsen due to workload separation requirements

Engineering Contradiction:
Improveresource utilization efficiencyVSAvoidworkload isolation security
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent segments workloads into isolated execution environments with enforced separation boundaries, allocating specific resource subsets to each workload while maintaining physical or virtual isolation mechanisms. This segmentation enables high resource utilization through shared infrastructure while ensuring security and reliability through enforced workload separation and access control.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20220318674A1Planet-scale, fully managed artificial intelligence infrastructure service
Publication Date: 2022.10.06 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20220318674A1 patent drawing
  • US20220318674A1 patent drawing
  • US20220318674A1 patent drawing

AI summary

The disclosure herein describes managing artificial intelligence (AI) workloads in a cloud infrastructure platform. A set of distributed infrastructure resources are integrated into the cloud infrastructure platform via native support interfaces. AI workloads are received from a plurality of tenants, wherein the AI workloads include training workloads and inferencing workloads and resource subsets of the set of distributed infrastructure resources are assigned to the received AI workloads. The received AI workloads are scheduled for execution on the assigned resource subsets and based on the scheduling of the AI workloads, they are executed on the assigned resource subsets. The described cloud infrastructure platform provides efficient, secure execution of AI workloads for many different tenants and enables the flexible use of a wide variety of both third-party and first-party infrastructure resources.