Distributed AI Operating System for Enterprise Data Center Resource Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI software lacks management functionalities for enterprise data centers, such as hardware, data, model, job, algorithm, and user management, and does not provide adequate installation and deployment support for AI applications, leading to inefficiencies and increased maintenance costs in multi-tenant environments.

Innovation Solution

A distributed computing system comprising a master machine and worker machines, with an API server, resource allocation module, and container scheduling module, that allows for the creation of virtual machines to perform AI and ML tasks across multiple physical machines, enabling efficient resource management and job scheduling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing AI software is used in multi-tenant data centers, then AI applications can run on available hardware, but the software lacks management functionalities for hardware, data, models, jobs, algorithms, and users, leading to operational inefficiencies

Engineering Contradiction:
Improvemanagement functionalitiesVSAvoidsoftware functionality
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal operating system that provides comprehensive management functionalities across multiple dimensions including hardware resources, data assets, machine learning models, job scheduling, algorithms, and user management. This multi-functional system replaces the need for separate specialized software components, enabling a single platform to handle all enterprise AI data center operations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The operating system is structured into distinct functional modules that independently manage different aspects of the data center ecosystem. Each module (hardware management, data management, model management, job scheduling, etc.) operates semi-independently, allowing the system to provide specialized functionality for each management domain while maintaining overall system integration through the core operating system architecture.

Inventive Principle:
Principle #1Segmentation

2Ease of manufacture

If AI applications are deployed without dedicated operating system support, then deployment is simpler, but installation and deployment support, as well as maintenance, become more difficult and increase total ownership cost

Engineering Contradiction:
Improvedeployment simplicityVSAvoidmaintenance difficulty
Core Design Contradiction:
Ease of manufactureVSEase of repair

Solution Approach 1:

The operating system incorporates automated self-service capabilities including self-installation, self-configuration, and self-maintenance functions. The system automatically manages resource allocation, job scheduling, and system optimization without requiring manual intervention, thereby simplifying deployment while simultaneously reducing maintenance complexity through automated monitoring and self-healing mechanisms.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The operating system performs preliminary configuration and setup actions during the installation phase, pre-configuring hardware resources, data storage structures, model repositories, and job scheduling parameters. This advance preparation eliminates the need for complex post-deployment configuration and reduces maintenance burden by establishing optimal system states before production use begins.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If virtual machines are distributed across multiple worker machines, then resource utilization and scalability improve, but system complexity and coordination overhead increase

Engineering Contradiction:
Improveresource utilizationVSAvoidsystem coordination
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The operating system serves as an intermediary layer between the physical hardware infrastructure and the virtual machine workloads. It manages the complexity of distributing virtual machines across multiple worker machines by providing centralized scheduling, resource allocation, and coordination mechanisms. This intermediary architecture enables efficient resource utilization while abstracting the coordination complexity from individual components.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements dynamic resource allocation and virtual machine placement strategies that adapt to changing workload demands and resource availability in real-time. The operating system continuously monitors system state and dynamically adjusts virtual machine distribution, resource provisioning, and job scheduling to optimize productivity while managing coordination overhead through adaptive rather than static configurations.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10782988B2Operating system for distributed enterprise artificial intelligence programs on data centers and the clouds
Publication Date: 2020.09.22 PETUUM INC
  • US10782988B2 patent drawing
  • US10782988B2 patent drawing
  • US10782988B2 patent drawing

AI summary

A system including a master machine and a plurality of worker machines is disclosed. The master machine includes, for example, an API server configured to receive a job description; a resource allocation module configured to determine a number of virtual machines required to perform a job based on the job description; a container scheduling module configured to create a container containing the number of virtual machines required to perform the job, wherein at least two of the virtual machines in the container resides on different worker machines, and wherein each of the virtual machines is configured to run a same application to perform the job.