Kubernetes Resource Scheduler for Deep Learning Frameworks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current deep learning frameworks require manual resource management and task scheduling, which can be complex and time-consuming, especially when using multiple frameworks, as developers need to manage computational resources and network policies across server clusters, limiting efficiency and flexibility.

Innovation Solution

A method and apparatus for scheduling resources using a Kubernetes platform, which involves querying deep learning job statuses at intervals, submitting resource requests based on specific statuses, and automating resource allocation and reclamation, allowing for seamless initiation and completion of deep learning training tasks across a uniform resource pool.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual resource management and task scheduling are used for deep learning frameworks, then developers can control resource allocation, but the process becomes complex and time-consuming

Engineering Contradiction:
Improveresource management complexityVSAvoidautomatic resource scheduling
Core Design Contradiction:
Ease of operationVSExtent of automation

Solution Approach 1:

The patent introduces a Kubernetes platform as an intermediary between developers and physical machine resources. The platform includes an API server that receives resource requests, a scheduler that automatically allocates resources, and a controller that manages the entire lifecycle of deep learning training tasks. This intermediary system automates the complex manual processes while maintaining developer control through standardized interfaces.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables self-service automation where the Kubernetes platform automatically queries job statuses, submits resource requests, allocates computational resources, and reclaims resources without requiring manual developer intervention. The platform monitors training task progress and autonomously manages the complete resource lifecycle, freeing developers from repetitive management tasks.

Inventive Principle:
Principle #25Self-service

2Productivity

If developers manage resources manually across multiple deep learning frameworks, then framework flexibility is maintained, but efficiency is limited

Engineering Contradiction:
Improvetraining task efficiencyVSAvoidtime for resource management
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The Kubernetes platform provides a universal resource management system that supports multiple deep learning frameworks through a unified interface. The platform can manage resources for different frameworks simultaneously, eliminating the need for separate management processes for each framework. This multi-functional approach consolidates resource management tasks and improves overall productivity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs preliminary actions by pre-configuring resource pools and establishing scheduling policies before training tasks are submitted. The Kubernetes platform pre-allocates computational resources and sets up the infrastructure needed for various deep learning frameworks, so that when training tasks are submitted, resources are already prepared and available, reducing setup time and improving efficiency.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If a uniform resource pool is used for multiple deep learning frameworks, then resource allocation is simplified, but specific framework requirements may not be met

Engineering Contradiction:
Improveframework support flexibilityVSAvoidresource management system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The Kubernetes platform implements local quality by allowing different resource configurations and policies for different frameworks or job types within the uniform resource pool. Each deep learning framework can have customized resource requirements, scheduling policies, and allocation strategies tailored to its specific needs, while still operating from the same underlying resource pool. This enables the system to maintain simplicity at the pool level while providing customization at the framework level.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11762697B2Method and apparatus for scheduling resource for deep learning framework
Publication Date: 2023.09.19 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11762697B2 patent drawing
  • US11762697B2 patent drawing
  • US11762697B2 patent drawing

AI summary

The present disclosure discloses a method and apparatus for scheduling a resource for a deep learning framework. The method can comprise: querying statuses of all deep learning job objects from a Kubernetes platform at a predetermined interval; and submitting, in response to finding from the queried deep learning job objects a deep learning job object having a status conforming to a resource request submission status, a resource request to the Kubernetes platform to schedule a physical machine where the Kubernetes platform is located to initiate a deep learning training task. The method can completely automate the allocation and release on the resource of the deep learning training task.