A high-performance computing cluster resource scheduling method and system based on TR-DQN

By using a two-level neural network based on TR-DQN and a deep reinforcement learning method, the dynamic adaptability and efficiency problems in resource scheduling of high-performance computing clusters are solved, achieving efficient resource utilization and rapid response.

CN117591273BActive Publication Date: 2026-07-21WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN UNIV
Filing Date
2023-11-06
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Traditional high-performance computing cluster resource scheduling algorithms cannot adapt to dynamic changes in cluster status and nodes, resulting in job starvation and high computing resource requirements, leading to low scheduling efficiency.

Method used

We employ a two-level neural network based on TR-DQN and a deep reinforcement learning method. We use a window access waiting queue and combine task priority and node information for scheduling. We design a specific reward function to optimize resource utilization and response time.

Benefits of technology

It achieves adaptive scheduling of high-performance computing cluster resources, reduces job starvation, and improves resource utilization and task response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117591273B_ABST
    Figure CN117591273B_ABST
Patent Text Reader

Abstract

The application discloses a kind of high-performance computing cluster resource scheduling method and system based on TR-DQN, first user submits task request, all requests enter waiting queue and wait for scheduling;Then the priority of submitting task is calculated, and the waiting queue is reordered;Then the node information and task information of cluster are collected and processed, and the processed data is input into TR-DQN model for scheduling;Finally, after task scheduling is completed, it enters corresponding node operation.TR-DQN model combines the characteristics of high-performance computing cluster scheduling into deep reinforcement learning, and introduces two-level neural network structure, the first neural network is used to select tasks for immediate execution or reserved execution, and the second neural network is used to select tasks for backfilling, which can improve the resource utilization of the cluster, reduce the waiting time of the task, and quickly adapt to changes in the cluster load environment. In addition, it can also minimize the problem of work starvation in the cluster.
Need to check novelty before this filing date? Find Prior Art