一种面向GPU系统深度学习推理的能效感知自适应调度方法及系统
By leveraging reinforcement learning and NVIDIA's MPS technology, the batch size and frequency of GPU deep learning inference tasks are dynamically scheduled, solving the problems of high GPU power consumption and latency response, and achieving maximum energy efficiency and fast response under different loads.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HARBIN INST OF TECH
- Filing Date
- 2023-09-30
- Publication Date
- 2026-07-17
AI Technical Summary
In existing technologies, GPU deep learning inference tasks face the problems of high energy consumption and latency response, especially under high load conditions where it is difficult to effectively schedule batch size and GPU frequency to maximize energy efficiency.
An energy-efficient adaptive scheduling method based on reinforcement learning is adopted. By combining NVIDIA's MPS technology and transfer learning with an asynchronous execution strategy and an energy-efficient adaptive scheduler, batch size and GPU frequency are dynamically coordinated. The model is trained using reinforcement learning algorithms to reduce energy consumption and meet latency requirements.
It maximizes the energy efficiency of GPU inference under different load conditions, and can adaptively select the most suitable batch size and GPU frequency to reduce energy consumption and meet latency requirements, thereby improving the response speed and efficiency of inference tasks.
Smart Images

Figure CN117667336B_ABST