Deep learning inference performance interference-aware gpu resource configuration method
By establishing a performance prediction model and mathematical optimization problem, the GPU resource allocation was optimized, which solved the performance interference problem when multiple DNNs share resources, achieving cost minimization and performance guarantee, and improving resource utilization efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- EAST CHINA NORMAL UNIV
- Filing Date
- 2022-03-24
- Publication Date
- 2026-07-21
AI Technical Summary
When multiple DNN inference workloads share GPU resources, performance interference can lead to unmet user service level objectives and high inference costs. Existing technologies cannot effectively predict and optimize resource allocation.
By establishing a performance prediction model, obtaining hardware and load parameters, constructing a mathematical optimization problem, optimizing GPU resource configuration to minimize costs, and ensuring the performance SLO of the DNN inference load, the iGniter framework is used for resource allocation and load placement.
It significantly reduces user costs, improves resource utilization efficiency, and reduces additional overhead caused by performance interference while ensuring DNN inference performance.
Smart Images

Figure CN115237586B_ABST