用于人工智能模型的GPU资源动态调度方法及装置
By using a task-aware mechanism and Kubernetes management, GPU resources are dynamically scheduled, solving the problems of low resource utilization and high cost under the dedicated GPU allocation method. This enables flexible GPU resource management, improves resource utilization, and reduces user costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING SILICONFLOW TECHNOLOGY CO LTD
- Filing Date
- 2025-02-07
- Publication Date
- 2026-07-17
AI Technical Summary
The existing dedicated GPU resource allocation method results in low resource utilization, high user costs, and insufficient flexibility, making it impossible to dynamically adjust GPU resource usage according to task requirements.
The system identifies the GPU resource requirements of AI model tasks through a task-aware mechanism, monitors GPU load in real time, dynamically mounts and unmounts GPU resources, and manages them using custom resource definitions and controllers in the Kubernetes cluster.
It improves the utilization of GPU resources, reduces user costs, enhances the flexibility and applicability of the system, and supports a variety of custom recycling strategies.
Smart Images

Figure CN120066782B_ABST