用于人工智能模型的GPU资源动态调度方法及装置

By using a task-aware mechanism and Kubernetes management, GPU resources are dynamically scheduled, solving the problems of low resource utilization and high cost under the dedicated GPU allocation method. This enables flexible GPU resource management, improves resource utilization, and reduces user costs.

CN120066782BActive Publication Date: 2026-07-17BEIJING SILICONFLOW TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING SILICONFLOW TECHNOLOGY CO LTD
Filing Date
2025-02-07
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

The existing dedicated GPU resource allocation method results in low resource utilization, high user costs, and insufficient flexibility, making it impossible to dynamically adjust GPU resource usage according to task requirements.

Method used

The system identifies the GPU resource requirements of AI model tasks through a task-aware mechanism, monitors GPU load in real time, dynamically mounts and unmounts GPU resources, and manages them using custom resource definitions and controllers in the Kubernetes cluster.

Benefits of technology

It improves the utilization of GPU resources, reduces user costs, enhances the flexibility and applicability of the system, and supports a variety of custom recycling strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066782B_ABST
    Figure CN120066782B_ABST
Patent Text Reader

Abstract

本申请涉及一种用于人工智能模型的GPU资源动态调度方法及装置。该方法包括:通过任务感知机制,识别人工智能模型任务对GPU的资源需求;在所述人工智能模型任务的执行过程中,实时监测GPU负载情况;根据所述GPU负载情况确定GPU的资源需求以动态挂载GPU资源;通过所述GPU资源执行所述人工智能模型任务,得到输出结果;在满足GPU回收策略时,动态卸载所述GPU资源。本申请涉及的用于人工智能模型的GPU资源动态调度方法及装置,能够实现GPU资源的分时共享,显著提高了GPU利用率,降低了用户使用成本,同时支持多种自定义回收策略,增强了系统的灵活性和适用性。
Need to check novelty before this filing date? Find Prior Art