Single-gpu multi-model efficient deployment method based on dynamic video memory management
By implementing hierarchical memory management and collaborative scheduling, the problems of low memory utilization and fragmentation in single-GPU multi-model deployments are solved, achieving efficient reuse of memory resources and stability of task scheduling, and adapting to load changes in different scenarios.
CN122412136APending Publication Date: 2026-07-17
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Filing Date
- 2026-04-21
- Publication Date
- 2026-07-17
Smart Images

Figure CN122412136A_ABST
Abstract
本发明公开了基于动态显存管理的单GPU多模型高效部署方法,属于GPU资源调度与深度学习模型部署技术领域,本发明通过分层显存池动态管理与多模型协同调度的深度融合,首先构建基于推理生命周期感知的分层显存池动态管理机制,对单GPU全局显存进行逻辑分层划分,结合模型计算图张量生命周期实现匹配式显存分配与异步回收,解决现有技术显存利用率低、碎片化严重的问题;在此基础上构建优先级‑显存预测双维度时隙化协同调度算法,通过轻量化模型预测任务显存与推理时长,结合优先级构建调度代价函数,实现显存资源的时空复用与高优先级任务QoS保障。
Need to check novelty before this filing date? Find Prior Art