Single-gpu multi-model efficient deployment method based on dynamic video memory management

By implementing hierarchical memory management and collaborative scheduling, the problems of low memory utilization and fragmentation in single-GPU multi-model deployments are solved, achieving efficient reuse of memory resources and stability of task scheduling, and adapting to load changes in different scenarios.

CN122412136APending Publication Date: 2026-07-17

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Filing Date
2026-04-21
Publication Date
2026-07-17

Smart Images

  • Figure CN122412136A_ABST
    Figure CN122412136A_ABST
Patent Text Reader

Abstract

本发明公开了基于动态显存管理的单GPU多模型高效部署方法,属于GPU资源调度与深度学习模型部署技术领域,本发明通过分层显存池动态管理与多模型协同调度的深度融合,首先构建基于推理生命周期感知的分层显存池动态管理机制,对单GPU全局显存进行逻辑分层划分,结合模型计算图张量生命周期实现匹配式显存分配与异步回收,解决现有技术显存利用率低、碎片化严重的问题;在此基础上构建优先级‑显存预测双维度时隙化协同调度算法,通过轻量化模型预测任务显存与推理时长,结合优先级构建调度代价函数,实现显存资源的时空复用与高优先级任务QoS保障。
Need to check novelty before this filing date? Find Prior Art