一种算力服务器集群的智能运维监测方法及平台

By analyzing the historical operation records of the computing server cluster, the GPU memory access latency and task processing throughput stability are quantified, specified patterns are identified and adaptation values ​​are corrected, solving the problem of accumulated memory access latency fluctuations during GPU switching or migration in existing technologies, and improving the operational stability and resource utilization efficiency of the computing server cluster.

CN121958020BActive Publication Date: 2026-07-17ANHUI SHARETRONIC DATA TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANHUI SHARETRONIC DATA TECHNOLOGY CO LTD
Filing Date
2026-01-07
Publication Date
2026-07-17

Smart Images

  • Figure CN121958020B_ABST
    Figure CN121958020B_ABST
Patent Text Reader

Abstract

本发明适用于算力服务器集群运维与GPU资源调度技术领域,提供了一种算力服务器集群的智能运维监测方法及平台,所述方法包括:针对存在指定模式的候选GPU,确定其短周期波动程度对应的转折阈值,获取该候选GPU对应的最新历史任务样本的短周期波动程度,并根据最新短周期波动程度相对于转折阈值的偏离程度生成修正因子,利用修正因子对该候选GPU的适配性值进行下调修正。本发明避免了对正常GPU的过度干预,同时有效降低了在处理显存占用需求相对稳定且执行时长较长、并对持续算力稳定性敏感的计算任务时,因隐性运行波动累积而引发中后期吞吐劣化的风险,提升了算力服务器集群调度决策的稳定性、准确性和整体运行效率。
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • CN120469797A

  • CN120631598A