一种算力服务器集群的智能运维监测方法及平台
By analyzing the historical operation records of the computing server cluster, the GPU memory access latency and task processing throughput stability are quantified, specified patterns are identified and adaptation values are corrected, solving the problem of accumulated memory access latency fluctuations during GPU switching or migration in existing technologies, and improving the operational stability and resource utilization efficiency of the computing server cluster.
CN121958020BActive Publication Date: 2026-07-17ANHUI SHARETRONIC DATA TECHNOLOGY CO LTD
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI SHARETRONIC DATA TECHNOLOGY CO LTD
- Filing Date
- 2026-01-07
- Publication Date
- 2026-07-17
Smart Images

Figure CN121958020B_ABST
Abstract
本发明适用于算力服务器集群运维与GPU资源调度技术领域,提供了一种算力服务器集群的智能运维监测方法及平台,所述方法包括:针对存在指定模式的候选GPU,确定其短周期波动程度对应的转折阈值,获取该候选GPU对应的最新历史任务样本的短周期波动程度,并根据最新短周期波动程度相对于转折阈值的偏离程度生成修正因子,利用修正因子对该候选GPU的适配性值进行下调修正。本发明避免了对正常GPU的过度干预,同时有效降低了在处理显存占用需求相对稳定且执行时长较长、并对持续算力稳定性敏感的计算任务时,因隐性运行波动累积而引发中后期吞吐劣化的风险,提升了算力服务器集群调度决策的稳定性、准确性和整体运行效率。
Need to check novelty before this filing date? Find Prior Art
Citation Information
Patent Citations
CN120469797A
CN120631598A