Method for monitoring running state of ai server and related equipment
By collecting and analyzing the operational status data of the AI server's host side and accelerated computing side, cross-component relationships are established to locate the source of resource interference, solving the problem of difficulty in locating GPU performance anomalies in existing technologies and improving diagnostic efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
- Filing Date
- 2026-04-08
- Publication Date
- 2026-07-03
AI Technical Summary
Existing GPU monitoring tools cannot identify the root cause of abnormal GPU performance, especially in multi-tenant shared GPU infrastructures, where performance fluctuations caused by resource contention on the host side are difficult to locate quickly.
Collect operational status data from the host side and accelerated computing side of the AI server, establish a cross-component related dataset through time correlation processing, perform status analysis, identify abnormal operational status data, and analyze the types of resource interference.
It improves the accuracy and efficiency of root cause diagnosis of GPU performance anomalies and supports the operation and maintenance management and performance optimization of multi-tenant GPU infrastructure.
Smart Images

Figure CN122332170A_ABST