Method for monitoring running state of ai server and related equipment

By collecting and analyzing the operational status data of the AI ​​server's host side and accelerated computing side, cross-component relationships are established to locate the source of resource interference, solving the problem of difficulty in locating GPU performance anomalies in existing technologies and improving diagnostic efficiency and accuracy.

CN122332170APending Publication Date: 2026-07-03CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
Filing Date
2026-04-08
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing GPU monitoring tools cannot identify the root cause of abnormal GPU performance, especially in multi-tenant shared GPU infrastructures, where performance fluctuations caused by resource contention on the host side are difficult to locate quickly.

Method used

Collect operational status data from the host side and accelerated computing side of the AI ​​server, establish a cross-component related dataset through time correlation processing, perform status analysis, identify abnormal operational status data, and analyze the types of resource interference.

Benefits of technology

It improves the accuracy and efficiency of root cause diagnosis of GPU performance anomalies and supports the operation and maintenance management and performance optimization of multi-tenant GPU infrastructure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122332170A_ABST
    Figure CN122332170A_ABST
Patent Text Reader

Abstract

This invention provides a method and related equipment for monitoring the operational status of an AI server. The method includes: collecting operational status data from the host side and the accelerated computing side of the AI ​​server to obtain a raw operational status dataset; performing time correlation processing on the data from both sides based on time information to obtain a cross-component correlated dataset; performing status analysis on the cross-component correlated dataset to identify abnormal operational status data; and performing correlation parsing on the abnormal operational status data to determine the type of resource interference and obtain monitoring results. By establishing cross-layer data correlation between the host side and the accelerated computing side, this invention can locate the source of host-side resource interference causing abnormal accelerated computing performance, improving the accuracy and efficiency of root cause diagnosis and effectively supporting the operation and maintenance management and performance optimization of multi-tenant GPU infrastructure.
Need to check novelty before this filing date? Find Prior Art