Monitoring method and device for AI computing power card
By obtaining specified indicator data from the heterogeneous AI computing card interface and performing alarm judgment, the problem of effective monitoring of heterogeneous AI computing cards is solved, the container scheduling strategy is optimized, and the computing efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202510956876.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-11-14
AI Technical Summary
In complex computing resource environments, there is a lack of effective regulatory means in existing technologies to effectively manage and monitor heterogeneous AI computing cards to meet the application scenarios of model training and inference, as well as the special needs of large models and edge computing.
By obtaining data on different specified indicators from the heterogeneous AI computing power card interface, using the alarm rules configured by the set indicators to make alarm judgments, outputting alarm information, and performing fine-grained monitoring of the containers that call the AI computing power card, the container scheduling strategy is optimized.
It enables effective monitoring of heterogeneous AI computing cards, timely detection of anomalies, optimization of container scheduling strategies, improvement of overall computing efficiency, avoidance of large-scale computing interruptions, prediction of expansion needs, and improvement of resource utilization.
Smart Images

Figure CN120950327A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of resource operation and maintenance monitoring technology, specifically to a monitoring method and device for AI computing power cards. Background Technology
[0002] AI computing power is mainly used in multiple application scenarios such as machine learning platforms, intelligent language and speech, and IoT platforms. It must not only meet the application scenarios of model training and inference, but also support special needs such as large models and edge computing.
[0003] Due to the explosive growth of large-scale models, users need to procure different AI computing power, which in turn introduces more large-scale model frameworks and business scenarios. Faced with extremely scarce computing resources and complex computing usage scenarios, effectively managing different AI computing cards becomes crucial. Given that heterogeneous AI computing power must not only meet the application scenarios of model training and inference but also support specialized needs such as large-scale models and edge computing, how to effectively manage heterogeneous AI computing cards is an unresolved issue. Summary of the Invention
[0004] The main objective of this invention is to provide a monitoring method for AI computing cards to address the shortcomings in related technologies.
[0005] To achieve the above objectives, according to a first aspect of the present invention, a monitoring method for AI computing cards is provided, comprising obtaining data of different specified indicators from heterogeneous AI computing card interfaces, wherein different AI computing cards correspond to different specified indicators; performing alarm judgment on the data using alarm rules configured based on the set indicators; and outputting alarm information.
[0006] Optionally, the heterogeneous AI computing card interface can obtain data for different specified indicators, including:
[0007] For NPU-type computing cards, the following data is collected: number of AI processors, real-time receive rate of AI processor network port, real-time send rate of AI processor network port, AI processor network port Link status, AI processor network health status, AI processor error code, AI processor name and ID, AI processor health status, AI processor power consumption, AI processor temperature, AI processor DDR memory usage, total AI processor DDR memory, AI processor HBM memory usage, total AI processor HBM memory, AI processor AI Core utilization, current AI Core frequency, AI processor process information, memory usage of processes, AI processor voltage, NPU-Exporter version information, total NPU memory size with container information, NPU memory usage with container information, and NPU utilization with container information.
[0008] Optionally, for DCU type computing cards, data can be collected on DCU temperature, DCU power consumption, DCU chip frequency, DCU power consumption limit, DCU utilization rate, DCU memory usage, total DCU memory, and containers that call the DCU.
[0009] Optionally, for Nvidia DCGM, GPU utilization, memory bandwidth utilization, encoder utilization, frame buffer, number of frame buffers used, SM clock frequency, memory clock frequency, SM application clock frequency, memory application clock frequency, slow throughput root cause, memory temperature, GPU temperature, power, and energy consumed since driver loading.
[0010] Optionally, after collecting data on set indicators for different types of computing power cards, the method includes: inputting the data into a standardized model to map the data of different AI computing power cards to standard data.
[0011] Optionally, after collecting data on set indicators for different types of computing power cards, the data of containers calling different AI computing power cards in a specified dimension is determined; based on the data in the specified dimension, the resource requirements of different tasks are determined, and then the container scheduling strategy is optimized.
[0012] Optionally, the method may further include obtaining data of a first set indicator from a specified physical node interface; or obtaining data of a second set indicator from a storage cluster interface; or obtaining data of a third set indicator from a service interface.
[0013] According to a second aspect of the present invention, a monitoring device for AI computing power cards is provided, comprising a data acquisition interface for acquiring specified indicator data from a heterogeneous computing power card interface and collecting the indicator data; collecting data of set indicators for different types of computing power cards; and performing alarm judgment on the indicator data based on alarm rules configured according to the set indicators, and outputting alarm information.
[0014] By performing fine-grained monitoring of containers that utilize AI computing power cards, analyzing the resource requirements of different tasks, and optimizing container scheduling strategies, overall computing efficiency can be improved.
[0015] According to a third aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing the computer to perform the method described in any one of the first aspects.
[0016] According to a fourth aspect of the present invention, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the at least one processor to perform the method described in any implementation of the first aspect.
[0017] This embodiment describes a method and apparatus for monitoring AI computing cards. The method includes acquiring data on different specified indicators from heterogeneous AI computing card interfaces, where different AI computing cards correspond to different specified indicators; using alarm rules configured based on the set indicators to perform alarm judgment on the data, and outputting alarm information. This achieves the monitoring of different AI computing cards. Attached Figure Description
[0018] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 This is a flowchart of a monitoring method for AI computing power cards according to an embodiment of the present invention;
[0020] Figure 2 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0021] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of the invention described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0023] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0024] According to embodiments of the present invention, a monitoring method for AI computing power cards is provided, such as... Figure 1 As shown, steps 101 to 102 are included below:
[0025] Step 101: Obtain data for different specified indicators from the heterogeneous AI computing power card interface, where different AI computing power cards correspond to different specified indicators.
[0026] In this step, heterogeneous computing cards refer to different types of AI computing cards that acquire different metric data from the Exporter component. For example, data can be obtained from the Exporter component using the Prometheus monitoring and alerting toolkit, where each Exporter component converts the different metric data into data that Prometheus can recognize. Based on the collected data, further monitoring and alerting can be implemented.
[0027] Step 102: Use alarm rules configured based on set indicators to make alarm judgments on the data and output alarm information.
[0028] In this step, alarm rules can be defined based on different indicators, thereby enabling the immediate detection of abnormal situations in heterogeneous AI computing cards and timely response and handling.
[0029] As an optional implementation method in this embodiment, after collecting data on set indicators for different types of computing power cards, the data of containers calling different AI computing power cards in a specified dimension is determined; based on the data in the specified dimension, the resource requirements of different tasks are determined, and then the container scheduling strategy is optimized.
[0030] In this optional implementation, fault risks can be detected in advance based on the indicator information of different AI computing cards, and risky cards can be replaced to avoid large-scale computing interruptions due to single card failures. At the same time, capacity can be predicted to provide a basis for capacity expansion. Subsequent data analysis can be performed to find the root cause of low scheduling efficiency and optimize the scheduling logic.
[0031] Therefore, by monitoring key indicators such as the temperature, power consumption, utilization rate, and memory usage of the AI computing card, the containers that call the AI computing card can be monitored in a fine-grained manner, the resource requirements of different tasks can be analyzed, and the container scheduling strategy can be optimized, thereby improving the overall computing efficiency.
[0032] As an optional implementation of this embodiment, data on different specified indicators obtained from the heterogeneous AI computing card interface includes: for NPU type computing cards, the number of AI processors, the real-time receiving rate of the AI processor network port, the real-time sending rate of the AI processor network port, the Link status of the AI processor network port, the network health status of the AI processor, the error code of the AI processor, the name and ID of the AI processor, the health status of the AI processor, the power consumption of the AI processor, the temperature of the AI processor, the amount of DDR memory used by the AI processor, the total amount of DDR memory of the AI processor, the amount of HBM memory used by the AI processor, the total HBM memory of the AI processor, the AI Core utilization rate of the AI processor, the current frequency of the AI Core of the AI processor, the process information of the AI processor, the amount of memory used by the process, the voltage of the AI processor, the NPU-Exporter version information, the total size of NPU memory with container information, the memory used by NPU with container information, and the NPU utilization rate with container information.
[0033] In this optional implementation, refer to Table 1:
[0034] Table 1
[0035]
[0036]
[0037] As an optional implementation method in this embodiment, the following data is collected for DCU type computing cards: DCU temperature, DCU power consumption, DCU chip frequency, DCU power consumption limit, DCU utilization rate, DCU memory usage, total DCU memory, and containers that call the DCU.
[0038] For specific metrics and metric descriptions in this optional implementation method, please refer to Table 2:
[0039] Table 2
[0040]
[0041]
[0042] As an optional implementation of this embodiment, the following data is collected for Nvidia DCGM: GPU utilization, memory bandwidth utilization, encoder utilization, frame buffer, number of frame buffers used, SM clock frequency, memory clock frequency, SM application clock frequency, memory application clock frequency, slow throughput root cause, memory temperature, GPU temperature, power, and energy consumed since driver loading.
[0043] In this optional implementation method, please refer to the indicators and indicator descriptions shown in Table 3.
[0044]
[0045]
[0046] As an optional implementation of this embodiment, after collecting data on set indicators for different types of computing power cards, the method includes: inputting the data into a standardized model to map the data of different AI computing power cards into standard data.
[0047] In this optional implementation, a standardized monitoring indicator system is established in advance, and a customized analysis algorithm for time-series multidimensional data of different AI computing cards is used to calculate standardized monitoring data such as utilization rate, power consumption, and fault risk at any time stamp in real time.
[0048] Taking fault risk as an example, when mapping data from different AI computing cards to standard data, taking an NPU-type computing card as an example, the fault risk index at a certain timestamp is calculated as follows:
[0049] High risk = Network health status is unhealthy || Processor health status is unhealthy || Processor temperature has consistently been above 70 degrees Celsius in the past 120 seconds || AI Core utilization has consistently been above 90% in the past 120 seconds || NPU utilization with container information has consistently been above 90% in the past 120 seconds.
[0050] Different computing power cards can establish different standardized monitoring indicator systems. After determining a certain dimension indicator, alarm rules can be configured to further realize alarm judgment.
[0051] As an optional implementation of this embodiment, the method further includes obtaining data of a first set indicator from a specified physical node interface; or obtaining data of a second set indicator from a storage cluster interface; or obtaining data of a third set indicator from a service interface.
[0052] In this optional implementation, physical node status monitoring includes not only the aforementioned AI computing card, but also node CPU / memory usage, disk I / O usage / rate / latency, system load, network card traffic, temperature, and fan status.
[0053] Storage cluster status monitoring includes storage cluster health status, storage cluster capacity statistics, storage cluster IOPS / bandwidth, data disk SMART information, disk temperature, and disk bad sector health.
[0054] Control service status monitoring: Real-time status monitoring of critical operating system services, microservice control services, cloud service control services, and automation center services.
[0055] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0056] According to an embodiment of the present invention, a monitoring device for AI computing power cards is also provided, including a data acquisition interface for acquiring specified indicator data from a heterogeneous computing power card interface and collecting the indicator data; collecting data of set indicators for different types of computing power cards; and performing alarm judgment on the indicator data based on alarm rules configured according to the set indicators, and outputting alarm information.
[0057] As an optional implementation of this embodiment, data on different specified indicators obtained from the heterogeneous AI computing card interface includes: for NPU type computing cards, the number of AI processors, the real-time receiving rate of the AI processor network port, the real-time sending rate of the AI processor network port, the Link status of the AI processor network port, the network health status of the AI processor, the error code of the AI processor, the name and ID of the AI processor, the health status of the AI processor, the power consumption of the AI processor, the temperature of the AI processor, the amount of DDR memory used by the AI processor, the total amount of DDR memory of the AI processor, the amount of HBM memory used by the AI processor, the total HBM memory of the AI processor, the AI Core utilization rate of the AI processor, the current frequency of the AI Core of the AI processor, the process information of the AI processor, the amount of memory used by the process, the voltage of the AI processor, the NPU-Exporter version information, the total size of NPU memory with container information, the memory used by NPU with container information, and the NPU utilization rate with container information.
[0058] As an optional implementation method in this embodiment, the following data is collected for DCU type computing cards: DCU temperature, DCU power consumption, DCU chip frequency, DCU power consumption limit, DCU utilization rate, DCU memory usage, total DCU memory, and containers that call the DCU.
[0059] As an optional implementation of this embodiment, the following data is collected for Nvidia DCGM: GPU utilization, memory bandwidth utilization, encoder utilization, frame buffer, number of frame buffers used, SM clock frequency, memory clock frequency, SM application clock frequency, memory application clock frequency, information on the reason for the constant slowdown, memory temperature, GPU temperature, power, and energy consumed since driver loading.
[0060] As an optional implementation of this embodiment, after collecting data on set indicators for different types of computing power cards, the method includes: inputting the data into a standardized model to map the data of different AI computing power cards into standard data.
[0061] As an optional implementation method in this embodiment, after collecting data on set indicators for different types of computing power cards, the data of containers calling different AI computing power cards in a specified dimension is determined; based on the data in the specified dimension, the resource requirements of different tasks are determined, and then the container scheduling strategy is optimized.
[0062] As an optional implementation of this embodiment, the method further includes obtaining data of a first set indicator from a specified physical node interface; or obtaining data of a second set indicator from a storage cluster interface; or obtaining data of a third set indicator from a service interface.
[0063] According to embodiments of the present invention, the present invention also provides an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the methods described in any of the above embodiments.
[0064] According to embodiments of the present invention, the present invention also provides a readable storage medium storing computer instructions that enable a computer to perform the methods described in any of the above embodiments when executed.
[0065] According to embodiments of the present invention, the present invention also provides a computer program product that, when executed by a processor, can implement the methods described in any of the above embodiments.
[0066] Figure 2A schematic block diagram of an example electronic device 300 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices.
[0067] like Figure 2 As shown, the electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. The RAM 303 may also store various programs and data required for the operation of the electronic device 300. The computing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0068] Multiple components in electronic device 300 are connected to I / O interface 305, including: input unit 306, such as keyboard, mouse, etc.; output unit 307, such as various types of displays, speakers, etc.; storage unit 308, such as disk, optical disk, etc.; and communication unit 309, such as network card, modem, wireless transceiver, etc. Communication unit 309 allows electronic device 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0069] The computing unit 301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as the object matching method. For example, in some embodiments, the object matching method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 300 via ROM 302 and / or communication unit 309. When the computer program is loaded into RAM 303 and executed by the computing unit 301, one or more steps of the methods described above may be performed.
[0070] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0071] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0072] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
Claims
1. A monitoring method for AI computing power cards, characterized in that, include: Data for different specified indicators are obtained from the interface of heterogeneous AI computing power cards, where different AI computing power cards correspond to different specified indicators; The alarm rules configured based on set indicators are used to determine the alarm status of the data and output alarm information.
2. The monitoring method for AI computing power cards according to claim 1, characterized in that, Data for different specified metrics obtained from heterogeneous AI computing power card interfaces includes: For NPU-type computing cards, the following data is collected: number of AI processors, real-time receive rate of AI processor network port, real-time send rate of AI processor network port, AI processor network port Link status, AI processor network health status, AI processor error code, AI processor name and ID, AI processor health status, AI processor power consumption, AI processor temperature, AI processor DDR memory usage, total AI processor DDR memory, AI processor HBM memory usage, total AI processor HBM memory, AI processor AICore utilization, current AI Core frequency, AI processor process information, memory usage of processes, AI processor voltage, NPU-Exporter version information, total NPU memory size with container information, NPU memory usage with container information, and NPU utilization with container information.
3. The monitoring method for AI computing power cards according to claim 1, characterized in that, For DCU type computing cards, collect data on DCU temperature, DCU power consumption, DCU chip frequency, DCU power consumption limit, DCU utilization rate, DCU memory usage, total DCU memory, and containers that call the DCU.
4. The monitoring method for AI computing power cards according to claim 1, characterized in that, For Nvidia DCGM, collect GPU utilization, memory bandwidth utilization, encoder utilization, frame buffer, frame buffer usage, SM clock frequency, memory clock frequency, SM application clock frequency, memory application clock frequency, slow throughput root cause, memory temperature, GPU temperature, power consumption, and energy consumed since driver loading.
5. The monitoring method for AI computing power cards according to claim 4, characterized in that, After collecting and setting data for different types of computing power cards, the methods include: The data is input into a standardized model to map the data from different AI computing cards to standard data.
6. The monitoring method for AI computing power cards according to claim 5, characterized in that, After collecting data on the set indicators for different types of computing power cards, determine the data in the specified dimensions of the containers that call different AI computing power cards; Based on the data of the specified dimensions, the resource requirements of different tasks are determined, and the container scheduling strategy is then optimized.
7. The monitoring method for AI computing power cards according to claim 1, characterized in that, The method also includes obtaining data for a first set metric from the interface of a specified physical node; Alternatively, data for the second set metric can be obtained from the storage cluster interface; Alternatively, data for a third-party set metric can be obtained from the service interface.
8. A monitoring device for AI computing power cards, characterized in that, include: The data acquisition interface is used to obtain specified indicator data from the heterogeneous computing power card interface and to collect the indicator data. Collect and set data for different types of computing power cards; Based on the alarm rules configured for the set indicators, the system performs alarm judgment on the indicator data and outputs alarm information.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the method according to any one of claims 1-7.
10. An electronic device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the at least one processor to perform the method according to any one of claims 1-7.