A GPU card monitoring method, monitoring system and related devices

By receiving monitoring requests and using the Prometheus monitoring system to obtain the UUID information and operation resource information of the target GPU card, the problem of real-time monitoring of non-integrated card service operation information is solved, and efficient utilization of GPU resources is achieved.

CN114860536BActive Publication Date: 2025-05-06ZHENGZHOU YUNHAI INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210428502.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-22
Publication Date
2025-05-06
Estimated Expiration
2042-04-22

AI Technical Summary

Technical Problem

In the cluster, GPU resources are not fully utilized because the prior art cannot monitor service operation information on non-integrated cards in real time, resulting in users being unable to regulate services in real time.

Method used

By receiving monitoring requests, the Prometheus monitoring system is used to obtain the currently running minimum resource service unit, determine the UUID information of the target GPU card where the target minimum resource service unit is located, and obtain the GPU operation resource information, including GPU utilization, memory usage and core proportion.

Benefits of technology

Real-time monitoring of non-integrated card resource services is realized, and users can regulate in real time based on the service operation information to ensure that GPU resources are maximized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114860536B_ABST
    Figure CN114860536B_ABST
Patent Text Reader

Abstract

The present application provides a monitoring method for a GPU card, comprising: receiving a monitoring request; using the Prometheus monitoring system to obtain the currently running minimum resource service unit; determining the target minimum resource service unit applied by the service corresponding to the monitoring request, and obtaining the UUID information of the target GPU card where the target minimum resource service unit is located; obtaining the GPU operation resource information of the target GPU card according to the UUID information; the GPU operation resource information includes at least one of GPU utilization, video memory usage, and core share. The present application can monitor the resource consumption of non-whole cards, thereby filling the gap in monitoring non-whole card resource services, making it easier for users to control the GPU operation status in real time. The present application also provides a monitoring system, a computer-readable storage medium, and an electronic device for a GPU card, which have the above-mentioned beneficial effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of server information monitoring, and in particular to a GPU card monitoring method, system and related devices. Background Art

[0002] In a cluster, GPU resources are very valuable. If you can only deploy whole-card services, the number of services that can be deployed will be very limited. For example, if there is only one GPU card, you can only deploy one whole-card service. This will make the GPU underutilized. To solve this problem, fine-grained services that are not deployed with whole-card resources can now be used in k8s clusters, that is, multiple GPU services run on one card. In this case, GPU resources will be utilized to the maximum extent.

[0003] However, due to the use of non-whole card deployment services, only hardware information can be obtained, but no corresponding services can be obtained. For clusters that want to deploy non-whole card resource services, it is impossible to confirm the service operation information in real time, and this part of information is very important to users, so that they can make real-time adjustments based on the service operation information. Summary of the invention

[0004] The purpose of this application is to provide a GPU card monitoring method, monitoring system, computer-readable storage medium and electronic device, which can monitor service operation information on non-entire cards.

[0005] In order to solve the above technical problems, the present application provides a GPU card monitoring method, and the specific technical solution is as follows:

[0006] receiving monitoring requests;

[0007] Use the Prometheus monitoring system to obtain the currently running minimum resource service unit;

[0008] Determine the target minimum resource service unit applied by the service corresponding to the monitoring request, and obtain the UUID information of the target GPU card where the target minimum resource service unit is located;

[0009] The GPU operation resource information of the target GPU card is obtained according to the UUID information; the GPU operation resource information includes at least one of GPU utilization, video memory usage and core ratio.

[0010] Optionally, also include:

[0011] Obtaining CPU resource information and storage resource information of the service;

[0012] After obtaining the operating resource information of the target GPU card according to the UUID information, the method further includes:

[0013] The service operation status information including the CPU resource information, the storage resource information and the GPU operation resource information is output.

[0014] Optionally, use the Prometheus monitoring system to obtain the currently running minimum resource service unit, including:

[0015] Enter the preset query statement in the Prometheus monitoring system to obtain the currently running minimum resource service unit.

[0016] Optionally, determining a target minimum resource service unit applied by the service corresponding to the monitoring request includes:

[0017] Parse the non-whole card resource service list to determine the service name and scenario included in the service corresponding to the monitoring request;

[0018] All target minimum resource service units corresponding to the service are determined according to the service name and the scenario.

[0019] Optionally, obtaining the UUID information of the target GPU card where the target minimum resource service unit is located includes:

[0020] Determine the unit information of the target minimum resource service unit by using the CoreV1 interface;

[0021] The UUID information of the target GPU card is determined by analyzing the environment variables of the target minimum resource service unit according to the unit information.

[0022] Optionally, acquiring the GPU operation resource information of the target GPU card according to the UUID information includes:

[0023] Enter the UUID information and execute the nvidia-smi -L command to determine the correspondence between the UUID information and the GPU card number;

[0024] Determine the target GPU card according to the GPU card number;

[0025] The nvidia-smi command is used to collect the GPU operation resource information of the target GPU card.

[0026] The present application provides a monitoring system for a GPU card, comprising:

[0027] A request receiving module, used for receiving monitoring requests;

[0028] The service unit acquisition module is used to use the Prometheus monitoring system to obtain the currently running minimum resource service unit;

[0029] A UUID determination module is used to determine the target minimum resource service unit applied by the service corresponding to the monitoring request, and obtain the UUID information of the target GPU card where the target minimum resource service unit is located;

[0030] The information collection module is used to obtain the GPU operation resource information of the target GPU card according to the UUID information; the GPU operation resource information includes at least one of GPU utilization, video memory usage and core ratio.

[0031] Optionally, also include:

[0032] A service resource information collection module, used to obtain the CPU resource information and storage resource information of the service;

[0033] The operation information output module is used to output the service operation status information including the CPU resource information, the storage resource information and the GPU operation resource information.

[0034] The present application also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method described above are implemented.

[0035] The present application also provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps of the above-mentioned method when calling the computer program in the memory.

[0036] The present application provides a monitoring method for a GPU card, comprising: receiving a monitoring request; obtaining a currently running minimum resource service unit using a Prometheus monitoring system; determining a target minimum resource service unit applied by a service corresponding to the monitoring request, and obtaining UUID information of a target GPU card where the target minimum resource service unit is located; obtaining GPU operating resource information of the target GPU card according to the UUID information; the GPU operating resource information includes at least one of GPU utilization, video memory usage, and core share.

[0037] This application can monitor the resource consumption of non-whole cards. After receiving a monitoring request, it can obtain the UUID information of the target GPU card where the target minimum resource service unit is located, and then obtain the corresponding GPU operation resource information. This application can not only monitor non-whole cards, but also whole card resources, thus filling the gap in non-whole card resource service monitoring and facilitating users to control the GPU operation status in real time.

[0038] The present application also provides a GPU card monitoring system, a computer-readable storage medium and an electronic device, which have the above-mentioned beneficial effects and are not described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0040] Figure 1 A flowchart of a GPU card monitoring method provided in an embodiment of the present application;

[0041] Figure 2 A schematic diagram of the structure of a monitoring system for a GPU card provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0043] Please refer to Figure 1 , Figure 1 A flowchart of a GPU card monitoring method provided in an embodiment of the present application, the method comprising:

[0044] S101: receiving a monitoring request;

[0045] This step receives a monitoring request, and there is no limitation on how to receive the monitoring request. The monitoring request may include a resource monitoring request for a service. For a service, it may include one or more minimum resource service units (also called "pods"). The creation of non-whole-card resource services enables the sharing of GPU cards, that is, multiple services are allowed to run on the same GPU card, that is, different services may run on the same GPU card. When creating a non-whole-card resource service, it will compare all GPU cards with remaining resources for random scheduling.

[0046] S102: Use the Prometheus monitoring system to obtain the currently running minimum resource service unit;

[0047] This step aims to obtain the currently running minimum resource service unit. It requires that the kube-state-metrics monitoring component has been deployed in the environment. In this way, you can enter the preset query statement in the Prometheus monitoring system to obtain the currently running minimum resource service unit. One feasible way is to obtain the metric of kube_pod_labels of the kube-state-metrics component through the Prometheus monitoring system, and use it to obtain the minimum resource service unit of non-whole card resource services. In the process of obtaining the minimum resource service unit, you can directly pull out all the minimum resource service units in the current environment.

[0048] S103: Determine the target minimum resource service unit applied by the service corresponding to the monitoring request, and obtain UUID information of the target GPU card where the target minimum resource service unit is located;

[0049] This step is to determine the target minimum resource service unit used by the service in the monitoring request, and determine the UUID (Universally Unique Identifier) ​​information of the target GPU card where the target minimum resource service unit is located. The UUID information is the unique identification code of the GPU card. There is no limitation on the type of UUID information used for the GPU card, as long as it can be used as the unique identification code of the GPU card.

[0050] The target minimum resource service unit for the service application determined in this step is usually at least one. If it is not found, it indicates that the service may be abnormal. If the service is not abnormal, this step should be able to query all the target minimum resource service units used by the service. The target minimum resource service unit in this step indicates the minimum resource service unit used by the service to be queried.

[0051] When determining the target minimum resource service unit, a feasible execution method can parse the non-whole card resource service list, determine the service name and scenario contained in the service corresponding to the monitoring request, and then determine all the target minimum resource service units corresponding to the service according to the service name and scenario. That is, the monitoring request can contain the scene and service name that need to be determined. Since the same service name may exist in different scenes, the service that needs to be monitored can be uniquely determined with the help of the scene and service name. In this execution method, the non-whole card resource service list can be sorted out in advance, so that after receiving the monitoring request, the service can be directly found according to the non-whole card resource service list, and the target minimum resource service unit corresponding to the service can be determined, thereby reducing the time for querying the target minimum resource service unit corresponding to the service and improving the query efficiency.

[0052] When determining the UUID information of the target GPU card, you can first use the CoreV1 interface to determine the unit information of the target minimum resource service unit, and then analyze the environment variables of the target minimum resource service unit based on the unit information to determine the UUID information of the target GPU card. Specifically, you can use the CoreV1 interface read_namespaced_pod of the k8s cluster to obtain the specific information of the minimum resource service unit, and you can obtain the card UUID where the minimum resource service unit is located by analyzing the environment variable "NVIDIA_VISIBLE_DEVICES" of the minimum resource service unit. If this field cannot be obtained, the status of this service may be abnormal.

[0053] S104: Obtaining GPU operation resource information of the target GPU card according to the UUID information;

[0054] After determining the UUID information, it is equivalent to that the GPU information is uniquely determined, and the GPU operating resource information of the target GPU card can be directly obtained. The GPU operating resource information may include at least one of the GPU utilization, video memory usage and core ratio, which can be a combination of any several of them.

[0055] In this step, you can execute the nvidia-smi-L command in the container to obtain the correspondence between the card UUID and the card number (for example, the card number is 0, and the UUID is in the format of GPU-001234-89090-c67677). The nvidia-smi command and the card number are used to parse the resource usage information of the card where the minimum resource service unit is located, including the GPU number, GPU model, temperature, power, GPU usage, GPU video memory usage, GPU video memory total amount, etc. The GPU number, GPU model, temperature, power, non-whole card resource service and whole card data are consistent. The upper limit and lower limit of GPU card allocation for non-whole card resource services are known. For example, the lower limit of GPU card allocation for a service is 0.1, and the upper limit is 0.2. Then, it is stipulated that the GPU video memory allocation of this service can be obtained by multiplying the upper limit allocation ratio by the total GPU video memory, that is, the total video memory of this service. For GPU usage, it is necessary to obtain all the minimum resource service units that have been running on this card, and calculate the GPU usage of each minimum resource service unit. For example, if two minimum resource service units are running on this card, the ratio of the first minimum resource service unit is 0.1-0.2, and the ratio of the second minimum resource service unit is 0.3-0.4. Then, based on the upper limit, the usage ratio of the first minimum resource service unit is 0.2 / (0.2+0.4)=1 / 3, and the ratio of the second is 2 / 3. Through nvidia-smi, the information of the currently running processes is obtained. If two minimum resource service units occupy the GPU at the same time, the calculation can be performed according to the ratio. If only one of the minimum resource service units occupies the GPU, the GPU usage of the card is the GPU usage of the minimum resource service unit. Similarly, the GPU memory usage will also be calculated in this way.

[0056] For users, it is necessary to intuitively obtain the core ratio and video memory ratio of the card, that is, what is the ratio of each service to the card and what is the video memory ratio. Through the above operations, the minimum resource service unit information running on the GPU has been obtained. The minimum resource service unit GPU ratio of the same service on the same card is added together to obtain the core ratio of the service on the card. For example, service 1 creates two minimum resource service units with an allocation upper limit of 0.2. If both minimum resource service units run on GPU card 1, then the core ratio of service 1 is 0.2+0.2=0.4. If the two minimum resource service units run on different GPU cards, then the core ratio of service 1 on the two GPU cards is 0.2 respectively.

[0057] In a feasible execution method, while executing this step or this embodiment, the CPU resource information and storage resource information of the service can also be obtained. After executing this step, the service operation status information including the CPU resource information, storage resource information and GPU operation resource information can be output, so that the user can clearly know all the operation information of the current service, not limited to the operation information of the GPU card.

[0058] This application can monitor the resource consumption of non-whole cards. After receiving the monitoring request, it can obtain the UUID information of the target GPU card where the target minimum resource service unit is located, and then obtain the corresponding GPU operation resource information. This application can not only monitor non-whole cards, but also monitor the resources of whole cards, thus filling the gap in monitoring the resource services of non-whole cards, making it easier for users to control the GPU operation status in real time.

[0059] A monitoring system for a GPU card provided in an embodiment of the present application is introduced below. The monitoring system for a GPU card described below and the monitoring method for a GPU card described above can refer to each other.

[0060] See also Figure 2 , Figure 2 A schematic diagram of the structure of a monitoring system for a GPU card provided in an embodiment of the present application. The monitoring system for a GPU card provided in the present application includes:

[0061] A request receiving module, used for receiving monitoring requests;

[0062] The service unit acquisition module is used to use the Prometheus monitoring system to obtain the currently running minimum resource service unit;

[0063] A UUID determination module is used to determine the target minimum resource service unit applied by the service corresponding to the monitoring request, and obtain the UUID information of the target GPU card where the target minimum resource service unit is located;

[0064] The information collection module is used to obtain the GPU operation resource information of the target GPU card according to the UUID information; the GPU operation resource information includes at least one of GPU utilization, video memory usage and core ratio.

[0065] Based on the above embodiments, as a preferred embodiment, it also includes:

[0066] A service resource information collection module, used to obtain the CPU resource information and storage resource information of the service;

[0067] The operation information output module is used to output the service operation status information including the CPU resource information, the storage resource information and the GPU operation resource information.

[0068] The present application also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed, the steps provided in the above embodiment can be implemented. The storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0069] The present application also provides an electronic device, which may include a memory and a processor, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps provided in the above embodiment may be implemented. Of course, the electronic device may also include various network interfaces, power supplies and other components.

[0070] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the system provided in the embodiment, since it corresponds to the method provided in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.

[0071] Specific examples are used herein to illustrate the principles and implementation methods of the present application, and the description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

[0072] It should also be noted that, in this specification, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.

Claims

1. A GPU card monitoring method, characterized in that: include: Receive a monitoring request, the monitoring request including a scene and a service name; Using the Prometheus monitoring system to obtain the currently running minimum resource service unit, including: inputting a preset query statement in the Prometheus monitoring system to obtain the currently running minimum resource service unit, including: obtaining the metric of kube_pod_labels of the kube-state-metrics component through the Prometheus monitoring system, and using the metric to obtain the minimum resource service unit of the non-whole card resource service; wherein the non-whole card resource service is used to realize the sharing of GPU cards, and multiple services run on the same GPU card; Determine the target minimum resource service unit applied by the service corresponding to the monitoring request, and obtain the UUID information of the target GPU card where the target minimum resource service unit is located, including: parsing the non-whole card resource service list to determine the service name and scenario included in the service corresponding to the monitoring request; determining all target minimum resource service units corresponding to the service according to the service name and the scenario; determining the unit information of the target minimum resource service unit by using the CoreV1 interface; and determining the UUID information of the target GPU card by analyzing the environment variables of the target minimum resource service unit according to the unit information; Obtain the GPU operating resource information of the target GPU card according to the UUID information; the GPU operating resource information includes at least one of GPU utilization, video memory usage and core share; add the minimum resource service unit GPU share of the same service on the same card to obtain the core share of the service on the card.

2. The monitoring method according to claim 1, characterized in that: Also includes: Obtaining CPU resource information and storage resource information of the service; After obtaining the operating resource information of the target GPU card according to the UUID information, the method further includes: The service operation status information including the CPU resource information, the storage resource information and the GPU operation resource information is output.

3. The monitoring method according to claim 1, characterized in that: Acquiring the GPU operation resource information of the target GPU card according to the UUID information includes: Enter the UUID information and execute the nvidia-smi -L command to determine the correspondence between the UUID information and the GPU card number; Determine the target GPU card according to the GPU card number; The nvidia-smi command is used to collect the GPU operation resource information of the target GPU card.

4. A GPU card monitoring system, characterized in that: include: A request receiving module, used to receive a monitoring request, wherein the monitoring request includes a scene and a service name; A service unit acquisition module is used to use the Prometheus monitoring system to obtain the currently running minimum resource service unit, including: inputting a preset query statement in the Prometheus monitoring system to obtain the currently running minimum resource service unit, including: obtaining the metric of kube_pod_labels of the kube-state-metrics component through the Prometheus monitoring system, and using the metric to obtain the minimum resource service unit of the non-whole card resource service; wherein the non-whole card resource service is used to realize the sharing of GPU cards, and multiple services run on the same GPU card; A UUID determination module is used to determine the target minimum resource service unit applied by the service corresponding to the monitoring request, and obtain the UUID information of the target GPU card where the target minimum resource service unit is located, including: parsing a non-whole card resource service list to determine the service name and scenario included in the service corresponding to the monitoring request; determining all target minimum resource service units corresponding to the service according to the service name and the scenario; determining the unit information of the target minimum resource service unit using the CoreV1 interface; and determining the UUID information of the target GPU card according to the unit information by analyzing the environment variables of the target minimum resource service unit; The information collection module is used to obtain the GPU operation resource information of the target GPU card according to the UUID information; the GPU operation resource information includes at least one of GPU utilization, video memory usage and core ratio; the minimum resource service unit GPU ratio of the same service on the same card is added to obtain the core ratio of the service on the card.

5. The monitoring system according to claim 4, characterized in that: Also includes: A service resource information collection module, used to obtain the CPU resource information and storage resource information of the service; The operation information output module is used to output the service operation status information including the CPU resource information, the storage resource information and the GPU operation resource information.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the GPU card monitoring method according to any one of claims 1 to 3 are implemented.

7. An electronic device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the GPU card monitoring method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Container GPU resource monitoring system in container cluster environment

    CN114281647A