Heterogeneous computing power task monitoring method and device, computer equipment, readable storage medium and program product

By acquiring monitoring index data of heterogeneous computing power tasks, summarizing and aggregating them at the task granularity, and identifying and destroying abnormal tasks, the problem of low utilization of heterogeneous computing power resources was solved, resource utilization was improved and costs were reduced.

CN121880122APending Publication Date: 2026-04-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2024-10-11
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

The utilization rate of heterogeneous computing resources is low, and these resources are wasted due to their occupation. Existing technologies cannot effectively identify and handle abnormal tasks.

Method used

By acquiring monitoring index data during heterogeneous computing power tasks, the index data is summarized and aggregated according to task granularity, and abnormal tasks are identified and eliminated.

Benefits of technology

This improved the utilization rate of heterogeneous computing resources, reduced business computing costs, and achieved cost reduction and efficiency improvement in business operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121880122A_ABST
    Figure CN121880122A_ABST
Patent Text Reader

Abstract

The invention relates to a heterogeneous computing power task monitoring method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the steps of obtaining monitoring index data collected in the process that different computing power containers execute heterogeneous computing power tasks; summarizing the monitoring index data according to the task granularity to obtain aggregated index data corresponding to each heterogeneous computing power task; performing exception identification processing based on the aggregation index data, and determining an exception task in the heterogeneous computing power tasks; and searching, killing and destroying the abnormal task. Monitoring index data of heterogeneous computing power tasks are aggregated according to task granularity, heterogeneous computing power resources occupied by the abnormal tasks are activated through identification and searching, killing and destroying processing of the abnormal tasks, the utilization rate of the heterogeneous computing power resources is increased, and therefore unnecessary expenditure of service computing power cost is reduced, and the service efficiency is improved. And the effects of reducing cost and increasing efficiency of business are comprehensively realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for monitoring heterogeneous computing power tasks. Background Technology

[0002] With the development of computer technology, heterogeneous computing technology has emerged, which involves using different types of processors or computing devices to complete computational tasks. This technology includes, but is not limited to, Central Processing Units (CPUs), Graphics Processing Units (GPUs), Field-Programmable Gate Arrays (FPGAs), and Application-Specific Integrated Circuits (ASICs), each specializing in different computational tasks such as vector operations, matrix calculations, and spatial operations. The advantage of heterogeneous computing lies in its ability to select the most suitable computing device for computation based on the needs of different tasks, thereby improving the efficiency and performance of the entire system. With the rise of large-scale artificial intelligence models, the scale of heterogeneous computing has increased significantly, and the importance of its stability is constantly increasing.

[0003] As for the computing task status of heterogeneous multi-GPU systems, the current judgment is mainly based on the computing status of the business itself. If the status is in the running state, it means that the business is still running and needs to be actively exited before switching to other tasks for execution. If an abnormality occurs during the task execution process, or if the training data is not prepared in time during the business's computing operation, although the computing task is still in the running state, the execution of the computing is invalid. In this case, the heterogeneous computing resources are occupied and wasted, resulting in low utilization of heterogeneous computing resources. Summary of the Invention

[0004] Therefore, it is necessary to provide a heterogeneous computing power task monitoring method, device, computer equipment, computer-readable storage medium, and computer program product that can effectively improve the utilization efficiency of heterogeneous resource computing power and address the above-mentioned technical problems.

[0005] Firstly, this application provides a method for monitoring heterogeneous computing power tasks, including:

[0006] Acquire monitoring index data collected during the execution of heterogeneous computing tasks by different computing power containers;

[0007] The monitoring index data are aggregated and processed according to the task granularity to obtain the aggregated index data corresponding to each heterogeneous computing power task.

[0008] Anomaly identification processing is performed based on the aggregated index data to determine the abnormal tasks in the heterogeneous computing power tasks.

[0009] Perform a detection and destruction process on the abnormal task.

[0010] Secondly, this application also provides a heterogeneous computing power task monitoring device, comprising:

[0011] The indicator acquisition module is used to acquire monitoring indicator data collected during the execution of heterogeneous computing tasks by different computing power containers;

[0012] The indicator aggregation module is used to summarize and process the monitoring indicator data according to the task granularity to obtain the aggregated indicator data corresponding to each heterogeneous computing power task.

[0013] An abnormal task identification module is used to perform abnormal identification processing based on the aggregated index data to identify abnormal tasks in the heterogeneous computing power tasks.

[0014] The abnormal task detection and removal module is used to perform detection, removal, and destruction processing on the abnormal tasks.

[0015] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0016] Acquire monitoring index data collected during the execution of heterogeneous computing tasks by different computing power containers;

[0017] The monitoring index data are aggregated and processed according to the task granularity to obtain the aggregated index data corresponding to each heterogeneous computing power task.

[0018] Anomaly identification processing is performed based on the aggregated index data to determine the abnormal tasks in the heterogeneous computing power tasks.

[0019] Perform a detection and destruction process on the abnormal task.

[0020] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0021] Acquire monitoring index data collected during the execution of heterogeneous computing tasks by different computing power containers;

[0022] The monitoring index data are aggregated and processed according to the task granularity to obtain the aggregated index data corresponding to each heterogeneous computing power task.

[0023] Anomaly identification processing is performed based on the aggregated index data to determine the abnormal tasks in the heterogeneous computing power tasks.

[0024] Perform a detection and destruction process on the abnormal task.

[0025] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:

[0026] Acquire monitoring index data collected during the execution of heterogeneous computing tasks by different computing power containers;

[0027] The monitoring index data are aggregated and processed according to the task granularity to obtain the aggregated index data corresponding to each heterogeneous computing power task.

[0028] Anomaly identification processing is performed based on the aggregated index data to determine the abnormal tasks in the heterogeneous computing power tasks.

[0029] Perform a detection and destruction process on the abnormal task.

[0030] The aforementioned heterogeneous computing power task monitoring method, device, computer equipment, computer-readable storage medium, and computer program product acquire monitoring indicator data collected during the execution of heterogeneous computing power tasks by different computing power containers. This data forms the basis for heterogeneous computing power task monitoring, analysis, and processing. The monitoring indicator data is then aggregated at the task granularity to obtain aggregated indicator data corresponding to each heterogeneous computing power task. This aggregates monitoring indicator data collected from different computing power containers and devices under the name of the heterogeneous computing power task for analysis. Anomaly identification processing is performed based on the aggregated indicator data to identify abnormal tasks within the heterogeneous computing power tasks. These abnormal tasks are then detected and destroyed. By confirming and destroying abnormal tasks during execution, the occupied heterogeneous card resources are released, improving the utilization rate of heterogeneous computing power resources. This application aggregates monitoring index data of heterogeneous computing power tasks by task granularity, and revitalizes the heterogeneous computing power resources occupied by these abnormal tasks by identifying, detecting and destroying abnormal tasks, thereby improving the utilization rate of heterogeneous computing power resources, reducing unnecessary expenditures on business computing power costs, and comprehensively achieving the effect of reducing business costs and increasing efficiency. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is an application environment diagram of a heterogeneous computing power task monitoring method in one embodiment;

[0033] Figure 2 This is a flowchart illustrating a heterogeneous computing power task monitoring method in one embodiment;

[0034] Figure 3 This is a schematic diagram of the field names included in a database table design in one embodiment;

[0035] Figure 4 This is a logic block diagram of the process for identifying abnormal tasks in one embodiment;

[0036] Figure 5 This is a schematic diagram illustrating how heterogeneous computing power task monitoring provides support for model-related computational tasks in one embodiment.

[0037] Figure 6 This is a flowchart of a heterogeneous computing power task monitoring method in one embodiment;

[0038] Figure 7 This is a flowchart illustrating the abnormal task identification and evaluation process in one embodiment;

[0039] Figure 8 This is a flowchart illustrating the heterogeneous computing power task monitoring method in another embodiment;

[0040] Figure 9 This is a structural block diagram of a heterogeneous computing power task monitoring device in one embodiment;

[0041] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0043] The heterogeneous computing power task monitoring method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, monitoring server 102 communicates with heterogeneous computing server 104 via a network. A data storage system can store the data that the heterogeneous computing server 104 needs to process. The data storage system can be integrated onto the heterogeneous computing server 104, or it can be placed in the cloud or on other network servers. Monitoring server 102 can collect various indicators from the heterogeneous computing server 104. When monitoring computing tasks, monitoring server 102 first obtains monitoring indicator data collected during the execution of heterogeneous computing tasks by different computing containers from the heterogeneous computing server 104; it then summarizes the monitoring indicator data according to task granularity to obtain aggregated indicator data corresponding to each heterogeneous computing task; based on the aggregated indicator data, it performs anomaly identification processing to determine abnormal tasks in the heterogeneous computing tasks; and it performs detection and destruction processing on the abnormal tasks to achieve control over the heterogeneous computing server 104. Monitoring server 102 and heterogeneous computing server 104 can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers providing cloud computing services.

[0044] In one exemplary embodiment, such as Figure 2 As shown, a method for monitoring heterogeneous computing power tasks is provided, which can be applied to... Figure 1 Taking monitoring server 102 as an example, the explanation includes the following steps 201 to 207. Wherein:

[0045] Step 201: Obtain monitoring index data collected during the execution of heterogeneous computing tasks by different computing power containers.

[0046] In this context, a computing power container refers to a container that executes heterogeneous computing power tasks. A container is a lightweight and portable virtualization technology that packages an application and its dependencies into a single, portable unit. The container contains the necessary elements for the successful execution of the application, such as code, environment variables, processes, runtime environment, and software dependencies. Similar to shipping containers used in the transportation industry to isolate different goods for transport, containers isolate applications and their dependencies to run in different environments. In practice, an environment for executing heterogeneous computing power tasks can be deployed within the computing power container, enabling it to perform these tasks. Heterogeneous computing power tasks refer to computational tasks implemented using heterogeneous computing power devices, such as vector operations, matrix calculations, and spatial operations. The advantage of heterogeneous computing power lies in its ability to select the most suitable computing device for computation based on the needs of different tasks, thereby improving the efficiency and performance of the entire system. In specific embodiments, the heterogeneous computing power tasks of this application can be training tasks for various types of machine learning models, such as mixed-mode models, visual models, speech models, game models, medical models, and natural language processing models. At this point, heterogeneous computing power tasks include model pre-training, model fine-tuning, and model enhancement. Monitoring metrics data are collected through agent clients deployed on each computing power container. The collected data includes monitoring metrics representing the heterogeneous computing status during task execution. This data, along with task information, is stored in a database. During monitoring and analysis, the monitoring metrics data collected during the execution of heterogeneous computing power tasks by different computing power containers can be directly retrieved from the database. Specifically, the collected monitoring metrics data can include various utilization information and communication bandwidth information.

[0047] For example, during the execution of heterogeneous computing tasks by the heterogeneous computing server 104, it may be contaminated by residual dirty data tasks, thereby affecting the efficiency of heterogeneous computing power flow. Evaluating the computational status of heterogeneous computing tasks is an effective way to identify dirty data tasks. By identifying and eliminating dirty data tasks, the performance of the entire heterogeneous computing system can be effectively improved. To monitor and evaluate the computational status of heterogeneous computing tasks, it is necessary to monitor the device and task status during the execution of the heterogeneous computing tasks and obtain the data collected by the computing container during the execution of the heterogeneous computing tasks as the basis for analysis. In a specific embodiment, a data collection agent client can be pre-deployed on the computing container executing the heterogeneous computing tasks. Then, during the execution of the heterogeneous computing tasks, these agent clients can collect various types of monitoring indicator data, which will be transmitted to the monitoring server 102 via the network. The monitoring server 102 then performs specific monitoring data analysis based on this monitoring indicator data. In a specific embodiment, this application is applied to the status monitoring of various machine learning model pre-training processes. At this point, after a new model training task is issued, the model training task can be split into various containers for execution, and monitoring metrics can be collected during the model pre-training process.

[0048] Step 203: Summarize and process the monitoring index data according to the task granularity to obtain the aggregated index data corresponding to each heterogeneous computing power task.

[0049] Among these, summarizing monitoring indicator data by task granularity refers to aggregating monitoring indicator data from different containers and devices under the same task within the collected monitoring indicator data for analysis. Aggregated indicator data is then used to characterize various types of monitoring indicator data for a single heterogeneous computing power task.

[0050] For example, since heterogeneous computing tasks vary in scale, some large computing tasks may be deployed on different computing containers and devices. Since the objective of this application is to identify dirty data tasks within heterogeneous computing tasks, it is necessary to aggregate the indicator data of the same heterogeneous computing task for analysis in order to identify tasks with abnormal dirty data. The result is aggregated indicator data. During the aggregation process, each monitoring indicator data can be aggregated individually, for example, by averaging, to obtain aggregated indicator data for a single heterogeneous computing task. In one embodiment, the field names included in the database table design for storing the monitoring indicator data can be specifically referred to... Figure 3As shown, "Time" refers to the point in time when heterogeneous computing power metrics are collected. This time-based identification avoids overlap in metric collection. "Task Name" refers to the name of the heterogeneous computing component; this metric is essential for the coordination of multi-dimensional metrics. "Container Name" refers to the name of the heterogeneous computing power container, identifying each container within the task. "Container Startup Time" refers to the time the container starts running; this time also represents the task's start and running time, and its duration can serve as an auxiliary indicator for assessing task status. The five-tuple information represents the five monitored metrics collected. During data aggregation, the five-tuple information for the same task can be aggregated based on the task name in this table to obtain the aggregated metric data.

[0051] Step 205: Perform anomaly identification processing based on aggregated index data to identify abnormal tasks in heterogeneous computing power tasks.

[0052] Step 207: Perform a detection and removal process on the abnormal task.

[0053] Anomaly identification and processing refers to identifying heterogeneous computing tasks with abnormal computational states based on monitoring indicator data. These abnormal tasks are specifically dirty data tasks, which waste the heterogeneous computing resources they occupy, thereby affecting the efficiency of heterogeneous computing resource circulation. By detecting and destroying these abnormal tasks, it is beneficial to the overall circulation of computing power in the heterogeneous resource pool and to optimize the development quality of heterogeneous computing tasks.

[0054] For example, by aggregating and summarizing, aggregated indicator data corresponding to each heterogeneous computing task is obtained. Anomaly analysis can then be performed on these aggregated indicator data. The anomaly analysis process can be discussed in different cases, such as identifying anomalies in single-card tasks and multi-card tasks. The identification process mainly analyzes these heterogeneous computing tasks based on the aggregated indicator data obtained from the database, thereby outputting a list of anomalous tasks. In a specific embodiment, a fallback utilization threshold can be used to filter out dirty tasks with excessively low utilization. Then, tasks with utilization rates above the threshold are further analyzed based on their specific circumstances. For example, the analysis can be performed separately for single-card tasks and multi-card tasks. For single-card tasks, dirty data tasks can be identified based on the utilization rate of the single card's computing unit and memory utilization. For multi-card tasks, the analysis is based on the communication bandwidth of the heterogeneous cards. After identifying anomalous tasks, these anomalous tasks are also saved and then aggregated for further processing and destruction.

[0055] The aforementioned heterogeneous computing task monitoring method acquires monitoring indicator data collected during the execution of heterogeneous computing tasks by different computing power containers. This data forms the basis for heterogeneous computing task monitoring and analysis. The monitoring indicator data is then aggregated at the task granularity to obtain aggregated indicator data for each heterogeneous computing task. This aggregated data from different computing power containers and devices is analyzed under the name of the heterogeneous computing task. Anomaly identification is performed based on the aggregated indicator data to identify abnormal tasks within the heterogeneous computing task pool. These abnormal tasks are then eliminated and destroyed. By confirming and eliminating abnormal tasks during execution, the occupied heterogeneous card resources are released, improving the utilization rate of heterogeneous computing resources. This application aggregates monitoring indicator data for heterogeneous computing tasks at the task granularity and, through the identification and elimination of abnormal tasks, revitalizes the heterogeneous computing resources occupied by these abnormal tasks, improving their utilization rate and reducing unnecessary expenditures on business computing costs, thus achieving a comprehensive effect of cost reduction and efficiency improvement.

[0056] In an exemplary embodiment, the method further includes: deploying metric acquisition software on each computing power container; collecting task information and monitoring metric data of heterogeneous computing power tasks executed by the computing power containers in a timed polling manner using the metric acquisition software; and summarizing the task information and monitoring metric data into an metric data table. Step 201 includes: reading the monitoring metric data collected during the execution of heterogeneous computing power tasks by different computing power containers from the metric data table.

[0057] The indicator acquisition software refers to the agent client program deployed within the computing power container, which can collect various monitoring indicator data required for subsequent analysis during the execution of heterogeneous computing power tasks. Regarding the handling of timed polling, timed polling is a common communication mechanism in which the client sends requests to the server at predetermined time intervals, and the server returns the corresponding data upon receiving the requests. This mechanism is often used in application scenarios that require periodic data updates. Specifically, this application can implement the monitoring of heterogeneous computing power tasks in a periodic manner. In this case, within each period, the server responsible for monitoring and analysis will access the indicator acquisition software through timed polling to collect the monitoring indicator data for the current period for analysis. The task information and monitoring indicator data summarized in the indicator data table can be found in [reference needed]. Figure 3 As shown.

[0058] For example, this application can collect monitoring indicator data from various computing power containers by deploying indicator collection software. Before data collection, the pre-written indicator collection software can be deployed to each computing power container in the system. When the computing power containers run heterogeneous computing power tasks, this indicator collection software will perform the corresponding data collection function. The server responsible for monitoring and analysis will request data from the indicator collection software on the computing power containers in a periodic polling manner according to the monitoring and analysis cycle. The indicator collection software will return the task information of the heterogeneous computing power tasks currently being executed on the computing power containers, and also return the collected monitoring indicator data. This data will form a data structure such as... Figure 3 The same data structure is used, and the data is aggregated into a database table for analysis. During analysis, the server reads monitoring indicator data collected during the execution of heterogeneous computing tasks by different computing containers directly from the database's indicator data table. And for... Figure 3 In this scenario where monitoring indicator data is stored according to a data table, step 203 specifically includes: classifying the monitoring indicator data based on task information to obtain monitoring indicator data for different tasks; and aggregating the monitoring indicator data for different tasks to obtain aggregated indicator data corresponding to each heterogeneous computing power task. That is, for the indicator data table stored in the database, the monitoring indicator data is classified based on the preceding task information, thereby grouping monitoring indicator data for the same task together to obtain monitoring indicator data for different tasks. Then, these monitoring indicator data for different tasks are aggregated to obtain aggregated indicator data corresponding to each heterogeneous computing power task. This application can aggregate indicators based on task names and then perform evaluation and judgment based on the aggregated indicator values. In this embodiment, by deploying indicator collection software, various types of indicator data are collected in a timed polling manner, thereby effectively ensuring the accuracy and efficiency of data collection, improving the success rate of identifying abnormal tasks, and ensuring monitoring effectiveness.

[0059] In an exemplary embodiment, step 205 includes: determining the heterogeneous computing power utilization rate of each heterogeneous computing power task based on aggregated index data; and identifying heterogeneous computing power tasks with heterogeneous computing power utilization rates lower than the utilization rate threshold as abnormal tasks.

[0060] For example, in the process of anomaly identification, the heterogeneous computing power utilization rate of each heterogeneous computing power task can be determined based on aggregated indicator data, and anomaly identification can be performed based on a preset utilization rate threshold. This utilization rate threshold is a value that the business computing needs to achieve as a safety net. The purpose of setting this value is to ensure the basic use of heterogeneous resources. This utilization rate threshold can be set by the operation and maintenance personnel of the monitored heterogeneous computing power server platform based on the cluster being operated. According to the set threshold, if it is lower than the threshold, the task name is recorded in the list of abnormal tasks; if the heterogeneous computing power utilization rate is higher than the threshold, it needs to be analyzed in conjunction with other types of aggregated indicator data. In a specific embodiment, for the process of determining the heterogeneous computing power utilization rate of each heterogeneous computing power task based on aggregated indicator data, the heterogeneous card utilization rate parameter corresponding to each heterogeneous computing power task can be determined based on the aggregated indicator data; the heterogeneous card utilization rate parameter corresponding to each heterogeneous computing power task is averaged to determine the heterogeneous computing power utilization rate of each heterogeneous computing power task. By averaging, the overall processing status of current tasks can be effectively analyzed, thereby filtering out heterogeneous computing tasks whose overall utilization does not meet processing requirements. For example, in a specific embodiment, heterogeneous computing task A is allocated to containers 1, 2, and 3 for execution. The heterogeneous card utilization rate of the subtask executed on container 1 is 45%, that on container 2 is 52%, and that on container 3 is 38%. Therefore, the corresponding heterogeneous computing power utilization rate for task A is (45% + 52% + 38%) / 3 = 45%. In this embodiment, based on a preset utilization threshold, abnormal heterogeneous computing power tasks are filtered out, thus eliminating tasks whose utilization rate fails to meet basic business requirements. This ensures efficient subsequent processing by effectively identifying abnormal tasks.

[0061] In an exemplary embodiment, step 205 further includes: identifying the task type of the heterogeneous computing power task when the heterogeneous computing power utilization rate of the heterogeneous computing power task is equal to or higher than the utilization rate threshold; and identifying abnormal tasks in the heterogeneous computing power task based on aggregated index data and task type.

[0062] For example, the task types of heterogeneous computing power tasks can specifically include single-card tasks and multi-card tasks. A single-card task is a heterogeneous computing task implemented using a single heterogeneous card, while a multi-card task requires multiple heterogeneous computing cards to work together for computation. When the heterogeneous computing power utilization rate of a heterogeneous computing power task is equal to or higher than the utilization rate threshold, abnormal tasks can be identified and processed by combining the task type and aggregated indicator data.

[0063] For heterogeneous computing tasks involving single-card architectures, after identifying the single-card tasks within the heterogeneous computing task landscape, the utilization rate of computing units and memory for each single-card task can be determined based on aggregated metric data. Single-card tasks with computing unit utilization or memory utilization below a threshold are identified as anomalous tasks. Computing unit utilization indicates the utilization rate of computing units, while memory utilization represents the memory utilization rate on the heterogeneous card. When evaluating single-card tasks, if the aggregated metric data indicates that the utilization rate of the card's heterogeneous computing units or memory is very low, for example, below 5%, it indicates that the task is not actually performing heterogeneous computing and is considered a "dirty data" task, thus classifying it as an anomalous task.

[0064] For the anomaly analysis of multi-GPU tasks, after identifying multi-GPU tasks within heterogeneous computing power tasks, the communication bandwidth between each multi-GPU task is first determined based on aggregated indicator data. Then, multi-GPU tasks with communication bandwidth below a certain threshold are identified as anomalous tasks. Specifically, communication bandwidth can include NVLink communication traffic and RDMA (Remote Direct Memory Access) communication traffic. NVLink communication traffic refers to the NVLink communication bandwidth between heterogeneous GPUs, while RDMA communication traffic refers to the high-speed RDMA communication bandwidth between heterogeneous devices. In other words, for multi-GPU heterogeneous computing tasks—that is, tasks running on multiple heterogeneous GPUs (possibly on the same machine or across different machines)—the evaluation criteria are based on the NVLink communication bandwidth or RDMA network bandwidth between GPUs. If the bandwidth value is very low, for example, bandwidth utilization is below 5%, it indicates that the multi-GPU computing is in an abnormal state and needs to be marked as a dirty task. Therefore, multi-GPU tasks with communication bandwidth below the threshold are identified as anomalous tasks. In one embodiment, the overall process for identifying abnormal tasks can be referred to Figure 4As shown. By utilizing heterogeneous card utilization and combining it with subsequent single-card and multi-card task analysis, various types of heterogeneous computing tasks can be accurately identified. In this embodiment, heterogeneous computing tasks are classified according to specific computing task conditions, and single-card and multi-card tasks are analyzed separately based on the classification results to identify heterogeneous computing power tasks that exhibit abnormalities under each type, thereby effectively improving the accuracy of abnormal task identification. In summary, this application integrates abnormal tasks in the running state into an automated processing flow, alleviating the manpower input of heterogeneous computing power operation and maintenance personnel, avoiding the manual process of screening and cleaning abnormal computing tasks, and helping to reduce operation and maintenance costs. From an operational perspective, this application evaluates heterogeneous computing tasks by aggregating multi-dimensional computing power indicators and automates the execution of processing, which is conducive to the overall flow of computing power in the heterogeneous resource pool and helps to optimize the development quality of heterogeneous computing tasks.

[0065] In an exemplary embodiment, the method further includes: obtaining a model computing task processing request; parsing the model computing task processing request to obtain a model computing task to be processed; splitting the model computing task to obtain a heterogeneous computing power task; and allocating the heterogeneous computing power task to different computing power containers.

[0066] For example, this application can be applied to the import process of heterogeneous computing tasks in various types of model computing task processing, such as model pre-training, model fine-tuning, and model enhancement. Upon receiving a corresponding model computing task processing request, the request can be parsed to determine the current model computing task to be processed, and then processed according to the nature of the task. If the model computing task is small, a corresponding single-card type heterogeneous computing task can be directly generated and allocated to a single heterogeneous card in the computing power container for processing. If the computing scale is large, the model computing task can be split as needed to obtain corresponding heterogeneous computing tasks. These tasks can be allocated to the same computing power container and executed by multiple heterogeneous cards within the container, or they can be allocated to different containers and executed by heterogeneous cards in different containers. The specific process can be referred to... Figure 5As shown, for models such as hybrid models, visual models, speech models, game models, medical models, and natural language processing models at the product layer, the computational tasks they generate are imported into computing power containers for monitoring and collaborative evaluation. By collecting and aggregating indicators of the heterogeneous computational tasks executed in the computing power containers, abnormal tasks can be identified using monitoring indicator data. These abnormal tasks are then eliminated and destroyed, ensuring efficiency and accuracy. In one embodiment, this application is used to monitor the training tasks of natural language processing models. In this case, the corresponding initial natural language processing model needs to be deployed within the container first. Then, various types of text data for model training are acquired, and different batches of model training samples are constructed based on this text data, generating model training computational tasks corresponding to each container. The model training samples are then input into different initial natural language processing models for model training. The combined model training computational tasks executed by multiple containers constitute the processing task specified in the model computational task processing request. In this embodiment, by splitting and allocating the received model computation tasks, the computation tasks to be executed can be effectively distributed to various computing power containers for processing, ensuring the efficiency and accuracy of computation task allocation and processing.

[0067] In an exemplary embodiment, step 207 includes: generating an abnormal task data table corresponding to the abnormal task; saving the abnormal task data table to the task database; querying the task database in a polling manner based on a preset task processing cycle to obtain the abnormal task data table within the current task processing cycle; and performing detection and destruction processing on the abnormal tasks within the current task processing cycle according to the abnormal task data table.

[0068] For example, the process of detecting and destroying abnormal tasks can begin by organizing and analyzing the data of these abnormal tasks, and then searching for and destroying these abnormal tasks according to the processing cycle, processing multiple abnormal tasks at once to improve processing efficiency. First, an abnormal task data table needs to be generated for each abnormal task. If the abnormal task data table has already been generated, it can be updated based on the information of the abnormal task. The abnormal task data table needs to record task information including the time the task was marked as an abnormal task, the task name, and the time the task started running from the start timer. The abnormal task data table is saved to the task database. The cycled processing specifically refers to querying the task database in a polling manner at fixed time nodes in each task processing cycle, such as the node about to switch to the next task processing cycle, to obtain the abnormal task data table for the current task processing cycle. Then, the abnormal tasks recorded in the abnormal task data table are detected and destroyed. In one specific embodiment, the process of detecting and destroying abnormal tasks requires a final list confirmation by the operations and maintenance personnel responsible for task management to prevent certain important personnel from being eliminated. Before performing the detection and destruction process on the abnormal task data table within the current task processing cycle, a page generation process can be performed based on the abnormal task information in the abnormal task data table to obtain a confirmation page. The confirmation page can include the task name of the identified abnormal task and various related information, and then be pushed to the operations and maintenance personnel. Upon receiving the data table modification message based on the confirmation page, the operations and maintenance personnel can directly update the abnormal task data table to obtain an updated task data table. Based on the updated task data table, the detection and destruction process is then performed on the abnormal tasks within the current task processing cycle. This auxiliary confirmation method improves the accuracy of the detection and destruction process. In this embodiment, by generating an abnormal task data table and then implementing the detection and destruction of abnormal tasks in a polling manner, the efficiency and accuracy of the detection and destruction process can be effectively improved.

[0069] In one embodiment, the complete steps of the heterogeneous computing power task monitoring method of this application include: acquiring model computing task processing requests; parsing model computing task processing requests to obtain model computing tasks to be processed; splitting the model computing tasks to obtain heterogeneous computing power tasks, and allocating the heterogeneous computing power tasks to different computing power containers. Deploying indicator acquisition software on each computing power container; collecting task information and monitoring indicator data of the heterogeneous computing power tasks executed by the computing power containers in a timed polling manner using the indicator acquisition software; summarizing the task information and monitoring indicator data into an indicator data table; reading the monitoring indicator data collected during the execution of heterogeneous computing power tasks by different computing power containers from the indicator data table. Based on the task information, classifying the monitoring indicator data to obtain monitoring indicator data for different tasks; aggregating the monitoring indicator data for different tasks to obtain aggregated indicator data corresponding to each heterogeneous computing power task. Based on the aggregated indicator data, determining the heterogeneous card utilization parameter corresponding to each heterogeneous computing power task; averaging the heterogeneous card utilization parameter corresponding to each heterogeneous computing power task to determine the heterogeneous computing power utilization rate of each heterogeneous computing power task. Heterogeneous computing power tasks with utilization rates below a utilization threshold are identified as anomalous tasks. When the heterogeneous computing power utilization rate of a heterogeneous computing power task is equal to or higher than the utilization threshold, the task type of the heterogeneous computing power task is identified; single-card tasks within the heterogeneous computing power task are identified; based on aggregated indicator data, the computing unit utilization rate and memory utilization rate of each single-card task are determined; single-card tasks with computing unit utilization rate or memory utilization rate below the utilization threshold are identified as anomalous tasks. Multi-card tasks within the heterogeneous computing power task are identified; based on aggregated indicator data, the communication bandwidth traffic between various multi-card tasks is determined; multi-card tasks with communication bandwidth traffic below the traffic threshold are identified as anomalous tasks. Generate an abnormal task data table corresponding to the abnormal tasks; save the abnormal task data table to the task database; query the task database in a polling manner based on a preset task processing cycle to obtain the abnormal task data table within the current task processing cycle; generate a page based on the abnormal task information in the abnormal task data table to obtain a detection and removal confirmation page; push the detection and removal confirmation page; upon receiving a data table modification message based on the detection and removal confirmation page, update the abnormal task data table to obtain an updated task data table; and perform detection and removal processing on the abnormal tasks within the current task processing cycle based on the updated task data table.

[0070] This application also provides an application scenario, which is illustrated by taking the above-mentioned heterogeneous computing power task monitoring method as an example. The heterogeneous computing power task monitoring method specifically includes:

[0071] When users build a heterogeneous computing platform to handle various types of machine learning model-related computational tasks, some of these tasks are abnormal due to the presence of dirty data. For example, if the business computation fails or training data is not prepared in time, the computation, though still running, becomes invalid. This results in wasted heterogeneous computing resources, with the cost still borne by the business – a lose-lose situation for both. Traditional solutions cannot automate scenarios involving abnormal environments or runtime states in heterogeneous computing tasks. Therefore, the heterogeneous computing task monitoring system described in this application collects various monitoring metrics during task execution and performs overall data analysis based on these metrics.

[0072] This solution focuses on constructing multi-dimensional monitoring indicators for collaborative heterogeneous computing power to assess whether the status of heterogeneous computing services is abnormal. The key aspect is to eliminate and destroy tasks with dirty data to revitalize occupied heterogeneous resources. The flowchart is as follows: Figure 6 As shown, heterogeneous computing tasks are broken down into multiple heterogeneous computing containers for execution. Afterward, metrics are collected. Specifically, a collection agent client program is deployed on each computing container to perform periodic polling and collection of metrics. In addition to monitoring metric data, the collected metrics also include task information from the computing containers. This information is used to summarize the metrics at the task granularity. The collected metric data is then stored in a database table.

[0073] Once the collected metrics are aggregated into the database, a periodically triggered evaluation program can perform assessments and filtering based on the metric data recorded in the database, outputting a list of abnormal tasks. The workflow of the evaluation program can be found in [reference needed]. Figure 7As shown, after retrieving the metric data from the database table, the metric can be aggregated directly based on the task name corresponding to the metric data, grouping metrics with the same task name together. Then, evaluation and judgment are performed based on the aggregated metric values. First, the platform operations and maintenance personnel will configure the utilization threshold. The purpose of setting this value is to ensure the basic use of heterogeneous resources. This value can be evaluated and set by the platform operations and maintenance personnel based on the cluster being operated. Then, based on the set thresholds, the card utilization of each heterogeneous computing task is evaluated. If the card utilization is lower than the threshold, the task name is recorded in the dirty task list. If the card utilization of a heterogeneous computing task is higher than the threshold, the resource specifications of the heterogeneous computing task are evaluated separately. For a single-card heterogeneous computing task, i.e., the heterogeneous computing task runs on a single heterogeneous card, the utilization of the heterogeneous computing unit or memory of that card is evaluated to see if it is very low, for example, below 5%. This indicates that the task is not actually performing heterogeneous computing and is considered a dirty data task. For a multi-card heterogeneous computing task, i.e., the heterogeneous computing task runs on multiple cards, which may be on the same machine or across multiple machines, the evaluation indicators need to be performed through the inter-card NVLink communication bandwidth value or RDMA computing network bandwidth value. If the bandwidth value is very low, for example, the bandwidth utilization is below 5%, this indicates that the multi-card computing is in an abnormal state and needs to be marked as a dirty task. Once a task is identified as dirty, it is recorded in the database for removal and subsequent backtracking analysis. The task information recorded in the database table includes the time the task was marked as dirty, the task name, and the time the task started running from its inception timer. An abnormal task data table is saved to the task database. Based on a preset task processing cycle, the task database is queried in a polling manner to obtain the abnormal task data table for the current processing cycle. Based on the abnormal task data table, abnormal tasks within the current processing cycle are removed and destroyed. During this process, operations personnel can also manually confirm the final generated abnormal task list. Based on the abnormal task information in the abnormal task data table, a confirmation page is generated and pushed to the system. Upon receiving a data table modification message based on the confirmation page, the abnormal task data table is updated to obtain an updated task data table. The updated task data table is then used to search for abnormal tasks, ensuring the accuracy of anomaly identification and task removal.

[0074] In one embodiment, the heterogeneous computing power task monitoring method of this application can be specifically referred to Figure 8As shown, after the platform operations and maintenance personnel begin implementing the solution, they first need to configure the minimum utilization threshold required for heterogeneous card computing. This threshold is derived from the operations and maintenance personnel's assessment based on the cluster's operational status. Then, using a round-robin approach, they read the metric data from the database at the task granularity, aggregate it to obtain the task utilization, and then perform a judgment between the utilization and the threshold. If the utilization is below the threshold, it is marked as dirty data, and the task is then entered into the database table for round-robin detection and destruction. If the utilization is above the threshold, it is determined whether the task is single-card or multi-card. If it is single-card, it is checked whether the utilization of the computing unit or video memory is very low; if so, it is marked as dirty data. If it is a multi-card task, it is checked whether the communication bandwidth of NVLink or RDMA is very low; if so, the marking operation is performed. If the metrics for single-card or multi-card tasks are normal, the process of reading the task's metric data from the database table in a round-robin manner continues, performing the task status assessment process. Finally, the identified abnormal tasks are recorded in the data table, and dirty data tasks are detected and destroyed using a round-robin approach.

[0075] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0076] Based on the same inventive concept, this application also provides a heterogeneous computing power task monitoring device for implementing the heterogeneous computing power task monitoring method described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more heterogeneous computing power task monitoring device embodiments provided below can be found in the limitations of the heterogeneous computing power task monitoring method above, and will not be repeated here.

[0077] In one exemplary embodiment, such as Figure 9 As shown, a heterogeneous computing power task monitoring device is provided, comprising:

[0078] The indicator acquisition module 902 is used to acquire monitoring indicator data collected during the execution of heterogeneous computing tasks by different computing power containers.

[0079] The indicator aggregation module 904 is used to summarize and process the monitoring indicator data according to the task granularity to obtain the aggregated indicator data corresponding to each heterogeneous computing power task.

[0080] The abnormal task identification module 906 is used to perform abnormal identification processing based on aggregated index data to identify abnormal tasks in heterogeneous computing power tasks.

[0081] The abnormal task detection and removal module 908 is used to perform detection, removal, and destruction of abnormal tasks.

[0082] The aforementioned heterogeneous computing power task monitoring device acquires monitoring indicator data collected during the execution of heterogeneous computing power tasks by different computing power containers. This data forms the basis for heterogeneous computing power task monitoring and analysis. The monitoring indicator data is then aggregated at the task granularity to obtain aggregated indicator data for each heterogeneous computing power task. This aggregates monitoring indicator data collected from different computing power containers and devices under the heterogeneous computing power task name for analysis. Anomaly identification is performed based on the aggregated indicator data to identify abnormal tasks within the heterogeneous computing power tasks. These abnormal tasks are then eliminated and destroyed. By confirming and eliminating abnormal tasks during execution, the occupied heterogeneous card resources are released, improving the utilization rate of heterogeneous computing power resources. This application aggregates monitoring indicator data of heterogeneous computing power tasks at the task granularity and, through the identification and elimination of abnormal tasks, revitalizes the heterogeneous computing power resources occupied by these abnormal tasks, improving the utilization rate of heterogeneous computing power resources and reducing unnecessary expenditures on business computing power costs, thereby achieving a comprehensive effect of cost reduction and efficiency improvement.

[0083] In one embodiment, a data acquisition module is further included, used for: deploying indicator acquisition software on each computing power container; collecting task information and monitoring indicator data of heterogeneous computing power tasks executed by the computing power containers in a timed polling manner through the indicator acquisition software; and summarizing the task information and monitoring indicator data into an indicator data table. The indicator acquisition module 902 is specifically used in the indicator data table to read the monitoring indicator data collected during the execution of heterogeneous computing power tasks by different computing power containers.

[0084] In one embodiment, the indicator aggregation module 904 is configured to: classify the monitoring indicator data based on task information to obtain monitoring indicator data for different tasks; and aggregate the monitoring indicator data for different tasks to obtain aggregated indicator data corresponding to each heterogeneous computing power task.

[0085] In one embodiment, the abnormal task identification module 906 is used to: determine the heterogeneous computing power utilization rate of each heterogeneous computing power task based on aggregated index data; and identify heterogeneous computing power tasks with heterogeneous computing power utilization rates lower than the utilization rate threshold as abnormal tasks.

[0086] In one embodiment, the abnormal task identification module 906 is used to: determine the heterogeneous card utilization parameter corresponding to each heterogeneous computing power task based on aggregated index data; and perform averaging processing on the heterogeneous card utilization parameter corresponding to each heterogeneous computing power task to determine the heterogeneous computing power utilization rate of each heterogeneous computing power task.

[0087] In one embodiment, the abnormal task identification module 906 is configured to: identify the task type of the heterogeneous computing power task when the heterogeneous computing power utilization rate of the heterogeneous computing power task is equal to or higher than the utilization rate threshold; and identify abnormal tasks in the heterogeneous computing power task based on aggregated index data and task type.

[0088] In one embodiment, the abnormal task identification module 906 is used to: identify single-card tasks in heterogeneous computing power tasks; determine the computing unit utilization rate and memory utilization rate of each single-card task based on aggregated index data; and identify single-card tasks with computing unit utilization rate or memory utilization rate lower than the utilization rate threshold as abnormal tasks.

[0089] In one embodiment, the abnormal task identification module 906 is used to: identify multi-card tasks in heterogeneous computing power tasks; determine the communication bandwidth traffic between each multi-card task based on aggregated index data; and identify multi-card tasks whose communication bandwidth traffic is lower than the traffic threshold as abnormal tasks.

[0090] In one embodiment, the system further includes a task allocation module, configured to: obtain a model computing task processing request; parse the model computing task processing request to obtain a model computing task to be processed; split the model computing task to obtain a heterogeneous computing power task; and allocate the heterogeneous computing power task to different computing power containers.

[0091] In one embodiment, the abnormal task detection and elimination module 908 performs the following steps: generates an abnormal task data table corresponding to the abnormal task; saves the abnormal task data table to the task database; queries the task database in a polling manner based on a preset task processing cycle to obtain the abnormal task data table within the current task processing cycle; and performs detection and elimination processing on the abnormal tasks within the current task processing cycle according to the abnormal task data table.

[0092] In one embodiment, the abnormal task detection and elimination module 908: generates a page based on the abnormal task information in the abnormal task data table to obtain a detection and elimination confirmation page; pushes the detection and elimination confirmation page; upon receiving a data table modification message based on the detection and elimination confirmation page, updates the abnormal task data table to obtain an updated task data table; and performs detection and elimination destruction processing on abnormal tasks within the current task processing cycle according to the updated task data table.

[0093] Each module in the aforementioned heterogeneous computing power task monitoring device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0094] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10 The computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores data related to heterogeneous computing task monitoring. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a heterogeneous computing task monitoring method.

[0095] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0096] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0097] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0098] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the steps in the above method embodiments.

[0099] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0100] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0101] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0102] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for monitoring heterogeneous computing power tasks, characterized in that, The method includes: Acquire monitoring index data collected during the execution of heterogeneous computing tasks by different computing power containers; The monitoring index data are aggregated and processed according to the task granularity to obtain the aggregated index data corresponding to each heterogeneous computing power task. Anomaly identification processing is performed based on the aggregated index data to determine the abnormal tasks in the heterogeneous computing power tasks. Perform a detection and removal process on the abnormal task.

2. The method according to claim 1, characterized in that, The method further includes: Deploy metrics collection software on each computing container; The aforementioned indicator acquisition software collects task information and monitoring indicator data of heterogeneous computing power tasks executed by the computing power container in a timed polling manner. The task information and the monitoring indicator data are summarized into an indicator data table; The monitoring index data collected during the execution of heterogeneous computing tasks by different computing power containers includes: Read the monitoring index data collected during the execution of heterogeneous computing tasks by different computing power containers from the index data table.

3. The method according to claim 2, characterized in that, The process of summarizing the monitoring index data by task granularity to obtain aggregated index data corresponding to each heterogeneous computing power task includes: Based on the task information, the monitoring indicator data is classified and processed to obtain monitoring indicator data for different tasks. The monitoring index data of different tasks are aggregated to obtain the aggregated index data corresponding to each heterogeneous computing power task.

4. The method according to claim 1, characterized in that, The anomaly identification process based on the aggregated index data to determine the abnormal tasks in the heterogeneous computing power tasks includes: Based on the aggregated index data, the heterogeneous computing power utilization rate of each heterogeneous computing power task is determined; Heterogeneous computing power tasks whose utilization rate is lower than the utilization rate threshold are identified as abnormal tasks.

5. The method according to claim 4, characterized in that, The process of determining the heterogeneous computing power utilization rate of each heterogeneous computing power task based on the aggregated index data includes: Based on the aggregated index data, the heterogeneous card utilization parameters corresponding to each heterogeneous computing power task are determined; The heterogeneous card utilization rate parameter corresponding to each heterogeneous computing power task is averaged to determine the heterogeneous computing power utilization rate of each heterogeneous computing power task.

6. The method according to claim 4, characterized in that, The step of identifying anomalous tasks in the heterogeneous computing power tasks based on the aggregated index data further includes: When the heterogeneous computing power utilization rate of a heterogeneous computing power task is equal to or higher than the utilization rate threshold, the task type of the heterogeneous computing power task is identified. Based on the aggregated index data and the task type, abnormal tasks in the heterogeneous computing power tasks are identified.

7. The method according to claim 6, characterized in that, The task types include single-card tasks; The step of identifying abnormal tasks in the heterogeneous computing power tasks based on the aggregated index data and the task type includes: Identify single-card tasks in the heterogeneous computing power tasks; Based on the aggregated index data, the computing unit utilization and video memory utilization of each single-card task are determined; Single-card tasks with computing unit utilization or video memory utilization below the utilization threshold are identified as abnormal tasks.

8. The method according to claim 6, characterized in that, The task types include multi-card tasks; The step of identifying abnormal tasks in the heterogeneous computing power tasks based on the aggregated index data and the task type includes: Identify multi-GPU tasks within the heterogeneous computing power tasks; Based on the aggregated index data, the communication bandwidth traffic between each multi-card task is determined; Multi-card tasks with communication bandwidth traffic below the traffic threshold are identified as abnormal tasks.

9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Obtain the model computation task processing request; Parse the model computation task processing request to obtain the model computation task to be processed; The model computation task is split into heterogeneous computing power tasks, and these heterogeneous computing power tasks are allocated to different computing power containers.

10. The method according to any one of claims 1 to 9, characterized in that, The process of detecting and destroying the abnormal task includes: Generate the abnormal task data table corresponding to the abnormal task; Save the abnormal task data table to the task database; Based on a preset task processing cycle, the task database is queried in a polling manner to obtain the abnormal task data table within the current task processing cycle. Based on the abnormal task data table, abnormal tasks within the current task processing cycle are detected and destroyed.

11. The method according to claim 10, characterized in that, Before performing the detection and destruction process on the abnormal task data table within the current task processing cycle, the method further includes: Based on the abnormal task information in the abnormal task data table, a page generation process is performed to obtain a detection confirmation page. The notification will push the confirmation page for the virus removal. Upon receiving a data table modification message based on the detection confirmation page, the abnormal task data table is updated to obtain an updated task data table. The step of performing detection and destruction processing on abnormal tasks within the current task processing cycle based on the abnormal task data table includes: Based on the updated task data table, abnormal tasks within the current task processing cycle are detected and destroyed.

12. A heterogeneous computing power task monitoring device, characterized in that, The device includes: The indicator acquisition module is used to acquire monitoring indicator data collected during the execution of heterogeneous computing tasks by different computing power containers; The indicator aggregation module is used to summarize and process the monitoring indicator data according to the task granularity to obtain the aggregated indicator data corresponding to each heterogeneous computing power task. An abnormal task identification module is used to perform abnormal identification processing based on the aggregated index data to identify abnormal tasks in the heterogeneous computing power tasks. The abnormal task detection and removal module is used to perform detection, removal, and destruction processing on the abnormal tasks.

13. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.

15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.