Computing power-based reasoning task monitoring and alarming method and device for intelligent computing center
By obtaining task type information, setting monitoring indicators and alarm conditions in the intelligent computing center, and monitoring and sending alarm information in real time, the task abnormality caused by insufficient computing power is solved, ensuring the stability and efficiency of the inference task.
Patent Information
- Application Number
- CN202510454867.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-25
AI Technical Summary
In the intelligent computing center, inference tasks based on computing power may have insufficient computing power, resulting in task lag, delay or interruption, data abnormalities and model performance degradation, resulting in serious consequences. The existing technology lacks an effective monitoring and alarm mechanism.
By obtaining task type information of the target inference task, determining monitoring indicators and alarm triggering conditions, monitoring indicator data in real time, and sending alarm information of preset information templates to the user when the alarm triggering conditions are met, including email, SMS, instant messaging application push or API interface call.
It realizes the stable and efficient operation of the inference tasks based on computing power of the intelligent computing center, timely discovers and handles exceptions, and avoids data leakage and business paralysis.
Smart Images

Figure CN120371658A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure, and particularly relates to a method and device for monitoring and alarming inference tasks based on computing power in an intelligent computing center. Background Technique
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, to mainly provide the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.
[0004] An "intelligent computing center" includes but is not limited to an "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.
[0006] "Computing power" is the core of an "intelligent computing center" and an "intelligent computing center". It is the ability of a computer device or a computing / data center to process information. It is the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement. It is the computing ability to process information data and achieve the output of a target result. It is a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0007] An "inference task based on computing power" refers to a task that uses computing power to process and analyze data to draw conclusions or make inferences.
[0008] Reasoning tasks based on computing power can use computing power to analyze and process data, and play an important role in many fields, such as assisting in the diagnosis of diseases in medical care, realizing intelligent monitoring in transportation, and predicting risks in the financial field. It can improve decision-making efficiency and accuracy and promote the intelligent development of various industries. However, during the execution of reasoning tasks, insufficient computing power may cause task jams, delays, or even interruptions, or data anomalies and model performance degradation may occur, resulting in errors in reasoning results. If they cannot be processed immediately, they may lead to serious consequences such as data leakage and business paralysis, affecting the normal development and development of related businesses. Therefore, since the emergence of intelligent computing centers, in order to ensure the smooth execution of reasoning tasks based on computing power, how to implement monitoring and alarming of reasoning tasks based on computing power in intelligent computing centers has become a technical problem that needs to be solved urgently. Summary of the invention
[0009] The present invention provides a method and device for monitoring and giving an alarm for reasoning tasks based on computing power of an intelligent computing center, so as to solve the problem of monitoring and giving an alarm for reasoning tasks based on computing power of an intelligent computing center.
[0010] To solve the above problems, the present invention is achieved as follows:
[0011] In a first aspect, the present invention provides a method for monitoring and alarming inference tasks based on computing power in an intelligent computing center, comprising:
[0012] Step S1, obtaining task type information of the target reasoning task;
[0013] Step S2: determining the monitoring indicators and alarm triggering conditions of the target reasoning task according to the task type information of the target reasoning task;
[0014] Step S3: monitoring the indicator data generated during the operation of the target reasoning task based on the monitoring indicator;
[0015] Step S4: When the indicator data meets the alarm triggering condition, send alarm information to the user terminal based on a preset information template.
[0016] In one embodiment, the alarm triggering condition includes an alarm threshold, and step S2 includes:
[0017] Step S21, determining the monitoring index of the target reasoning task according to the task type information of the target reasoning task;
[0018] Step S22: acquiring a plurality of historical indicator data of the target reasoning task in a historical period according to the monitoring indicator;
[0019] Step S23, calculating the mean and standard deviation of the plurality of historical indicator data;
[0020] Step S24: Determine the alarm threshold for the target inference task based on the mean and standard deviation of the multiple historical metric data.
[0021] In one embodiment, step S4 includes:
[0022] Step S41: When the metric data continuously meets the alarm trigger condition for a preset duration, send an alarm message to the user terminal based on a preset information template corresponding to the target inference task through a target communication method, where the target communication method includes at least one of email, SMS, instant messaging application push, or application programming interface (API) call.
[0023] In one embodiment, the monitoring metric includes at least one of the request count of the target inference task and the computing power resource utilization rate of the target inference task.
[0024] In one embodiment, the preset information template includes the sender of the alarm message, the recipient of the alarm message, an alarm identifier, an alarm time, and alarm content.
[0025] In a second aspect, the present invention further provides a monitoring and alarming device for inference tasks based on computing power in an intelligent computing center, including:
[0026] A first acquisition module, configured to acquire the task type information of the target inference task;
[0027] A first determination module, configured to determine the monitoring metric and the alarm trigger condition of the target inference task according to the task type information of the target inference task;
[0028] A first monitoring module, configured to monitor the metric data generated during the operation of the target inference task based on the monitoring metric;
[0029] A first sending module, configured to send an alarm message to the user terminal based on a preset information template when the metric data meets the alarm trigger condition.
[0030] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps in the monitoring and alarming method for inference tasks based on computing power in the intelligent computing center as described in the first aspect above are implemented.
[0031] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the monitoring and alarming method for inference tasks based on computing power in the intelligent computing center as described in the first aspect above are implemented.
[0032] In a fifth aspect, the present invention further provides a computer program product, including computer instructions, which when executed by a processor, implement the steps in the method for monitoring and alarming inference tasks based on computing power of the intelligent computing center as described in the first aspect above.
[0033] In the present invention, task type information of a target inference task is obtained; according to the task type information of the target inference task, monitoring metrics and alarm trigger conditions of the target inference task are determined; based on the monitoring metrics, metric data generated during the running of the target inference task is monitored; and when the metric data meets the alarm trigger conditions, an alarm message is sent to the user side based on a preset information template. In this way, through type recognition of the target inference task, setting of monitoring metrics and alarm conditions, real-time monitoring, and timely alarming, a complete monitoring and management system for inference tasks based on computing power of the intelligent computing center is formed, which can effectively ensure the stable and efficient operation of the target inference task. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] To more clearly illustrate the technical solutions of the present invention, the drawings required for the description of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0035] Figure 1 is a flowchart of a method for monitoring and alarming inference tasks based on computing power of an intelligent computing center provided by the present invention;
[0036] Figure 2 is a test schematic diagram of a model service provided by the present invention;
[0037] Figure 3 is a configuration schematic diagram of alarm trigger conditions provided by the present invention;
[0038] Figure 4 is a template schematic diagram of an alarm message provided by the present invention;
[0039] Figure 5 is a structural diagram of a device for monitoring and alarming inference tasks based on computing power of an intelligent computing center provided by the present invention;
[0040] Figure 6 is a structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0042] The "computing power" as described in the present invention refers to: the ability of a computer device or a computing / data center to process information, which is the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, and is the computing ability to output a target result by processing information data. It is a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0043] The "computational power" (Computational Power, CP) as described in the present invention refers to: the ability of a data center server to process data and output results, which is a comprehensive indicator for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 .
[0044] The "carrying capacity" (Network Power, NP) as described in the present invention refers to: the performance of the data transmission ability of computing power facilities, which is a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., and involves network transmission inside and between data centers, and is a comprehensive indicator for measuring network transmission scheduling ability.
[0045] The "storage power" (Storage Power, SP) as described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage ability of a data center, including external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), and the commonly used measurement unit for performance is the number of read / write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB). The disaster recovery ratio is an important manifestation of security and reliability.
[0046] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure integrating information computing power, network carrying capacity, and data storage capacity, which can realize the centralized computing, storage, transmission, and application of information.
[0047] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0048] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.
[0049] The "general computing power" described in the present invention refers to: the computing power provided by servers based on Central Processing Unit (CPU) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0050] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform deployed on a large scale based on dedicated chips such as Graphics Processing Unit (GPU), Field Programmable Gate Array (FPGA), and Application Specific Integrated Circuit (ASIC), such as natural language processing, machine vision, etc.
[0051] The "super computing power" described in the present invention mainly refers to: the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0052] The "intelligent computing center" described in the present invention refers to: a facility that mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0053] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".
[0054] The "Intelligent Computing Center" described in the present invention, namely the artificial intelligence computing center, is a type of computing infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.
[0055] The "Computing Power Center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, and having computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0056] The "Supercomputing Center" described in the present invention refers to the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0057] The "Computing Power Resources" described in the present invention refer to technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0058] The "Inference Task Based on Computing Power" described in the present invention refers to a task of using computing power to process and analyze data to draw conclusions or make inferences.
[0059] In the existing technology, reasoning tasks based on computing power can use computing power to analyze and process data, and play an important role in many fields, such as assisting in the diagnosis of diseases in medical treatment, realizing intelligent monitoring in transportation, and predicting risks in the financial field. It can improve decision-making efficiency and accuracy and promote the intelligent development of various industries. However, in the process of executing reasoning tasks, insufficient computing power may cause task jams, delays, or even interruptions, or data anomalies and model performance degradation may occur, resulting in errors in reasoning results. If they cannot be processed immediately, serious consequences such as data leakage and business paralysis may occur, affecting the normal development and development of related businesses. Therefore, since the emergence of intelligent computing centers, in order to ensure the smooth execution of reasoning tasks based on computing power, how to realize monitoring and alarming of reasoning tasks based on computing power in intelligent computing centers has become a technical problem that needs to be solved urgently. In order to realize monitoring and alarming of the reasoning tasks based on computing power of the intelligent computing center, in the present invention, the task type information of the target reasoning task is obtained; the monitoring index and alarm triggering condition of the target reasoning task are determined according to the task type information of the target reasoning task; the index data generated by the target reasoning task during operation is monitored based on the monitoring index; when the index data meets the alarm triggering condition, the alarm information is sent to the user end based on the preset information template. In this way, by identifying the type of target reasoning task, setting monitoring indicators and alarm conditions, real-time monitoring and timely alarming, a complete monitoring and management system for reasoning tasks based on computing power of the intelligent computing center is formed, which can effectively ensure the stable and efficient operation of the target reasoning task.
[0060] For details, see Figure 1 , Figure 1 It is a flow chart of a method for monitoring and alarming a reasoning task based on computing power in an intelligent computing center provided by the present invention. Figure 1 As shown, the following steps are included:
[0061] Step S1, obtaining task type information of the target reasoning task;
[0062] In this step, the intelligent computing center needs to clarify the specific type of the target reasoning task, such as image recognition reasoning task, natural language processing reasoning task, predictive analysis reasoning task, etc. This is the basis for a series of subsequent operations. Only by accurately knowing the task type can subsequent processing be carried out in a targeted manner.
[0063] Step S2: determining the monitoring indicators and alarm triggering conditions of the target reasoning task according to the task type information of the target reasoning task;
[0064] In this step, based on the task type information obtained in step S1, appropriate monitoring metrics are determined for the target inference task, such as inference time, accuracy rate, resource utilization (usage of CPU, memory, GPU, etc.). At the same time, set the conditions under which an alarm will be triggered. For example, an alarm is triggered when the inference accuracy rate is lower than a certain threshold, the resource utilization exceeds 80%, etc. Different task types focus on different key points. By specifically determining the monitoring metrics and alarm trigger conditions, the running status of the task can be monitored more accurately. For example, for an inference task with high real-time requirements, the inference time is mainly monitored, and an alarm is triggered once the set time limit is exceeded, so as to timely discover and solve problems and ensure the normal operation of the task and service quality. Reasonable alarm trigger conditions can notify relevant personnel in a timely manner when problems occur and prevent the problems from expanding.
[0065] Step S3: Monitor the metric data generated during the running of the target inference task based on the monitoring metrics;
[0066] In this step, during the running of the target inference task, relevant data generated by the task is collected and recorded in real time according to the monitoring metrics determined in step S2. For example, continuously monitor the change in inference accuracy rate, the fluctuation of resource usage, etc. Exemplarily, for a target inference task of image recognition, assume the determined monitoring metrics are inference accuracy rate, average inference time, and GPU utilization rate. During the running of the task, the corresponding data needs to be monitored according to these metrics. For example, after processing 1000 test images, it is found that 850 images are correctly recognized, then the inference accuracy rate at this time is 85%. As the task continues to run, this accuracy rate data will be continuously updated to observe whether the performance of the model is stable or has a changing trend.
[0067] Step S4: When the metric data meets the alarm trigger conditions, send an alarm message to the user terminal based on a preset information template.
[0068] In this step, when the metric data monitored in step S3 reaches or exceeds the alarm trigger conditions set in step S2, the intelligent computing center will send an alarm message to the user terminal according to the pre-set information template to notify the user that an abnormal situation has occurred in the task running. The information template usually includes the basic information of the task, the metrics that trigger the alarm, and the specific abnormal situation, etc.
[0069] In the present invention, task type information of a target inference task is obtained; according to the task type information of the target inference task, monitoring metrics and alarm trigger conditions of the target inference task are determined; based on the monitoring metrics, metric data generated during the operation of the target inference task is monitored; and in the case where the metric data meets the alarm trigger conditions, an alarm message is sent to a user terminal based on a preset information template. In this way, through type recognition, monitoring metric and alarm condition setting, real-time monitoring, and timely alarming of the target inference task, a complete task monitoring and management system is formed, which can effectively ensure the stable and efficient operation of the target inference task.
[0070] In one embodiment, the alarm trigger condition includes an alarm threshold, and step S2 includes:
[0071] Step S21, according to the task type information of the target inference task, determine the monitoring metrics of the target inference task;
[0072] Step S22, according to the monitoring metrics, obtain multiple historical metric data of the target inference task in a historical period;
[0073] Step S23, calculate the mean and standard deviation of the multiple historical metric data;
[0074] Step S24, determine the alarm threshold of the target inference task based on the mean and standard deviation of the multiple historical metric data.
[0075] In the above embodiment, different types of target inference tasks have different characteristics and requirements, and aspects that need to be concerned about are also different. For example, for a sentiment analysis inference task in natural language processing, metrics such as accuracy rate and recall rate may be focused on; for a target detection inference task in a video stream, metrics such as frame rate and mAP (mean average precision) of the model may be more concerned about. By clarifying the task type, monitoring metrics that can most reflect the operation status and performance of the task can be specifically selected, providing a basis for subsequent monitoring and analysis.
[0076] After determining the monitoring metrics, relevant metric data of the target inference task in the past period of time needs to be collected. These historical data can reflect the performance of the task in the normal operation state and various possible situations. For example, collect the daily accuracy rate data of the sentiment analysis inference task in the past week, or the hourly frame rate data of the video target detection inference task in the past month, etc. Through the analysis of historical data, the operation rules and change trends of the task can be understood.
[0077] The mean is the average of all historical metric data, which represents the average level of the task within the historical period. The standard deviation reflects the degree of dispersion of the data, that is, the fluctuation of the data relative to the mean. Taking the accuracy rate of the sentiment analysis inference task as an example, the calculated mean allows us to know the average accuracy of the task in history, while the standard deviation can tell us the fluctuation range of the accuracy rate. If the standard deviation is small, it indicates that the accuracy rate is relatively stable; if the standard deviation is large, it means that the accuracy rate fluctuates greatly and there may be some unstable factors that need further analysis.
[0078] The alarm threshold is the boundary value used to determine whether the task has an anomaly. Generally speaking, the alarm threshold can be reasonably set according to the mean and the standard deviation. For example, the mean plus a certain multiple of the standard deviation can be used as the upper alarm threshold, and the mean minus a certain multiple of the standard deviation can be used as the lower alarm threshold. The thresholds set in this way can comprehensively consider the historical performance of the task and the fluctuation of the data. When the current metric data exceeds this range, it means that the task may have an anomaly and an alarm needs to be sent in a timely manner.
[0079] In this embodiment, determining the monitoring metrics based on the task type and then determining the alarm threshold in combination with the statistical characteristics of the historical data can more accurately reflect the actual operation situation and the normal fluctuation range of the task. Compared with simply setting a fixed alarm threshold, this method can avoid false alarms or missed alarms caused by unreasonable threshold settings and improve the accuracy and reliability of the alarm.
[0080] In one embodiment, step S4 includes:
[0081] Step S41, when the metric data continuously reaches the alarm trigger condition for a preset duration, send an alarm message to the user terminal based on a preset information template corresponding to the target inference task through a target communication method, where the target communication method includes at least one of email, SMS, instant messaging application push, or application programming interface (API) interface call.
[0082] In the above embodiment, when monitoring the target inference task, the obtained metric data (such as inference accuracy rate, resource utilization rate, etc.) does not immediately trigger an alarm once it reaches the previously set alarm trigger condition (such as the accuracy rate is lower than a certain threshold, the resource utilization rate exceeds a certain percentage, etc.). Instead, these metric data need to continuously meet the alarm trigger condition, and the continuous time needs to reach the pre-set duration. For example, if the pre-set alarm trigger condition is that the CPU utilization rate exceeds 80%, and the preset duration is 10 minutes, then only when the CPU utilization rate continuously exceeds 80% for 10 minutes will the subsequent alarm operation be executed.
[0083] Each target inference task has a pre-set information template corresponding to it. This template stipulates the format and content of the alarm information. For example, it will include information such as the name of the task, the abnormal metric, the specific value of the metric, and the time when the alarm is triggered. Using the pre-set information template can ensure the standardization and integrity of the alarm information, enabling users to quickly and clearly understand the specific situation of the task abnormality.
[0084] After the above-mentioned conditions are met, the system will send the alarm information to the user side through specific communication methods. The target communication methods mentioned here include at least one of email, SMS, instant messaging application push, or application programming interface (API) call. For example, detailed alarm information can be sent to the user's email in the form of an email; the user can also be quickly notified of the task abnormality through SMS; or messages can be pushed through instant messaging applications (such as WeChat, Enterprise WeChat, etc.) so that users can receive alarm prompts in a timely manner; in addition, the alarm information can be integrated into other relevant systems or applications through API calls for further processing or notifying relevant personnel.
[0085] In this embodiment, using the pre-set information template to send alarm information makes the alarm content have a unified format and specification. Users can quickly understand the key information of the task abnormality without spending time sorting out and analyzing non-standard alarm content. This helps users more efficiently judge the severity of the problem and take corresponding handling measures, improving the efficiency of problem handling.
[0086] In one embodiment, the monitoring metrics include at least one of the request count of the target inference task and the computing power resource utilization rate of the target inference task.
[0087] In the above embodiment, the request count of the target inference task refers to the number of requests issued for the target inference task. For example, in the inference application of a machine learning model, multiple clients may send requests to the server to require the model to perform inference calculations. The request count here is the statistics of these requests, which can reflect the frequency of the inference task being called.
[0088] The computing power resource utilization rate of the target inference task refers to the proportion of the computing power resources used when executing the target inference task in the total available computing power resources. For example, the server has certain computing power resources such as CPU and GPU. When the target inference task runs, it will occupy a part of these resources for data processing and model inference. This metric is used to measure the occupancy of the computing power resources by the task.
[0089] In one embodiment, the preset information template includes the sender of the alarm information, the recipient of the alarm information, an alarm identifier, an alarm time, and alarm content.
[0090] In the above embodiment, the sender: clarifies the source of the alarm information, enabling the recipient to clearly know which system, department, or specific module sent the information;
[0091] The recipient: designates the recipient of the alarm information, which can be an individual, such as the email address or mobile phone number of a specific project leader or operation and maintenance personnel; it can also be a group, such as the work group of a project team or the notification group of an operation and maintenance team. This ensures that the alarm information can be accurately conveyed to the relevant personnel or groups that need to pay attention to the problem.
[0092] The alarm identifier: is a classification identifier for the alarm information, used to distinguish different types of alarms. For example, there may be an alarm identifier for a decrease in the accuracy of the inference task, an alarm identifier for excessive resource occupancy, etc. Through the alarm identifier, the recipient can quickly identify the nature and general category of the alarm, facilitating the prioritized handling of important or urgent alarms.
[0093] The alarm time: records the specific time when the alarm occurred, which is very important for analyzing the chronological order of problem occurrence, investigating possible related events, and evaluating the timeliness of the problem. For example, through the alarm time, it can be determined whether the problem occurred during the business peak period or after a specific operation.
[0094] The alarm content: details the specific situation of the alarm, including the name of the target inference task, the monitored metrics with abnormal values, and possible cause analysis, etc. The alarm content can provide the recipient with sufficient information to understand the essence of the problem, so as to take corresponding solutions.
[0095] In this embodiment, using the preset information template containing the sender, recipient, alarm identifier, alarm time, and alarm content has significant benefits. Clearly defining the sender allows for tracing the information source and strengthening the sense of responsibility; accurately positioning the recipient ensures that the information reaches the relevant personnel directly and avoids omission. The alarm identifier facilitates quickly distinguishing alarm categories and prioritizing the handling of urgent and critical issues. The alarm time provides a time clue to assist in troubleshooting the root cause of the problem according to the event sequence. The rich alarm content enables the recipient to comprehensively grasp the problem without additional inquiries, quickly respond and solve it, greatly improving the alarm handling efficiency and ensuring the stable operation of the system.
[0096] Actually, when testing the monitoring and alarm function of the inference task based on computing power in the intelligent computing center, the user can first deploy an open-source large model in the intelligent computing center, specify the computing power specifications, and configure the model service not to allow the system to automatically adjust computing resources.
[0097] See Figure 2 Figure 2 , continuously send requests to the model service so that the model service performs an inference task based on the request and thus uses computing power resources.
[0098] See Figure 3 Figure 3 , set an alarm trigger condition in advance based on the inference task of the model service. Exemplarily, when the GPU occupancy rate of the model service exceeds 20%, an alarm message can be sent to the user side. The template of the alarm message can be seen in Figure 4 Figure 4 , and the alarm message can be sent in the form of an email.
[0099] During testing, a stress testing tool can be used to gradually increase the concurrency of model service requests until the GPU occupancy exceeds 20%.
[0100] Please see Figure 5 Figure 5 , Figure 5 Figure 5 is the structural diagram of an inference task monitoring and alarming device based on computing power in the intelligent computing center provided by the present invention. As Figure 5 Figure 5 shown, the inference task monitoring and alarming device 500 based on computing power in the intelligent computing center includes:
[0101] A first acquisition module 501, configured to acquire the task type information of the target inference task;
[0102] A first determination module 502, configured to determine the monitoring metrics and alarm trigger conditions of the target inference task according to the task type information of the target inference task;
[0103] A first monitoring module 503, configured to monitor the metric data generated during the operation of the target inference task based on the monitoring metrics;
[0104] A first sending module 504, configured to send an alarm message to the user side based on a preset information template when the metric data meets the alarm trigger condition.
[0105] In one embodiment, the alarm trigger condition includes an alarm threshold, and the first determination module 602 includes:
[0106] A first determination unit, configured to determine the monitoring metrics of the target inference task according to the task type information of the target inference task;
[0107] A first acquisition unit, configured to acquire multiple historical metric data of the target inference task in a historical period according to the monitoring metrics;
[0108] A first calculation unit, configured to calculate the mean and standard deviation of the multiple historical metric data;
[0109] A second determination unit, configured to determine an alarm threshold for the target inference task based on the mean and standard deviation of the multiple historical metric data.
[0110] In one embodiment, the first sending module 504 includes:
[0111] A first sending unit, configured to, when the metric data continuously reaches the alarm trigger condition for a preset duration, send an alarm message to a user terminal through a target communication method based on a preset information template corresponding to the target inference task, where the target communication method includes at least one of email, short message, instant messaging application push, or application programming interface (API) interface call.
[0112] In one embodiment, the monitored metric includes at least one of the number of requests for the target inference task and the utilization rate of computing power resources for the target inference task.
[0113] In one embodiment, the preset information template includes the sender of the alarm message, the recipient of the alarm message, an alarm identifier, an alarm time, and alarm content.
[0114] The inference task monitoring and alarming device for an intelligent computing center based on computing power provided by the present invention can implement each process of the above-mentioned inference task monitoring and alarming method for an intelligent computing center based on computing power. The technical features correspond one by one and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.
[0115] It should be noted that the inference task monitoring and alarming device for an intelligent computing center based on computing power in the present invention can be a device, or a component, an integrated circuit, or a chip in an electronic device.
[0116] The present invention further provides an electronic device. Refer to Figure 6 , Figure 6 is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. The electronic device includes a memory 601, a processor 602, and a program or instruction running on the memory 601. When the program or instruction is executed by the processor 602, it can implement Figure 1 any step in the corresponding inference task monitoring and alarming method embodiment of the intelligent computing center based on computing power and achieve the same beneficial effects. They will not be elaborated here.
[0117] Among them, the processor 602 can be a CPU, an ASIC, an FPGA, or a GPU.
[0118] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above-mentioned inference task monitoring and alarming method embodiment of the intelligent computing center based on computing power can be completed by hardware related to program instructions, and the program can be stored in a readable medium.
[0119] The present invention also provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above Figure 1 The corresponding steps in the embodiment of the monitoring and alarm method for the reasoning task of the intelligent computing center based on computing power can achieve the same technical effect. To avoid repetition, they are not repeated here. The storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.
[0120] The present invention also provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the above Figure 1 The corresponding intelligent computing center has each process of the embodiment of the reasoning task monitoring and alarm method based on computing power, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0121] The terms "first", "second" etc. in the present invention are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. In addition, the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusions, for example, the process, method, system, product or equipment comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment. In addition, "and / or" is used in the present application to represent at least one of the connected objects, such as A and / or B and / or C, which means to include 7 situations including single A, single B, single C, and A and B all exist, B and C all exist, A and C all exist, and A, B and C all exist.
[0122] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0123] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of the various embodiments of the present application.
[0124] The embodiments of the present application are described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. An inference task monitoring and alarming method based on computing power in an intelligent computing center, characterized in that, Including: Step S1, obtaining the task type information of the target inference task; Step S2, determining the monitoring metrics and alarm triggering conditions of the target inference task according to the task type information of the target inference task; Step S3, monitoring the metric data generated during the operation of the target inference task based on the monitoring metrics; Step S4, when the metric data meets the alarm triggering conditions, sending an alarm message to the user terminal based on a preset information template.
2. The method according to claim 1, characterized in that, The alarm triggering conditions include an alarm threshold, and the step S2 includes: Step S21, determining the monitoring metrics of the target inference task according to the task type information of the target inference task; Step S22, obtaining multiple historical metric data of the target inference task in a historical period according to the monitoring metrics; Step S23, calculating the mean and standard deviation of the multiple historical metric data; Step S24, determining the alarm threshold of the target inference task based on the mean and standard deviation of the multiple historical metric data.
3. The method according to claim 1, wherein The step S4 includes: Step S41, when the metric data continuously reaches the alarm triggering conditions for a preset duration, sending an alarm message to the user terminal through a target communication method based on a preset information template corresponding to the target inference task, and the target communication method includes at least one of email, SMS, instant messaging application push, or application programming interface (API) interface call.
4. The method according to any one of claims 1 to 3, characterized in that The monitoring metrics include at least one of the request number of the target inference task and the computing power resource utilization rate of the target inference task.
5. The method according to any one of claims 1 to 3, characterized in that, The preset information template includes the sender of the alarm message, the recipient of the alarm message, an alarm identifier, an alarm time, and alarm content.
6. An inference task monitoring and alarming device based on computing power in an intelligent computing center, characterized in that, Including: A first acquisition module, configured to obtain the task type information of the target inference task; A first determination module, configured to determine the monitoring metrics and alarm triggering conditions of the target inference task according to the task type information of the target inference task; A first monitoring module, configured to monitor the metric data generated during the operation of the target inference task based on the monitoring metrics; A first sending module, configured to send an alarm message to the user terminal based on a preset information template when the metric data meets the alarm triggering conditions.
7. The device according to claim 6, characterized in that, The alarm triggering conditions include an alarm threshold, and the first determination module includes: A first determination unit, configured to determine the monitoring metrics of the target inference task according to the task type information of the target inference task; A first acquisition unit, configured to obtain multiple historical metric data of the target inference task in a historical period according to the monitoring metrics; A first calculation unit, configured to calculate the mean and standard deviation of the multiple historical metric data; A second determination unit, configured to determine the alarm threshold of the target inference task based on the mean and standard deviation of the multiple historical metric data.
8. An electronic device, characterized in that, Including: A processor, a memory, and a program stored on the memory and executable on the processor. When the program is executed by the processor, the steps of the inference task monitoring and alarm method based on computing power of the intelligent computing center as described in any one of claims 1 to 5 are implemented.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the inference task monitoring and warning method of the intelligent computing center based on computing power as described in any one of claims 1 to 5 are implemented.
10. A computer program product, characterized in that, It includes computer instructions, and when the computer instructions are executed by a processor, the steps of the inference task monitoring and warning method of the intelligent computing center based on computing power as described in any one of claims 1 to 5 are implemented.