Network monitoring and alarming method and device for computing power operation task state of intelligent computing center

Through the visualization and alarm mechanism of the operation task status of the intelligent computing center's computing power, the problem of lack of network monitoring warnings in the existing technology is solved, and real-time monitoring and efficiency improvement of abnormal situations is achieved.

CN120358125APending Publication Date: 2025-07-22DATACANVAS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510391451.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The lack of network monitoring warnings on the status of computing power operation tasks in the intelligent computing center in the prior art, resulting in a decrease in the execution efficiency of computing power operation tasks.

Method used

By visualizing the state of the computing power operation task of the intelligent computing center that is monitored in real time, a visual interface of the operation status, performance status and traffic status is generated, and alarm information is displayed on the interface, including communication abnormalities, throughput abnormalities, average delay rate abnormalities, packet loss rate abnormalities and traffic volatility abnormalities, etc.

Benefits of technology

Real-time monitoring and abnormal alarms of the status of computing power operation tasks are realized, reducing the reduction in task execution efficiency in abnormal situations and improving task execution efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358125A_ABST
    Figure CN120358125A_ABST
Patent Text Reader

Abstract

The invention provides a network monitoring and warning method and device for the computing power operation task state of an intelligent computing center, and relates to the technical field of computing power infrastructure, and the method comprises the steps: S1, carrying out the visualization of the computing power operation task state, monitored in real time, of the intelligent computing center, and obtaining a visual interface, the computing power operation task state comprises an operation state, a performance state and a flow state, and the visual interface comprises an operation state interface, a performance state interface and a flow state interface; and S2, when any one of the computing power operation task states is abnormal, displaying alarm information on the corresponding visual interface. Through network monitoring of the computing power operation task state of the intelligent computing center, an alarm is given for an abnormal condition in real time, the condition that the execution efficiency of the computing power operation task is reduced due to the fact that any state in the computing power operation task state is abnormal is reduced, and therefore the execution efficiency of the computing power operation task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure, and particularly relates to a method and device for network monitoring and warning of the status of computing power operation tasks in an intelligent computing center. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes but is not limited to the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that is based on artificial intelligence theory, adopts an artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", and is the ability of computer devices or computing / data centers to process information. It is the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement. It is the computing ability to achieve the output of the target result by processing information data. It is a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0007] At present, as the core in the era of artificial intelligence (AI), the intelligent computing center undertakes ultra-large-scale computing power operation tasks such as model training and inference. The network environment of the intelligent computing center has high dynamicity and high bandwidth requirements. For AI model training, parameter synchronization between nodes needs to be performed frequently, and network jitter or congestion directly affects the training efficiency; edge inference tasks have strict requirements for real-time performance and need to ensure low-latency data transmission. However, since the emergence of intelligent computing centers, there has been a lack of network monitoring and warning of the status of computing power operation tasks in the prior art, resulting in a decline in the execution efficiency of computing power operation tasks.

[0008] It can be seen that it is an urgent problem to be solved to perform network monitoring and warning on the status of computing power operation tasks in the intelligent computing center. Summary of the Invention

[0009] An embodiment of the present invention provides a method and device for network monitoring and alarming of the computing power operation task status in an intelligent computing center, so as to solve the problem in the related art that the lack of network monitoring and warning of the computing power operation task status in the intelligent computing center leads to a decrease in the execution efficiency of the computing power operation task.

[0010] To solve the above problems, the present invention is implemented as follows:

[0011] In a first aspect, an embodiment of the present invention provides a method for network monitoring and alarming of the computing power operation task status in an intelligent computing center, including:

[0012] Step S1: Visualize the computing power operation task status of the intelligent computing center monitored in real time to obtain a visualization interface, where the computing power operation task status includes an operation status, a performance status, and a traffic status, and the visualization interface includes an operation status interface, a performance status interface, and a traffic status interface;

[0013] Step S2: When any status in the computing power operation task status is abnormal, display an alarm message on the corresponding visualization interface.

[0014] In one embodiment, before the step S2, it further includes:

[0015] Step S3: Monitor the network link to determine whether there is a first abnormal information, where the first abnormal information includes communication abnormality, and the communication abnormality indicates a link interruption in the network layer;

[0016] Step S4: When there is the first abnormal information, determine that the operation status is abnormal.

[0017] In one embodiment, before the step S2, it further includes:

[0018] Step S5: Monitor the network link to determine whether there is a second abnormal information, where the second abnormal information includes at least one of throughput abnormality, average delay rate abnormality, and packet loss rate abnormality. Among them, the throughput abnormality means that the amount of data successfully transmitted per unit time is less than a first threshold, the average delay rate abnormality means that the delay rate of the average time of data from the sending end to the receiving end is greater than a second threshold, and the packet loss rate abnormality means that the proportion of data lost during transmission is greater than a third threshold;

[0019] Step S6: When there is the second abnormal information, determine that the performance status is abnormal.

[0020] In one embodiment, before the step S2, it further includes:

[0021] Step S7: Monitor the traffic within a preset time interval to determine whether there is a third abnormal message. The third abnormal message includes an abnormal traffic volatility, which means that the volatility of the traffic within the preset time interval is greater than a fourth threshold value.

[0022] Step S8: When there is the third abnormal message, determine that the traffic status is abnormal.

[0023] In one embodiment, the step S2 includes at least one of the following:

[0024] Step S21: When the operating status is abnormal, display the alarm message corresponding to the first abnormal level on the operating status interface according to the first abnormal level of the operating status. The first abnormal level is a level determined according to the influence factors corresponding to each link interruption in the network layer.

[0025] Step S22: When the performance status is abnormal, display the alarm message corresponding to the second abnormal level on the performance status interface according to the second abnormal level of the performance status. The second abnormal level is a level determined according to the deviation degree between the obtained data throughput, average delay rate, and / or packet loss rate and their respective threshold values.

[0026] Step S23: When the traffic status is abnormal, display the alarm message corresponding to the third abnormal level on the traffic status interface according to the third abnormal level of the traffic status. The third abnormal level is a level determined according to the deviation degree between the traffic volatility within the obtained preset time interval and its corresponding threshold value.

[0027] In one embodiment, it further includes:

[0028] Step S9: When receiving test data, obtain the visualization result of the test data in the visualization interface. The test data is data used to simulate any abnormal state in the computing power operation task state.

[0029] Step S10: When the visualization result indicates that the alarm message corresponding to the test data is not displayed in the visualization interface, send a preset indication message, where the preset indication message is used to indicate that the function of displaying the alarm message in the visualization interface is abnormal.

[0030] In a second aspect, an embodiment of the present invention further provides an intelligent computing center computing power operation task state network monitoring and alarming device, including:

[0031] A visualization module, configured to visualize the operation task status of the computing power in the intelligent computing center monitored in real time, so as to obtain a visualization interface, where the operation task status of the computing power includes an operation status, a performance status, and a traffic status, and the visualization interface includes an operation status interface, a performance status interface, and a traffic status interface;

[0032] An alarm module, configured to display alarm information on the corresponding visualization interface when any one of the operation task statuses of the computing power is abnormal.

[0033] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps in the method for network monitoring and alarming of the operation task status of the computing power in the intelligent computing center as described in the first aspect above are implemented.

[0034] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method for network monitoring and alarming of the operation task status of the computing power in the intelligent computing center as described in the first aspect above are implemented.

[0035] In a fifth aspect, the present invention further provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the steps in the method for network monitoring and alarming of the operation task status of the computing power in the intelligent computing center as described in the first aspect above are implemented.

[0036] In the embodiment of the present invention, the operation task status of the computing power in the intelligent computing center monitored in real time is visualized to obtain a visualization interface; when any one of the operation task statuses of the computing power is abnormal, alarm information is displayed on the corresponding visualization interface. Through network monitoring of the operation task status of the computing power in the intelligent computing center, abnormal situations are alarmed in real time, reducing the situation where the execution efficiency of the operation task of the computing power decreases due to any abnormal status in the operation task status of the computing power, thereby improving the execution efficiency of the operation task of the computing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 is a flowchart of a method for network monitoring and alarming of the operation task status of the computing power in the intelligent computing center provided by the embodiment of the present invention;

[0039] Figure 2 It is a structural diagram of a network monitoring and warning device for the operation task status of computing power in an intelligent computing center provided by an embodiment of the present invention;

[0040] Figure 3 It is a structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0042] The "computing power" described in the present invention refers to: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to output a target result through processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0043] The "computational power" (Computational Power, CP) described in the present invention refers to: the ability of a data center server to process data and output results, a comprehensive index to measure the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 .

[0044] The "carrying capacity" (Network Power, NP) described in the present invention refers to: the performance of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission inside and between data centers, and is a comprehensive index to measure network transmission scheduling ability.

[0045] The "Storage Power (SP)" described in the present invention refers to the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and built-in storage devices in servers. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0046] The "computing power infrastructure" described in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage power, and can realize the centralized computing, storage, transmission, and application of information.

[0047] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0048] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.

[0049] The "general computing power" described in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0050] The "intelligent computing power" described in the present invention refers to a computing platform that is scaled up for various artificial intelligence innovation applications based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing and machine vision.

[0051] The "super computing power" described in the present invention mainly refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.

[0052] The "Intelligent Computing Center" described in the present invention refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0053] The "Intelligent Computing Center" described in the present invention includes, but is not limited to, the "Intelligent Computing Center".

[0054] The "Intelligent Computing Center" described in the present invention, namely the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.

[0055] The "Computing Power Center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity, and IT software and hardware devices, and having computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0056] The "Supercomputing Center" described in the present invention refers to the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0057] The "Computing Power Resources" described in the present invention refers to technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0058] The "model" described in the present invention includes, but is not limited to, "large language models" and "multimodal large models".

[0059] The "Large Language Model" described in the present invention refers to a large language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language, trained through a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0060] The "Multimodal Large Models" described in the present invention refer to models trained by jointly using multimodal information such as text, images, videos, and audio, including but not limited to multimodal large language models.

[0061] The "computing power operation task" described in the present invention refers to specific workloads or jobs executed on computing power resources that require a certain amount of computing power support, usually involving scenarios such as complex data processing, numerical calculations, model training, or simulation.

[0062] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for network monitoring and alarming of the computing power operation task status in an intelligent computing center provided by an embodiment of the present invention. As Figure 1 shown, it includes the following steps:

[0063] Step S1: Visualize the real-time monitored status of the computing power operation task in the intelligent computing center to obtain a visualization interface. The computing power operation task status includes an operation status, a performance status, and a traffic status. The visualization interface includes an operation status interface, a performance status interface, and a traffic status interface;

[0064] In this step, the computing power operation task status described in the present invention includes but is not limited to an operation status, a performance status, and a traffic status; correspondingly, the visualization interface described in the present invention includes but is not limited to an operation status interface, a performance status interface, and a traffic status interface.

[0065] Exemplarily, toolkits such as Zabbix and Nagios can be configured to achieve real-time monitoring of the operation status; toolkits such as iPerf, Prometheus, and Grafana can be configured to achieve real-time monitoring of the performance status; toolkits such as sFlow, NetFlow, and Wireshark can be configured to achieve real-time monitoring of the traffic status.

[0066] By visualizing the real-time monitored status of the computing power operation task in the intelligent computing center to obtain a visualization interface, the operation status, performance status, and traffic status can be intuitively grasped. Among them, the operation status may include communication information; the performance status may include information such as data throughput, average latency rate, and packet loss rate; the traffic status may include information such as traffic volatility. Through the display of the operation status interface, it is convenient for operation and maintenance personnel to master real-time network operation information; through the display of the performance status interface, it is convenient for operation and maintenance personnel to master real-time network performance information; through the display of the traffic status interface, it is convenient for operation and maintenance personnel to master real-time network congestion information.

[0067] Step S2: When any of the computing power operation task states is abnormal, display an alarm message on the corresponding visualization interface.

[0068] In this step, when any of the computing power operation task states is abnormal, an alarm message can be displayed on the corresponding visualization interface, so as to realize the early warning of the execution efficiency of the computing power operation task. The operation and maintenance personnel can take relevant measures according to the corresponding alarm message to reduce the situation where the execution efficiency of the computing power operation task decreases due to any abnormal state in the computing power operation task state.

[0069] Among them, when any of the computing power operation task states is abnormal, it can be considered that the device associated with this state has a fault.

[0070] In an example, when the operation state is abnormal, the corresponding alarm message can be displayed on the operation state interface; when the performance state is abnormal, the corresponding alarm message can be displayed on the performance state interface; when the traffic state is abnormal, the corresponding alarm message can be displayed on the traffic state interface. In this way, by displaying the alarm messages corresponding to different computing power operation task states on different visualization interfaces, the monitoring and alarming of the fault information (that is, when any state is abnormal) are realized, so that the execution efficiency of the computing power operation task can be warned, and the situation where the execution efficiency of the computing power operation task decreases due to any abnormal state in the computing power operation task state can be reduced.

[0071] In another example, the visualization interface can also include a fault state interface. By displaying the fault state interface, it is convenient for the operation and maintenance personnel to simultaneously master the abnormal situations of each computing power operation task state. For example, when the operation state is abnormal, the corresponding alarm message can be displayed on the fault state interface; when the performance state is abnormal, the corresponding alarm message can be displayed on the fault state interface; when the traffic state is abnormal, the corresponding alarm message can be displayed on the fault state interface. In this way, by displaying the alarm messages corresponding to different computing power operation task states on the same visualization interface, the monitoring and alarming of the fault information (that is, when any state is abnormal) are realized, so that the execution efficiency of the computing power operation task can be warned, and the situation where the execution efficiency of the computing power operation task decreases due to any abnormal state in the computing power operation task state can be reduced.

[0072] In the embodiment of the present invention, the computing power operation task state of the intelligent computing center monitored in real time is visualized to obtain a visualization interface; when any of the computing power operation task states is abnormal, an alarm message is displayed on the corresponding visualization interface. Through the network monitoring of the computing power operation task state of the intelligent computing center, the abnormal situation is alarmed in real time, and the situation where the execution efficiency of the computing power operation task decreases due to any abnormal state in the computing power operation task state is reduced, thereby improving the execution efficiency of the computing power operation task.

[0073] In one embodiment, step S2 includes at least one of the following:

[0074] Step S21: When the operating state is abnormal, according to the first abnormal level of the operating state, display the alarm information corresponding to the first abnormal level on the operating state interface, where the first abnormal level is a level determined according to the influence factors corresponding to each link interruption in the network layer;

[0075] Step S22: When the performance state is abnormal, according to the second abnormal level of the performance state, display the alarm information corresponding to the second abnormal level on the performance state interface, where the second abnormal level is a level determined according to the deviation degree between the obtained data throughput, average latency rate, and / or packet loss rate and their respective thresholds;

[0076] Step S23: When the traffic state is abnormal, according to the third abnormal level of the traffic state, display the alarm information corresponding to the third abnormal level on the traffic state interface, where the third abnormal level is a level determined according to the deviation degree between the obtained traffic volatility within a preset time interval and its corresponding threshold.

[0077] In one embodiment, the influence factors corresponding to each link in the network layer are set respectively, and different influence factors represent the influence degree when the link is interrupted. For example, when the first link is interrupted, it will cause the network to be paralyzed, and when the second link is interrupted, it will cause network latency. In other words, the influence degree when the first link is interrupted is greater than that when the second link is interrupted, so the influence factor of the first link is set to be greater than that of the second link. In this way, when the operating state is abnormal, first, according to the influence factors corresponding to each link interruption in the network layer, determine the first abnormal level of the operating state, that is, determine the influence degree when the link is interrupted at this time; then display the alarm information corresponding to the first abnormal level on the operating state interface. For example, when the first abnormal level indicates a minor abnormality, the relevant link can be marked in yellow and a prompt message can be displayed. When the first abnormal level indicates a moderate abnormality, it can be marked in orange. When the first abnormal level indicates a severe abnormality, it can be marked in red and an alarm box can be popped up.

[0078] In another embodiment, thresholds for data throughput, average latency rate, and packet loss rate are set respectively. For example, the data throughput threshold is set to 100 Mbps, the average latency rate threshold is set to 10 ms, and the packet loss rate threshold is set to 1%. Calculate the degree of deviation between the obtained data throughput, average latency rate, and packet loss rate and their corresponding thresholds. For example, if the current data throughput is 80 Mbps, the degree of deviation is (100 - 80) / 100 = 20%. Determine the second anomaly level according to the degree of deviation, which can also be divided into mild anomaly, moderate anomaly, and severe anomaly. When there is an anomaly in the performance state, display the corresponding warning information on the performance state interface according to the second anomaly level. The display method can refer to the display of warning information for abnormal operating states.

[0079] Among them, when the degrees of deviation between the data throughput, average latency rate, and packet loss rate and their respective corresponding thresholds are all different. For example, when the degree of deviation of the current data throughput is between 10% - 20% it is a mild anomaly, when the degree of deviation of the current average latency rate is between 10% - 20% it is a moderate anomaly, and when the degree of deviation of the current packet loss rate exceeds 5% it is a severe anomaly. In one example, at this time, the warning information corresponding to the levels of data throughput, average latency rate, and packet loss rate can be displayed in different areas of the performance state interface respectively. In another example, at this time, the data throughput, average latency rate, and packet loss rate with the largest degree of deviation can be used as the performance state anomaly for warning.

[0080] In yet another embodiment, a threshold for the traffic volatility within a preset time interval is set. For example, the traffic volatility threshold is set to 20%. Calculate the degree of deviation between the obtained traffic volatility within the preset time interval and its corresponding threshold. Determine the third anomaly level according to the degree of deviation, which is divided into mild anomaly, moderate anomaly, and severe anomaly. When there is an anomaly in the traffic state, display the corresponding warning information on the traffic state interface according to the third anomaly level.

[0081] Among them, to calculate the traffic volatility, it is first necessary to obtain network traffic data. Data collection can be completed with the help of network monitoring devices (such as switches, routers), network management systems, or professional traffic monitoring software. In actual operation, network traffic data needs to be collected at preset time intervals, for example, once every 5 minutes or once every 10 minutes. The traffic volatility reflects the change range of network traffic within a preset time interval. By comparing the traffic data at two adjacent time points, the traffic volatility can be calculated. When there is an anomaly in the traffic state, a warning message is sent to the administrator. The warning message can be displayed on the traffic state interface or sent to the administrator by means of text messages, emails, instant messaging tools, etc.

[0082] After determining the anomaly level based on the degree of deviation, corresponding alarm messages are generated according to different anomaly levels. For example: If it is a minor anomaly, the alarm message can be "The traffic status is slightly abnormal, and the current traffic volatility deviates from the threshold to a small extent. Please pay attention." If it is a moderate anomaly, the alarm message can be "The traffic status is moderately abnormal, and the current traffic volatility deviates from the threshold to a large extent. It is necessary to check in time." If it is a severe anomaly, the alarm message can be "The traffic status is severely abnormal, and the current traffic volatility deviates from the threshold to a very large extent, which may cause network congestion. Please handle it immediately!" Thus, congestion alarms based on traffic volatility thresholds and anomaly levels are achieved.

[0083] In one embodiment, before the step S2, it further includes:

[0084] Step S3: Monitor the network link to determine whether there is a first anomaly message, where the first anomaly message includes communication anomalies, and the communication anomaly indicates a link interruption in the network layer;

[0085] Step S4: When there is the first anomaly message, determine that the operating state is abnormal.

[0086] In this embodiment, exemplarily, the port status of the switch / router can be queried regularly; the link connectivity can be detected using ping, traceroute, or a dedicated tool (such as Ipsla); the link layer data packets can be captured through an optical splitter, and if there is no traffic for 30 consecutive seconds, an alarm is triggered to indicate a communication anomaly. By monitoring the network link, it is determined whether there is a first anomaly message.

[0087] When it is determined that there is a link interruption in the network layer, it can be determined that the operating state of the intelligent computing center is abnormal. At this time, the following measures can be taken: Send an alarm message to the administrator by email, text message, system pop-up window, etc., informing that the operating state is abnormal and elaborating on the abnormal data. The abnormal data can include the time of link interruption, the number of the interrupted link, the network device associated with the interrupted link, etc. For example, the content of the email sent is as follows: "The operating state of the intelligent computing center is abnormal: Link A connection is interrupted. Please handle it in time." The network link of the intelligent computing center can be monitored in real time to determine whether an abnormal situation occurs.

[0088] In one embodiment, before the step S2, it further includes:

[0089] Step S5: Monitor the network link to determine whether there is a second abnormal information, where the second abnormal information includes at least one of throughput abnormality, average delay rate abnormality, and packet loss rate abnormality. Among them, the throughput abnormality means that the amount of data successfully transmitted per unit time is less than the first threshold, the average delay rate abnormality means that the delay rate of the average time for data from the sending end to the receiving end is greater than the second threshold, and the packet loss rate abnormality means that the proportion of lost data during transmission is greater than the third threshold;

[0090] Step S6: When there is the second abnormal information, determine that the performance state is abnormal.

[0091] In this embodiment, the lower limit value of the data throughput can be determined according to the actual requirements of the network and the service level agreement (SLA) as the first threshold. For example, if it is required that the network bandwidth reaches at least 100 Mbps, that is, the amount of data successfully transmitted per unit time reaches at least 100 Mbps, then the first threshold can be set to 100 Mbps. The upper limit value of the average delay rate is set as the second threshold according to the tolerance of the service to network delay. For example, for real-time services, the average delay rate should be controlled within 50 ms, then the second threshold can be set to 50 ms. According to the stability of the network and service requirements, the upper limit value of the packet loss rate is set as the third threshold. For example, for data transmission services, the packet loss rate should be controlled within 1%, then the third threshold can be set to 1%.

[0092] Compare the collected data throughput, average delay rate, and packet loss rate with the corresponding thresholds respectively. If the data throughput is less than the first threshold, it is determined that there is a throughput abnormality; if the average delay rate is greater than the second threshold, it is determined that there is an average delay rate abnormality; if the packet loss rate is greater than the third threshold, it is determined that there is a packet loss rate abnormality.

[0093] The second abnormal information includes at least one of throughput abnormality, average delay rate abnormality, and packet loss rate abnormality. Therefore, when it is determined in step S5 that there is at least one of throughput abnormality, average delay rate abnormality, or packet loss rate abnormality, it can be determined that the performance state of the intelligent computing center is abnormal. At this time, an alarm message can be sent to the administrator in various ways (such as email, text message, system pop-up window, etc.), detailing the type of performance state abnormality (throughput abnormality, average delay rate abnormality, or packet loss rate abnormality) and the relevant index data. For example, the content of the email can be: "The performance state of the intelligent computing center is abnormal: the data throughput is 80 Mbps, less than the first threshold of 100 Mbps, please handle it in time".

[0094] In one embodiment, before the step S2, it further includes:

[0095] Step S7: Monitor the traffic for a preset time interval to determine whether there is a third abnormal message, where the third abnormal message includes an abnormal traffic volatility, and the abnormal traffic volatility means that the volatility of the traffic within the preset time interval is greater than a fourth threshold value;

[0096] Step S8: When there is the third abnormal message, determine that the traffic status is abnormal.

[0097] In this embodiment, a reasonable upper limit of traffic volatility can be set as the fourth threshold value according to the historical traffic data of the network, service characteristics, and network carrying capacity of the network. For example, by analyzing the traffic data in the past period of time, it is found that the traffic volatility usually fluctuates between 10% - 15%, then the fourth threshold value can be set to 20%. The obtained traffic volatility is compared with the fourth threshold value. If the traffic volatility is greater than the fourth threshold value, it is determined that there is an abnormal traffic volatility, that is, there is a third abnormal message.

[0098] Among them, when there is an abnormal traffic volatility, there may be a traffic peak in a short period of time. This causes the network bandwidth to be occupied in large amounts instantaneously, resulting in a sharp increase in bandwidth utilization. For example, during the peak period of inference task requests, the traffic suddenly surges, which may fill up the network bandwidth in a short time. While during the traffic trough period, a large amount of bandwidth will be idle, causing resource waste. During the traffic peak, the amount of data in the network increases sharply, and the waiting time for data packets to queue for processing at network nodes (such as routers, switches) becomes longer, resulting in a significant increase in network latency. In addition, when the network traffic suddenly increases and exceeds the processing capacity of network devices, the network devices may discard some data packets because they are unable to process too many data packets in time, resulting in an increase in the packet loss rate. For applications with high requirements for data transmission reliability, such as file transfer, database synchronization, etc., packet loss will cause data transmission errors and require retransmission, thereby reducing the transmission efficiency and even possibly resulting in data loss. Therefore, continuous monitoring of the traffic status is a key link to ensure the stable operation of the network and improve service reliability. Through network monitoring of the computing power operation task status of the intelligent computing center, real-time alarms for traffic anomalies are generated, reducing the situation where traffic congestion occurs in the computing power operation task execution due to abnormal traffic status in the computing power operation task status, thereby improving the execution efficiency of the computing power operation task.

[0099] In one embodiment, it further includes:

[0100] Step S9: When receiving test data, obtain the visualization result of the test data in the visualization interface, where the test data is data used to simulate the occurrence of any abnormal status in the computing power operation task status;

[0101] Step S10: When the visualization result indicates that the warning information corresponding to the test data is not displayed on the visualization interface, send a preset indication message, which is used to indicate that the function of displaying warning information on the visualization interface is abnormal.

[0102] In this embodiment, an abnormal situation is simulated. By inputting test data and then checking the display effect of the visualization interface, the processing ability of the system in the face of abnormal data is verified. If it is found that the visualization interface does not display the warning information as expected, it indicates that there may be an abnormality in the function of displaying warning information in the system. At this time, a preset indication message needs to be sent to remind relevant personnel. The specific steps are as follows:

[0103] The test data is used to simulate data that causes any abnormality in the computing power operation task status. These data need to be generated according to the previously set abnormal rules. For example: If the abnormal rule for the running state sets that the CPU utilization rate exceeding 90% is abnormal, the CPU utilization rate in the test data can be set to 95%. If the abnormal rule for the performance state sets that the average latency rate exceeding 50 ms is abnormal, the average latency rate in the test data can be set to 60 ms. If the abnormal rule for the traffic state sets that the traffic volatility rate exceeding 20% is abnormal, the traffic volatility rate in the test data can be set to 25%.

[0104] Wait for the test data to be processed and displayed on the visualization interface. The screenshot of the visualization interface, data display content, etc. can be obtained through an automated script or manual operation as the visualization result. For example, using automated testing tools such as Selenium to simulate a user opening the visualization interface and then capturing the interface content. If the corresponding warning information is not found in the visualization result, it is determined that the function of displaying warning information is abnormal.

[0105] Once the function is determined to be abnormal, a preset indication message can be sent. The preset indication message can be in the form of an email, text message, internal system message, etc. For example, send it to the system administrator by email, and the email content can include the detailed information of the test data, the screenshot of the visualization result, and a prompt such as "There may be an abnormality in the function of displaying warning information. Please check!"

[0106] Visualize the operation task status of the computing power in the intelligent computing center in real time to obtain a visualization interface; when any status in the computing power operation task status is abnormal, display an alarm message on the corresponding visualization interface. Through network monitoring of the computing power operation task status in the intelligent computing center, alarm for abnormal situations in real time, reduce the situation where the execution efficiency of the computing power operation task decreases due to any abnormal status in the computing power operation task status, thereby improving the execution efficiency of the computing power operation task. In addition, by simulating the scenario of abnormal status with test data and then viewing the display effect of the visualization interface, the processing ability of the system in the face of abnormal data is verified, and the timeliness and accuracy of the alarm are improved.

[0107] Please refer to Figure 2 , Figure 2 FIG. is a structural diagram of a network monitoring and alarming device for the operation task status of the computing power in the intelligent computing center provided by an embodiment of the present invention. As Figure 2 shown, the network monitoring and alarming device 200 for the operation task status of the computing power in the intelligent computing center includes:

[0108] A visualization module 201, configured to visualize the operation task status of the computing power in the intelligent computing center monitored in real time to obtain a visualization interface, where the computing power operation task status includes an operation status, a performance status, and a traffic status, and the visualization interface includes an operation status interface, a performance status interface, and a traffic status interface;

[0109] An alarm module 202, configured to display an alarm message on the corresponding visualization interface when any status in the computing power operation task status is abnormal.

[0110] In one embodiment, it further includes:

[0111] A first monitoring module, configured to monitor the network link to determine whether there is a first abnormal information, where the first abnormal information includes communication abnormality, and the communication abnormality indicates a link interruption in the network layer;

[0112] A first determination module, configured to determine that the operation status is abnormal when there is the first abnormal information.

[0113] In one embodiment, it further includes:

[0114] A second monitoring module, configured to monitor the network link to determine whether there is a second abnormal information, where the second abnormal information includes at least one of throughput abnormality, average delay rate abnormality, and packet loss rate abnormality. Among them, the throughput abnormality means that the amount of data successfully transmitted per unit time is less than a first threshold, the average delay rate abnormality means that the delay rate of the average time of data from the sending end to the receiving end is greater than a second threshold, and the packet loss rate abnormality means that the proportion of lost data during transmission is greater than a third threshold;

[0115] A second determination module, configured to determine that there is an abnormality in the performance state when there is the second abnormal information.

[0116] In one embodiment, it further includes:

[0117] A third monitoring module, configured to monitor the traffic within a preset time interval to determine whether there is third abnormal information, where the third abnormal information includes an abnormal traffic volatility, and the abnormal traffic volatility indicates that the volatility of the traffic within the preset time interval is greater than a fourth threshold;

[0118] A third determination module, configured to determine that there is an abnormality in the traffic state when there is the third abnormal information.

[0119] In one embodiment, the alarm module includes at least one of the following:

[0120] A first sub-alarm module, configured to, when there is an abnormality in the operating state, display alarm information corresponding to the first abnormal level of the operating state on the operating state interface according to the first abnormal level of the operating state, where the first abnormal level is a level determined according to influence factors corresponding to interruptions of each link in the network layer;

[0121] A second sub-alarm module, configured to, when there is an abnormality in the performance state, display alarm information corresponding to the second abnormal level of the performance state on the performance state interface according to the second abnormal level of the performance state, where the second abnormal level is a level determined according to the deviation degree between the obtained data throughput, average latency rate, and / or packet loss rate and their respective thresholds;

[0122] A third sub-alarm module, configured to, when there is an abnormality in the traffic state, display alarm information corresponding to the third abnormal level of the traffic state on the traffic state interface according to the third abnormal level of the traffic state, where the third abnormal level is a level determined according to the deviation degree between the obtained traffic volatility within a preset time interval and its corresponding threshold.

[0123] In one embodiment, it further includes:

[0124] An acquisition module, configured to, when receiving test data, acquire a visualization result of the test data in the visualization interface, where the test data is data used to simulate an abnormality in any state of the computing power operation task state;

[0125] A sending module, configured to send preset indication information when the visualization result indicates that alarm information corresponding to the test data is not displayed in the visualization interface, where the preset indication information is used to indicate an abnormality in the function of displaying alarm information in the visualization interface.

[0126] The network monitoring and warning device for the computing power operation task status of the intelligent computing center provided by the embodiments of the present invention can implement each process of the above-mentioned network monitoring and warning method for the computing power operation task status of the intelligent computing center. The technical features correspond one by one and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.

[0127] It should be noted that the network monitoring and warning device for the computing power operation task status of the intelligent computing center in the embodiments of the present invention can be a device, or a component, an integrated circuit, or a chip in an electronic device.

[0128] The embodiments of the present invention also provide an electronic device. Refer to Figure 3 , Figure 3 which is a schematic structural diagram of an electronic device provided by the embodiments of the present invention. The electronic device includes a memory 301, a processor 302, and a program or instruction running on the memory 301. When the program or instruction is executed by the processor 302, it can implement Figure 1 any step in the corresponding embodiment of the network monitoring and warning method for the computing power operation task status of the intelligent computing center and achieve the same beneficial effects, which will not be elaborated here.

[0129] Among them, the processor 302 can be a CPU, an ASIC, an FPGA, or a GPU.

[0130] Those of ordinary skill in the art can understand that all or part of the steps to implement the embodiments of the above-mentioned network monitoring and warning method for the computing power operation task status of the intelligent computing center can be completed by hardware related to program instructions, and the program can be stored in a readable medium.

[0131] The embodiments of the present invention also provide a readable storage medium. A computer program is stored on the readable storage medium. When the computer program is executed by a processor, it can implement the above Figure 1 any step in the corresponding embodiment of the network monitoring and warning method for the computing power operation task status of the intelligent computing center and can achieve the same technical effects. To avoid repetition, they will not be elaborated here. The storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0132] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the above Figure 1 each process of the corresponding embodiment of the network monitoring and warning method for the computing power operation task status of the intelligent computing center and can achieve the same technical effects. To avoid repetition, they will not be elaborated here.

[0133] In the embodiments of the present invention, terms such as "first" and "second" are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units is not necessarily limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to such process, method, product, or device. In addition, in this application, the use of "and / or" means at least one of the connected objects. For example, A and / or B and / or C means including the 7 cases of A alone, B alone, C alone, A and B existing together, B and C existing together, A and C existing together, and A, B, and C existing together.

[0134] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising such element.

[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of the various embodiments of this application.

[0136] The above describes the embodiments of this application in conjunction with the drawings, but this application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of this application, those of ordinary skill in the art can also make many forms without departing from the purpose of this application and the scope protected by the claims, and all belong to the protection scope of this application.

Claims

1. A method for network monitoring and alarming of the computing power operation task status in an intelligent computing center, characterized in that, Including: Step S1: Visualize the operation task status of the computing power of the intelligent computing center in real time to obtain a visualization interface. The operation task status of the computing power includes an operation status, a performance status, and a traffic status. The visualization interface includes an operation status interface, a performance status interface, and a traffic status interface. Step S2: When any of the statuses in the operation task status of the computing power is abnormal, display an alarm message on the corresponding visualization interface.

2. The method according to claim 1, wherein Before the step S2, it further includes: Step S3: Monitor the network link to determine whether there is a first abnormal message. The first abnormal message includes communication abnormality, and the communication abnormality indicates a link interruption in the network layer. Step S4: When there is the first abnormal message, determine that the operation status is abnormal.

3. The method according to claim 1, wherein Before the step S2, it further includes: Step S5: Monitor the network link to determine whether there is a second abnormal message. The second abnormal message includes at least one of throughput abnormality, average delay rate abnormality, and packet loss rate abnormality. Among them, the throughput abnormality means that the amount of data successfully transmitted per unit time is less than a first threshold, the average delay rate abnormality means that the delay rate of the average time for data from the sending end to the receiving end is greater than a second threshold, and the packet loss rate abnormality means that the proportion of data lost during transmission is greater than a third threshold. Step S6: When there is the second abnormal message, determine that the performance status is abnormal.

4. The method according to claim 1, characterized in that, Before the step S2, it further includes: Step S7: Monitor the traffic in a preset time interval to determine whether there is a third abnormal message. The third abnormal message includes traffic volatility abnormality, and the traffic volatility abnormality means that the volatility of the traffic in the preset time interval is greater than a fourth threshold. Step S8: When there is the third abnormal message, determine that the traffic status is abnormal.

5. The method according to any one of claims 1 to 4, characterized in that, The step S2 includes at least one of the following: Step S21: When the operation status is abnormal, display the alarm message corresponding to the first abnormal level on the operation status interface according to the first abnormal level of the operation status. The first abnormal level is a level determined according to the influence factors corresponding to each link interruption in the network layer. Step S22: When the performance status is abnormal, display the alarm message corresponding to the second abnormal level on the performance status interface according to the second abnormal level of the performance status. The second abnormal level is a level determined according to the deviation degree between the obtained data throughput, average delay rate, and / or packet loss rate and their respective corresponding thresholds. Step S23: When the traffic status is abnormal, display the alarm message corresponding to the third abnormal level on the traffic status interface according to the third abnormal level of the traffic status. The third abnormal level is a level determined according to the deviation degree between the obtained traffic volatility in the preset time interval and its corresponding threshold.

6. The method according to any one of claims 1 to 4, characterized in that, It further includes: Step S9: When test data is received, obtain the visualization result of the test data in the visualization interface, where the test data is data used to simulate data that causes any abnormality in the computing power operation task status. Step S10: When the visualization result indicates that the warning information corresponding to the test data is not displayed in the visualization interface, send a preset indication information, where the preset indication information is used to indicate that the function of displaying warning information in the visualization interface is abnormal.

7. An intelligent computing center computing power operation task status network monitoring and alarming device, characterized in that, Comprising: A visualization module, configured to visualize the computing power operation task status of the intelligent computing center monitored in real time to obtain a visualization interface, where the computing power operation task status includes an operation status, a performance status, and a traffic status, and the visualization interface includes an operation status interface, a performance status interface, and a traffic status interface; An alarm module, configured to display alarm information in the corresponding visualization interface when any status in the computing power operation task status is abnormal.

8. An electronic device, characterized in that, Comprising: A processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, the steps of the method for network monitoring and alarming of the computing power operation task status of the intelligent computing center according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method for network monitoring and alarming of the computing power operation task status of the intelligent computing center according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that, Including computer instructions, where when the computer instructions are executed by a processor, the steps of the method for network monitoring and alarming of the computing power operation task status of the intelligent computing center according to any one of claims 1 to 6 are implemented.