Computing power resource real-time monitoring method and device of intelligent computing center

By monitoring computing resources in real time in the intelligent computing center and displaying resource indicator data on the operation and maintenance interface, the problem that the intelligent computing center cannot timely obtain the status of computing resources, and the operation and maintenance efficiency of computing resources is improved.

CN120216294APending Publication Date: 2025-06-27DATACANVAS LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510369398.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The intelligent computing center cannot promptly understand the current status of computing power resources, resulting in a long collection and feedback cycle for computing power resource monitoring information, which in turn reduces the operation and maintenance efficiency of computing power resources.

Method used

By monitoring resource information in computing resources in real time, generating resource indicator data, and displaying these data on the operation and maintenance interface, so that users can observe and grasp resource indicators and their changing trends in real time.

Benefits of technology

Real-time monitoring and management of computing power resources is realized, the operation and maintenance efficiency of computing power resources is significantly improved, and the efficiency and reliability of resource utilization are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216294A_ABST
    Figure CN120216294A_ABST
Patent Text Reader

Abstract

The invention provides a computing power resource real-time monitoring method and device for an intelligent computing center, and relates to the technical field of intelligent computing centers, intelligent computing centers and computing power infrastructure, and the method comprises the steps: S1, carrying out the real-time monitoring of resource information in computing power resources under the condition that a computing power operation task is executed based on the computing power resources, and carrying out the real-time monitoring of the resource information in the computing power resources; resource monitoring information of the computing power resources in the process of executing the computing power operation task is obtained, and the resource monitoring information comprises computing resource information, storage resource information and network resource information; s2, generating resource index data of the computing power resource based on resource monitoring information; and S3, displaying the resource index data based on the operation and maintenance interface of the computing power resource. According to the method, a user can grasp the resource index data and the change trend thereof in real time, and the operation and maintenance efficiency of computing power resources can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and particularly relates to a method and device for real-time monitoring of computing power resources in an intelligent computing center. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.

[0003] An "intelligent computing center" refers to a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power and intelligent computing power. An intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0004] The "intelligent computing center" includes but is not limited to the "intelligent computing center".

[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.

[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers". It is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] Currently, in the process of providing computing power services to users, an intelligent computing center cannot timely obtain the computing power resource data of computing power resources at the current moment, resulting in a very long collection and feedback cycle of computing power resource monitoring information, and further leading to a very low operation and maintenance efficiency of computing power resources. Summary of the Invention

[0008] Embodiments of the present invention provide a method and device for real-time monitoring of computing power resources in an intelligent computing center to solve the problem of very low operation and maintenance efficiency of computing power resources in the prior art.

[0009] To solve the above problems, the present invention is implemented as follows:

[0010] In a first aspect, embodiments of the present invention provide a method for real-time monitoring of computing power resources in an intelligent computing center, including:

[0011] Step S1: When performing a computing power operation task based on computing power resources, monitor the resource information in the computing power resources in real time to obtain the resource monitoring information of the computing power resources during the execution of the computing power operation task, where the resource monitoring information includes computing resource information, storage resource information, and network resource information;

[0012] Step S2: Generate resource metric data of the computing power resources based on the resource monitoring information;

[0013] Step S3: Display the resource metric data on the operation and maintenance interface of the computing power resources.

[0014] In one embodiment, the resource monitoring information is the resource information of the computing power resources at a first time point, and the method further includes:

[0015] Step S4: Input the resource metric data into a pre-trained estimation model for estimation to obtain resource metric estimation information, where the resource metric estimation information is used to represent the resource metric estimation data of the computing power resources within a first time period, and the first time period is the time period between the first time point and a second time point, and the second time point is a time point after the first time point;

[0016] Step S5: Display the resource metric estimation information on the operation and maintenance interface of the computing power resources.

[0017] In one embodiment, the resource metric data includes computing resource metric data, storage resource metric data, and network resource metric data;

[0018] Among them, the computing resource metric data is used to indicate the performance of the computing device corresponding to the computing power resources, the storage resource metric data is used to indicate the performance of the storage device corresponding to the computing power resources, and the network resource metric data is used to indicate the performance of the network device corresponding to the computing power resources.

[0019] In one embodiment, the computing resource metric data includes cluster resource utilization rate and load information, the storage resource metric data includes storage capacity usage and disk throughput, and the network resource metric data includes the status information of the network device, network traffic information, and network latency.

[0020] In one embodiment, the method further includes at least one of the following:

[0021] Step S6: Generate a first prompt message when the cluster resource utilization rate is greater than a first threshold;

[0022] Step S7: Generate a second prompt message when the storage capacity usage is greater than a second threshold;

[0023] Step S8, when the network delay is greater than a third threshold, generate a third prompt message.

[0024] In one embodiment, step S3 includes:

[0025] Step S31, generate a visualization chart based on the resource metric data, where the visualization chart is used to represent the resource metric data of the computing power resources during the execution of the computing power operation task;

[0026] Step S32, display the visualization chart based on the operation and maintenance interface of the computing power resources.

[0027] In a second aspect, an embodiment of the present invention further provides a real-time monitoring device for computing power resources of an intelligent computing center, including:

[0028] A monitoring module, configured to, when executing a computing power operation task based on computing power resources, perform real-time monitoring on the resource information in the computing power resources to obtain resource monitoring information of the computing power resources during the execution of the computing power operation task, where the resource monitoring information includes computing resource information, storage resource information, and network resource information;

[0029] A generating module, configured to generate resource metric data of the computing power resources based on the resource monitoring information;

[0030] A first display module, configured to display the resource metric data based on the operation and maintenance interface of the computing power resources.

[0031] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps in the real-time monitoring method for computing power resources of the intelligent computing center described in the first aspect above are implemented.

[0032] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the real-time monitoring method for computing power resources of the intelligent computing center described in the first aspect above are implemented.

[0033] In a fifth aspect, the present invention further provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the steps in the real-time monitoring method for computing power resources of the intelligent computing center described in the first aspect above are implemented.

[0034] In an embodiment of the present invention, when performing a computing power operation task based on computing power resources, the resource information in the computing power resources is monitored in real time to obtain resource monitoring information of the computing power resources during the execution of the computing power operation task, where the resource monitoring information includes computing resource information, storage resource information, and network resource information; resource index data of the computing power resources is generated based on the resource monitoring information; and the resource index data is displayed based on an operation and maintenance interface of the computing power resources. In this way, by monitoring the resource information in the computing power resources in real time and displaying the resource index data based on the operation and maintenance interface of the computing power resources, the user can observe the resource index data on the operation and maintenance interface of the computing power resources, and further, the user can grasp the resource index data and its change trend in real time, thereby greatly improving the operation and maintenance efficiency of the computing power resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0036] Figure 1 is a flowchart of a method for real-time monitoring of computing power resources of an intelligent computing center provided by an embodiment of the present invention;

[0037] Figure 2 is one of the schematic diagrams of the computing resource index data of a computing power resource provided by an embodiment of the present invention;

[0038] Figure 3 is another schematic diagram of the computing resource index data of a computing power resource provided by an embodiment of the present invention;

[0039] Figure 4 is one of the schematic diagrams of the storage resource index data of a computing power resource provided by an embodiment of the present invention;

[0040] Figure 5 is another schematic diagram of the storage resource index data of a computing power resource provided by an embodiment of the present invention;

[0041] Figure 6 is a schematic diagram of the network resource index data of a computing power resource provided by an embodiment of the present invention;

[0042] Figure 7 is a structural diagram of a device for real-time monitoring of computing power resources of an intelligent computing center provided by an embodiment of the present invention;

[0043] Figure 8 is a structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0045] The "computing power" described in the present invention refers to: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of a target result through processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.

[0046] The "computational power" (Computational Power, CP) described in the present invention refers to: the ability of a data center server to process data and achieve result output, a comprehensive index for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 +CP 智能 +CP 超级 .

[0047] The "network power" (NP) described in the present invention refers to: the manifestation of the data transmission ability of computing power facilities, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission within and between data centers, and a comprehensive index for measuring network transmission scheduling ability.

[0048] The "Storage Power (SP)" described in the present invention refers to the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and built-in storage devices of servers. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read / write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.

[0049] The "computing power infrastructure" described in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage power, and can realize the centralized computing, storage, transmission, and application of information.

[0050] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0051] The "computing power" described in the present invention includes general computing power, intelligent computing power, and super computing power.

[0052] The "general computing power" described in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0053] The "intelligent computing power" described in the present invention refers to a computing platform that is scaled up and deployed for various artificial intelligence innovation applications based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing and machine vision.

[0054] The "super computing power" described in the present invention mainly refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.

[0055] The "Intelligent Computing Center" described in the present invention refers to a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.

[0056] The "Intelligent Computing Center" described in the present invention includes, but is not limited to, the "Intelligent Computing Center".

[0057] The "Intelligent Computing Center" described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.

[0058] The "Computing Power Center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0059] The "Supercomputing Center" described in the present invention refers to, that is, the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.

[0060] The "Computing Power Resources" described in the present invention refers to technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.

[0061] The "Computing Power Operation Task" described in the present invention refers to a specific workload or job executed on computing power resources and requiring a certain amount of computing power support, usually involving scenarios such as complex data processing, numerical calculation, model training, or simulation.

[0062] The "Visual Monitoring" described in the present invention refers to converting the real-time performance data of computing, storage, and network devices corresponding to computing power resources into perceivable charts or dashboards through an interactive interface, which can be used to help users quickly understand the status of computing power resources.

[0063] In the prior art, during the process of providing computing power services to users, an intelligent computing center is unable to timely obtain the computing power resource data of computing power resources at the current moment, resulting in a long collection and feedback cycle of computing power resource monitoring information, and further leading to very low operation and maintenance efficiency of computing power resources. In the embodiments of the present invention, by real-time monitoring the resource information in the computing power resources and displaying the resource index data based on the operation and maintenance interface of the computing power resources, users can observe the resource index data on the operation and maintenance interface of the computing power resources. Furthermore, users can grasp the resource index data and its change trend in real time, thereby being able to greatly improve the operation and maintenance efficiency of computing power resources.

[0064] Specifically, please refer to Figure 1 , Figure 1 which is a flowchart of a method for real-time monitoring of computing power resources of an intelligent computing center provided by an embodiment of the present invention. As Figure 1 shown, it includes the following steps:

[0065] Step S1: When a computing power operation task is executed based on the computing power resources, real-time monitor the resource information in the computing power resources to obtain the resource monitoring information of the computing power resources during the execution of the computing power operation task, where the resource monitoring information includes computing resource information, storage resource information, and network resource information.

[0066] In this step, the above-mentioned computing power operation task can be a specific workload or job that is executed on the computing power resources and requires a certain amount of computing power support, and usually involves scenarios such as complex data processing, numerical calculation, model training, or simulation. For example, the above-mentioned computing power operation task can be a model inference task deployed on the computing power resources, such as simultaneously identifying the categories of a large number of objects using a trained image classification model, or it can also be a GPU cluster training an image recognition model, etc.

[0067] The above-mentioned resource monitoring information can be the original monitoring information collected by real-time monitoring, that is, the original performance data collected from the hardware or software layer, and can include fine-grained underlying indicators. For example, the utilization rate of a single CPU core, the number of read and write operations of each disk, etc.

[0068] Step S2: Generate the resource index data of the computing power resources based on the resource monitoring information.

[0069] In this step, the above-mentioned resource metric data can be high-level metrics obtained by aggregating, calculating, and normalizing the original monitoring data. For example, the overall CPU utilization rate of the cluster, the average value of disk throughput, etc. Further, the above-mentioned resource metric data can also be the overall performance value of the system, which can be used to characterize the overall performance of the computing power resources at a certain moment. By integrating multi-dimensional performance data into a single value, the specific value range can be from 0 to 1, and it can be expressed in the form of a percentage. This application does not make specific limitations on this.

[0070] Step S3: Display the resource metric data on the operation and maintenance interface of the computing power resources.

[0071] In this step, the above-mentioned operation and maintenance interface can be an interactive tool for monitoring, managing, and maintaining the computing power resources. Specifically, it can be a visualization platform or software interface for centrally displaying the status data of the computing devices, storage devices, and network devices corresponding to the computing power resources.

[0072] It can be understood that displaying the resource metric data on the operation and maintenance interface of the computing power resources can achieve visual monitoring of the computing power resources.

[0073] In the embodiment of the present invention, when performing a computing power operation task based on the computing power resources, the resource information in the computing power resources is monitored in real time to obtain the resource monitoring information of the computing power resources during the execution of the computing power operation task, where the resource monitoring information includes computing resource information, storage resource information, and network resource information; generating the resource metric data of the computing power resources based on the resource monitoring information; and displaying the resource metric data on the operation and maintenance interface of the computing power resources. In this way, by monitoring the resource information in the computing power resources in real time and displaying the resource metric data on the operation and maintenance interface of the computing power resources, users can observe the resource metric data on the operation and maintenance interface of the computing power resources, and then users can master the resource metric data and its change trend in real time, thereby greatly improving the operation and maintenance efficiency of the computing power resources.

[0074] In one embodiment, the resource monitoring information is the resource information of the computing power resources at the first time point, and the method further includes:

[0075] Step S4: Input the resource metric data into a pre-trained estimation model for estimation to obtain resource metric estimation information, where the resource metric estimation information is used to characterize the resource metric estimation data of the computing power resources within the first time period, and the first time period is the time period between the first time point and the second time point, and the second time point is a time point after the first time point;

[0076] Step S5: Display the resource metric estimation information on the operation and maintenance interface of the computing power resources.

[0077] Specifically, the above-mentioned first time point may refer to the time point of the resource information collected in real time currently, that is, the current time point. The above-mentioned pre-trained estimation model can be used to estimate the resource usage trend within a future time period, that is, the first time period, based on the current resource metric data, so as to obtain the estimated resource metric data at future moments. Specifically, the estimation model may be a time series estimation model based on machine learning, such as Long Short-Term Memory (LSTM), etc.

[0078] It should be noted that the display method of the above-mentioned resource metric estimation information on the operation and maintenance interface based on the computing power resources may be different from the method of displaying the resource metric data. For example, through the distinction between solid lines and dashed lines, the solid line is used to represent historical or real-time data, representing the occurred state, and the dashed line is used to estimate the trend, reflecting the possibility and uncertainty of future changes.

[0079] In the above embodiment, by inputting the resource metric data into the pre-trained estimation model for estimation, obtaining the resource metric estimation information, and displaying the resource metric estimation information based on the operation and maintenance interface of the computing power resources, the user can quickly grasp the future change trend of the resource metric data, and can actively adjust the strategy according to the resource metric estimation information, thereby further improving the operation and maintenance efficiency of the computing power resources.

[0080] In one embodiment, the resource metric data includes computing resource metric data, storage resource metric data, and network resource metric data;

[0081] Among them, the computing resource metric data is used to indicate the performance of the computing device corresponding to the computing power resources, the storage resource metric data is used to indicate the performance of the storage device corresponding to the computing power resources, and the network resource metric data is used to indicate the performance of the network device corresponding to the computing power resources.

[0082] Specifically, the above-mentioned computing devices may include CPUs, GPUs, and server clusters, etc., the above-mentioned storage devices may include hard disks, Solid State Drives (SSDs), and cloud storage services, etc., and the above-mentioned network devices may include routers and switches, etc.

[0083] In the above embodiment, by subdividing the resource metric data into three core resource dimension data of computing resource metric data, storage resource metric data, and network resource metric data, the user can obtain the relevant real-time data of the corresponding resource dimension according to actual needs, and can adjust the corresponding resources according to the metrics of different dimensions.

[0084] In one embodiment, the computing resource metric data includes cluster resource utilization rate and load information, the storage resource metric data includes storage capacity usage and disk throughput, and the network resource metric data includes the status information of the network device, network traffic information, and network latency.

[0085] Specifically, the above-mentioned cluster resource utilization rate can be used to measure the comprehensive utilization rate of the entire computing cluster, including the overall occupancy of resources such as CPU, memory, and GPU. The above-mentioned load information can be used to reflect the balance and real-time pressure status of each node or task allocation. The above-mentioned storage capacity usage can be used to indicate the used space of the storage device, and the above-mentioned disk throughput can be used to measure the disk read and write performance, such as the number of input / output (IO) operations per second and the data transfer rate. The above-mentioned status information of the network device can be used to reflect the operating health of network hardware (such as routers and switches), including whether it is online and the port status. The above-mentioned network traffic information can be used to indicate the data transfer volume and bandwidth utilization rate per unit time. The above-mentioned network latency can be the time required for data to travel back and forth between devices and is used to reflect the network response speed.

[0086] In the above embodiment, by subdividing the metrics of computing resource metric data, storage resource metric data, and network resource metric data, the accuracy of real-time monitoring of computing power resources can be improved.

[0087] In one embodiment, the method further includes at least one of the following:

[0088] Step S6: Generate a first prompt message when the cluster resource utilization rate is greater than a first threshold;

[0089] Step S7: Generate a second prompt message when the storage capacity usage is greater than a second threshold;

[0090] Step S8: Generate a third prompt message when the network latency is greater than a third threshold.

[0091] Specifically, the above-mentioned first prompt message can be used to indicate that the cluster resource utilization rate exceeds the limit, the above-mentioned second prompt message can be used to indicate that the storage capacity usage exceeds the limit, and the above-mentioned third prompt message can be used to indicate that the network latency exceeds the limit.

[0092] In the above embodiment, by generating a first prompt message when the cluster resource utilization rate is greater than a first threshold, generating a second prompt message when the storage capacity usage is greater than a second threshold, and generating a third prompt message when the network latency is greater than a third threshold, automated alarm based on preset thresholds can be realized, which helps users quickly respond to problems and can further improve the operation and maintenance efficiency of computing power resources.

[0093] In one embodiment, step S3 includes:

[0094] Step S31: Generate a visualization chart based on the resource metric data, where the visualization chart is used to characterize the resource metric data of the computing power resources during the execution of the computing power operation task;

[0095] Step S32: Display the visualization chart based on the operation and maintenance interface of the computing power resources.

[0096] Specifically, the above visualization chart can be a time series line chart, a bar chart, etc., and can be used to display the change trend of the resource metric data of the computing power resources over time during the execution of the computing power operation task.

[0097] Exemplarily, Figure 2 and Figure 3 are schematic diagrams of the computing resource metric data of a computing power resource provided by an embodiment of the present invention. Among them, Figure 2 it is specifically a schematic diagram of the overall total load and the overall average CPU usage rate, Figure 3 it is specifically a schematic diagram of the total GPU usage rate, Figure 4 and Figure 5 are schematic diagrams of the storage resource metric data of a computing power resource provided by an embodiment of the present invention. Among them, Figure 4 it is specifically a schematic diagram of the overall total disk and the overall average disk usage rate, Figure 5 it is specifically a schematic diagram of the overall total memory and the overall average memory usage rate, Figure 6 is a schematic diagram of the network resource metric data of a computing power resource provided by an embodiment of the present invention, specifically a schematic diagram of the overall network bandwidth usage.

[0098] In the above embodiment, a visualization chart is generated through the resource metric data, and the visualization chart is displayed based on the operation and maintenance interface of the computing power resources, so that users can intuitively obtain and master the resource metric data of the computing power resources during the execution of the computing power operation task, and then can adjust the computing power resources in a timely manner based on the visualization chart, thereby further improving the operation and maintenance efficiency of the computing power resources.

[0099] Please refer to Figure 7 , Figure 7 is a structural diagram of a real-time monitoring device for the computing power resources of an intelligent computing center provided by an embodiment of the present invention. As Figure 7 shown, the real-time monitoring device 700 for the computing power resources of the intelligent computing center includes:

[0100] The monitoring module 701 is used to, when a computing power operation task is executed based on computing power resources, monitor the resource information in the computing power resources in real time, and obtain resource monitoring information of the computing power resources during the execution of the computing power operation task, where the resource monitoring information includes computing resource information, storage resource information, and network resource information;

[0101] The generating module 702 is used to generate resource index data of the computing power resources based on the resource monitoring information;

[0102] The first display module 703 is used to display the resource index data based on the operation and maintenance interface of the computing power resources.

[0103] In one embodiment, the resource monitoring information is the resource information of the computing power resources at a first time point, and the real-time monitoring device for the computing power resources of the intelligent computing center further includes:

[0104] The estimating module is used to input the resource index data into a pre-trained estimating model for estimation to obtain resource index estimation information, where the resource index estimation information is used to represent the resource index estimation data of the computing power resources within a first time period, and the first time period is the time period between the first time point and a second time point, and the second time point is a time point after the first time point;

[0105] The second display module is used to display the resource index estimation information based on the operation and maintenance interface of the computing power resources.

[0106] In one embodiment, the resource index data includes computing resource index data, storage resource index data, and network resource index data;

[0107] Among them, the computing resource index data is used to indicate the performance of the computing device corresponding to the computing power resources, the storage resource index data is used to indicate the performance of the storage device corresponding to the computing power resources, and the network resource index data is used to indicate the performance of the network device corresponding to the computing power resources.

[0108] In one embodiment, the computing resource index data includes cluster resource utilization rate and load information, the storage resource index data includes storage capacity usage and disk throughput, and the network resource index data includes the status information of the network device, network traffic information, and network latency.

[0109] In one embodiment, the real-time monitoring device for the computing power resources of the intelligent computing center further includes at least one of the following:

[0110] The first generating module is used to generate a first prompt message when the cluster resource utilization rate is greater than a first threshold;

[0111] A second generation module, configured to generate a second prompt message when the storage capacity usage is greater than a second threshold;

[0112] A third generation module, configured to generate a third prompt message when the network delay is greater than a third threshold.

[0113] In one embodiment, the first display module 703 includes:

[0114] A generation unit, configured to generate a visualization chart based on the resource metric data, where the visualization chart is used to characterize the resource metric data of the computing power resources during the execution of the computing power operation task;

[0115] A display unit, configured to display the visualization chart based on the operation and maintenance interface of the computing power resources.

[0116] The real-time monitoring device for computing power resources of the intelligent computing center provided by the embodiments of the present invention can implement each process of the above-mentioned real-time monitoring method for computing power resources of the intelligent computing center. The technical features correspond one by one and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0117] It should be noted that the real-time monitoring device for computing power resources of the intelligent computing center in the embodiments of the present invention can be a device, or a component, an integrated circuit, or a chip in an electronic device.

[0118] The embodiments of the present invention also provide an electronic device. Refer to Figure 8 , Figure 8 is a schematic structural diagram of an electronic device provided by the embodiments of the present invention. The electronic device includes a memory 801, a processor 802, and a program or instruction running on the memory 801. When the program or instruction is executed by the processor 802, it can implement Figure 1 any step in the corresponding embodiment of the real-time monitoring method for computing power resources of the intelligent computing center and achieve the same beneficial effects, which will not be elaborated here.

[0119] Among them, the processor 802 can be a CPU, an ASIC, an FPGA, or a GPU.

[0120] Those of ordinary skill in the art can understand that all or part of the steps of implementing the embodiments of the above-mentioned real-time monitoring method for computing power resources of the intelligent computing center can be completed by hardware related to program instructions, and the program can be stored in a readable medium.

[0121] The embodiments of the present invention also provide a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the above-mentioned Figure 1Any step in the embodiment of the real-time monitoring method for the computing power resources of the corresponding intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. The storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0122] The present invention also provides a computer program product, including computer instructions, which when executed by a processor, implement each process of the above-mentioned Figure 1 corresponding embodiment of the real-time monitoring method for the computing power resources of the intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0123] The terms "first", "second", etc. in the embodiments of the present invention are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. In addition, the terms "comprising" and "having" and any of their variations are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In addition, the use of "and / or" in this application means at least one of the connected objects. For example, A and / or B and / or C means including A alone, B alone, C alone, and A and B both exist, B and C both exist, A and C both exist, and A, B, and C all exist, a total of 7 cases.

[0124] It should be noted that in this article, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not clearly listed, or also includes elements inherent to this process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or device comprising that element.

[0125] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of the various embodiments of the present application.

[0126] The embodiments of the present application are described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A real-time monitoring method for computing resources of an intelligent computing center, characterized in that: include: Step S1: When executing a computing power operation task based on computing power resources, monitor the resource information in the computing power resources in real time to obtain resource monitoring information of the computing power resources in the process of executing the computing power operation task, wherein the resource monitoring information includes computing resource information, storage resource information and network resource information; Step S2: Generate resource indicator data of the computing power resources based on the resource monitoring information; Step S3: Display the resource indicator data based on the operation and maintenance interface of the computing power resources.

2. The method according to claim 1, characterized in that The resource monitoring information is resource information of the computing resource at a first time point, and the method further includes: Step S4: input the resource indicator data into a pre-trained estimation model for estimation to obtain resource indicator estimation information, wherein the resource indicator estimation information is used to characterize the resource indicator estimation data of the computing power resource within a first time period, the first time period being a time period between the first time point and a second time point, and the second time point being a time point after the first time point; Step S5: Display the resource indicator estimation information based on the operation and maintenance interface of the computing power resources.

3. The method according to claim 1, characterized in that The resource indicator data includes computing resource indicator data, storage resource indicator data and network resource indicator data; Among them, the computing resource indicator data is used to indicate the performance of the computing device corresponding to the computing power resources, the storage resource indicator data is used to indicate the performance of the storage device corresponding to the computing power resources, and the network resource indicator data is used to indicate the performance of the network device corresponding to the computing power resources.

4. The method according to claim 3, characterized in that The computing resource indicator data includes cluster resource utilization and load information, the storage resource indicator data includes storage capacity usage and disk throughput, and the network resource indicator data includes status information of the network device, network traffic information and network latency.

5. The method according to claim 4, characterized in that The method further comprises at least one of the following: Step S6: when the cluster resource utilization is greater than a first threshold, generate first prompt information; Step S7: if the storage capacity usage is greater than a second threshold, generate a second prompt message; Step S8: When the network delay is greater than a third threshold, generate a third prompt message.

6. The method according to any one of claims 1 to 5, characterized in that The step S3 comprises: Step S31: generating a visualization chart based on the resource indicator data, wherein the visualization chart is used to represent the resource indicator data of the computing power resource in the process of executing the computing power operation task; Step S32: Display the visualization chart based on the operation and maintenance interface of the computing power resources.

7. A real-time monitoring device for computing resources of an intelligent computing center, characterized in that: include: A monitoring module is used to monitor the resource information in the computing power resources in real time when the computing power operation task is executed based on the computing power resources, and obtain the resource monitoring information of the computing power resources in the process of executing the computing power operation task, wherein the resource monitoring information includes computing resource information, storage resource information and network resource information; A generation module, used to generate resource indicator data of the computing power resources based on resource monitoring information; The first display module is used to display the resource indicator data based on the operation and maintenance interface of the computing power resources.

8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for real-time monitoring of computing resources of an intelligent computing center as described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for real-time monitoring of computing resources of an intelligent computing center according to any one of claims 1 to 6.

10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the real-time monitoring method for computing power resources of an intelligent computing center as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Computing power resource network storage method and device of intelligent computing center cloud platform

    CN120475039A