Visual monitoring method and device for computing power resource system performance of intelligent computing center
By obtaining and visualizing the system performance data of computing power resources in the intelligent computing center, the problem that users cannot directly understand the performance of computing power resources is solved, and automated monitoring and efficient operation and maintenance of computing power resources are realized.
Patent Information
- Application Number
- CN202510369347.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-27
AI Technical Summary
When the intelligent computing center provides computing power services to users, users cannot directly understand the current computing power resource system performance, and the relevant performance data needs to be manually retrieved, resulting in low operation and maintenance efficiency.
When executing computing power running tasks, a collection of system data of computing power resources is obtained, including system performance data at multiple moments, such as CPU usage, GPU usage, memory usage and disk usage. Based on these data, visual charts are generated to characterize the system performance data of computing power resources within the target time period.
The automated collection and visual presentation of computing power resource system performance data is realized, making the system behavior intuitive and easy to understand. Users can quickly master the system performance data of computing power resource and adjust resources in a timely manner, thereby greatly improving operation and maintenance efficiency.
Smart Images

Figure CN120216301A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure, and particularly relates to a method and device for visually monitoring the performance of a computing power resource system of an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power and intelligent computing power. The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", which is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0007] Currently, in the process of providing computing power services to users, users of intelligent computing centers cannot directly obtain the current performance of the computing power resource system, and relevant performance data also needs to be manually retrieved, which is inefficient and cumbersome, and thus leads to a very low operation and maintenance efficiency of computing power resources. Summary of the Invention
[0008] The present invention provides a method and device for visually monitoring the performance of a computing power resource system of an intelligent computing center to solve the problem of very low operation and maintenance efficiency of computing power resources.
[0009] To solve the above technical problems, the present invention is implemented as follows:
[0010] In a first aspect, the present invention provides a method for visually monitoring the performance of a computing power resource system of an intelligent computing center, including:
[0011] Step S1: During the execution of the computing power operation task, obtain the system data set of the computing power resources, where the system data set includes the system performance data of the computing power resources at multiple moments within the target time period, and the system performance data includes at least one of the central processing unit (CPU) usage rate, GPU usage rate, memory usage rate, and disk usage rate;
[0012] Step S2: Based on the system data set, generate at least one first visualization chart, which is used to represent the system performance data of the computing power resources within the target time period.
[0013] In one embodiment, the method further includes:
[0014] Step S3: Based on the system data set, determine the overall performance of the computing power resources at each of the multiple moments, and obtain a plurality of system overall performance values, where the plurality of system overall performance values correspond one-to-one to the multiple moments;
[0015] Step S4: Generate a second visualization chart based on the plurality of system overall performance values, which is used to represent the system overall performance values of the computing power resources within the target time period.
[0016] In one embodiment, the system performance data includes a plurality of performance index data, and step S3 includes:
[0017] Step S31: Determine a plurality of weight values corresponding one-to-one to the plurality of performance index data;
[0018] Step S32: Perform weighted calculation on the plurality of performance index data in the target system performance data and the plurality of weight values to obtain the system overall performance value corresponding to the first moment;
[0019] Wherein, the target system performance data is any one of the system performance data of the multiple moments, and the first moment is the moment corresponding to the target system performance data among the multiple moments.
[0020] In one embodiment, step S1 includes:
[0021] Step S11: Determine the multiple moments within the target time period based on a preset time interval, where the time interval between any two adjacent moments among the multiple moments is the preset time interval;
[0022] Step S12: Monitor the computing power resources according to the multiple moments to obtain the system performance data corresponding to the multiple moments respectively.
[0023] In one embodiment, when the overall performance value of the target system is less than a preset threshold, the second visualization chart includes marker information corresponding to the second moment;
[0024] Wherein, the overall performance value of the target system is any one of the overall performance values of the multiple systems; the second moment is the moment corresponding to the overall performance value of the target system among the multiple moments; the marker information is used to characterize that there is an abnormality in the computing power resources at the second moment; the preset threshold is the threshold corresponding to the second moment.
[0025] In one embodiment, before step S4, the method further includes:
[0026] Step S6: Input multiple historical overall system performance values into a pre-trained model for calculation to obtain multiple preset thresholds corresponding one by one to the multiple moments;
[0027] Wherein, the historical overall system performance value is the overall system performance value of the computing power resources before the target time period.
[0028] In a second aspect, the present invention provides a visualization monitoring device for the system performance of computing power resources in an intelligent computing center, including:
[0029] An acquisition module, configured to acquire a set of system data of the computing power resources during the execution of the computing power operation task, wherein the set of system data includes system performance data of the computing power resources at multiple moments within a target time period, and the system performance data includes at least one of the usage rates of the central processing unit (CPU), GPU, memory, and disk;
[0030] A first generation module, configured to generate at least one first visualization chart based on the set of system data, and the first visualization chart is used to characterize the system performance data of the computing power resources within the target time period.
[0031] In a third aspect, the present invention provides an electronic device, including: a processor, a memory, and a program stored on the memory and executable on the processor, and when the program is executed by the processor, it implements the steps of the visualization monitoring method for the system performance of computing power resources in an intelligent computing center as described in the first aspect above.
[0032] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the visualization monitoring method for the system performance of computing power resources in an intelligent computing center as described in the first aspect above.
[0033] Fifth aspect, the present invention provides a computer program product, including computer instructions, which when executed by a processor, implement the steps of the method for visual monitoring of the computing power resource system performance of the intelligent computing center as described in the first aspect above.
[0034] In an embodiment of the present invention, during the execution of the computing power operation task, a system data set of the computing power resources is obtained, where the system data set includes system performance data of the computing power resources at multiple moments within a target time period, and the system performance data includes at least one of the central processing unit (CPU) usage rate, GPU usage rate, memory usage rate, and disk usage rate; based on the system data set, at least one first visualization chart is generated, and the first visualization chart is used to represent the system performance data of the computing power resources within the target time period. In this way, by collecting the system performance data of the computing power resources at multiple moments within the target time period and presenting it to the user in the form of a visualization chart, the automatic collection and visualization presentation of the system performance data of the computing power resources are realized, making the system behavior intuitive and easy to understand. The user can quickly master the system performance data of the computing power resources within the target time period, and the user can intuitively obtain the change situation of the system performance. Furthermore, the computing power resources can be adjusted in a timely manner based on the visualization chart, thereby greatly improving the operation and maintenance efficiency of the computing power resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0036] Figure 1 is a flowchart of a method for visual monitoring of the computing power resource system performance of an intelligent computing center provided by an embodiment of the present invention;
[0037] Figure 2 is a schematic diagram of a system data set provided by an embodiment of the present invention;
[0038] Figure 3 is a schematic diagram of the overall total memory and overall average memory usage rate of a computing power resource provided by an embodiment of the present invention;
[0039] Figure 4 is a schematic diagram of the overall total disk and overall average disk usage rate of a computing power resource provided by an embodiment of the present invention;
[0040] Figure 5 is a schematic diagram of the overall total load and overall average CPU usage rate of a computing power resource provided by an embodiment of the present invention;
[0041] Figure 6 It is a schematic diagram of the total CPU usage rate of computing power resources provided by an embodiment of the present invention;
[0042] Figure 7 It is a schematic diagram of the overall network bandwidth usage of computing power resources provided by an embodiment of the present invention;
[0043] Figure 8 It is a schematic diagram of the overall system performance of computing power resources provided by an embodiment of the present invention;
[0044] Figure 9 It is a structural diagram of a visualization monitoring device for the system performance of computing power resources in an intelligent computing center provided by an embodiment of the present invention;
[0045] Figure 10 It is a structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0046] Next, the technical solutions in the present invention will be clearly and completely described in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0047] Next, the technical terms involved in the present invention will be briefly described first.
[0048] The "computing power" described in the present invention refers to: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of a target result by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.
[0049] The "computational power" (Computational Power, CP) described in the present invention refers to: the ability of a data center server to process data and achieve result output, a comprehensive index to measure the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 .
[0050] The "Network Power (NP)" described in the present invention refers to: It is an indication of the data transmission capacity of computing power facilities, and is a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling ability.
[0051] The "Storage Power (SP)" described in the present invention refers to: It is the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, and includes external storage devices such as storage arrays and server-internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0052] The "computing power infrastructure" described in the present invention refers to: A new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize centralized computing, storage, transmission, and application of information.
[0053] The "new type of information infrastructure" described in the present invention refers to: It mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0054] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.
[0055] The "general computing power" described in the present invention refers to: The computing ability provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0056] The "intelligent computing power" described in the present invention refers to: For various artificial intelligence innovation applications, a computing platform is deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, etc.
[0057] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0058] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms mainly for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0059] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".
[0060] The "intelligent computing center" described in the present invention, namely the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.
[0061] The "computing power center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, transportation power, and storage power, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0062] The "supercomputing center" described in the present invention refers to: namely the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0063] The "computing power resources" described in the present invention refers to: technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0064] The "computing power operation task" described in the present invention refers to: a specific workload or job executed on computing power resources and requiring a certain amount of computing power support, usually involving scenarios such as complex data processing, numerical calculation, model training, or simulation.
[0065] The "visual monitoring" described in the present invention refers to: converting the original monitoring data into intuitive forms such as charts, dashboards or heat maps, and visually presenting the system performance data through a graphical interface, so that users can quickly and clearly understand the system performance status.
[0066] In the prior art, during the process of an intelligent computing center providing computing power services to users, users cannot directly obtain the current performance of the computing power resource system, and the relevant performance data also needs to be manually retrieved, which is inefficient and cumbersome, resulting in a very low operation and maintenance efficiency of the computing power resources. In the embodiments of the present invention, by collecting the system performance data of the computing power resources at multiple moments within a target time period and presenting it to users in the form of visual charts, the automatic collection and visual presentation of the system performance data of the computing power resource system are realized, making the system behavior intuitive and easy to understand. Users can quickly master the system performance data of the computing power resources within the target time period, and users can intuitively obtain the change situation of the system performance. Furthermore, users can make timely adjustments to the computing power resources based on the visual charts, thereby being able to greatly improve the operation and maintenance efficiency of the computing power resources.
[0067] Specifically, please refer to Figure 1 , Figure 1 which is a flowchart of a method for visual monitoring of the system performance of the computing power resources of an intelligent computing center provided by an embodiment of the present invention. As Figure 1 shown, it includes the following steps:
[0068] Step S1: During the execution of the computing power operation task, obtain the system data set of the computing power resources, where the system data set includes the system performance data of the computing power resources at multiple moments within a target time period, and the system performance data includes at least one of the central processing unit (CPU) usage rate, GPU usage rate, memory usage rate, and disk usage rate.
[0069] In this step, the above-mentioned computing power operation task can be a specific workload or job that is executed on the computing power resources and requires a certain amount of computing power support, and usually involves scenarios such as complex data processing, numerical calculation, model training, or simulation. For example, the above-mentioned computing power operation task can be a model inference task deployed on the computing power resources, such as simultaneously identifying the categories of a large number of objects using a trained image classification model, or it can be a GPU cluster training an image recognition model, etc.
[0070] The above-mentioned system performance data can be a set of quantitative indicators used to reflect the resource usage situation, efficiency, and health status of the computing power resources during operation, and specifically can include at least one performance index data. For example, the central processing unit (CPU) usage rate, GPU usage rate, memory usage rate, disk usage rate, network bandwidth usage, load average, and throughput, etc.
[0071] The above target time period can be a predefined monitoring cycle. For example, the target time period can be set to monitor the system performance of the computing power resources in real time. In this case, the target time period includes the current time point and the past time period. The target time period can also be set to the past 24 hours, the peak period of the computing power operation task, and the comparison period before and after the computing power task runs, etc.
[0072] The system performance data at the above multiple moments can be collected regularly, such as once a minute, to form a time series data set. Figure 2 It is a schematic diagram of a system data set provided by an embodiment of the present invention. As Figure 2 shown, the system data set includes the system performance data of the computing power resources at time t1, time t2, and time t3 respectively. Among them, the system performance data includes CPU usage rate, GPU usage rate, memory occupancy rate, etc. Time t4 represents a future time.
[0073] Step S2: Based on the system data set, generate at least one first visualization chart, and the first visualization chart is used to represent the system performance data of the computing power resources within the target time period.
[0074] In this step, the above first visualization chart can be used to display the system performance data of the computing power resources within the target time period and the change trend over time. The types of the above first visualization chart include time series line chart, bar chart, resource utilization stacked chart, etc. Among them, the time series line chart and bar chart can be used to display the change trend of the system performance indicators of the computing power resources over time within the target time period, and the resource utilization stacked chart can be used to compare the occupancy ratios of different resource types in the time dimension.
[0075] It can be understood that the above first visualization chart can be a simple trend chart of a single indicator or a composite chart of multiple indicators superimposed or compared. When the above first visualization chart is a single - indicator trend chart, multiple first visualization charts can be generated based on the system data set, and the system performance data in the multiple first visualization charts is different. For example, the first visualization chart can be a line chart showing the change of CPU usage rate over time; when the above first visualization chart includes multiple indicators, only one first visualization chart can be generated based on the system data set, and the first visualization chart includes multiple system performance data. For example, the first visualization chart can be a line chart with the superposition of CPU usage rate and the number of cores.
[0076] Exemplarily, Figure 3 It is a schematic diagram of the overall total memory and overall average memory usage rate of a computing power resource provided by an embodiment of the present invention. Figure 4 It is a schematic diagram of the overall total disk and overall average disk usage rate of a computing power resource provided by an embodiment of the present invention.Figure 5 It is a schematic diagram of the overall total load of computing power resources and the overall average CPU usage rate provided by an embodiment of the present invention. Figure 6 It is a schematic diagram of the total CPU usage rate of computing power resources provided by an embodiment of the present invention. Figure 7 It is a schematic diagram of the overall network bandwidth usage of computing power resources provided by an embodiment of the present invention.
[0077] In an embodiment of the present invention, during the execution of a computing power operation task, a system data set of computing power resources is obtained, where the system data set includes system performance data of the computing power resources at multiple moments within a target time period, and the system performance data includes at least one of the central processing unit (CPU) usage rate, GPU usage rate, memory usage rate, and disk usage rate; based on the system data set, at least one first visualization chart is generated, and the first visualization chart is used to represent the system performance data of the computing power resources within the target time period. In this way, by collecting the system performance data of the computing power resources at multiple moments within the target time period and presenting it to the user in the form of a visualization chart, the automatic collection and visualization presentation of the system performance data of the computing power resources are realized, making the system behavior intuitive and easy to understand. The user can quickly master the system performance data of the computing power resources within the target time period, and the user can intuitively obtain the change situation of the system performance. Furthermore, the computing power resources can be adjusted in a timely manner based on the visualization chart, thereby greatly improving the operation and maintenance efficiency of the computing power resources.
[0078] In one embodiment, the method further includes:
[0079] Step S3: Based on the system data set, determine the overall performance of the computing power resources at each of the multiple moments to obtain multiple system overall performance values, where the multiple system overall performance values correspond one-to-one to the multiple moments.
[0080] In this step, the above system overall performance value is used to represent the overall performance of the computing power resources at a certain moment. By integrating multi-dimensional system performance data into a single value corresponding to each moment, the specific value range can be 0 - 1, and it can be expressed in the form of a percentage. The present application does not make specific limitations on this.
[0081] Step S4: Generate a second visualization chart based on the multiple system overall performance values, and the second visualization chart is used to represent the system overall performance values of the computing power resources within the target time period.
[0082] In this step, the above second visualization chart can be a time series line chart, a bar chart, etc., and can be used to display the change trend of the system overall performance value over time within the target time period. Exemplarily, Figure 8It is a schematic diagram of the overall system performance of computing power resources provided by an embodiment of the present invention.
[0083] In the above embodiment, the overall system performance value of the computing power resources at each moment is determined through the system data set of the computing power resources and presented to the user in the form of a visualization chart, so that the user can directly obtain and master the overall system performance value of the computing power resources within the target time period without having to retrieve and calculate again, and thus can adjust the computing power resources in a timely manner based on the visualization chart, thereby further improving the operation and maintenance efficiency of the computing power resources.
[0084] In one embodiment, the system performance data includes multiple performance index data, and the step S3 includes:
[0085] Step S31: Determine a plurality of weight values corresponding one by one to the multiple performance index data;
[0086] Step S32: Perform weighted calculation on the multiple performance index data in the target system performance data and the multiple weight values to obtain the overall system performance value corresponding to the first moment;
[0087] Wherein, the target system performance data is any one of the system performance data at the multiple moments, and the first moment is the moment corresponding to the target system performance data among the multiple moments.
[0088] Specifically, the above weight values can be used to reflect the contribution degree of each performance index data index to the overall system performance, and the specific value ranges from 0 to 1. It can be understood that the above weight values can adjust the weights according to task requirements or resource importance. For example, in the AI training scenario, the weight of GPU usage rate can be set relatively high.
[0089] Exemplarily, assume that a certain computing power operation task needs to monitor three performance indexes: CPU usage rate, GPU utilization rate, and memory occupancy rate. Among them, the current moment value of CPU usage rate is 80%, the weight is 30%, the current moment value of GPU utilization rate is 75%, the weight is 50%, and the current moment value of memory occupancy rate is 60%, the weight is 20%. It can be calculated that the current moment value of the overall system performance value = (80 * 0.3) + (75 * 0.5) + (60 * 0.2) = 73.5.
[0090] In the above embodiment, by assigning weights to key system performance indexes and performing weighted calculation, the overall system performance value at each moment is automatically generated. The user can directly grasp the overall state and the change trend of the overall state of the computing power resources without manually retrieving and calculating relevant data, thereby further improving the operation and maintenance efficiency of the computing power resources.
[0091] In one embodiment, the step S1 includes:
[0092] Step S11: Determine the multiple moments within the target time period based on a preset time interval, where, among the multiple moments, the time interval between any two adjacent moments is the preset time interval;
[0093] Step S12: Monitor the computing power resources according to the multiple moments to obtain the system performance data corresponding to the multiple moments respectively.
[0094] It can be understood that the determination of the above preset time interval can be combined with specific scenario requirements and system limitations. The length of the preset time interval will directly affect the accuracy of data collection and related resource consumption.
[0095] Specifically, the process of monitoring the computing power resources according to the multiple moments can be: within the target time period, a monitoring point is generated every preset time interval. For example, if the preset interval is 5 minutes and the time period is 60 minutes, then a total of 12 moments are generated. At each generated moment point, a monitoring instruction is triggered to call the system monitoring tool, and then the data of the current state of the computing power resources is obtained.
[0096] In the above embodiment, by sampling at a fixed time interval, the performance of the computing power resources within the target time period is systematically monitored, so that the generated time series data can be used to draw the first visualization chart and directly calculate the overall system performance value of the computing power resources, providing good data support for the subsequent generation of the visualization chart, thereby making the generated visualization chart more accurate.
[0097] In one embodiment, when the overall target system performance value is less than a preset threshold, the second visualization chart includes marking information corresponding to a second moment;
[0098] Wherein, the overall target system performance value is any one of the multiple overall system performance values; the second moment is the moment corresponding to the overall target system performance value among the multiple moments; the marking information is used to characterize that there is an abnormality in the computing power resources at the second moment; the preset threshold is the threshold corresponding to the second moment.
[0099] Specifically, the above preset threshold can be a critical value set according to business requirements or historical data. When the overall system performance value is lower than the preset threshold, it indicates that there may be an abnormality in the computing power resources, such as insufficient computing power resources or service failures, etc. The above second moment can be the specific time point when the overall system performance value first or continuously falls below the preset threshold.
[0100] The above marking information can be marked by symbols or text in the visualization chart, explicitly reminding the user at the corresponding time point, i.e., the second moment, that there may be an abnormality in the computing power resources at this moment. Among them, the symbols can be red triangles, warning icons, etc., and the text markings can include prompt instructions, reasons for abnormalities, or corresponding suggestions, etc.
[0101] In the above embodiment, by comparing the current overall system performance value with the preset threshold at this moment, when the overall system performance value is lower than the preset threshold, this moment is marked, so that when the overall system performance of the computing power resources drops suddenly, the abnormal time point can be quickly located through the marking information, and the reason can be analyzed in combination with other monitoring tools to solve the abnormality, thereby ensuring the stable operation of the computing power resources.
[0102] In one embodiment, before step S4, the method further includes:
[0103] Step S6: Input multiple historical overall system performance values into a pre-trained model for calculation to obtain multiple preset thresholds corresponding to the multiple moments one by one;
[0104] Among them, the historical overall system performance value is the overall system performance value of the computing power resources before the target time period.
[0105] Specifically, the above historical overall system performance value can be the overall system performance value before the target time period. Exemplarily, assuming that the current monitored overall system performance value is for "2025-03-05", then the historical overall system performance value can include the overall system performance value from "2025-02-04 to 2025-03-04".
[0106] The above pre-trained model can be used to calculate the threshold for future moments based on the historical overall system performance value. The model can specifically be a time series calculation model such as a Long Short-Term Memory (LSTM) neural network, etc., or a statistical regression model.
[0107] It should be noted that for each moment within the target time period, the model will generate a dynamic threshold instead of a fixed value. Exemplarily, if the monitoring period, i.e., the target time period, is one day of "2025-03-05" (each hour is a moment), then the model will output 24 different thresholds, and each threshold corresponds to the lower limit of the normal overall system performance value for a certain hour.
[0108] In the above embodiment, by dynamically generating a preset threshold corresponding to a specific moment through a machine learning model, the adaptability of monitoring the system performance of computing power resources can be improved, and the judgment criteria can be automatically adjusted according to historical laws, thereby being able to more accurately identify the abnormal moments that may exist in the computing power resources.
[0109] Please refer to Figure 9 , Figure 9 which is a structural diagram of a visualization monitoring device for the computing power resource system performance of an intelligent computing center provided by an embodiment of the present invention. As Figure 9 shown, the visualization monitoring device 900 for the computing power resource system performance of the intelligent computing center includes:
[0110] An acquisition module 901, configured to acquire a system data set of computing power resources during the execution of a computing power operation task, where the system data set includes system performance data of the computing power resources at multiple moments within a target time period, and the system performance data includes at least one of the central processing unit (CPU) usage rate, GPU usage rate, memory usage rate, and disk usage rate;
[0111] A first generation module 902, configured to generate at least one first visualization chart based on the system data set, where the first visualization chart is used to represent the system performance data of the computing power resources within the target time period.
[0112] In one embodiment, the visualization monitoring device 900 for the computing power resource system performance of the intelligent computing center further includes:
[0113] A determination module, configured to determine the overall performance of the computing power resources at each of the multiple moments based on the system data set, and obtain a plurality of system overall performance values, where the plurality of system overall performance values correspond one-to-one to the multiple moments;
[0114] A second generation module, configured to generate a second visualization chart based on the plurality of system overall performance values, where the second visualization chart is used to represent the system overall performance values of the computing power resources within the target time period.
[0115] In one embodiment, the system performance data includes a plurality of performance index data, and the determination module includes:
[0116] A first determination unit, configured to determine a plurality of weight values corresponding one-to-one to the plurality of performance index data;
[0117] A calculation unit, configured to perform weighted calculation on the plurality of performance index data in the target system performance data and the plurality of weight values to obtain the system overall performance value corresponding to the first moment;
[0118] where the target system performance data is any one of the system performance data of the multiple moments, and the first moment is the moment corresponding to the target system performance data among the multiple moments.
[0119] In one embodiment, the acquisition module includes:
[0120] A second determination unit, configured to determine the multiple moments within the target time period based on a preset time interval, wherein, among the multiple moments, the time interval between any two adjacent moments is the preset time interval;
[0121] A monitoring unit, configured to monitor the computing power resources according to the multiple moments, and obtain system performance data respectively corresponding to the multiple moments.
[0122] In one embodiment, when the overall performance value of the target system is less than a preset threshold, the second visualization chart includes marking information corresponding to a second moment;
[0123] Wherein, the overall performance value of the target system is any one of the multiple overall performance values of the system; the second moment is the moment corresponding to the overall performance value of the target system among the multiple moments; the marking information is used to characterize that there is an abnormality in the computing power resources at the second moment; the preset threshold is the threshold corresponding to the second moment.
[0124] In one embodiment, the visualization monitoring device 900 for the system performance of the computing power resources of the intelligent computing center further includes:
[0125] A calculation module, configured to input multiple historical overall performance values of the system into a pre-trained model for calculation, and obtain multiple preset thresholds corresponding one by one to the multiple moments;
[0126] Wherein, the historical overall performance value of the system is the overall performance value of the computing power resources before the target time period.
[0127] The visualization monitoring device for the system performance of the computing power resources of the intelligent computing center provided by the embodiments of the present invention can implement each process of the above-mentioned visualization monitoring method for the system performance of the computing power resources of the intelligent computing center. The technical features correspond one by one and can achieve the same technical effects. To avoid repetition, they are not described here again.
[0128] It should be noted that the visualization monitoring device for the system performance of the computing power resources of the intelligent computing center in the embodiments of the present invention can be a device, or a component, an integrated circuit, or a chip in an electronic device.
[0129] The embodiments of the present invention further provide an electronic device. Refer to Figure 10 , Figure 10 is a schematic structural diagram of an electronic device provided by the embodiments of the present invention. The electronic device includes a memory 1001, a processor 1002, and a program or instruction stored in the memory 1001 and running on the processor 1002. When the program or instruction is executed by the processor 1002, it can implement Figure 1Any steps in the embodiments of the method for visual monitoring of the computing power resource system performance of the corresponding intelligent computing center and the same beneficial effects are not elaborated here.
[0130] Among them, the processor 1002 can be a CPU, ASIC, FPGA or GPU.
[0131] Those of ordinary skill in the art can understand that all or part of the steps of implementing the embodiments of the method for visual monitoring of the computing power resource system performance of the intelligent computing center can be completed by hardware related to program instructions, and the program can be stored in a readable medium.
[0132] The embodiments of the present invention also provide a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement any step in the embodiments of the above Figure 1 corresponding method for visual monitoring of the computing power resource system performance of the intelligent computing center, and can achieve the same technical effects. To avoid repetition, it is not elaborated here. The storage medium such as a Read-Only Memory (ROM), Random Access Memory (RAM), magnetic disk or optical disc, etc.
[0133] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, it implements each process in the embodiments of the above Figure 1 corresponding method for visual monitoring of the computing power resource system performance of the intelligent computing center, and can achieve the same technical effects. To avoid repetition, it is not elaborated here.
[0134] The terms "first", "second", etc. in the embodiments of the present invention are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In addition, in this application, "and / or" is used to represent at least one of the connected objects. For example, A and / or B and / or C represents 7 cases including A alone, B alone, C alone, A and B existing together, B and C existing together, A and C existing together, and A, B, and C existing together.
[0135] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including such element.
[0136] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of this application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of the various embodiments of this application.
[0137] The embodiments of this application are described above in conjunction with the accompanying drawings. However, this application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this application, those of ordinary skill in the art can also make many forms without departing from the purpose of this application and the scope protected by the claims, and all of them fall within the protection scope of this application.
Claims
1. A method for visually monitoring the performance of a computing resource system in an intelligent computing center, characterized in that: include: Step S1: in the process of executing the computing power operation task, a system data set of the computing power resources is obtained, wherein the system data set includes system performance data of the computing power resources at multiple times within a target time period, and the system performance data includes at least one of a central processing unit CPU usage rate, a GPU usage rate, a memory usage rate, and a disk usage rate; Step S2: Based on the system data set, generate at least one first visualization chart, where the first visualization chart is used to represent the system performance data of the computing power resources within the target time period.
2. The method according to claim 1, characterized in that The method further comprises: Step S3: determining the overall performance of the computing resource at each of the multiple moments based on the system data set, and obtaining multiple system overall performance values, wherein the multiple system overall performance values correspond to the multiple moments one by one; Step S4: Generate a second visualization chart based on the multiple system overall performance values, where the second visualization chart is used to represent the system overall performance value of the computing power resource within the target time period.
3. The method according to claim 2, characterized in that The system performance data includes a plurality of performance indicator data, and the step S3 includes: Step S31: determining a plurality of weight values corresponding one-to-one to the plurality of performance indicator data; Step S32: performing weighted calculation on the multiple performance indicator data in the target system performance data and the multiple weight values to obtain the overall system performance value corresponding to the first moment; The target system performance data is any one of the system performance data at the multiple moments, and the first moment is a moment among the multiple moments corresponding to the target system performance data.
4. The method according to claim 1, characterized in that: The step S1 comprises: Step S11: determining the multiple moments within the target time period based on a preset time interval, wherein the time interval between any two adjacent moments in the multiple moments is the preset time interval; Step S12: Monitor the computing resources at the multiple moments to obtain system performance data corresponding to the multiple moments respectively.
5. The method according to claim 2, characterized in that: When the overall performance value of the target system is less than a preset threshold, the second visualization chart includes marking information corresponding to the second moment; Among them, the target system overall performance value is any one of the multiple system overall performance values; the second moment is the moment among the multiple moments corresponding to the target system overall performance value; the marking information is used to characterize that there is an abnormality in the computing power resource at the second moment; and the preset threshold is the threshold corresponding to the second moment.
6. The method according to claim 5, characterized in that Before step S4, the method further includes: Step S6: inputting a plurality of historical system overall performance values into a pre-trained model for calculation to obtain a plurality of preset thresholds corresponding to the plurality of moments; The historical system overall performance value is the system overall performance value of the computing power resource before the target time period.
7. A visual monitoring device for computing resource system performance of an intelligent computing center, characterized in that: include: An acquisition module is used to acquire a system data set of computing resources during the execution of a computing power operation task, wherein the system data set includes system performance data of the computing resources at multiple times within a target time period, and the system performance data includes at least one of a CPU usage rate, a GPU usage rate, a memory usage rate, and a disk usage rate; The first generating module is used to generate at least one first visualization chart based on the system data set, where the first visualization chart is used to represent the system performance data of the computing power resource within the target time period.
8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for visually monitoring the performance of a computing resource system of an intelligent computing center as described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for visually monitoring the performance of a computing resource system of an intelligent computing center according to any one of claims 1 to 6.
10. A computer program product, characterized in that It comprises computer instructions, which, when executed by a processor, implement the steps of the method for visually monitoring the performance of a computing resource system of an intelligent computing center as described in any one of claims 1 to 6.
Citation Information
Cited By
Multi-dimensional performance index-based computing power metering method and device for intelligent computing center cloud platform
CN121277792A