Computing power resource operation data visualization method and device of intelligent computing center
By building a visual interface in the intelligent computing center and automatically analyzing the root cause, the problem of users needing to frequently switch interfaces to obtain information is solved, the operation and maintenance efficiency and resource management capabilities are improved, and the global system state awareness is realized.
Patent Information
- Application Number
- CN202510440777.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the monitoring of the intelligent computing center relies on multiple independent tools, resulting in users needing to frequently switch interfaces to obtain complete information, and the correlation analysis of key indicators is inefficient.
By deploying monitoring components in the intelligent computing center, obtaining the operating index values, building a visual interface, combining predefined threshold comparison, automatically analyzing the root cause and positioning the resource dimension with the highest correlation, forming a global system visualization cognition.
It has achieved improvement in operation and maintenance efficiency, quickly identified abnormal indicators, reduced manual investigation time, optimized resource management capabilities, and improved the correlation analysis efficiency of key indicators.
Smart Images

Figure CN120336601A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and particularly relates to a method and device for visualizing the operation data of computing power resources of an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, to mainly provide the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios for artificial intelligence deep learning model development, model training, and model inference, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that is based on artificial intelligence theory, adopts an artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", is the ability of a computer device or a computing / data center to process information, is the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, is the computing ability to achieve the output of a target result through processing information data, is a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0007] At present, with the rapid development of information technology, the real-time monitoring and resource management of intelligent computing centers have become the core requirements for ensuring business continuity and operation and maintenance efficiency. In the prior art, system monitoring usually relies on scattered local tools or traditional monitoring platforms, such as a data display system based on multi-window split screens, a static chart analysis tool, and a report generation scheme based on a fixed template. These technical solutions generally adopt a hierarchical menu operation and multi-module switching mode, obtain resource status information through periodic data polling or batch processing methods, and present historical data trends in basic forms such as tables and line charts. Since the emergence of intelligent computing centers, due to relying on multiple independent tools to implement monitoring functions in different dimensions, users need to frequently switch interfaces to obtain complete information, resulting in low efficiency in the correlation analysis of key indicators. Therefore, the problem that it is difficult for intelligent computing centers to form a global system visualization cognition is an urgent problem to be solved. Summary of the Invention
[0008] An embodiment of the present invention provides a method and device for visualizing the operation data of computing power resources in an intelligent computing center, so as to solve the problem in the prior art that multiple independent tools are relied on to implement monitoring functions in different dimensions, and users need to frequently switch interfaces to obtain complete information, resulting in low efficiency in the correlation analysis of key indicators.
[0009] To solve the above problems, the present invention is implemented as follows:
[0010] In a first aspect, an embodiment of the present invention provides a method for visualizing the operation data of computing power resources in an intelligent computing center, including:
[0011] Step S1: Based on the monitoring components pre-deployed in the intelligent computing center, obtain the values corresponding to at least one operation indicator of the intelligent computing center;
[0012] Step S2: Construct a visualization interface according to the values corresponding to the at least one operation indicator, the predefined thresholds corresponding to the at least one operation indicator, and the comparison results corresponding to the at least one operation indicator; wherein, the comparison results are determined according to the values corresponding to the operation indicators and the predefined thresholds corresponding to the operation indicators;
[0013] Step S3: When it is detected that the target comparison result in the visualization interface shows an anomaly, determine the root cause analysis result of the anomaly shown by the target comparison result according to the real-time values and historical values corresponding to the target operation indicator collected by the monitoring component, and the comparison results corresponding to the at least one operation indicator include the target comparison result;
[0014] Step S4: According to the root cause analysis result, display the resource dimension with the highest correlation with the anomaly shown by the target comparison result in the visualization interface.
[0015] In one embodiment, the step S1 includes at least one of the following:
[0016] Step S11: Based on the monitoring components pre-deployed in at least one first target node, obtain the values corresponding to at least one operation indicator in the at least one first target node;
[0017] Step S12: Monitor at least one second target node based on the Prometheus service discovery mechanism, and obtain the values corresponding to at least one operation indicator in the at least one second target node;
[0018] Wherein, the first target node is a node in the intelligent computing center with a fixed Internet Protocol (IP) address or Domain Name System (DNS) endpoint, and the second target node is a node in the intelligent computing center with a dynamically changing IP address or DNS endpoint.
[0019] In one embodiment, the visualization interface includes at least one of a sub-dashboard view, a cross-level metric comparison chart, a cross-level association analysis channel, and a pop-up comparison chart;
[0020] The step S2 includes at least one of the following:
[0021] Step S21: Based on the label or directory level in the intelligent computing center selected by the user, construct at least one sub-dashboard view associated with the label or the directory level; wherein, each sub-dashboard view corresponds to a resource type or service module in the intelligent computing center, and the sub-dashboard view is used to display the values corresponding to multiple running metrics associated with the resource type or the service module;
[0022] Step S22: In response to a first operation of the user on a target service module in the intelligent computing center, construct a cross-level metric comparison chart associated with the target service module; wherein, the cross-level metric comparison chart includes a composite view of a time series line chart and a heat map, and the cross-level metric comparison chart is used to synchronously present the real-time load and historical trend of multiple running metrics associated with the target service module;
[0023] Step S23: In response to a second operation of the user on a target node in the intelligent computing center, construct a cross-level association analysis channel; wherein, the cross-level association analysis channel is used to display the values corresponding to the running metrics of at least one associated node having a data interaction relationship with the target computing node;
[0024] Step S24: Construct a pop-up comparison chart in the visualization interface; wherein, the pop-up comparison chart is used to present the correlation data of the running metrics of the target computing node and the at least one associated node in the time dimension and the space dimension, and the correlation data includes at least one of the utilization rate of the graphics processing unit (GPU), the occupancy rate of the central processing unit (CPU) cache, and the cross-node communication delay.
[0025] In one embodiment, the step S3 includes:
[0026] Step S31: In the case where it is monitored that the target comparison result shows an anomaly, obtain the historical values within a preset time period before the target comparison result shows an anomaly;
[0027] Step S32: Extract at least one candidate metric having a predefined dependency relationship with the target running metric from other running metrics collected by the monitoring component, construct a multi-dimensional feature matrix for the time series of the at least one candidate metric, and calculate the correlation degree value of the residual term between each candidate metric and the target running metric;
[0028] Step S33: Determine at least one resource dimension for which the target comparison result shows an anomaly and the contribution degree ranking of the at least one resource dimension according to the correlation degree value and the logical distance weights between nodes in the preset resource dependency topology graph;
[0029] Step S34: Screen out the resource dimensions whose confidence scores exceed the preset threshold according to the contribution degree ranking of the at least one resource dimension to obtain the root cause analysis result.
[0030] In one embodiment, the step S4 includes:
[0031] Step S41: When the user selects any target resource dimension marked as the root cause in the root cause analysis result, display the dependency relationship link of the target resource dimension in the visualization interface:
[0032] Step S42: Determine the shortest path corresponding to the dependency relationship link according to the physical connection or task scheduling logic in the node topology model of the intelligent computing center;
[0033] Step S43: Display the shortest path on the visualization interface and parallelly display the index trend curve of the target resource dimension on the shortest path in the form of a time axis in the visualization interface.
[0034] In one embodiment, the method further includes:
[0035] Step S5: Perform image processing on the visualization interface for the node corresponding to the target operation index or the associated communication link in the intelligent computing center;
[0036] Wherein, the image processing includes performing visual enhancement processing on the display effect of the visualization interface.
[0037] In a second aspect, an embodiment of the present invention further provides a visualization device for the computing power resource operation data of an intelligent computing center, including:
[0038] An index acquisition module, configured to acquire the values corresponding to at least one operation index of the intelligent computing center based on a monitoring component pre-deployed in the intelligent computing center;
[0039] An interface construction module, configured to construct a visualization interface according to the values corresponding to the at least one operation index, the predefined thresholds corresponding to the at least one operation index, and the comparison results corresponding to the at least one operation index; wherein, the comparison result is determined according to the value corresponding to the operation index and the predefined threshold corresponding to the operation index;
[0040] A result determination module, configured to, when it is detected that the display of the target comparison result in the visualization interface is abnormal, determine a root cause analysis result of the abnormal display of the target comparison result according to the real-time value and historical value of the target operation index collected by the monitoring component, where the comparison result corresponding to the at least one operation index includes the target comparison result;
[0041] An interface display module, configured to display, in the visualization interface, a resource dimension with the highest correlation with the abnormal display of the target comparison result according to the root cause analysis result.
[0042] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps in the method for visualizing the operation data of the computing power resources of the intelligent computing center as described in the first aspect above are implemented.
[0043] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method for visualizing the operation data of the computing power resources of the intelligent computing center as described in the first aspect above are implemented.
[0044] In a fifth aspect, the present invention further provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the steps in the method for visualizing the operation data of the computing power resources of the intelligent computing center as described in the first aspect above are implemented.
[0045] In the embodiments of the present invention, by collecting the operation index data of the intelligent computing center and constructing a visualization interface in combination with threshold comparison, the intuitive presentation of the abnormal state is realized. When an abnormality is detected, the root cause is automatically analyzed and the resource dimension with the highest correlation is located, forming a closed-loop management from monitoring, early warning to fault root cause diagnosis. Thus, the embodiments of the present invention improve the operation and maintenance efficiency, quickly identify abnormal indicators through automatic comparison and visual display, reduce the manual troubleshooting time, realize accurate fault location, and directly point to the core resource dimension of the problem based on the comparative analysis of real-time and historical data, improve the correlation analysis efficiency of indicators, optimize the resource management ability, improve the correlation analysis efficiency of key indicators, and form a global system state awareness. Description of the Drawings
[0046] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a flowchart of a method for visualizing the operation data of computing power resources in an intelligent computing center provided by an embodiment of the present invention;
[0048] Figure 2 It is one of the schematic diagrams of the visualization interface provided by an embodiment of the present invention;
[0049] Figure 3 It is the second of the schematic diagrams of the visualization interface provided by an embodiment of the present invention;
[0050] Figure 4 It is a schematic diagram of a device for visualizing the operation data of computing power resources in an intelligent computing center provided by an embodiment of the present invention;
[0051] Figure 5 It is a structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] The "computing power" as described in the present invention refers to: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to process information data and output a target result, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, which is mainly provided to the society through computing power infrastructure.
[0054] The "computational power" (Computational Power, CP) as described in the present invention refers to: the ability of a data center server to process data and output a result, which is a comprehensive indicator for measuring the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 .
[0055] The "Network Power (NP)" described in the present invention refers to: the manifestation of the data transmission capacity of computing power facilities, which is a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., and involves network transmission within and between data centers. It is a comprehensive indicator for measuring network transmission scheduling ability.
[0056] The "Storage Power (SP)" described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage ability of a data center, including external storage devices such as storage arrays and internal storage devices of servers. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0057] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize centralized computing, storage, transmission, and application of information.
[0058] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0059] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.
[0060] The "general computing power" described in the present invention refers to: the computing ability provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0061] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform is deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, etc.
[0062] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0063] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms mainly for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0064] The "intelligent computing center" described in the present invention includes, but is not limited to, the "intelligent computing center".
[0065] The "intelligent computing center" described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.
[0066] The "intelligence center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, transportation power, and storage power, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0067] The "supercomputing center" described in the present invention refers to: that is, the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0068] The "computing power resources" described in the present invention refers to: technologies and facilities required for the development of the digital society with information computing, transmission, storage, and application capabilities, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0069] The "models" described in the present invention include, but are not limited to, "large language models" and "multi-modal large models".
[0070] The "Large Language Model" described in the present invention refers to a large language model (LLM), which is a language model with a relatively large number of parameters. It aims to understand and generate human language, is trained with a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0071] The "Multimodal Large Models" described in the present invention refers to a model that jointly trains multimodal information such as text, images, videos, and audio, including but not limited to multimodal large language models.
[0072] Specifically, please refer to Figure 1 , Figure 1 which is a flowchart of a method for visualizing the operation data of computing power resources in an intelligent computing center provided by an embodiment of the present invention. As Figure 1 shown, it includes the following steps:
[0073] Step S1: Based on the monitoring components pre-deployed in the intelligent computing center, obtain the values corresponding to at least one operation metric of the intelligent computing center.
[0074] It can be understood that the monitoring components can be software / hardware tools (such as Prometheus, Zabbix, etc.) pre-deployed in the intelligent computing center, which can monitor the operation status of resources such as servers, networks, and storage in real time through the Application Programming Interface (API) or sensors. For example, obtain the values of metrics such as CPU usage rate, memory occupancy rate, and disk I / O rate. These data provide a basic basis for subsequent analysis.
[0075] Through the above steps, the embodiments of the present invention can achieve automated collection of full-link metrics in the intelligent computing center, ensure data timeliness, and avoid omissions or delays in manual statistics.
[0076] Step S2: Construct a visualization interface according to the values corresponding to the at least one operation metric, the predefined thresholds corresponding to the at least one operation metric, and the comparison results corresponding to the at least one operation metric; wherein, the comparison results are determined according to the values corresponding to the operation metrics and the predefined thresholds corresponding to the operation metrics.
[0077] In the above steps, the values corresponding to at least one operating metric obtained previously can be compared with the preset thresholds corresponding to at least one operating metric. For example, when the CPU usage rate exceeds 80%, an alarm is triggered. Thus, comparison results such as "normal / abnormal" are generated. Subsequently, the status of the metrics can be displayed on the visualization interface in the form of a dashboard, line chart, heat map, etc. For example, the exceeded metrics are marked in red, and the difference between the specific value and the threshold is marked. In this way, this step converts complex data into intuitive graphics to help the operation and maintenance personnel quickly locate the abnormal points. For example, the memory usage rate of a certain server suddenly soars to 95%, shortening the problem discovery time.
[0078] Step S3: When it is detected that the target comparison result in the visualization interface shows abnormality, based on the real-time value and historical value of the target operating metric collected by the monitoring component, determine the root cause analysis result of the abnormality shown by the target comparison result. The comparison results corresponding to the at least one operating metric include the target comparison result.
[0079] Specifically, when the visualization interface detects that a target metric continuously exceeds the standard, such as a CPU anomaly alarm, the intelligent computing center traces the root cause of the anomaly based on real-time data (current CPU load) and historical trends (load curve in the past 24 hours) through algorithms (such as statistical analysis, machine learning models). For example, if it is found that the memory occupancy of a certain process suddenly increases during the fault period, then this process is determined to be the root cause.
[0080] Step S4: According to the root cause analysis result, display the resource dimension with the highest correlation with the abnormality shown by the target comparison result in the visualization interface.
[0081] In the embodiment of the present invention, according to the root cause analysis result obtained in the previous steps, the resource dimension most relevant to the abnormality is highlighted in the visualization interface. For example, if the CPU anomaly is caused by the occupation of a specific virtual machine, then "Server A-VM01" is marked on the interface and its detailed metrics (such as process list, network traffic) are associated. Thus, this embodiment can convert abstract alarms into specific operation guides, directly target the highly correlated resources (such as restarting the faulty virtual machine) to quickly handle problems, improving the efficiency of correlation analysis of metrics. Optimize the resource management ability.
[0082] In one embodiment, step S1 includes at least one of the following:
[0083] Step S11: Based on the monitoring components pre-deployed on at least one first target node, obtain the values corresponding to at least one operating metric in the at least one first target node;
[0084] Step S12: Monitor at least one second target node based on the Prometheus service discovery mechanism, and obtain the values corresponding to at least one running metric in the at least one second target node;
[0085] Among them, the first target node is a node in the intelligent computing center with a fixed Internet Protocol (IP) address or Domain Name System (DNS) endpoint, and the second target node is a node in the intelligent computing center with a dynamically changing IP address or DNS endpoint.
[0086] It should be noted that the above first target node can be a physical server, a dedicated virtual machine, etc. with a fixed IP address or a DNS endpoint. The second target node can be dynamic resources such as Kubernetes Pods and cloud platform auto-scaling group instances. The monitoring component can be an Exporter tool such as Node Exporter and GPU DCGM Exporter. The service discovery mechanism can be understood as methods such as Kubernetes API, Consul registry, and AWS EC2 tags.
[0087] In some embodiments, monitoring components are pre-deployed on the first target node. These monitoring components, that is, Exporters, continuously collect local resource metrics (such as CPU usage, memory occupancy, disk I / O, GPU temperature, etc.). Specifically, Prometheus can directly specify the endpoint address of the first target node through a static configuration file, such as http: / / 192.0.2.1:9100 / metrics, etc., and actively request to obtain the values corresponding to the running metrics according to a preset pulling frequency (such as every 15 seconds). Exemplarily, if a server with a fixed IP of server-01.example.com deploys Node Exporter, the target endpoint can be manually added to the Prometheus configuration file to regularly collect its CPU load.
[0088] In the above embodiments, the static configuration is applicable to scenarios where the node addresses are stable (such as core database servers and physical machine clusters), providing the simplest monitoring solution. It does not rely on a complex dynamic discovery mechanism, the configuration is direct and reliable, and it is especially suitable for small or non-scalable environments to ensure stable and complete collection of key node metrics.
[0089] In some other embodiments, for the second target node, Prometheus detects and manages the target endpoints in real time through the service discovery mechanism. Exemplarily, in a Kubernetes cluster, Prometheus is configured to obtain information about all the smallest scheduling units (Pods) with specific labels from the kube-apiserver and automatically generate corresponding collection tasks. When a new Pod is started, the service discovery module automatically adds its IP address to the monitoring list. If a certain node fails or is destroyed, the target is automatically removed. Taking the cloud environment as an example, if instances are marked with AWS EC2 labels (such as Prometheus = "true"), Prometheus can dynamically obtain the public / private IPs of these instances through the AWS API and pull their Exporter data.
[0090] In this way, the service discovery mechanism adopted in the above embodiments solves the problems of dynamic node addresses not being fixed and the number changing frequently. For example, in a cloud native environment, an auto-scaling group may create or destroy multiple computing nodes at any time, and traditional static configurations cannot adapt to such changes in a timely manner. Through service discovery, Prometheus can seamlessly track the metrics of all target nodes, ensure that the monitoring coverage is always synchronized with the actual resource status, reduce the manual maintenance cost, and improve the scalability of large-scale distributed systems.
[0091] In one embodiment, the visualization interface includes at least one of the following: a sub-dashboard view, a cross-hierarchy metric comparison chart, a cross-hierarchy correlation analysis channel, and a pop-up comparison chart;
[0092] The step S2 includes at least one of the following:
[0093] Step S21: Based on the label or directory hierarchy selected by the user in the intelligent computing center, construct at least one sub-dashboard view associated with the label or the directory hierarchy; wherein each sub-dashboard view corresponds to a resource type or service module in the intelligent computing center, and the sub-dashboard view is used to display the values of multiple running metrics associated with the resource type or the service module;
[0094] Step S22: In response to a first operation of the user on a target service module in the intelligent computing center, construct a cross-hierarchy metric comparison chart associated with the target service module; wherein the cross-hierarchy metric comparison chart includes a composite view of a time series line chart and a heat map, and the cross-hierarchy metric comparison chart is used to synchronously present the real-time load and historical trend of multiple running metrics associated with the target service module;
[0095] Step S23: In response to the user's second operation on the target node in the intelligent computing center, construct a cross-hierarchy association analysis channel; wherein, the cross-hierarchy association analysis channel is used to display the values corresponding to the operation metrics of at least one associated node having a data interaction relationship with the target computing node.
[0096] Step S24: Construct a pop-up comparison chart in the visualization interface; wherein, the pop-up comparison chart is used to present the correlation data of the operation metrics of the target computing node and the at least one associated node in the time dimension and the space dimension, and the correlation data includes at least one of the utilization rate of the graphics processing unit (GPU), the occupancy rate of the central processing unit (CPU) cache, and the cross-node communication delay.
[0097] In a specific embodiment, the label of the intelligent computing center selected by the user can be the label of a GPU cluster, a database module, etc. The selection of the above directory hierarchy can be like the selection of East China Data Center → Server Group A. The intelligent computing center can automatically screen the associated resource types or service modules and dynamically generate corresponding sub dashboards. Exemplarily, if the user selects the label "AI Training Platform", the sub dashboard view will display the operation metrics of all GPU nodes under this platform, such as GPU utilization rate, memory bandwidth, number of processes, etc., as well as the corresponding threshold comparison results (i.e., normal or abnormal). In addition, in the visualization interface, each sub dashboard can customize the layout by dragging components, such as adding a line chart to display the trend of the GPU video memory occupancy rate.
[0098] In this way, this embodiment can achieve on-demand aggregation of information, that is, the user can focus on specific resources or service modules without traversing all the data (for example, only view the disk I / O metrics of the "Storage Cluster"). It can also achieve hierarchical management of the visualization interface, support multi-level directory navigation (such as Data Center → Computer Room → Server Group), and adapt to the complex topology of large computing centers. It can also achieve the reuse of standardized templates, that is, the predefined dashboard templates can be quickly copied to similar resources, such as applying the same monitoring metric configuration to all "Kubernetes Cluster" nodes.
[0099] In yet another embodiment, when the user performs a first operation, which can be, for example, clicking on a target service module, such as clicking on the video streaming processing service, the intelligent computing center can automatically generate a composite view. Specifically, it includes: a time-series line chart that can display the real-time load of the core metrics of the module, such as the current CPU usage rate of 95%, and can overlay the historical load trend, such as the average value in the past 7 days; a heat map that horizontally compares multiple metrics by node or time period. For example, the disk latency of all backend servers is visualized with different shades of color to quickly locate the performance bottleneck nodes. Exemplarily, in the video stream service module, the line chart can show that the current peak load of the encoding GPU is 85%, while the heat map below presents the fluctuation trend of the GPU memory occupancy rate during the early morning hours.
[0100] Thus, the above-mentioned step S22 realizes the spatio-temporal correlation analysis, enabling the user to quickly compare the real-time state with the historical baseline, and also realizes the cross-verification of multiple metrics, that is, observing multiple dimensions through the same chart simultaneously. In addition, it can also perform anomaly pattern recognition, that is, the heat map can intuitively expose periodic fluctuations or persistent problems of local nodes.
[0101] In yet another embodiment, when the user performs a second operation, which can be, for example, selecting a target node (such as "server node-07") and triggering an operation, the intelligent computing center automatically identifies the associated nodes that communicate with it directly or indirectly through a pre-stored resource interaction relationship graph (based on logs or configuration information). Exemplarily, data flow analysis can be performed. If node-07 is the main database node, the associated nodes may include slave databases, cache servers, and front-end application servers. Metric synchronization display can also be performed, that is, in the channel, the system will display the core metrics of these nodes side by side. For example, the CPU usage rate of node-07, the latency of its slave database, and the request success rate of the front end, which can help the user determine whether the anomaly is caused by upstream and downstream problems.
[0102] Thus, the above-mentioned step 23 can achieve end-to-end tracing, break resource isolation islands, and reveal cross-level dependency relationships. For example, whether the database performance degradation is caused by a sudden increase in front-end traffic. By comparing the metrics of associated nodes, local node failures can be distinguished from external interferences. For example, it can be distinguished that the network latency is not a problem of node-07 itself, but the bandwidth of its switch neighbor node-08 is full.
[0103] In some specific embodiments, when the user needs to deeply analyze the target node and its associated nodes, a pop-up window is triggered on the visualization interface. For example, when double-clicking on a certain metric, the intelligent computing center can generate a dynamic comparison chart. Specifically, the dynamic comparison chart can display the correlation in the time dimension. That is, it can use a line chart to superimpose and display the change curves of the GPU utilization rate of the target node and the memory occupancy rate of the associated cache server over time, and calculate the Pearson correlation coefficient between the two. The dynamic comparison chart can also achieve spatial dimension analysis. That is, it can display the cross-node communication delay through a heat matrix. For example, the delay from node-07 to node-12 is 3 times higher than the average value, or use a scatter plot to correlate the CPU cache occupancy rate with the process response time. In addition, the dynamic comparison chart can also perform causality verification. That is, it can intuitively present the strong and weak correlations between metrics. For example, it can prove that the sudden drop in GPU utilization rate is indeed caused by excessive cache occupancy due to a memory leak in a certain process.
[0104] In the embodiments of the present invention, multi-dimensional insights into the intelligent computing center are realized. From the single dashboard to the layer-by-layer drill-down design of the pop-up window, a complete diagnostic process of global overview, module focus, correlation analysis, and detail verification is achieved. By using composite charts and automated correlation analysis, manual information splicing is reduced. For example, there is no need to manually compare the metric files of two independent nodes. In addition, the embodiments of the present invention can also adapt to various complex architectures and are applicable to multi-level resource scenarios such as cloud-native and hybrid clouds. For example, it can monitor physical machines, virtual machines, and containerized services simultaneously.
[0105] Exemplarily, when a GPU cluster performance anomaly occurs in the intelligent computing center, the user can quickly locate the problem through the following steps: Select the "GPU cluster" sub-dashboard through S21 and find that the video memory utilization rate of a certain node exceeds the standard; Use S23 to trigger correlation analysis and identify a sharp increase in the communication delay between this node and the storage server; Compare the delay and IO metrics of the two in the S24 pop-up window, and finally confirm that the overload of the storage service causes the GPU waiting time to extend. Thus, through the embodiments of the present invention, the operation and maintenance efficiency can be significantly improved, the manual maintenance cost can be reduced, and the scalability of large-scale distributed systems can be enhanced.
[0106] In one embodiment, step S3 includes:
[0107] Step S31: In the case where it is monitored that the target comparison result shows an anomaly, obtain the historical values within a preset time period before the target comparison result shows an anomaly;
[0108] Step S32: Extract at least one candidate metric that has a predefined dependency relationship with the target running metric from other running metrics collected by the monitoring component, construct a multi-dimensional feature matrix for the time series of the at least one candidate metric, and calculate the correlation degree value of the residual term between each candidate metric and the target running metric;
[0109] Step S33: Determine at least one resource dimension for which the target comparison result shows an anomaly and the contribution degree ranking of the at least one resource dimension according to the correlation degree value and the logical distance weights between nodes in the preset resource dependency topology graph;
[0110] Step S34: Filter out the resource dimensions whose confidence scores exceed a preset threshold according to the contribution degree ranking of the at least one resource dimension to obtain the root cause analysis result.
[0111] In some specific embodiments, such as Figure 2 or Figure 3 shown, through multi-dimensional data association and dependency topology analysis, the root resource dimension and its contribution degree that cause the anomaly of the comparison result can be located. When it is detected that the comparison result of a certain target operation index (such as GPU memory usage rate) shows an anomaly, a preset time period can be set, such as 1 hour before the anomaly occurs, to capture the "before and after comparison" data of the anomaly occurrence. Extract the numerical sequence of this index within the specified time period from the historical database of the monitoring component to form a benchmark data set. Assume that the comparison result shows that the memory usage rate of GPU-A suddenly exceeds the threshold at 10:30. The intelligent computing center automatically obtains the time series of the memory usage rate of GPU-A in the past 1 hour.
[0112] Furthermore, according to the predefined dependency relationships, such as physical connections or task scheduling logics, other indicators that are potentially associated with the target operation index are filtered out. The memory usage rate of GPU-A may be affected by the following factors, including the load of the CPU cluster, network bandwidth, storage I / O rate, and a list of candidate indicators, such as [CPU load, network bandwidth, storage IOPS]. Thus, a multi-dimensional feature matrix can be constructed by merging the time series data of the candidate indicators and the target indicator into a matrix, where each column represents the numerical change of an indicator. The multi-dimensional feature matrix can be shown in the following table:
[0113] Timestamp GPU Memory CPU Load Network Bandwidth Storage IOPS 10:29 85% 60% 300Mbps 2000 10:30 96% 70% 150Mbps 4000
[0114] In the above steps, the "residual" (actual value - normal value) at the anomaly point of the target indicator can be calculated. For example, if the threshold is 90%, the residual is 96% - 90% = 6%. Quantitative analysis of the correlation degree between each candidate indicator and the target indicator can be performed using methods such as Pearson correlation coefficient, dynamic time warping (DTW) distance, and Granger causality test.
[0115] Specifically, according to the resource dependency topology graph, a logical distance weight can be assigned to the association between each candidate metric and the target metric. For example, if a candidate metric (such as network bandwidth) is directly connected to the target resource (GPU-A), the logical distance is 1. Resources with indirect dependencies (such as indirectly affecting video memory through a storage server) may be assigned a value of 2. Thus, the original association degree between the candidate metric and the target can be multiplied by the logical distance weight to obtain the comprehensive contribution score. Exemplarily, assuming the association degree of network bandwidth is -0.92 (strong negative correlation) and the logical distance is 1, the contribution degree can be obtained as 0.92×1 = 0.92. The association degree of storage IOPS is 0.67 and the logical distance is 2, so the contribution degree can be obtained as 0.67 / 2≈0.34.
[0116] It is worth mentioning that the resource dependency topology graph can be a model that graphically describes the logical or physical association relationships between different resources (such as hardware devices, software services, task processes, etc.). It regards each component in the intelligent computing center as a node and represents the dependency relationships between them as edges or arrows.
[0117] In this way, the contribution score can be converted into a confidence score, perhaps through normalization processing or weighted summation. For example, confidence = (contribution degree × weight factor) + other prior knowledge correction terms. Among them, a preset threshold is set (such as confidence ≥ 0.8), and only candidate metrics exceeding the threshold are retained. For example, the confidence of network bandwidth is 92% (passing the threshold), while the CPU load is 85% (not reaching the threshold and may be excluded due to insufficient logical distance weight). Thus, the final root cause analysis result is that the abnormal network bandwidth leads to a sharp increase in the GPU video memory usage rate.
[0118] When it is measured that the video memory usage rate of GPU-A reaches 96% (threshold 90%) at 10:30, data for the past 1 hour is obtained, and candidate metrics are extracted, including CPU load, network bandwidth, storage IOPS, etc. It can be calculated that the association degree shows that the network bandwidth drops suddenly by 50% and is highly negatively correlated with the abnormal GPU video memory (caused by resource competition); combined with the logical distance weight of the topology graph, it is determined that the contribution degree of network bandwidth is the highest (0.92). The confidence of 92% exceeds the threshold, and the final root cause is located as "insufficient network bandwidth".
[0119] In this way, the embodiment of the present invention avoids misjudging non-directly related factors through the combination of association degree and topology weight. For example, the CPU load may have an indirect relationship with the GPU video memory, but if it does not reach the threshold, it will not be selected as the root cause. In addition, it can also automatically filter out candidate metrics with low confidence, focus on the key problem path, dynamically adapt to complex scenarios, and flexibly support multi-level dependency relationships.
[0120] In one embodiment, the step S4 includes:
[0121] Step S41: When the user selects any target resource dimension marked as a root cause in the root cause analysis result, display the dependency link of the target resource dimension in the visualization interface:
[0122] Step S42: Determine the shortest path corresponding to the dependency link according to the physical connection or task scheduling logic in the node topology model of the intelligent computing center;
[0123] Step S43: Display the shortest path on the visualization interface, and parallelly display the index trend curve of the target resource dimension on the shortest path in the form of a timeline in the visualization interface.
[0124] In some specific embodiments, when the user selects a resource dimension "marked as a root cause" (such as insufficient video memory of a certain GPU node) in the root cause analysis result, the intelligent computing center can trigger the parsing of the following information: dependency relationship acquisition and link visualization construction. Among them, dependency relationship acquisition is to extract the dependency relationship between the target resource and other resources based on the resource topology model pre-constructed by the intelligent computing center. The dependency relationship can include physical connection (for example, the GPU node is connected to the storage server through a network card) and task scheduling logic (for example, a certain AI training task occupies multiple GPUs and CPU clusters at the same time).
[0125] Link visualization construction can be to convert these dependency relationships into displayable graphic elements. Exemplarily, if the target resource is GPU-A, its dependency link may include GPU-A, storage server-B (data reading and writing), network switch-C (communication relay), task scheduler-D (load distribution).
[0126] Furthermore, these links can be displayed on the visualization interface in the form of lines, arrows or topology diagrams, and the critical path may be marked with colors. Suppose the root cause selected by the user is insufficient video memory of GPU-A, the intelligent computing center will automatically display all resources directly connected to this node (such as storage server-B), and expand the indirectly dependent network devices, task schedulers, etc. to form a complete link diagram.
[0127] In a specific embodiment, the intelligent computing center internally maintains a node topology database, which contains the dependency relationships of all physical connections and task scheduling logics. A weight can be assigned to the relationship between each node, such as physical distance, network latency or task priority, etc. If the physical distance is concerned, the Dijkstra algorithm can be used to calculate the minimum number of hops. If dynamic factors (such as real-time network latency) need to be considered, the path search A* algorithm can be used in combination with the current state data; if the dependency relationship is an undirected graph and the weights are the same, it can be simplified to the Breadth-First Search (BFS).
[0128] Exemplarily, if the target resource is the video memory problem of GPU-A, the shortest path between it and storage server-B can be found. Among them, the topology model shows that there are two direct connection paths, namely Path 1: GPU-A, network switch-C (delay 0.5ms), storage B, with a total weight of 0.5ms. Path 2: GPU-A, access layer-D (delay 2ms), core network-E, storage B, with a total weight of 4ms. According to the weights, the intelligent computing center can select Path 1 as the shortest path.
[0129] Furthermore, the calculated shortest path can be highlighted in the main topology diagram, for example, connecting the nodes on the path with a red dashed line. Or, a timeline panel can be newly added on the side to integrate the metric data of all resources on the path within a specific time period. Metric trend curves can also be used for parallel display, and the intelligent computing center extracts the historical metric data of each node in the path from a time series database (such as Prometheus). For example, for GPU-A: video memory usage rate, temperature; for network switch-C: bandwidth occupancy; for storage B: IOPS (number of I / O operations per second).
[0130] In a specific embodiment, an AI training task suddenly freezes, and the root cause analysis points out that the video memory of GPU-A is insufficient. Then the process of displaying the dependency link can be that the intelligent computing center shows that there is a bottleneck in the path from storage B to GPU-A (for example, the bandwidth of switch-C is saturated). Thus, by quickly locating that the "network transit node" on the critical path may become a performance bottleneck, rather than only focusing on the GPU itself.
[0131] In one embodiment, the method further includes:
[0132] Step S5: Perform image processing on the visualization interface for the node corresponding to the target operation metric or the associated communication link in the intelligent computing center;
[0133] Among them, the image processing includes performing visual enhancement processing on the display effect of the visualization interface.
[0134] In some embodiments, the visualization interface can be visually enhanced through image processing to help users quickly identify key metric anomalies or node / link status changes. The above image processing can be to perform visual enhancement processing on the display effect of the visualization interface, and the visual enhancement effects can specifically include color gradients and icon blinking frequencies.
[0135] Among them, the color gradient levels can be mapped to the predefined threshold intervals corresponding to the target operation metrics. The color gradient associates the values of the target operation metrics with the preset threshold intervals, and visually shows the degree of deviation of the metrics from the normal range through color transition. For example, for the CPU usage rate from 0% to 100%, it can be designed as green (low load), yellow (medium load), and red (overload), with smooth transitions in between. In this way, in the visualization interface, without looking at the specific values, the color can convey the metric health, and distinguish minor anomalies from urgent problems through gradients. For example, a yellow warning allows observation, while a red one requires immediate handling.
[0136] The icon blinking frequency can be positively correlated with the severity of the abnormality in the target comparison result. For example, for a low-risk anomaly, the icon can blink slowly at 1 time per second. In the case of a high-risk anomaly, the icon can blink quickly at 5 times per second to attract attention. Specifically, it can be classified according to the severity level, and map the comparison results (such as "normal", "warning", "error") to the frequency. Thus, in the visualization interface, the user's attention can be quickly captured through high-frequency blinking to ensure that key anomalies are not overlooked, and the urgency can also be distinguished to avoid all alarms using the same frequency to interfere with operations (for example, gentle blinking for warnings and intense blinking for dangerous states).
[0137] Exemplarily, assume that the following problems occur in the GPU cluster of an intelligent computing center. The video memory occupancy rate of a certain GPU node reaches 92%. The icon in its topology diagram changes from green to orange-red and is shown as a red area in the line chart; the usage rate of the associated communication link (such as the bandwidth between this GPU and the storage server) is 85%, and it is marked with a yellow gradient background. This GPU node triggers a "critical" level alarm due to video memory overload, and its icon blinks quickly at a frequency of 2 times per second; the storage server shows a "warning" due to I / O latency problems and only blinks slowly at 1 time per second. Thus, the user can immediately locate the core faulty node through the color and high-frequency blinking, and combine the yellow gradient of the associated link to judge whether there are upstream and downstream impacts, and quickly decide whether to restart the GPU or expand the storage resources.
[0138] Please refer to Figure 4 , Figure 4 which is the structural diagram of a visualization device for the computing power resource operation data of an intelligent computing center provided by an embodiment of the present invention. As Figure 4 shown, the visualization device 200 for the computing power resource operation data of the intelligent computing center includes:
[0139] An index acquisition module 201, configured to obtain the values corresponding to at least one operation index of the intelligent computing center based on the monitoring components pre-deployed in the intelligent computing center;
[0140] An interface construction module 202, configured to construct a visualization interface according to the values corresponding to the at least one running metric, the predefined thresholds corresponding to the at least one running metric, and the comparison results corresponding to the at least one running metric; wherein, the comparison results are determined according to the values corresponding to the running metrics and the predefined thresholds corresponding to the running metrics;
[0141] A result determination module 203, configured to, when it is detected that the display of the target comparison result in the visualization interface is abnormal, determine a root cause analysis result of the abnormal display of the target comparison result according to the real-time value and the historical value of the target running metric collected by the monitoring component, and the comparison results corresponding to the at least one running metric include the target comparison result;
[0142] An interface display module 204, configured to display, in the visualization interface, the resource dimension with the highest association degree with the abnormal display of the target comparison result according to the root cause analysis result.
[0143] In one embodiment, the metric acquisition module 201 is specifically configured to perform at least one of the following:
[0144] Based on the monitoring components pre-deployed on at least one first target node, obtain the values corresponding to at least one running metric in the at least one first target node;
[0145] Based on the Prometheus service discovery mechanism, monitor at least one second target node, and obtain the values corresponding to at least one running metric in the at least one second target node;
[0146] Wherein, the first target node is a node in the intelligent computing center with a fixed Internet Protocol (IP) address or a Domain Name System (DNS) endpoint, and the second target node is a node in the intelligent computing center with a dynamically changing IP address or DNS endpoint.
[0147] In one embodiment, the visualization interface includes at least one of at least one sub-dashboard view, a cross-hierarchy metric comparison chart, a cross-hierarchy association analysis channel, and a pop-up comparison chart;
[0148] The interface construction module 202 is specifically configured to perform at least one of the following:
[0149] Based on the tags or directory hierarchies in the intelligent computing center selected by the user, construct at least one sub-dashboard view associated with the tags or the directory hierarchies; wherein, each sub-dashboard view corresponds to a resource type or a service module in the intelligent computing center, and the sub-dashboard view is used to display the values corresponding to multiple running metrics associated with the resource type or the service module;
[0150] In response to the user's first operation on the target service module in the intelligent computing center, construct a cross - level index comparison chart associated with the target service module; wherein, the cross - level index comparison chart includes a composite view of a time - series line chart and a heat map, and the cross - level index comparison chart is used to synchronously present the real - time load and historical trends of multiple operation indexes associated with the target service module.
[0151] In response to the user's second operation on the target node in the intelligent computing center, construct a cross - level association analysis channel; wherein, the cross - level association analysis channel is used to display the numerical values corresponding to the operation indexes of at least one associated node having a data interaction relationship with the target computing node.
[0152] Construct a pop - up comparison chart in the visualization interface; wherein, the pop - up comparison chart is used to present the correlation data of the operation indexes of the target computing node and the at least one associated node in the time dimension and the space dimension, and the correlation data includes at least one of the utilization rate of the graphics processing unit (GPU), the occupancy rate of the central processing unit (CPU) cache, and the cross - node communication delay.
[0153] In one embodiment, the result determination module 203 is specifically configured to:
[0154] When it is detected that the target comparison result shows an anomaly, obtain the historical values within a preset time period before the target comparison result shows an anomaly.
[0155] Extract at least one candidate index having a predefined dependency relationship with the target operation index from other operation indexes collected by the monitoring component, construct a multi - dimensional feature matrix for the time series of the at least one candidate index, and calculate the correlation degree values of the residual terms between each candidate index and the target operation index.
[0156] According to the correlation degree values and the logical distance weights between nodes in the preset resource dependency topology graph, determine at least one resource dimension for which the target comparison result shows an anomaly, and the contribution degree ranking of the at least one resource dimension.
[0157] According to the contribution degree ranking of the at least one resource dimension, filter out the resource dimensions whose confidence scores exceed the preset threshold to obtain the root cause analysis result.
[0158] In one embodiment, the interface display module 204 is specifically configured to:
[0159] When the user selects any target resource dimension marked as the root cause in the root cause analysis result, display the dependency relationship link of the target resource dimension in the visualization interface:
[0160] Determine the shortest path corresponding to the dependency link according to the physical connection or task scheduling logic in the node topology model of the intelligent computing center;
[0161] Display the shortest path on the visualization interface, and parallelly display the index trend curve of the target resource dimension on the shortest path in the form of a timeline in the visualization interface.
[0162] In one embodiment, the visualization device 200 for the computing power resource operation data of the intelligent computing center further includes:
[0163] An image processing module, configured to perform image processing on the visualization interface for the nodes corresponding to the target operation metrics or the associated communication links in the intelligent computing center;
[0164] Wherein, the image processing includes visually enhancing the display effect of the visualization interface.
[0165] The visualization device 200 for the computing power resource operation data of the intelligent computing center provided by the embodiments of the present invention can implement each process of the above Figure 1 shown embodiments of the visualization method for the computing power resource operation data of the intelligent computing center, with the technical features corresponding one by one and achieving the same technical effects. To avoid repetition, it will not be elaborated here.
[0166] It should be noted that the visualization device for the computing power resource operation data of the intelligent computing center in the embodiments of the present invention can be a device, or a component, integrated circuit, or chip in an electronic device.
[0167] The embodiments of the present invention further provide an electronic device. Refer to Figure 5 , Figure 5 is a schematic structural diagram of an electronic device provided by the embodiments of the present invention. The electronic device includes a memory 301, a processor 302, and a program or instruction stored on the memory 301 and running. When the program or instruction is executed by the processor 302, it can implement Figure 1 any step in the corresponding embodiment of the visualization method for the computing power resource operation data of the intelligent computing center and achieve the same beneficial effects, which will not be elaborated here.
[0168] Wherein, the processor 302 can be a CPU, ASIC, FPGA, or GPU.
[0169] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments of the visualization method for the computing power resource operation data of the intelligent computing center can be completed by hardware related to program instructions, and the program can be stored in a readable medium.
[0170] An embodiment of the present invention further provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement any of the steps in the above-mentioned Figure 1 embodiment of the method for visualizing the operation data of computing power resources of the corresponding intelligent computing center, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. The storage medium may be, for example, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc.
[0171] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement each process of the above-mentioned Figure 1 embodiment of the method for visualizing the operation data of computing power resources of the corresponding intelligent computing center, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0172] The terms "first", "second", etc. in the embodiments of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. In addition, in this application, "and / or" is used to represent at least one of the connected objects. For example, A and / or B and / or C represents 7 situations including A alone, B alone, C alone, A and B existing, B and C existing, A and C existing, and A, B, and C existing.
[0173] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not clearly listed, or also includes elements inherent to this process, method, article or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0174] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of various embodiments of the present application.
[0175] The embodiments of the present application are described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A visualization method for the operation data of computing power resources in an intelligent computing center, characterized in that, Including: Step S1: Based on the monitoring components pre-deployed in the intelligent computing center, obtain the values corresponding to at least one operation metric of the intelligent computing center; Step S2: Construct a visualization interface according to the values corresponding to the at least one operation metric, the predefined thresholds corresponding to the at least one operation metric, and the comparison results corresponding to the at least one operation metric; wherein, the comparison results are determined according to the values corresponding to the operation metrics and the predefined thresholds corresponding to the operation metrics; Step S3: When it is monitored that the target comparison result in the visualization interface shows an anomaly, determine the root cause analysis result of the anomaly shown by the target comparison result according to the real-time values and historical values corresponding to the target operation metric collected by the monitoring component, and the comparison results corresponding to the at least one operation metric include the target comparison result; Step S4: According to the root cause analysis result, display the resource dimension with the highest correlation with the anomaly shown by the target comparison result in the visualization interface.
2. The method according to claim 1, wherein The step S1 includes at least one of the following: Step S11: Based on the monitoring components pre-deployed in at least one first target node, obtain the values corresponding to at least one operation metric in the at least one first target node; Step S12: Monitor at least one second target node based on the Prometheus service discovery mechanism, and obtain the values corresponding to at least one operation metric in the at least one second target node; Wherein, the first target node is a node in the intelligent computing center with a fixed Internet Protocol (IP) address or Domain Name System (DNS) endpoint, and the second target node is a node in the intelligent computing center with a dynamically changing IP address or DNS endpoint.
3. The method according to claim 1, characterized in that, The visualization interface includes at least one of at least one sub-dashboard view, cross-hierarchy metric comparison chart, cross-hierarchy correlation analysis channel, and pop-up comparison chart; The step S2 includes at least one of the following: Step S21: Based on the tags or directory hierarchies selected by the user in the intelligent computing center, construct at least one sub-dashboard view associated with the tags or the directory hierarchies; wherein, each sub-dashboard view corresponds to a resource type or service module in the intelligent computing center, and the sub-dashboard view is used to display the values corresponding to multiple operation metrics associated with the resource type or the service module; Step S22: In response to the first operation of the user on the target service module in the intelligent computing center, construct a cross-hierarchy metric comparison chart; wherein, the cross-hierarchy metric comparison chart includes a composite view of a time series line chart and a heat map, and the cross-hierarchy metric comparison chart is used to synchronously present the real-time load and historical trends of multiple operation metrics associated with the target service module; Step S23: In response to the second operation of the user on the target node in the intelligent computing center, construct a cross-hierarchy correlation analysis channel; wherein, the cross-hierarchy correlation analysis channel is used to display the values corresponding to the operation metrics of at least one associated node having a data interaction relationship with the target computing node. Step S24: Construct a pop-up comparison chart in the visualization interface; wherein, the pop-up comparison chart is used to present the correlation data of the running metrics of the target computing node and the at least one associated node in the time dimension and the space dimension, and the correlation data includes at least one of the utilization rate of the graphics processing unit (GPU), the occupancy rate of the central processing unit (CPU) cache, and the cross-node communication latency.
4. The method according to any one of claims 1 to 3, characterized in that The step S3 includes: Step S31: When it is monitored that the target comparison result shows an anomaly, obtain the historical values within a preset time period before the target comparison result shows an anomaly; Step S32: Extract at least one candidate metric that has a predefined dependency relationship with the target running metric from other running metrics collected by the monitoring component, construct a multi-dimensional feature matrix for the time series of the at least one candidate metric, and calculate the correlation degree value of the residual term between each candidate metric and the target running metric; Step S33: According to the correlation degree value and the logical distance weights between nodes in the preset resource dependency topology graph, determine at least one resource dimension for which the target comparison result shows an anomaly, and the contribution degree ranking of the at least one resource dimension; Step S34: According to the contribution degree ranking of the at least one resource dimension, filter out the resource dimensions whose confidence scores exceed the preset threshold to obtain the root cause analysis result.
5. The method according to any one of claims 1 to 3, characterized in that, The step S4 includes: Step S41: When the user selects any target resource dimension marked as the root cause in the root cause analysis result, display the dependency relationship link of the target resource dimension in the visualization interface: Step S42: Determine the shortest path corresponding to the dependency relationship link according to the physical connection or task scheduling logic in the node topology model of the intelligent computing center; Step S43: Display the shortest path on the visualization interface, and parallelly display the metric trend curve of the target resource dimension on the shortest path in the form of a time axis in the visualization interface.
6. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Step S5: Perform image processing on the visualization interface for the node corresponding to the target running metric or the associated communication link in the intelligent computing center; Wherein, the image processing includes performing visual enhancement processing on the display effect of the visualization interface.
7. A visualization device for the operation data of computing power resources in an intelligent computing center, characterized in that, It includes: A metric acquisition module, configured to obtain the values corresponding to at least one running metric of the intelligent computing center based on a monitoring component pre-deployed in the intelligent computing center; An interface construction module, configured to construct a visualization interface according to the values corresponding to the at least one running metric, the predefined thresholds corresponding to the at least one running metric, and the comparison results corresponding to the at least one running metric; wherein, the comparison result is determined according to the value corresponding to the running metric and the predefined threshold corresponding to the running metric. A result determination module, configured to, when it is detected that the display of the target comparison result in the visualization interface is abnormal, determine a root cause analysis result for the abnormal display of the target comparison result according to the real-time value and the historical value of the target operation index collected by the monitoring component, where the comparison result corresponding to the at least one operation index includes the target comparison result; An interface display module, configured to display, in the visualization interface, a resource dimension with the highest correlation with the abnormal display of the target comparison result according to the root cause analysis result.
8. An electronic device, characterized in that, including: A processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, the steps of the method for visualizing the operation data of the computing power resources of the intelligent computing center according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method for visualizing the operation data of the computing power resources of the intelligent computing center according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that, including computer instructions, and when the computer instructions are executed by a processor, the steps of the method for visualizing the operation data of the computing power resources of the intelligent computing center according to any one of claims 1 to 6 are implemented.