Computing power resource node health monitoring method and device of intelligent computing center
By monitoring and calculating the computing power resource nodes of the intelligent computing center and generating visual charts, the problem that users cannot directly understand the health status of the computing nodes is solved, and efficient operation and maintenance of computing power resources is achieved.
Patent Information
- Application Number
- CN202510369375.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-27
AI Technical Summary
In the process of providing computing power services to users, the monitoring data related to computing nodes in computing power resources needs to be manually retrieved by users, resulting in users being unable to directly understand the health status of computing nodes, which is inefficient and cumbersome, which leads to very low operation and maintenance efficiency of computing power resources.
By monitoring multiple computing nodes in computing power resources, node status data, fault data and resource utilization rate are obtained, node health is calculated based on these data, and visual charts are generated to display node health information, so as to realize automated judgment and visual presentation of node health of computing power resources nodes.
The automated judgment and visual presentation of the health of computing power resource nodes is realized, allowing users to quickly grasp the health status of multiple computing nodes, and thus adjust computing power resources in a timely manner, significantly improving the operation and maintenance efficiency of computing power resources.
Smart Images

Figure CN120216293A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers, and computing power infrastructure, and particularly relates to a method and device for health monitoring of computing power resource nodes in an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power and intelligent computing power. The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and using an artificial intelligence computing architecture.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", and is the ability of computer devices or computing / data centers to process information. It is the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement. It is the computing ability to achieve the output of the target result by processing information data. It is a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0007] Currently, in the process of providing computing power services to users by intelligent computing centers, the monitoring data related to computing nodes in the computing power resources needs to be manually retrieved by users, so that users cannot directly know the health of the computing nodes in the current computing power resources, resulting in low efficiency and cumbersome work, and further leading to the problem of very low operation and maintenance efficiency of computing power resources. Summary of the Invention
[0008] Embodiments of the present invention provide a method and device for health monitoring of computing power resource nodes in an intelligent computing center, which are used to solve the problem of very low operation and maintenance efficiency of computing power resources.
[0009] To solve the above problems, the present invention is implemented as follows:
[0010] In a first aspect, embodiments of the present invention provide a method for health monitoring of computing power resource nodes in an intelligent computing center, including:
[0011] Step S1: Monitor multiple computing nodes in the computing resources to obtain multiple node data sets of the multiple computing nodes within the target time period. The node data set includes the node monitoring data of the corresponding computing node, and the node monitoring data includes at least one of the following: node status data, node failure data, and node resource utilization rate;
[0012] Step S2: Calculate the node health degree based on the multiple node data sets to obtain multiple node health information corresponding one by one to the multiple computing nodes. Among them, the node health information is used to characterize the node health status of the corresponding computing node within the target time period;
[0013] Step S3: Generate a visualization chart based on the node health information of the multiple computing nodes. The visualization chart is used to characterize the node health information of the multiple computing nodes within the target time period.
[0014] In one embodiment, step S1 includes:
[0015] Step S11: Determine multiple moments within the target time period based on a preset time interval. Among the multiple moments, the time interval between any two adjacent moments is the preset time interval;
[0016] Step S12: Monitor each computing node among the multiple computing nodes according to the multiple moments to obtain the node data set of each computing node within the target time period;
[0017] Among them, the node data set includes the node monitoring data of each computing node at the multiple moments, and the node health information includes the node health degree of each computing node at the multiple moments.
[0018] In one embodiment, when the target node health degree is greater than or equal to the first threshold, the visualization chart includes first marking information, and the first marking information is used to indicate that the health level of the target node health degree is the first health level;
[0019] When the target node health degree is less than the first threshold and the target node health degree is greater than or equal to the second threshold, the visualization chart includes second marking information, and the second marking information is used to indicate that the health level of the target node health degree is the second health level;
[0020] When the target node health degree is less than the second threshold, the visualization chart includes third marking information, and the third marking information is used to indicate that the health level of the target node health degree is the third health level;
[0021] Among them, the first health level is higher than the second health level, the second health level is higher than the third health level, the first threshold is greater than the second threshold, and the health degree of the target node is any one of the health degrees of each computing node at the multiple moments.
[0022] In one embodiment, the node monitoring data includes a plurality of node health index data, and step S2 includes:
[0023] Step S21: Determine a plurality of weight values corresponding one-to-one to the plurality of node health index data;
[0024] Step S22: Perform weighted calculation on the plurality of node health index data in the target node monitoring data and the plurality of weight values to obtain the health degree of the target computing node corresponding to the target moment;
[0025] Among them, the target node monitoring data is any one of the node monitoring data of the target computing node at the multiple moments, the target computing node is any one of the plurality of computing nodes, and the target moment is the moment corresponding to the target node monitoring data among the multiple moments.
[0026] In one embodiment, the first marking information, the second marking information, and the third marking information are color identification information in the visualization chart, and among the first marking information, the second marking information, and the third marking information, any two corresponding color identification information are different.
[0027] In one embodiment, the first marking information, the second marking information, and the third marking information are preset icons in the visualization chart, and among the first marking information, the second marking information, and the third marking information, any two corresponding preset icons are different.
[0028] In a second aspect, an embodiment of the present invention further provides a computing power resource node health monitoring device for an intelligent computing center, including:
[0029] A monitoring module, configured to monitor a plurality of computing nodes in the computing power resources to obtain a plurality of node data sets of the plurality of computing nodes respectively within a target time period, where the node data set includes node monitoring data of the corresponding computing node, and the node monitoring data includes at least one of the following: node status data, node failure data, and node resource utilization rate;
[0030] A calculation module, configured to calculate the node health based on the multiple node data sets, and obtain multiple node health information corresponding to the multiple computing nodes one by one, where the node health information is used to characterize the node health status of the corresponding computing node during the target time period;
[0031] A generation module, configured to generate a visualization chart based on the node health information of the multiple computing nodes, where the visualization chart is used to characterize the node health information of the multiple computing nodes during the target time period.
[0032] In a third aspect, the present invention further provides an electronic device, including a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps in the method for monitoring the health of computing power resource nodes of the intelligent computing center as described in the first aspect above.
[0033] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the method for monitoring the health of computing power resource nodes of the intelligent computing center as described in the first aspect above.
[0034] In a fifth aspect, the present invention further provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the steps in the method for monitoring the health of computing power resource nodes of the intelligent computing center as described in the first aspect above.
[0035] In the embodiments of the present invention, multiple computing nodes in the computing power resources are monitored to obtain multiple node data sets of the multiple computing nodes respectively during the target time period. The node data set includes node monitoring data of the corresponding computing node, and the node monitoring data includes at least one of the following: node status data, node failure data, and node resource utilization rate; calculate the node health based on the multiple node data sets to obtain multiple node health information corresponding to the multiple computing nodes one by one, where the node health information is used to characterize the node health status of the corresponding computing node during the target time period; generate a visualization chart based on the node health information of the multiple computing nodes, where the visualization chart is used to characterize the node health information of the multiple computing nodes during the target time period. In this way, through the multiple node data sets of the multiple computing nodes respectively during the target time period, multiple node health information corresponding to the multiple computing nodes one by one is determined and presented to the user in the form of a visualization chart, realizing the automatic judgment and visualization presentation of the health of computing power resource nodes, enabling the user to quickly and simultaneously master the node health status of multiple computing nodes, and then being able to timely adjust the computing power resources based on the visualization chart, thereby greatly improving the operation and maintenance efficiency of the computing power resources. Brief Description of the Drawings
[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0037] Figure 1 is a flowchart of a method for monitoring the health of computing power resource nodes in an intelligent computing center provided by an embodiment of the present invention;
[0038] Figure 2 is a schematic diagram of the node health of a computing node provided by an embodiment of the present invention;
[0039] Figure 3 is a schematic diagram of a node data set provided by an embodiment of the present invention;
[0040] Figure 4 is a structural diagram of a device for monitoring the health of computing power resource nodes in an intelligent computing center provided by an embodiment of the present invention;
[0041] Figure 5 is a structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0043] The "computing power" referred to in the present invention means: the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of a target result by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly providing services to society through computing power infrastructure.
[0044] The "Computational Power (CP)" described in the present invention refers to: the ability of a data center server to process data and output results, which is a comprehensive indicator for measuring the computing power of a data center and includes general computing power, supercomputing power, and intelligent computing power. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing power. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 + CP 智能 + CP 超级 。
[0045] The "Network Power (NP)" described in the present invention refers to: the performance of the data transmission capacity of computing power facilities, which is a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., and involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling ability.
[0046] The "Storage Power (SP)" described in the present invention refers to: the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon, which is a comprehensive indicator for measuring the data storage capacity of a data center and includes external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0047] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize centralized computing, storage, transmission, and application of information.
[0048] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0049] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and supercomputing power.
[0050] The "general computing power" described in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0051] The "intelligent computing power" described in the present invention refers to a computing platform that is scaled for various artificial intelligence innovation applications and is based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, and so on.
[0052] The "super computing power" described in the present invention mainly refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0053] The "intelligent computing center" described in the present invention refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0054] The "intelligent computing center" described in the present invention includes, but is not limited to, the "intelligent computing center".
[0055] The "intelligent computing center" described in the present invention, that is, the artificial intelligence computing center, is a type of computing power infrastructure that is based on artificial intelligence theory, adopts an artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.
[0056] The "computing power center" described in the present invention refers to a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, and has computing power, transportation power, and storage power, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0057] The "supercomputing center" described in the present invention, that is, the supercomputing data center, is a data center based on supercomputers or large-scale computing clusters, and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0058] The "computing power resources" described in the present invention refer to: technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0059] The "health of computing power resource nodes" described in the present invention refers to: an indicator that measures the comprehensive level of the operating status, availability, and performance of computing power resource nodes, and is the core decision-making basis for operating and scheduling computing power resources.
[0060] The "computing power operation task" described in the present invention refers to: specific workloads or jobs that are executed on computing power resources and require a certain amount of computing power support, usually involving scenarios such as complex data processing, numerical calculations, model training, or simulation.
[0061] In the prior art, during the process of an intelligent computing center providing computing power services to users, the monitoring data related to computing nodes in the computing power resources needs to be manually retrieved by the users, making it impossible for the users to directly know the health of the computing nodes in the current computing power resources. The efficiency is low and the work is cumbersome, thereby resulting in a very low operation and maintenance efficiency of the computing power resources. In the embodiments of the present invention, multiple node health information corresponding to multiple computing nodes is determined through multiple node data sets of multiple computing nodes within a target time period, and presented to the users in the form of a visual chart, realizing the automated judgment and visual presentation of the health of computing power resource nodes, enabling the users to quickly and simultaneously master the node health status of multiple computing nodes, and then being able to timely adjust the computing power resources based on the visual chart, thereby being able to greatly improve the operation and maintenance efficiency of the computing power resources.
[0062] Specifically, please refer to Figure 1 , Figure 1 which is a flowchart of a method for monitoring the health of computing power resource nodes of an intelligent computing center provided by an embodiment of the present invention. As Figure 1 shown, it includes the following steps:
[0063] Step S1: Monitor multiple computing nodes in the computing power resources to obtain multiple node data sets of the multiple computing nodes within a target time period. The node data set includes the node monitoring data of the corresponding computing node, and the node monitoring data includes at least one of the following: node status data, node failure data, and node resource utilization rate.
[0064] It is understandable that the monitoring of multiple computing nodes in the computing power resources can be carried out during the execution of the computing power operation task by the computing power resources. Among them, the computing power operation task can be a specific workload or job that is executed on the computing power resources and requires a certain amount of computing power support, and usually involves scenarios such as complex data processing, numerical calculation, model training, or simulation. For example, the above-mentioned computing power operation task can be a model inference task deployed on the computing power resources, such as simultaneously identifying the categories of a large number of objects using a trained image classification model, or it can also be training an image recognition model on a GPU cluster, etc.
[0065] In this step, the above-mentioned target time period can be a pre-defined monitoring period. For example, the target time period can be set to monitor multiple computing nodes in the computing power resources in real time. In this case, the target time period includes the current time point and the past time period. The target time period can also be set to the past 24 hours, the peak period of the computing power operation task, and the comparison period before and after the computing power task runs, etc.
[0066] The above-mentioned multiple node data sets can include the node monitoring data of each computing node in multiple computing nodes at one moment or multiple moments within the target time period. The above-mentioned node status data can be used to reflect the running status of the computing node at the corresponding moment. The above-mentioned node failure data can be used to indicate the type and scope of abnormal events that occur in the computing node, etc. The above-mentioned node resource utilization rate can be used to indicate the proportion of the hardware or software resources of the computing node that are actually used.
[0067] Step S2: Calculate the node health degree based on the multiple node data sets to obtain multiple node health information corresponding one-to-one to the multiple computing nodes, where the node health information is used to characterize the node health status of the corresponding computing node in the target time period.
[0068] In this step, the above-mentioned calculation of the node health degree based on the multiple node data sets can be understood as converting the multi-dimensional monitoring indicators in the node data sets into quantitative or structured node health information through a mathematical model or algorithm. Exemplarily, the above-mentioned node health information can be a digital description or a text description of the comprehensive health status of the computing node in the target time period. Among them, the digital description can be a health degree score, such as a decimal between 0 and 1, and the text description can be "healthy", "sub-healthy", or "faulty".
[0069] It should be noted that the above-mentioned node health information can include one node health degree or multiple node health degrees of the corresponding computing node in the target time period.
[0070] Step S3: Generate a visualization chart based on the node health information of the multiple computing nodes, where the visualization chart is used to characterize the node health information of the multiple computing nodes during the target time period.
[0071] In this step, the above visualization chart can be used to display the node health information of the multiple computing nodes during the target time period. Specifically, a heat map, a time series line chart, a bar chart, etc. can be used. Among them, as Figure 2 shown, the heat map can reflect the node health differences of different computing nodes at different time points through the brightness or color scale change of colors, and the time series line chart or bar chart can be used to display the node health change trend of the same computing node at different time points.
[0072] In the embodiment of the present invention, multiple computing nodes in the computing power resources are monitored to obtain multiple node data sets of the multiple computing nodes respectively during the target time period. The node data set includes the node monitoring data of the corresponding computing node, and the node monitoring data includes at least one of the following: node status data, node failure data, and node resource utilization rate; based on the multiple node data sets, node health degree calculation is performed to obtain multiple node health information corresponding to the multiple computing nodes one by one. Among them, the node health information is used to characterize the node health status of the corresponding computing node during the target time period; based on the node health information of the multiple computing nodes, a visualization chart is generated, and the visualization chart is used to characterize the node health information of the multiple computing nodes during the target time period. In this way, through multiple node data sets of multiple computing nodes respectively during the target time period, multiple node health information corresponding to the multiple computing nodes one by one is determined and presented to the user in the form of a visualization chart, realizing the automatic judgment and visualization presentation of the node health of the computing power resources, enabling the user to quickly and simultaneously master the node health status of multiple computing nodes, and then the computing power resources can be adjusted in time based on the visualization chart, thereby greatly improving the operation and maintenance efficiency of the computing power resources.
[0073] In one embodiment, step S1 includes:
[0074] Step S11: Determine multiple moments during the target time period based on a preset time interval. Among the multiple moments, the time interval between any two adjacent moments is the preset time interval;
[0075] Step S12: Monitor each computing node among the multiple computing nodes according to the multiple moments to obtain the node data set of each computing node during the target time period;
[0076] Among them, the node data set includes the node monitoring data of each computing node at the multiple moments, and the node health information includes the node health levels of each computing node at the multiple moments respectively.
[0077] It can be understood that the determination of the above preset time interval can be combined with specific scenario requirements and system limitations. The length of the preset time interval will directly affect the accuracy of data collection and the consumption of related resources.
[0078] Specifically, the process of monitoring each computing node among the multiple computing nodes according to the multiple moments respectively can be: within a target time period, a monitoring point is generated every preset time interval. For example, if the preset interval is 5 minutes and the time period is 60 minutes, then a total of 12 moments are generated. At each generated moment point, a monitoring instruction is triggered to call the system monitoring tool, so as to obtain the current node monitoring data of each computing node.
[0079] Exemplarily, Figure 3 is a schematic diagram of a node data set provided by an embodiment of the present invention. As Figure 3 shown, the node data set includes the node monitoring data of the first computing node, the second computing node, and the third computing node at moment t1, moment t2, and moment t3 respectively. Moment t4 represents the next moment after t3, that is, a future moment.
[0080] In the above embodiment, through the node data sets of each computing node within the target time period, the node health levels of each computing node at multiple moments are determined and presented to the user in the form of a visual chart, so that the user can directly obtain and master the change trend of the node health levels of multiple computing nodes within the target time period, and thus can quickly locate abnormal computing nodes and trace specific moments, thereby further improving the operation and maintenance efficiency of computing resources.
[0081] In one embodiment, when the target node health level is greater than or equal to the first threshold, the visual chart includes first marking information, and the first marking information is used to indicate that the health level of the target node health level is the first health level;
[0082] When the target node health level is less than the first threshold and the target node health level is greater than or equal to the second threshold, the visual chart includes second marking information, and the second marking information is used to indicate that the health level of the target node health level is the second health level;
[0083] When the target node health level is less than the second threshold, the visual chart includes third marking information, and the third marking information is used to indicate that the health level of the target node health level is the third health level;
[0084] Among them, the first health level is higher than the second health level, the second health level is higher than the third health level, the first threshold is greater than the second threshold, and the target node health degree is any node health degree among the node health degrees of each computing node at the multiple moments.
[0085] In the above embodiment, the node health degrees of the computing nodes are divided into three levels by the first threshold and the second threshold, so that the user can quickly master the state distribution of multiple computing nodes without manually parsing the data, thereby significantly shortening the fault handling time and improving the utilization rate of computing power resources.
[0086] In one embodiment, the node monitoring data includes multiple node health index data, and the step S2 includes:
[0087] Step S21: Determine multiple weight values corresponding one by one to the multiple node health index data;
[0088] Step S22: Perform weighted calculation on the multiple node health index data in the target node monitoring data and the multiple weight values to obtain the node health degree corresponding to the target computing node at the target moment;
[0089] Among them, the target node monitoring data is any one of the node monitoring data of the target computing node at the multiple moments, the target computing node is any one of the multiple computing nodes, and the target moment is the moment corresponding to the target node monitoring data among the multiple moments.
[0090] Specifically, the above weight values can be used to reflect the contribution degrees of the respective node health index data to the node health degree, and the specific values range from 0 to 1. It can be understood that the above weight values can adjust the weights according to the importance of the index categories. For example, by weight assignment, priority is given to node status data and node failure data, directly reflecting the availability and urgency of the computing node, and the resource utilization rate is used as a supplementary index.
[0091] It should be noted that multiple node health index data can be quantified first and then weighted calculation can be performed. For example, the node status data can include "normal", "warning" or "fault", and can be correspondingly converted into numerical values, such as 1, 0.5 and 0; the node failure data can be represented by binary values, with failure = 1 and no failure = 0; the resource utilization rate can include CPU usage rate and memory occupancy rate, etc.
[0092] In the above embodiments, by assigning weight values to the health indicator data of multiple nodes, over-reliance on indicator data of a single dimension is avoided, which can effectively help users evaluate the node health of computing nodes from multiple dimensions, thereby improving the accuracy of determining the node health.
[0093] In one embodiment, the first marking information, the second marking information, and the third marking information are color identification information in the visualization chart, and among the first marking information, the second marking information, and the third marking information, any two corresponding color identification information are different.
[0094] Exemplarily, the first marking information may be a green identifier, used to indicate that the corresponding computing node is in a normal state; the second marking information may be a yellow identifier, used to indicate that the corresponding computing node has potential risks and can be intervened in advance to avoid deterioration; the third marking information may be a red identifier, used to indicate a node failure that needs to be processed immediately.
[0095] In the above embodiments, by using different colors to identify different health levels, complex health degree values are converted into intuitive visual signals, enabling users to quickly grasp the state distribution of computing nodes, significantly improving the user's perception efficiency of the node health of computing nodes, and thereby further improving the operation and maintenance efficiency of computing power resources.
[0096] In one embodiment, the first marking information, the second marking information, and the third marking information are preset icons in the visualization chart, and among the first marking information, the second marking information, and the third marking information, any two corresponding preset icons are different.
[0097] Exemplarily, the first marking information may be a tick, the second marking information may be an exclamation mark, and the third marking information may be a cross.
[0098] In the above embodiments, by using different icons to identify different health levels, complex health degree values are converted into intuitive visual signals, enabling users to quickly grasp the state distribution of computing nodes, significantly improving the user's perception efficiency of the node health of computing nodes, and thereby further improving the operation and maintenance efficiency of computing power resources.
[0099] Please refer to Figure 4 , Figure 4 which is a structural diagram of a computing power resource node health monitoring device of an intelligent computing center provided by an embodiment of the present invention. As Figure 4 shown, the computing power resource node health monitoring device 400 of the intelligent computing center includes:
[0100] The monitoring module 401 is configured to monitor multiple computing nodes in the computing resources, and obtain multiple node data sets of the multiple computing nodes respectively within a target time period. The node data set includes node monitoring data of the corresponding computing node, and the node monitoring data includes at least one of the following: node status data, node failure data, and node resource utilization rate;
[0101] The calculation module 402 is configured to calculate the node health degree based on the multiple node data sets, and obtain multiple node health information corresponding to the multiple computing nodes one by one. Among them, the node health information is used to characterize the node health status of the corresponding computing node within the target time period;
[0102] The generation module 403 is configured to generate a visualization chart based on the node health information of the multiple computing nodes. The visualization chart is used to characterize the node health information of the multiple computing nodes within the target time period.
[0103] In one embodiment, the monitoring module 401 includes:
[0104] The first determination unit is configured to determine multiple moments within the target time period based on a preset time interval. Among the multiple moments, the time interval between any two adjacent moments is the preset time interval;
[0105] The monitoring unit is configured to monitor each computing node in the multiple computing nodes according to the multiple moments, and obtain the node data set of each computing node within the target time period;
[0106] Among them, the node data set includes the node monitoring data of each computing node at the multiple moments, and the node health information includes the node health degree of each computing node at the multiple moments.
[0107] In one embodiment, when the target node health degree is greater than or equal to the first threshold, the visualization chart includes first marking information, and the first marking information is used to indicate that the health level of the target node health degree is the first health level;
[0108] When the target node health degree is less than the first threshold and greater than or equal to the second threshold, the visualization chart includes second marking information, and the second marking information is used to indicate that the health level of the target node health degree is the second health level;
[0109] When the target node health degree is less than the second threshold, the visualization chart includes third marking information, and the third marking information is used to indicate that the health level of the target node health degree is the third health level;
[0110] Among them, the first health level is higher than the second health level, the second health level is higher than the third health level, the first threshold is greater than the second threshold, and the health degree of the target node is any one of the health degrees of each computing node at the multiple moments.
[0111] In one embodiment, the node monitoring data includes a plurality of node health index data, and the computing module 402 includes:
[0112] A second determination unit, configured to determine a plurality of weight values corresponding one-to-one to the plurality of node health index data;
[0113] A calculation unit, configured to perform weighted calculation on the plurality of node health index data in the target node monitoring data and the plurality of weight values to obtain the health degree of the target computing node corresponding to the target moment;
[0114] Among them, the target node monitoring data is any one of the node monitoring data of the target computing node at the multiple moments, the target computing node is any one of the plurality of computing nodes, and the target moment is the moment corresponding to the target node monitoring data among the multiple moments.
[0115] In one embodiment, the first marking information, the second marking information, and the third marking information are color identification information in the visualization chart, and among the first marking information, the second marking information, and the third marking information, any two corresponding color identification information are different.
[0116] In one embodiment, the first marking information, the second marking information, and the third marking information are preset icons in the visualization chart, and among the first marking information, the second marking information, and the third marking information, any two corresponding preset icons are different.
[0117] The computing power resource node health monitoring device of the intelligent computing center provided by the embodiments of the present invention can implement each process of the above-mentioned embodiments of the computing power resource node health monitoring method of the intelligent computing center, and the technical features correspond one-to-one and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0118] It should be noted that the computing power resource node health monitoring device in the embodiments of the present invention can be a device, or a component, an integrated circuit, or a chip in an electronic device.
[0119] The embodiments of the present invention further provide an electronic device. See Figure 5 , Figure 5It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device includes a memory 501, a processor 502, and a program or instruction running on the memory 501. When the program or instruction is executed by the processor 502, it can implement Figure 1 Any steps in the corresponding embodiment of the method for monitoring the health of computing power resource nodes in the intelligent computing center and achieve the same beneficial effects, which will not be elaborated here.
[0120] Among them, the processor 502 can be a CPU, ASIC, FPGA or GPU.
[0121] Those of ordinary skill in the art can understand that all or part of the steps of implementing the corresponding embodiment of the method for monitoring the health of computing power resource nodes in the intelligent computing center can be completed by hardware related to program instructions, and the program can be stored in a readable medium.
[0122] An embodiment of the present invention also provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it can implement the above Figure 1 Any steps in the corresponding embodiment of the method for monitoring the health of computing power resource nodes in the intelligent computing center and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. The storage medium such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc, etc.
[0123] The present invention also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, it implements the above Figure 1 Each process of the corresponding embodiment of the method for monitoring the health of computing power resource nodes in the intelligent computing center and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0124] The terms "first", "second", etc. in the embodiments of the present invention are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. In addition, in this application, "and / or" is used to represent at least one of the connected objects. For example, A and / or B and / or C represents 7 situations including A alone, B alone, C alone, A and B both exist, B and C both exist, A and C both exist, and A, B and C all exist.
[0125] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including such element.
[0126] From the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or a second terminal device, etc.) to execute the methods of various embodiments of this application.
[0127] The embodiments of this application are described above in conjunction with the accompanying drawings. However, this application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of this application and without departing from the purpose of this application and the scope protected by the claims, can also make many forms, all of which fall within the protection scope of this application.
Claims
1. A method for monitoring the health of computing resource nodes in an intelligent computing center, characterized in that: include: Step S1: monitor multiple computing nodes in the computing power resources to obtain multiple node data sets of the multiple computing nodes in a target time period, wherein the node data sets include node monitoring data of the corresponding computing nodes, and the node monitoring data includes at least one of the following: node status data, node fault data, and node resource utilization; Step S2, performing node health calculation based on the multiple node data sets to obtain multiple node health information corresponding to the multiple computing nodes, wherein the node health information is used to characterize the node health status of the corresponding computing node in the target time period; Step S3: Generate a visualization chart based on the node health information of the multiple computing nodes, where the visualization chart is used to represent the node health information of the multiple computing nodes within the target time period.
2. The method according to claim 1, characterized in that The step S1 comprises: Step S11: determining multiple moments within the target time period based on a preset time interval, wherein the time interval between any two adjacent moments in the multiple moments is the preset time interval; Step S12: monitoring each of the multiple computing nodes at the multiple time points to obtain a node data set of each computing node in a target time period; The node data set includes the node monitoring data of each computing node at the multiple time points, and the node health information includes the node health of each computing node at the multiple time points.
3. The method according to claim 2, characterized in that When the health of the target node is greater than or equal to a first threshold, the visualization chart includes first marking information, and the first marking information is used to indicate that the health level of the health of the target node is a first health level; When the health of the target node is less than the first threshold and the health of the target node is greater than or equal to the second threshold, the visualization chart includes second marking information, and the second marking information is used to indicate that the health level of the health of the target node is the second health level; When the health of the target node is less than the second threshold, the visualization chart includes third marking information, and the third marking information is used to indicate that the health level of the health of the target node is a third health level; Among them, the first health level is higher than the second health level, the second health level is higher than the third health level, the first threshold is greater than the second threshold, and the target node health is any node health of each computing node at the multiple moments.
4. The method according to claim 2, characterized in that The node monitoring data includes a plurality of node health indicator data, and the step S2 includes: Step S21: determining a plurality of weight values corresponding one-to-one to the plurality of node health indicator data; Step S22: performing weighted calculation on the multiple node health indicator data in the target node monitoring data and the multiple weight values to obtain the node health corresponding to the target computing node at the target time; Among them, the target node monitoring data is any one of the node monitoring data of the target computing node at the multiple moments, the target computing node is any one of the multiple computing nodes, and the target moment is the moment among the multiple moments corresponding to the target node monitoring data.
5. The method according to claim 3, characterized in that: The first marking information, the second marking information and the third marking information are color identification information in the visual chart, and any two corresponding color identification information in the first marking information, the second marking information and the third marking information are different.
6. The method according to claim 3, characterized in that The first marking information, the second marking information, and the third marking information are preset icons in the visual chart, and the preset icons corresponding to any two of the first marking information, the second marking information, and the third marking information are different.
7. A computing resource node health monitoring device for an intelligent computing center, characterized in that: include: A monitoring module is used to monitor multiple computing nodes in the computing power resources to obtain multiple node data sets of the multiple computing nodes in a target time period, wherein the node data sets include node monitoring data of the corresponding computing nodes, and the node monitoring data includes at least one of the following: node status data, node fault data, and node resource utilization; A calculation module, configured to perform node health calculation based on the plurality of node data sets to obtain a plurality of node health information corresponding to the plurality of computing nodes, wherein the node health information is used to characterize the node health status of the corresponding computing node in the target time period; A generation module is used to generate a visualization chart based on the node health information of the multiple computing nodes, where the visualization chart is used to represent the node health information of the multiple computing nodes within the target time period.
8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for monitoring the health of computing resource nodes of an intelligent computing center as described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for monitoring the health of computing resource nodes in an intelligent computing center as described in any one of claims 1 to 6.
10. A computer program product, characterized in that It comprises computer instructions, which, when executed by a processor, implement the steps of the method for health monitoring of computing resource nodes of an intelligent computing center as described in any one of claims 1 to 6.