Multi-dimensional fault sensing and alarming method and device for computing power node of intelligent computing center cloud platform

By deploying a containerized health status detection program on the computing nodes of the intelligent computing center cloud platform, the problem of low efficiency in fault detection of computing nodes is solved, multi-dimensional fault perception and automatic alarm are realized, and detection efficiency and accuracy are improved.

CN120658555APending Publication Date: 2025-09-16DATACANVAS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511013330.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In the existing technology, the efficiency of fault detection of computing nodes in intelligent computing centers is low, and it mainly relies on manual auxiliary detection, resulting in low detection efficiency.

Method used

By deploying health status detection programs on the computing nodes of the intelligent computing center cloud platform and using containerization to encapsulate preset detection programs and benchmark programs, multi-dimensional health status detection and fault perception of computing nodes can be achieved, including the use of mirrors and agent programs to automatically monitor and locate faults.

Benefits of technology

It improves the efficiency and accuracy of computing node fault detection, reduces manual intervention, and achieves rapid positioning and automated alarming of computing node faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120658555A_ABST
    Figure CN120658555A_ABST
Patent Text Reader

Abstract

The invention provides a computing power node multi-dimensional fault sensing alarm method and device of an intelligent computing center cloud platform, and relates to the technical field of intelligent computing centers, intelligent computing centers and computing power infrastructures. The method comprises the steps that S1, the health state of a computing power node of an intelligent computing center cloud platform is detected based on a first object to obtain health state information, the first object comprises a mirror image and / or an agent program, and the mirror image packages a health state detection program through a container and is deployed on the computing power node; the agent program deploys a health state detection program on a computing power node according to the instruction, wherein the health state detection program comprises a preset detection program and / or a preset benchmark test program; and S2, based on the health state information, determining whether the computing power node has a fault or not. According to the method and the device, the health state detection program is deployed on the computing power node through the containerization or agent program to detect the health state of the computing power node and determine whether the computing power node has the fault, so that the fault detection efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of intelligent computing centers, smart computing centers, computing power infrastructure and intelligent computing cloud technology, and in particular to a multi-dimensional fault perception and alarm method and device for computing power nodes of an intelligent computing center cloud platform. Background Art

[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged.

[0003] An "Intelligent Computing Center" is a facility that uses large-scale heterogeneous computing resources, including general-purpose and intelligent computing power, to provide the computing power, data, and algorithms required for AI applications (such as AI deep learning model development, model training, and model inference). The Intelligent Computing Center encompasses facilities, hardware, and software, and provides a full stack of capabilities, from bottom-level computing power to top-level application enablement.

[0004] “Intelligent Computing Center” includes but is not limited to “Smart Computing Center”.

[0005] "Intelligent Computing Center" refers to an artificial intelligence computing center. It is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0006] "Computing power" is the core of "intelligent computing center" and "intelligent computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform certain computing needs. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to society through computing power infrastructure.

[0007] Since the emergence of intelligent computing centers, hardware failures (such as GPU, CPU, memory, or disk corruption) and kernel issues (such as kernel deadlocks and file system corruption) have had a direct impact on cluster stability in Kubernetes cluster management. Currently, fault detection of computing nodes in intelligent computing centers is typically performed manually, resulting in low detection efficiency. Therefore, improving fault detection efficiency in computing nodes in intelligent computing centers is an urgent issue. Summary of the Invention

[0008] The present invention provides a multi-dimensional fault perception and alarm method and device for computing power nodes of an intelligent computing center cloud platform, which are used to solve the problem of how to improve the fault detection efficiency of computing power nodes.

[0009] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:

[0010] In a first aspect, the present invention provides a multi-dimensional fault perception and alarm method for computing nodes of an intelligent computing center cloud platform, comprising:

[0011] Step S1: Based on a prefabricated first object, the health status of a computing node of the intelligent computing center cloud platform is detected to obtain health status information of the computing node, wherein the first object includes at least one of an image and an agent program, the image encapsulates a health status detection program in a containerized deployment manner and is deployed on the computing node, the agent program deploys the health status detection program on the computing node according to instructions, and the agent program is further used to maintain the health status detection program in normal operation, and the health status detection program includes at least one of a preset detection program and a preset benchmark test program;

[0012] Step S2: Based on the health status information, determine whether the computing power node has a fault.

[0013] In a second aspect, the present invention provides a multi-dimensional fault perception and alarm device for computing nodes of an intelligent computing center cloud platform, comprising:

[0014] A detection module, configured to detect the health status of a computing node of a cloud platform of an intelligent computing center based on a prefabricated first object, and obtain health status information of the computing node, wherein the first object includes at least one of an image and an agent program, the image encapsulating a health status detection program in a containerized deployment manner and deploying it on the computing node, the agent program deploying the health status detection program on the computing node according to instructions, the agent program further configured to maintain the health status detection program in a normal operating state, and the health status detection program including at least one of a preset detection program and a preset benchmark test program;

[0015] The first determination module is used to determine whether the computing power node has a fault based on the health status information.

[0016] In a third aspect, the present invention provides a server comprising: a processor, a memory, and a program stored in the memory and runnable on the processor. When the program is executed by the processor, the steps of the multi-dimensional fault perception and alarm method for computing power nodes of the intelligent computing center cloud platform as described in the first aspect above are implemented.

[0017] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the multi-dimensional fault perception and alarm method for computing power nodes of the intelligent computing center cloud platform as described in the first aspect above are implemented.

[0018] In a fifth aspect, the present invention provides a computer program product comprising computer instructions, which, when executed by a processor, implement the steps of the multi-dimensional fault perception and alarm method for computing power nodes of the intelligent computing center cloud platform as described in the first aspect above.

[0019] In the present invention, a health status detection program is deployed on the computing power node of the intelligent computing center cloud platform through a containerized deployment method or an instruction-based deployment method. The intelligent computing center cloud platform can automatically detect whether there is a fault in the computing power node through the program, and can quickly locate the computing power node with a fault in the intelligent computing center cloud platform, and finally determine whether the computing power node has a fault, thereby improving the efficiency of computing power node fault detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0021] Figure 1 This is a flow chart of a multi-dimensional fault perception and alarm method for computing nodes of an intelligent computing center cloud platform of the present invention;

[0022] Figure 2 This is a schematic diagram of the architecture of the intelligent computing center cloud platform of the present invention;

[0023] Figure 3 This is a structural diagram of a multi-dimensional fault perception and alarm device for computing nodes of an intelligent computing center cloud platform of the present invention;

[0024] Figure 4 It is a structural diagram of the server of the present invention. DETAILED DESCRIPTION

[0025] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0026] First, the technical terms involved in the present invention are briefly explained below.

[0027] The "computing power" mentioned in the present invention refers to: the ability of computer equipment or computing / data centers to process information, the ability of computer hardware and software to work together to execute certain computing requirements, and the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.

[0028] The "computing power" (CP) mentioned in the present invention refers to: the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating-point operations performed per second (FLOPS, 1EFLOPS=10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP=CP 通用 +CP 智能 +CP 超级 .

[0029] The "carrying capacity" (Network Power, NP) mentioned in the present invention refers to: it is the performance of the data transmission capability of the computing power facility, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. It involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities.

[0030] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in terms of data storage capacity, performance, security and reliability, and environmental friendliness. It is a comprehensive indicator for measuring a data center's data storage capacity, encompassing both external storage devices such as storage arrays and internal server storage. Storage capacity is commonly measured in exabytes (EB, 1EB = 2^60 bytes), while performance is commonly measured in IOPS / TB (Input / Output Operations Per Second / TB). Disaster recovery ratio is a key indicator of security and reliability.

[0031] The "computing power infrastructure" mentioned in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, and can realize the centralized calculation, storage, transmission and application of information.

[0032] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G networks, fiber-optic broadband networks, backbone networks, international communication networks, satellite Internet, computing power infrastructure such as data centers, general computing power centers, intelligent computing centers, supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.

[0033] The "computing power" mentioned in the present invention includes: general computing power, intelligent computing power and super computing power.

[0034] The "general computing power" mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0035] The "intelligent computing power" mentioned in this invention refers to: a computing platform based on specialized chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various innovative artificial intelligence applications, such as natural language processing (NLP) and machine vision.

[0036] The "supercomputing power" mentioned in the present invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, genetic analysis, etc.

[0037] The "intelligent computing center" described in this article refers to a facility that provides the computing power, data, and algorithms required for artificial intelligence applications (such as AI deep learning model development, model training, and model inference) by utilizing large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center encompasses facilities, hardware, and software, and can provide a full stack of capabilities, from bottom-level computing power to top-level application enablement.

[0038] The "intelligent computing center cloud platform" mentioned in the present invention is referred to as "intelligent computing cloud", which refers to: a cloud computing platform that provides comprehensive services based on the hardware resources and software resources of the intelligent computing center.

[0039] The "intelligent computing center" mentioned in the present invention includes but is not limited to the "intelligent computing center".

[0040] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services and algorithm services required for artificial intelligence applications.

[0041] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0042] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide large-scale computing, storage and network services and other functions, and is widely used in application scenarios such as aerospace, oil exploration, climate modeling and genome sequencing.

[0043] The "computing resources" mentioned in the present invention refer to: technologies and facilities with information calculation, transmission, storage and application capabilities required for the development of a digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guarantee resources such as wind, fire, water, and electricity.

[0044] The "computing power node" mentioned in the present invention refers to the computing resources of the server / container that can process computing tasks.

[0045] The "computing power operation task" mentioned in the present invention refers to: a specific workload or job executed on computing power resources that requires a certain amount of computing power support, usually involving complex data processing, numerical calculations, model training or simulation scenarios.

[0046] The "health status information" mentioned in the present invention refers to: information used to characterize the health status of the computing power node, including but not limited to hardware status information (such as CPU status, GPU status, memory status, storage device status, etc.), software and system status information (operating system status, service status, etc.), network performance information, storage status information, file system status, etc.

[0047] The "health status detection program" mentioned in the present invention refers to: a program used to detect the health status of a computing power node.

[0048] The "pre-set detection program" mentioned in the present invention refers to: an automated detection tool or script pre-configured and deployed on the computing power node, which is used to automatically monitor the health status of the computing power node.

[0049] The "pre-set benchmark test program" mentioned in the present invention refers to: a program used to benchmark the performance of computing power nodes, such as GPU computing performance, network throughput, storage input / output (Input / Output, I / O) and other performance tests.

[0050] See also Figure 1 , Figure 1 This invention provides a multi-dimensional fault perception and alarm method for computing nodes of an intelligent computing center cloud platform. Figure 1 As shown, the method includes:

[0051] Step S1: Based on a prefabricated first object, the health status of a computing node of the intelligent computing center cloud platform is detected to obtain health status information of the computing node, wherein the first object includes at least one of an image and an agent program, the image encapsulates a health status detection program in a containerized deployment manner and is deployed on the computing node, the agent program deploys the health status detection program on the computing node according to instructions, and the agent program is further used to maintain the health status detection program in normal operation, and the health status detection program includes at least one of a preset detection program and a preset benchmark test program;

[0052] Step S2: Based on the health status information, determine whether the computing power node has a fault.

[0053] The prefabricated first object may be understood as an object of a detection program that is pre-configured and encapsulated through containerization, or a pre-configured detection program.

[0054] In some implementations, a health status detection program (including at least one of a preset detection program and a preset benchmark test program) is packaged as a container image and deployed on a computing node of the intelligent computing center cloud platform. The health status detection program may also be referred to as a fault detection program.

[0055] In some embodiments, according to the instructions of the main control unit (such as through a script or configuration file), an agent program including a health status detection program is installed on the computing power node. The agent program can control the health status detection program to be in normal operating state. When the health status detection program is in normal operating state, the detection program can be used to detect the health status of the computing power node.

[0056] For example, when the agent detects that the health status detection program has exited, it controls the program to re-run. The agent can also report the status of the health status detection program to the main control unit according to a preset period. If the main control unit does not receive the reporting information, it outputs an alarm prompt.

[0057] The health status detection program is used to detect the health status of computing power nodes. When deploying agent programs on multiple computing power nodes, it can improve deployment efficiency.

[0058] like Figure 2 As shown, the node fault detector is deployed on the computing power node. The node fault detector includes fault monitoring process A and fault monitoring process B, as well as a kernel monitor. When a fault is detected in the computing power node, the fault information is reported to the API server.

[0059] Among them, the preset detection program can be used to detect the health status of the hardware or system of the computing power node. For example, the preset detection program includes a script for detecting the GPU status of the computing power node and a script for detecting the network card status.

[0060] By running the preset detection program or agent program in the image, the health status of the computing power node can be detected, thereby perceiving the failure of the computing power node.

[0061] For example, a first detection program is run to obtain the temperature of the GPU (eg, 75° C.) and the memory usage (eg, 85%).

[0062] Run the second detection program to detect the speed (such as 1 Gbps) and packet loss rate (such as 0.05%) of the network card.

[0063] When running a benchmark program, you can perform benchmark tests to determine the health of computing nodes. For example, you can use the benchmark program to run high-load GPU computing tasks on one or more computing nodes to detect GPU performance degradation. If the benchmark test results do not meet the benchmark, it indicates that the computing node is faulty. You can further troubleshoot the fault location of the computing node, such as through binary tree querying, targeted troubleshooting, or other methods to locate the specific fault.

[0064] The preset detection program and the preset benchmark test program can be deployed as images or agents, respectively. For example, the preset detection program can be deployed in the image and the preset benchmark test program can be deployed in the agent, or vice versa. The preset detection program and the preset benchmark test program can also be deployed simultaneously in the image, or in the agent. When performing a health check on a computing power node, the corresponding test program can be selected for health check.

[0065] In some implementations, health detection can be performed using any of the above detection programs, or health status detection can be performed using the above two health status detection programs to determine whether the computing power node has a fault, thereby improving the detection accuracy.

[0066] For example, when the benchmark test program detects that the GPU performance of the computing power node has declined, the GPU is further tested through a preset test program; or the health status of the computing power node is tested through two test programs, and the health status information obtained is combined to comprehensively determine whether the computing power node has a fault.

[0067] In some embodiments, when an abnormality is found during the benchmark test on the node as a whole, if it is a single node, the single node is directly marked as faulty; if it is multiple nodes, a grouping strategy (such as binary grouping, topology grouping, device-by-device grouping) or other methods can be selected to divide the multiple nodes into multiple groups, and further benchmark tests are performed on each node group (which may include one or more nodes) until the faulty node is located.

[0068] The health status of the computing power node is detected by the above health status detection program to obtain the health status information of the computing power node. The health status information is used to determine whether the computing power node has a fault.

[0069] For example, the GPU temperature is detected to obtain a current temperature value. If the current temperature value exceeds a reference value, it indicates a GPU failure.

[0070] If the network card packet loss rate exceeds the baseline value, it indicates a network failure.

[0071] In some implementations, a comprehensive judgment can also be made using multiple indicators. For example, if the GPU memory usage exceeds the reference value for a duration greater than a preset duration and the temperature exceeds the reference value, it indicates a GPU failure.

[0072] In some implementations, the health status detection program may further include a hard disk health detection tool (such as smartmontools), a GPU status monitoring tool (such as nvidia-smi), etc., to better monitor the hardware status.

[0073] This invention deploys health status detection programs in various ways to achieve real-time monitoring of computing nodes, reducing manual intervention and improving monitoring effectiveness. Furthermore, containerized deployment allows for flexible expansion and facilitates the addition of detection programs. By using pre-set detection programs and benchmarking procedures, the accuracy of fault identification can be improved.

[0074] Optionally, step S1 includes:

[0075] Step S11: When the health status detection program includes the preset detection program, the health status information of the computing power node of the intelligent computing center cloud platform is collected based on the preset detection program in the prefabricated first object, and the health status information includes at least one of the hardware operating parameters and system performance parameters.

[0076] When a preset detection program is included in the image or agent program, the hardware health and system performance of the computing power node can be tested by running the preset detection program to obtain hardware operating parameters and system performance parameters.

[0077] For example, to support hardware fault awareness for servers, switches, and other hardware, hardware health information about the server can be accessed through integrated hardware monitoring tools (such as IPMI, iDRAC, and OpenBMC). These tools can provide detailed information about power supplies, fans, memory, hard drives, temperatures, and more.

[0078] In some implementations, a pre-set detection program can be used to periodically query hardware health information for real-time checks. If a power failure or memory error is detected, an alarm message is output.

[0079] In addition, you can also set specific detection conditions to automatically detect the health status of computing power nodes when the conditions are met. For example, set a first cycle, and automatically detect each corresponding period of the first cycle; when the health status score of the computing power node is lower than the preset value, check again after a preset interval or switch from the first cycle to a second cycle with a shorter interval than the first cycle; before starting a specific computing power task; when the computing power node health score is lower than the preset value, and so on.

[0080] Hardware operation parameter collection, for example, obtains GPU temperature, memory utilization, usage rate and other parameters; detects network card speed, packet loss rate, and link status.

[0081] System performance parameter collection, for example, obtains CPU usage, process resource usage; obtains memory usage, cache usage.

[0082] Based on the collected hardware operating parameters and system performance parameters, determine whether the computing power node is faulty.

[0083] By conducting multi-dimensional health status detection on computing power nodes, the accuracy of detection can be improved.

[0084] Optionally, the hardware operating parameters include at least one of the following:

[0085] GPU status information, the status information including at least one of GPU usage rate, temperature, and video memory utilization;

[0086] Working status of the GPU link;

[0087] GPU memory usage status;

[0088] The operating status of the network card;

[0089] File system status;

[0090] CPU status information;

[0091] System memory status information;

[0092] Storage device status information;

[0093] Network device status information;

[0094] Power status information;

[0095] The link status of the PCIE standard for fast peripheral interconnection.

[0096] Multi-dimensional fault detection can be performed on computing power nodes, including GPU status, GPU link, GPU memory, network card status, file system status, CPU status information, system memory status information, storage device status information, network device status information, power supply status information, and PCIE link status.

[0097] The GPU status includes, for example, GPU usage, temperature, and video memory utilization. You can periodically run commands (such as nvidia-smi) to monitor the GPU status. If GPU memory overload or device failure is detected, an alarm is triggered.

[0098] The working status detection of the GPU link can include whether the link connection is normal or whether the link speed is degraded. For example, monitoring the link status between the GPU and the host.

[0099] The usage status of GPU memory can be monitored using memory leak detection tools, for example, to detect whether the GPU memory usage is abnormally high.

[0100] You can check the network card's operating status and speed using tools such as ethtool or ip-s link. If the network card slows down due to a fault or driver issue, it will be automatically recorded and an alarm will be issued. You can also determine if the network card is faulty by checking for TX / RX errors or packet loss on the network interface.

[0101] The file system status can be checked regularly using file system consistency check tools (such as fsck) and hard disk health detection tools (such as smartctl).

[0102] If the file system fails to mount or a serious error occurs, the computing power node will be marked and a fault alarm will be triggered.

[0103] The CPU status is information used to reflect the CPU status, such as CPU usage, CPU temperature, CPU cache efficiency, CPU frequency, error count, etc.

[0104] The system memory status information is information indicating the system memory status, such as system memory usage, virtual memory usage, memory leaks, etc.

[0105] The storage device status information is information indicating the status of the storage device, for example, hard disk health status information (read and write errors, etc.), disk throughput, and the like.

[0106] The network device status information is information indicating the status of the network device, such as network interface error count, network delay, bandwidth utilization, switch port status, etc.

[0107] The power status information is information indicating the power status, such as power voltage, current, power, chassis temperature, etc.

[0108] The link status of the Peripheral Component Interconnect Express (PCIE) standard can be determined by obtaining link information or monitoring status changes to determine whether the link is faulty.

[0109] In addition, other parameters can also be tested.

[0110] Through the above preset detection program, multi-dimensional monitoring of hardware operating parameters such as GPU, network card, file system, CPU, system memory, storage device, network device, power supply, etc. can be achieved.

[0111] Optionally, step S1 includes:

[0112] Step S12: When the health status detection program includes the preset benchmark test program, a benchmark performance test is performed on the computing power node using the benchmark test program in the prefabricated first object to obtain performance information of the computing power node, where the performance information is used to characterize the health status of the computing power node.

[0113] The step S2 comprises:

[0114] Step S21: Based on the comparison result of the performance information of the computing power node and the benchmark performance baseline, determine whether the computing power node has a fault.

[0115] Deploy the image or agent program containing the benchmark program on the computing power node of the intelligent computing center cloud platform, and monitor the health of the computing power node by running the benchmark program.

[0116] Benchmark programs may include hardware performance tests or system stability tests, such as GPU computing stress tests, file system read and write performance tests, etc.

[0117] The above tests can generate performance information about the computing node, which can be used to reflect the health of the computing node. The performance information obtained from the benchmark performance test is compared with the benchmark performance baseline to determine whether the computing node is faulty. The benchmark performance baseline can be obtained in advance and is used to represent the baseline health of the computing node.

[0118] When it is determined based on the performance information obtained from the computing power node test that the computing power node does not reach the benchmark performance, it indicates that there is a fault in the computing power node.

[0119] Through the above method, the accuracy of health status or fault detection of computing power nodes can be improved.

[0120] Optionally, the computing power node is a computing power node group, and the computing power node group includes N computing power nodes, where N is an integer greater than 0; step S12 includes:

[0121] Step S121: When the health status detection program includes the preset benchmark test program, a benchmark performance test is performed on the computing power node group using the benchmark test program in the prefabricated first object to obtain overall performance information of the computing power node group, where the overall performance information is used to represent the overall health status of the computing power node group.

[0122] The step S21 includes:

[0123] Step S211: Comparing the overall performance information with a corresponding benchmark performance baseline, and determining whether the computing node group includes a faulty computing node based on the comparison result, wherein if the computing node group includes multiple computing nodes, the benchmark performance baseline is a first benchmark performance baseline; if the computing node group includes a single computing node, the benchmark performance baseline is a second benchmark performance baseline;

[0124] A computing node group may include one or more computing nodes. A benchmark performance test is performed based on the computing node group to obtain overall performance information of the computing node group.

[0125] If the computing node group only includes one computing node, the computing node is compared with the second benchmark performance baseline to determine whether the computing node is faulty. If the overall performance information reaches the benchmark performance baseline, it indicates that there is no fault; otherwise, it indicates that the computing node is faulty.

[0126] When a computing node group includes multiple computing nodes, the computing node group is compared with a first benchmark performance baseline to determine whether there are any faulty nodes in the computing node group. If the overall performance information reaches the benchmark performance baseline, there are no faults; otherwise, the computing node group includes a faulty computing node.

[0127] If N is greater than 1, that is, the computing power node group includes multiple computing power nodes, and the computing power node group includes a faulty computing power node, the computing power nodes can be further divided into groups. The computing power node group can be divided into two or three or more groups based on the size of N or the division method, and benchmark performance tests can be performed on each group of computing power nodes separately.

[0128] Obtain the overall performance information for each group of computing nodes and compare it to the corresponding benchmark performance baseline. Specifically, when each group has only one computing node, compare it to the second benchmark performance baseline; when each group has more than one computing node, compare it to the first benchmark performance baseline. Continue this process until the faulty node is located.

[0129] Determining whether a computing power node group has a fault by the above method can improve the efficiency of fault diagnosis.

[0130] Optionally, when N is greater than 2, after determining whether the computing power node group includes a faulty computing power node based on the comparison result, the method further includes:

[0131] Step S3: If it is determined that the computing power node group includes a faulty computing power node, the computing power node group is divided into M groups of computing power nodes, and a first benchmark performance test is performed on each of the M groups of computing power nodes to obtain a first performance test result, where M is an integer greater than 1 and M≤N;

[0132] Step S4: When it is determined according to the first performance test result that a first group of computing nodes in the M groups of computing nodes includes a faulty computing node, and the number of computing nodes in the first group of computing nodes is K, each computing node in the first group of computing nodes is combined with a computing node that does not have a fault in the M groups of computing nodes to obtain K groups of computing nodes, where K is an integer greater than 1 and K<N;

[0133] Step S5: performing a second benchmark performance test on each group of computing power nodes in the K groups of computing power nodes to obtain a second performance test result;

[0134] Step S6: When it is determined according to the second performance test result that the second group of computing power nodes in the K groups of computing power nodes includes a faulty computing power node, the faulty computing power node is determined based on the common computing power nodes in the first group of computing power nodes and the second group of computing power nodes.

[0135] If the overall performance results indicate that a computing node group includes a faulty computing node, the computing node group is divided into M groups. A benchmark performance test is performed on each group of computing nodes, and the test results are compared with the benchmark performance baseline. If the benchmark performance baseline is met, no computing nodes in the group are faulty. If the benchmark performance baseline is not met, the group includes a faulty computing node. For computing node groups that include a faulty computing node, each computing node is recombined with a healthy computing node. The combined computing nodes may include two or more computing nodes.

[0136] For example, a computing node group includes 10 computing nodes, and the benchmark test includes network speed and disk throughput (IO) testing. If the overall performance information obtained from the overall test does not meet the standards of the first benchmark performance baseline, it indicates that the computing node group contains a faulty computing node. The 10 computing nodes are then further divided into five groups, each with two computing nodes, and tested again. The above benchmark performance testing process and results are then performed separately for each group of computing nodes. If the performance information of the first group of computing nodes in the five groups does not meet the benchmark performance baseline, it indicates that a faulty computing node exists in that group. Simultaneously, a second group of computing nodes (any group of computing nodes without faults) is obtained whose performance information meets the benchmark performance baseline (indicating no faults). Computing nodes A and B in the first group are recombined with computing nodes C and D in the second group, respectively, to obtain two combined groups of computing nodes. The benchmark performance test is then re-performed on each of these two groups of computing nodes.

[0137] If the test results of the second group of computing power nodes including computing power node A do not reach the benchmark performance baseline during the benchmark performance test, and computing power node A is a computing power node shared by the first group of computing power nodes and the second group of computing power nodes, it indicates that computing power node A is faulty.

[0138] By dividing computing nodes into groups and conducting multiple benchmark performance tests, the effectiveness of computing node fault location can be improved. This approach also enables automated fault detection, reducing manual operations and improving fault detection efficiency.

[0139] Optionally, after step S6, the method further includes:

[0140] Step S7: When it is determined that the faulty computing node is the first computing node, a third benchmark performance test is performed on a third group of computing nodes using the prefabricated benchmark test program in the first object, where the third group of computing nodes is all computing nodes in the computing node group except the first computing node.

[0141] Step S8: When it is determined according to the test result of the third benchmark performance test that the third group of computing power nodes does not include a faulty computing power node, it is determined that the first computing power node is faulty.

[0142] When it is determined in the above manner that the faulty computing node is the first computing node, after removing the first computing node from the computing node group, all remaining computing nodes are combined again to obtain a third group of computing nodes.

[0143] The benchmark performance test is performed again on the third group of computing power nodes. If the benchmark performance test results of the third group of computing power nodes reach the benchmark performance baseline, it indicates that the first computing power node is a faulty computing power node.

[0144] Verification through the above method can improve the accuracy of diagnosis of faulty computing nodes.

[0145] Optionally, the benchmark performance baseline includes at least one of the following:

[0146] A first benchmark value obtained by performing a benchmark test on the computing power node under preset conditions;

[0147] A second benchmark value determined based on the configuration information of the computing power node and system performance parameters;

[0148] a third reference value generated based on health data of the computing power node within a preset historical period, wherein the third reference value is associated with a load of the computing power node;

[0149] In the case where the computing power node is a computing power node group including multiple computing power nodes, the fourth benchmark value is determined based on the node types of the multiple computing power nodes and the weight corresponding to each node type, and the node type is a type determined based on node parameters or topology information.

[0150] Among them, the first benchmark value is the performance parameter information obtained by the computing power node when performing a benchmark test under preset conditions, wherein the preset conditions can be ideal conditions, for example, the parameters obtained when the computing power node is tested under healthy conditions.

[0151] The second benchmark value can be a benchmark value determined in combination with configuration information (such as GPU model) and system performance parameters (such as CPU frequency). In addition, after the initial benchmark value is determined based on the configuration information and system performance parameters, the initial benchmark value can be corrected based on the first benchmark value to obtain the second benchmark value.

[0152] The third benchmark value can be generated based on the health data of the computing power node during a historical period. After a computing power node has been running for a long time, its performance may degrade. In some embodiments, the third benchmark value can be generated based on an ideal benchmark value obtained in a pre-test and the health data during the historical period, which can improve the accuracy of the judgment.

[0153] In some implementations, the third reference value may be determined in combination with the load of the computing power node.

[0154] For example, the system collects health data of computing nodes over a preset historical period, including load information, and associates this health data with the node's load status to differentiate performance under different load scenarios. Based on the statistical characteristics of the historical data (such as mean, median, and standard deviation), it generates corresponding third benchmark values ​​according to the load scenario. This allows the system to adapt to changes in the computing node's load and reduce misjudgment of faults due to short-term fluctuations.

[0155] In addition, a third benchmark value can be set based on the first benchmark value or the second benchmark value by comprehensively considering the historical operation data of the computing power node.

[0156] In some implementations, the above-mentioned benchmark values ​​may be adjusted based on user settings in combination with task requirements to adapt to actual needs.

[0157] In the case where the computing power node is a computing power node group, the node type to which each computing power node belongs is obtained, and the number of computing power nodes contained in each node type is obtained. The corresponding weight is determined based on the node type and the number of nodes, so as to calculate the benchmark performance value of the computing power node group, that is, the fourth benchmark value.

[0158] For example, a computing node group includes a first computing node and a second computing node. The first computing node has CPU parameters A1 and performance parameters A2, while the second computing node has CPU parameters B1 and performance parameters B2. Based on this computing node parameter information, the computing node group is divided into two node types. The baseline performance baseline for the first node type is baseline value A, while the baseline performance baseline for the second node type is baseline value B.

[0159] The weights of the benchmark values ​​of the first group of node types and the second group of node types are both 0.5, and the fourth benchmark value of the computing power node group corresponding to the first computing power node and the second computing power node is calculated based on the weights = A×0.5+B×0.5.

[0160] The above types can also be divided according to topology information or other methods. The number of computing power nodes of each topology type and the benchmark value corresponding to each topology type can be obtained, and the fourth benchmark value of the computing power node group can be calculated according to the above method.

[0161] The above approach can adapt to fault detection under different conditions and improve the accuracy of fault detection in some scenarios with high computing requirements.

[0162] Optionally, after step S2, the method further includes:

[0163] Step S5: When a fault occurs in the computing node, an alarm message is sent in a preset manner, where the alarm message includes the fault information of the computing node.

[0164] The preset method includes at least one of the following:

[0165] Sending the fault information to the monitoring system of the intelligent computing center cloud platform;

[0166] Sending the fault information to the event management system of the intelligent computing center cloud platform;

[0167] Sending the fault information to the administrator of the intelligent computing center cloud platform via email;

[0168] Sending the fault information to a second object associated with the intelligent computing center cloud platform through an instant messaging tool;

[0169] The fault information is sent to a third object associated with the intelligent computing center cloud platform through a cross-application communication technology.

[0170] In Kubernetes cluster management, computing node failures are usually invisible to the upstream layer of cluster management, resulting in Kubernetes being unable to perceive the actual health status of the computing nodes, which in turn causes computing tasks to run abnormally.

[0171] In this embodiment, when a computing power node fails, an alarm message is sent through any one or more of the following methods, and the alarm message carries the failure information of the computing power node.

[0172] In some implementations, when a hardware failure is detected, an alert can be sent by integrating with the event management system of the container orchestration system (Kubernetes), monitoring system (Prometheus, Alertmanager), and other tools. For example, Prometheus can be used to monitor the hardware status of the node and trigger an alert when a threshold is reached.

[0173] Fault information is sent to the cluster's event management system through Kubernetes' Events resource so that administrators and automated tools can respond to the fault.

[0174] In combination with the alarm system, the fault alarm can be sent to the administrator's email, and the fault information can be sent to the second party through the instant messaging tool, for example, the second party such as the team member corresponding to the intelligent computing center cloud platform;

[0175] Fault information can also be sent to a third party through cross-system technology. The third party can be a pre-set manager or other personnel so that timely action can be taken.

[0176] like Figure 3 As shown, Figure 3 The present invention provides a multi-dimensional fault perception and alarm device for computing nodes of an intelligent computing center cloud platform, comprising:

[0177] A detection module 301 is configured to detect the health status of a computing node of the intelligent computing center cloud platform based on a prefabricated first object to obtain health status information of the computing node, wherein the first object includes at least one of an image and an agent program, the image encapsulates a health status detection program in a containerized deployment manner and is deployed on the computing node, the agent program deploys the health status detection program on the computing node according to instructions, and is further configured to maintain the health status detection program in normal operation, and the health status detection program includes at least one of a preset detection program and a preset benchmark test program;

[0178] The first determination module 302 is used to determine whether the computing power node has a fault based on the health status information.

[0179] Optionally, the detection module is specifically configured to:

[0180] In the case where the health status detection program includes the preset detection program, the health status information of the computing power node of the intelligent computing center cloud platform is collected based on the preset detection program in the prefabricated first object, and the health status information includes at least one of the hardware operating parameters and system performance parameters.

[0181] Optionally, the hardware operating parameters include at least one of the following:

[0182] GPU status information, wherein the GPU status information includes at least one of GPU usage rate, temperature, and video memory utilization rate;

[0183] Working status of the GPU link;

[0184] GPU memory usage status;

[0185] The operating status of the network card;

[0186] File system status;

[0187] CPU status information;

[0188] System memory status information;

[0189] Storage device status information;

[0190] Network device status information;

[0191] Power status information;

[0192] The link status of the PCIE standard for fast peripheral interconnection.

[0193] Optionally, the detection module is specifically configured to, when the health status detection program includes the preset benchmark test program, perform a benchmark performance test on the computing power node using the benchmark test program in the prefabricated first object to obtain performance information of the computing power node, where the performance information is used to characterize the health status of the computing power node;

[0194] The first determination module is specifically configured to determine whether the computing power node has a fault based on a comparison result of the performance information of the computing power node with a benchmark performance baseline.

[0195] Optionally, the computing power node is a computing power node group, the computing power node group includes N computing power nodes, where N is an integer greater than 0; the determination and detection module is specifically configured to: when the health status detection program includes the preset benchmark test program, perform a benchmark performance test on the computing power node group using the benchmark test program in the prefabricated first object to obtain overall performance information of the computing power node group, where the overall performance information is used to characterize the overall health status of the computing power node group;

[0196] The first determination module is specifically used to compare the overall performance information with the corresponding benchmark performance baseline, and determine whether the computing power node group includes a faulty computing power node based on the comparison result, wherein, when the computing power node group includes multiple computing power nodes, the benchmark performance baseline is the first benchmark performance baseline, and when the computing power node group includes a single computing power node, the benchmark performance baseline is the second benchmark performance baseline.

[0197] Optionally, the device further comprises:

[0198] a first testing module configured to, when determining that the computing node group includes a faulty computing node, divide the computing node group into M groups of computing nodes, and perform a first benchmark performance test on each of the M groups of computing nodes to obtain a first performance test result, where M is an integer greater than 1 and M≤N;

[0199] a combining module, configured to, when it is determined according to the first performance test result that a first group of computing nodes in the M groups of computing nodes includes a faulty computing node and the number of computing nodes in the first group of computing nodes is K, combine each computing node in the first group of computing nodes with a computing node that does not have a fault in the M groups of computing nodes, to obtain K groups of computing nodes, where K is an integer greater than 1 and K<N;

[0200] A second testing module is configured to perform a second benchmark performance test on each group of computing power nodes in the K groups of computing power nodes to obtain a second performance test result;

[0201] The second determination module is used to determine the faulty computing nodes based on the common computing nodes in the first group of computing nodes and the second group of computing nodes when it is determined that the second group of computing nodes in the K groups of computing nodes includes the faulty computing nodes according to the second performance test results.

[0202] Optionally, the device further comprises:

[0203] a third testing module configured to, when determining that the faulty computing node is the first computing node, perform a third benchmark performance test on a third group of computing nodes using a prefabricated benchmark test program in the first object, the third group of computing nodes being all computing nodes in the computing node group except the first computing node;

[0204] The third determination module is used to determine that the first computing node has a fault when it is determined that the third group of computing nodes does not include a faulty computing node based on the test result of the third benchmark performance test.

[0205] Optionally, the benchmark performance baseline includes at least one of the following:

[0206] A first benchmark value obtained by performing a benchmark test on the computing power node under preset conditions;

[0207] A second benchmark value determined based on the configuration information of the computing power node and system performance parameters;

[0208] a third reference value generated based on health data of the computing power node within a preset historical period, wherein the third reference value is associated with a load of the computing power node;

[0209] In the case where the computing power node is a computing power node group including multiple computing power nodes, the fourth benchmark value is determined based on the node types of the multiple computing power nodes and the weight corresponding to each node type, and the node type is a type determined based on node parameters or topology information.

[0210] Optionally, the device further comprises:

[0211] A sending module, configured to send an alarm message in a preset manner when a fault occurs in the computing power node, wherein the alarm message includes fault information of the computing power node;

[0212] The preset method includes at least one of the following:

[0213] Sending the fault information to the monitoring system of the intelligent computing center cloud platform;

[0214] Sending the fault information to the event management system of the intelligent computing center cloud platform;

[0215] Sending the fault information to the administrator of the intelligent computing center cloud platform via email;

[0216] Sending the fault information to a second object associated with the intelligent computing center cloud platform through an instant messaging tool;

[0217] The fault information is sent to a third object associated with the intelligent computing center cloud platform through a cross-application communication technology.

[0218] The multi-dimensional fault perception and alarm device for computing power nodes of the intelligent computing center cloud platform provided by the present invention can realize the various processes of each embodiment of the multi-dimensional fault perception and alarm method for computing power nodes of the above-mentioned intelligent computing center cloud platform. The technical features correspond one to one and can achieve the same technical effects. To avoid repetition, they will not be repeated here.

[0219] It should be noted that the multi-dimensional fault perception and alarm device for computing power nodes of the intelligent computing center cloud platform in the present invention can be a device, or a component, integrated circuit, or chip in an electronic device.

[0220] Please refer to Figure 4 The present invention also provides a server 110, including a processor 111, a memory 112, and a computer program stored in the memory 112 and executable on the processor 111. When the computer program is executed by the processor 111, the various processes of the embodiment of the multi-dimensional fault perception and alarm method for computing power nodes of the above-mentioned intelligent computing center cloud platform are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0221] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the various processes of the embodiment of the multi-dimensional fault perception and alarm method for computing nodes of the intelligent computing center cloud platform, and can achieve the same technical effects. To avoid repetition, the details are not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0222] The present application also provides a computer program product including computer instructions, which, when executed by a processor, implement the above Figure 1 The various processes of the embodiment of the multi-dimensional fault perception and alarm method for computing power nodes of the intelligent computing center cloud platform shown in the figure can achieve the same technical effect. To avoid repetition, they will not be repeated here.

[0223] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0224] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0225] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.

Claims

1. A multi-dimensional fault perception and alarm method for computing nodes of an intelligent computing center cloud platform, characterized in that: include: Step S1: Based on a prefabricated first object, the health status of a computing node of the intelligent computing center cloud platform is detected to obtain health status information of the computing node, wherein the first object includes at least one of an image and an agent program, the image encapsulates a health status detection program in a containerized deployment manner and is deployed on the computing node, the agent program deploys the health status detection program on the computing node according to instructions, and the agent program is further used to maintain the health status detection program in normal operation, and the health status detection program includes at least one of a preset detection program and a preset benchmark test program; Step S2: Based on the health status information, determine whether the computing power node has a fault.

2. The method according to claim 1, characterized in that The step S1 comprises: Step S11: When the health status detection program includes the preset detection program, the health status information of the computing power node of the intelligent computing center cloud platform is collected based on the preset detection program in the prefabricated first object, and the health status information includes at least one of the hardware operating parameters and system performance parameters.

3. The method according to claim 2, characterized in that The hardware operating parameters include at least one of the following: GPU status information, wherein the GPU status information includes at least one of GPU usage rate, temperature, and video memory utilization rate; Working status of the GPU link; GPU memory usage status; The operating status of the network card; File system status; CPU status information; System memory status information; Storage device status information; Network device status information; Power status information; The link status of the PCIE standard for fast peripheral interconnection.

4. The method according to claim 1, wherein The step S1 comprises: Step S12: When the health status detection program includes the preset benchmark test program, a benchmark performance test is performed on the computing power node using the benchmark test program in the prefabricated first object to obtain performance information of the computing power node, where the performance information is used to characterize the health status of the computing power node. The step S2 comprises: Step S21: Based on the comparison result of the performance information of the computing power node and the benchmark performance baseline, determine whether the computing power node has a fault.

5. The method according to claim 4, characterized in that The computing power node is a computing power node group, and the computing power node group includes N computing power nodes, where N is an integer greater than 0; the step S12 includes: Step S121: When the health status detection program includes the preset benchmark test program, a benchmark performance test is performed on the computing power node group using the benchmark test program in the prefabricated first object to obtain overall performance information of the computing power node group, where the overall performance information is used to represent the overall health status of the computing power node group. The step S21 includes: Step S211: Compare the overall performance information with the corresponding benchmark performance baseline, and determine whether the computing power node group includes a faulty computing power node based on the comparison result, wherein, when the computing power node group contains multiple computing power nodes, the benchmark performance baseline is the first benchmark performance baseline, and when the computing power node group contains a single computing power node, the benchmark performance baseline is the second benchmark performance baseline.

6. The method according to claim 5, characterized in that When N is greater than 2, after determining whether the computing power node group includes a faulty computing power node based on the comparison result, the method further includes: Step S3: If it is determined that the computing power node group includes a faulty computing power node, the computing power node group is divided into M groups of computing power nodes, and a first benchmark performance test is performed on each of the M groups of computing power nodes to obtain a first performance test result, where M is an integer greater than 1 and M≤N; Step S4: When it is determined according to the first performance test result that a first group of computing nodes in the M groups of computing nodes includes a faulty computing node, and the number of computing nodes in the first group of computing nodes is K, each computing node in the first group of computing nodes is combined with a computing node that does not have a fault in the M groups of computing nodes to obtain K groups of computing nodes, where K is an integer greater than 1 and K<N; Step S5: performing a second benchmark performance test on each group of computing power nodes in the K groups of computing power nodes to obtain a second performance test result; Step S6: When it is determined according to the second performance test result that the second group of computing power nodes in the K groups of computing power nodes includes a faulty computing power node, the faulty computing power node is determined based on the common computing power nodes in the first group of computing power nodes and the second group of computing power nodes.

7. The method according to claim 6, characterized in that After step S6, the method further includes: Step S7: When it is determined that the faulty computing node is the first computing node, a third benchmark performance test is performed on a third group of computing nodes using the prefabricated benchmark test program in the first object, where the third group of computing nodes is all computing nodes in the computing node group except the first computing node. Step S8: When it is determined according to the test result of the third benchmark performance test that the third group of computing power nodes does not include a faulty computing power node, it is determined that the first computing power node is faulty.

8. The method according to any one of claims 4 to 7, characterized in that The benchmark performance baseline includes at least one of the following: A first benchmark value obtained by performing a benchmark test on the computing power node under preset conditions; A second benchmark value determined based on the configuration information of the computing power node and system performance parameters; a third reference value generated based on health data of the computing power node within a preset historical period, wherein the third reference value is associated with a load of the computing power node; In the case where the computing power node is a computing power node group including multiple computing power nodes, the fourth benchmark value is determined based on the node types of the multiple computing power nodes and the weight corresponding to each node type, and the node type is a type determined based on node parameters or topology information.

9. The method according to any one of claims 1 to 7, characterized in that After step S2, the method further includes: Step S9: When a fault occurs in the computing node, an alarm message is sent in a preset manner, where the alarm message includes the fault information of the computing node. The preset method includes at least one of the following: Sending the fault information to the monitoring system of the intelligent computing center cloud platform; Sending the fault information to the event management system of the intelligent computing center cloud platform; Sending the fault information to the administrator of the intelligent computing center cloud platform via email; Sending the fault information to a second object associated with the intelligent computing center cloud platform through an instant messaging tool; The fault information is sent to a third object associated with the intelligent computing center cloud platform through a cross-application communication technology.

10. A multi-dimensional fault perception and alarm device for computing nodes of an intelligent computing center cloud platform, characterized in that: include: A detection module, configured to detect the health status of a computing node of a cloud platform of an intelligent computing center based on a prefabricated first object, and obtain health status information of the computing node, wherein the first object includes at least one of an image and an agent program, the image encapsulating a health status detection program in a containerized deployment manner and deploying it on the computing node, the agent program deploying the health status detection program on the computing node according to instructions, the agent program further configured to maintain the health status detection program in a normal operating state, and the health status detection program including at least one of a preset detection program and a preset benchmark test program; The first determination module is used to determine whether the computing power node has a fault based on the health status information.

11. A server, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the multi-dimensional fault perception and alarm method for computing power nodes of an intelligent computing center cloud platform as described in any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the multi-dimensional fault perception and alarm method for computing power nodes of an intelligent computing center cloud platform as described in any one of claims 1 to 9.

13. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps of the multi-dimensional fault perception and alarm method for computing power nodes of an intelligent computing center cloud platform as described in any one of claims 1 to 9.