Computing node processing method and device, equipment and storage medium

By integrating single-node and inter-node network detection and optimization, the method improves the efficiency and reliability of computing node online processing within a cluster.

CN120321148APending Publication Date: 2025-07-15DAWNING INFORMATION IND (BEIJING) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510607901.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, the manual maintenance method before the computing node is launched is inefficient, resulting in a decrease in the operation reliability of the computing cluster.

Method used

Network detection of stand-alone dimensions and online dimensions, including IPMI network detection and IB network detection, combined with the configuration information of the computing node, stand-alone operation detection and online operation parameter optimization, and is processed online after the stand-alone operation test is passed.

Benefits of technology

It improves the efficiency of computing nodes going online, ensures the reliability and operational adaptability of computing nodes in the computing cluster, and improves the overall operational reliability of the computing cluster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321148A_ABST
    Figure CN120321148A_ABST
Patent Text Reader

Abstract

The invention relates to a computing node processing method and device, equipment and a storage medium. The method comprises the following steps: under the condition that a network detection event aiming at a computing node is monitored, performing single-machine-dimension network detection and online-dimension infinite bandwidth IB network detection on the computing node, and under the condition that the single-machine-dimension network detection is passed and the online-dimension IB network detection is not passed, according to configuration information of the computing node, determining whether the computing node passes the network detection according to the configuration information of the computing node; and single-machine operation detection is performed on the computing nodes, online operation parameters of the computing nodes are optimized, and online processing is performed on the computing nodes after the single-machine operation test is passed. By adopting the method, the online efficiency of the computational nodes can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and particularly to a method, apparatus, device, and storage medium for processing computing nodes. Background Art

[0002] With the continuous development of the computer field, the number of computing nodes in a computing cluster has also increased. To ensure the reliability of the operation of the computing cluster, generally, before a computing node goes online, the network where the computing node is located is detected, and in the case of detecting a network anomaly, the online process of the computing node is suspended to wait for the corresponding operation and maintenance personnel to process the computing node, and the subsequent online process is resumed after the processing is completed.

[0003] However, since the pre-online maintenance of each computing node is performed manually, the online efficiency of the computing nodes will be greatly reduced. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, apparatus, device, and storage medium for processing computing nodes that can improve the online efficiency of computing nodes for the above technical problems.

[0005] In a first aspect, the present application provides a method for processing computing nodes, including:

[0006] When a network detection event for a computing node is detected, perform network detection in a single-machine dimension and Infiniband (IB) network detection in an online dimension on the computing node;

[0007] When the network detection in the single-machine dimension passes and the IB network detection in the online dimension fails, perform single-machine operation detection on the computing node according to the configuration information of the computing node, and optimize the online operation parameters of the computing node;

[0008] After the single-machine operation test passes, perform an online process on the computing node.

[0009] In an embodiment of the present application, on the one hand, introducing IB network detection in the online dimension can ensure that after the computing node goes online, it is adapted to the operation of the computing cluster, ensuring the reliability of the operation of the computing node in the computing cluster; on the other hand, when the network detection in the single-machine dimension passes and the IB network detection in the online dimension fails, while optimizing the online operation parameters of the computing node, continue to perform the pre-online single-machine operation test, and after the single-machine operation test passes, the computing node will be put online regardless of whether the online operation parameters are optimized or not, which will greatly improve the online efficiency of the computing node on the basis of ensuring the online reliability of the computing node, thus ensuring the reliability of the operation of the computing cluster.

[0010] In one embodiment, a network detection event for a compute node is detected, including:

[0011] It is detected that the compute node is a node newly added to the compute cluster; or,

[0012] It is detected that the compute node has completed fault handling.

[0013] In the embodiments of the present application, by performing pre - online network detection and operation detection on both the compute nodes newly added to the compute cluster and the compute nodes that have completed fault handling, the reliability of the compute nodes after going online can be effectively guaranteed.

[0014] In one embodiment, optimizing the online operation parameters of the compute node includes:

[0015] According to the online abnormal parameters of the compute node in the IB network detection result, as well as the current global operation parameters and standard global operation parameters of the compute cluster in the online dimension, determining the online parameter adjustment strategy for the compute node;

[0016] Adopting the online parameter adjustment strategy to optimize the online operation parameters of the compute node.

[0017] In the embodiments of the present application, by determining the online parameter adjustment strategy for the compute node according to the online abnormal parameters, the current global operation parameters and the standard global operation parameters, the online operation parameters of the compute node can be optimized from the perspective of the global operation of the compute cluster, thereby ensuring the rationality of optimizing the online operation parameters.

[0018] In one embodiment, the method further includes:

[0019] After the compute node goes online, if the online operation parameters of the compute node have not been optimized yet, when there are unprocessed single - machine tasks in the compute cluster, allocate single - machine tasks to the compute node;

[0020] When there are no unprocessed single - machine tasks in the compute cluster and there are unprocessed online tasks, allocate online tasks to the compute node according to the task computation amount of the online tasks.

[0021] In the embodiments of the present application, after the compute node goes online and when the online operation parameters have not been optimized yet, by sending compute tasks to the compute node according to the type of compute tasks to be allocated in the compute cluster, the rationality of computing task allocation can be guaranteed.

[0022] In one embodiment, when the compute node is a node newly added to the compute cluster, the method further includes:

[0023] Select a target container image adapted to the computing node from each candidate container image according to the node type of the computing node;

[0024] Pull up the target container image in the computing node, and generate configuration information of the computing node according to the target container image.

[0025] In the embodiments of the present application, by pre-building each candidate container image and selecting a target container image adapted to the computing node from each candidate container image according to the node type of the computing node, it is not necessary to reconfigure the operating environment of the computing node. Only by pulling up the target container image in the computing node can the operating environment configuration be completed, which can effectively improve the efficiency of operating environment configuration.

[0026] In one of the embodiments, the method further includes:

[0027] After the computing node goes online, if it is detected that the computing node is running abnormally, obtain the running abnormal information of the computing node;

[0028] According to the hardware abnormal information in the running abnormal information, determine the abnormal hardware in the computing node, and process the abnormal hardware in the computing node according to the hardware abnormal information, the task information of the computing node, and the influence degree of the abnormal hardware on the processing task of the computing node; and / or,

[0029] According to the software abnormal information in the running abnormal information, update the software of the computing node, and update the configuration information of the computing node according to the software update situation.

[0030] In the embodiments of the present application, by adopting corresponding methods to process the abnormal hardware and / or software in the computing node according to the type of the running abnormal information, the flexibility and reliability of fault handling can be ensured.

[0031] In one of the embodiments, processing the abnormal hardware in the computing node according to the hardware abnormal information, the task information of the computing node, and the influence degree of the abnormal hardware on the processing task of the computing node includes:

[0032] Determine the repairability of the computing node according to the hardware abnormal information;

[0033] In the case where the repairability is repairable, select a replacement hardware from the alternative hardwares corresponding to the abnormal hardware according to the task information of the computing node and the influence degree of the abnormal hardware on the processing task of the computing node, and use the replacement hardware to replace the abnormal hardware in the computing node;

[0034] In the case where the repairability is not repairable, perform a decommissioning process on the computing node.

[0035] In an embodiment of the present application, by determining the repairability of a computing node based on hardware exception information and adopting different repair methods to process the computing nodes in different repair situations, the reliability of the computing node repair can be ensured.

[0036] In a second aspect, the present application further provides a computing node processing device, including:

[0037] A network detection module, configured to perform a single-machine dimension network detection and an infinite bandwidth InfiniBand (IB) network detection in the online dimension on the computing node when a network detection event for the computing node is detected;

[0038] An operation detection module, configured to perform a single-machine operation detection on the computing node and optimize the online operation parameters of the computing node according to the configuration information of the computing node when the single-machine dimension network detection passes and the IB network detection in the online dimension fails;

[0039] A node online module, configured to perform an online process on the computing node after the single-machine operation test passes.

[0040] In a third aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0041] When a network detection event for the computing node is detected, perform a single-machine dimension network detection and an infinite bandwidth InfiniBand (IB) network detection in the online dimension on the computing node;

[0042] When the single-machine dimension network detection passes and the IB network detection in the online dimension fails, perform a single-machine operation detection on the computing node and optimize the online operation parameters of the computing node according to the configuration information of the computing node;

[0043] After the single-machine operation test passes, perform an online process on the computing node.

[0044] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0045] When a network detection event for the computing node is detected, perform a single-machine dimension network detection and an infinite bandwidth InfiniBand (IB) network detection in the online dimension on the computing node;

[0046] When the single-machine dimension network detection passes and the IB network detection in the online dimension fails, perform a single-machine operation detection on the computing node and optimize the online operation parameters of the computing node according to the configuration information of the computing node;

[0047] After the single - machine operation test passes, the computing node is put online.

[0048] In a fifth aspect, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps:

[0049] In the case of detecting a network detection event for a computing node, perform network detection in the single - machine dimension and infinite - bandwidth InfiniBand (IB) network detection in the online dimension for the computing node;

[0050] In the case where the network detection in the single - machine dimension passes and the IB network detection in the online dimension fails, perform single - machine operation detection on the computing node according to the configuration information of the computing node, and optimize the online operation parameters of the computing node;

[0051] After the single - machine operation test passes, the computing node is put online.

[0052] The above - mentioned computing node processing method, device, equipment, and storage medium, by detecting a network detection event for the computing node, perform network detection in the single - machine dimension and infinite - bandwidth InfiniBand (IB) network detection in the online dimension for the computing node, and in the case where the network detection in the single - machine dimension passes and the IB network detection in the online dimension fails, perform single - machine operation detection on the computing node according to the configuration information of the computing node, and optimize the online operation parameters of the computing node, and then after the single - machine operation test passes, put the computing node online. Compared with the related technology where each computing node is maintained by manually maintaining an operation and maintenance form, adopting the above - mentioned method, on the one hand, introducing the IB network detection in the online dimension can ensure that the computing node is adapted to the operation of the computing cluster after going online, ensuring the reliability of the computing node running in the computing cluster; on the other hand, in the case where the network detection in the single - machine dimension passes and the IB network detection in the online dimension fails, while optimizing the online operation parameters of the computing node, continue to perform the single - machine operation test before going online, and after the single - machine operation test passes, the computing node will be put online regardless of whether the online operation parameters are optimized or not, which will greatly improve the efficiency of the computing node going online on the basis of ensuring the reliability of the computing node going online, thus ensuring the reliability of the computing cluster operation. Description of the Drawings

[0053] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for describing the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0054] Figure 1 is a schematic flowchart of a computing node processing method in an embodiment;

[0055] Figure 2 is a schematic diagram of the full-cycle state of a computing node in an embodiment;

[0056] Figure 3 is a schematic flowchart of optimizing online operation parameters in an embodiment;

[0057] Figure 4 is a schematic flowchart of node configuration in an embodiment;

[0058] Figure 5 is a schematic flowchart of fault handling in an embodiment;

[0059] Figure 6 is a schematic flowchart of abnormal hardware repair in an embodiment;

[0060] Figure 7 is a schematic flowchart of a computing node processing method in another embodiment;

[0061] Figure 8 is a structural block diagram of a computing node processing device in an embodiment;

[0062] Figure 9 is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0063] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0064] With the continuous development of the computer field, the number of computing nodes in a computing cluster also increases. To ensure the reliability of the operation of the computing cluster, generally, each computing node is monitored by manually maintaining an operation and maintenance table. That is, when it is detected that the operation of any computing node in the computing cluster is abnormal, the abnormality is registered in the operation and maintenance table and waiting for the corresponding operation and maintenance personnel to handle the abnormality.

[0065] However, since the efficiency of maintaining each computing node by manually maintaining the operation and maintenance table is low, this will reduce the reliability of the operation of the computing cluster.

[0066] Based on this, in an exemplary embodiment, a computing node processing method is provided. This method can be executed by a computer device, which can be a server or a terminal device with relatively powerful computing functions, such asFigure 1 As shown, the method specifically includes the following steps:

[0067] S101, when a network detection event for a computing node is detected, perform network detection in the single-machine dimension and InfiniBand (IB) network detection in the online dimension on the computing node.

[0068] Among them, a computing node is the basic storage and computing unit in a computing cluster; a network detection event is an event for detecting the network status of a computing node; network detection in the single-machine dimension is detection only for the network status where a single computing node is located. Further, network detection in the single-machine dimension may include, but is not limited to, detection of the management network where the computing node is located and detection of the Intelligent Platform Management Interface (IPMI) network.

[0069] InfiniBand (IB) network is an interconnection technology dedicated to high-performance computing on the server side, with extremely high throughput and extremely low latency, used for data interconnection between computers (such as replication, distributed work, etc.). The IB network is also used as a direct or switched interconnection between a server and a storage system, as well as an interconnection between storage systems, and communication between a server and a network, and is widely used in fields such as data centers and HPC high-performance storage. Further, IB network detection is detection for the IB network status where the computing node is located.

[0070] In an alternative embodiment, before the computing node goes online, to ensure the reliability of the computing node operation, a network detection event for the computing node will be generated; then, when the management node detects a network detection event for the computing node, various network detection methods can be used to perform network detection in the single-machine dimension and IB network detection in the online dimension on the network where the computing node is located, respectively.

[0071] Exemplarily, for IPMI network detection in the single-machine dimension, it can be detected whether the Internet Protocol (IP) address of the IPMI interface can be accessed by the management network; verify whether the IPMI network parameters (IP address, subnet mask, gateway, etc.) conform to the plan; confirm whether the Baseboard Management Controller (BMC) is running normally; check whether there is unauthorized access or security vulnerability, etc.

[0072] For the management network detection in the single-machine dimension, it is possible to detect whether the physical link and logical communication of the management network are normal; verify whether the parameter configurations such as IP address, subnet, gateway, DNS, etc. comply with the plan; confirm whether management services such as the Secure Shell (SSH) and IPMI are running normally; check whether firewall rules, encryption protocols, and access control policies are compliant; and check network latency, bandwidth utilization rate, and packet loss rate, etc.

[0073] For the Infiniband (IB) network detection in the online dimension, it is possible to verify the physical connection status of IB cables, optical modules, and switches; verify the configuration of the Subnet Manager (SM) and the rationality of partition division; test whether the bandwidth, latency, and packet loss rate meet the expectations; and locate anomalies such as link oscillation and SM failures.

[0074] S102. When the network detection in the single-machine dimension passes and the IB network detection in the online dimension fails, perform single-machine operation detection on the compute node according to the configuration information of the compute node, and optimize the online operation parameters of the compute node.

[0075] Among them, the so-called configuration information of the compute node may include, but is not limited to, configuration information in various dimensions such as hardware, software, network, and security; the so-called single-machine operation detection is a test on whether the compute node itself can operate normally, that is, to determine whether the compute node meets the online requirements; the so-called online operation parameters are the operation parameters related to the interaction of the compute node with other compute nodes in the compute cluster.

[0076] It can be understood that for the online of the compute node, only the network detection result in the single-machine dimension will affect whether the compute node can go online. That is, when the network detection in the single-machine dimension passes, the compute node can go online normally, and when the network detection in the single-machine dimension fails, the compute node cannot go online normally. At this time, it is necessary to optimize the operation parameters of the compute node in the single-machine dimension so that the compute node meets the online requirements.

[0077] Furthermore, since the IB network detection in the online dimension only affects the operation effect of the compute node in the cluster and does not affect the online of the compute node, therefore, regardless of whether the IB network detection passes or not, the compute node can continue with the subsequent single-machine operation detection before going online; furthermore, in order to ensure the operation effect of the compute node in the compute cluster after going online, the operation situation of the compute node in the compute cluster can be adjusted simultaneously when performing the single-machine operation detection on the compute node.

[0078] In one implementation, when the network detection at the single-machine level passes, if the IB network detection passes, the single-machine operation detection of the computing node is performed only according to the configuration information of the computing node; if the IB network detection fails, while performing the single-machine operation detection of the computing node according to the configuration information of the computing node, it is also necessary to optimize the online operation parameters of the computing node.

[0079] Exemplarily, performing the single-machine operation detection on the computing node can be: performing configuration detection on the baseline of the computing node according to the configuration information of the computing node. For example, batch scanning the systems, account permissions, databases, weak passwords, level protection compliance configurations, etc. associated with the computing node. Or, detecting each entity Agent in the computing node. Among them, Agent is an entity that can perceive the environment and make autonomous decisions.

[0080] Optimizing the online operation parameters of the computing node can be: optimizing parameters such as the performance of the central processing unit (CPU), memory bandwidth (memory read / write bandwidth and latency, etc.), storage input / output IO, and the utilization rate of the graphics processing unit (GPU) of the computing node.

[0081] S103, after the single-machine operation test passes, perform an online processing on the computing node.

[0082] Optionally, after the single-machine operation test of the computing node passes, an online processing can be performed on the computing node so that the computing node can operate normally in the computing cluster.

[0083] It can be understood that since the single-machine operation detection and the optimization of the online operation parameters are carried out simultaneously, therefore, if the optimization of the online operation parameters has not been completed before the computing node goes online, after the computing node goes online, the online operation situation of the computing node will continue to be adjusted and the online operation situation will be monitored in real time.

[0084] In the above method for processing computing nodes, when a network detection event for a computing node is detected, network detection in the single-machine dimension and Infiniband (IB) network detection in the online dimension are performed on the computing node. When the network detection in the single-machine dimension passes and the IB network detection in the online dimension fails, according to the configuration information of the computing node, single-machine operation detection is performed on the computing node, and the online operation parameters of the computing node are optimized. After the single-machine operation test passes, the computing node is put online. Compared with the related art where each computing node is maintained by manually maintaining an operation and maintenance table, adopting the above method, on the one hand, introducing IB network detection in the online dimension can ensure that the computing node is adapted to the operation of the computing cluster after going online, ensuring the reliability of the computing node when running in the computing cluster. On the other hand, when the network detection in the single-machine dimension passes and the IB network detection in the online dimension fails, while optimizing the online operation parameters of the computing node, the single-machine operation test before going online continues, and after the single-machine operation test passes, the computing node will be put online regardless of whether the online operation parameters are optimized. This greatly improves the efficiency of putting the computing node online while ensuring the reliability of putting the computing node online, thus ensuring the reliability of the operation of the computing cluster.

[0085] To ensure the reliability of the operation of the computing node, based on the above embodiments, in the embodiments of the present application, multiple optional ways to detect a network detection event for a computing node are provided. Specifically, it is detected that the computing node is a newly added node to the computing cluster; or, it is detected that the computing node has completed fault handling.

[0086] In an optional implementation manner, since a computing node newly added to the computing cluster has not undergone network detection processing, therefore, when it is detected that the computing node is a newly added node to the computing cluster, it is determined that a network detection event for the computing node is detected.

[0087] It can be understood that in the related art, after a computing node undergoes fault handling, the computing node is directly put online again, lacking verification of the reliability of the operation of the computing node. Therefore, in another optional implementation manner, to ensure the reliability of the operation of the computing node after it is put online again after completing fault handling, when it is detected that the computing node has completed fault handling, it can be determined that a network detection event for the computing node is detected.

[0088] Furthermore, when the network detection of the computing node in the single-machine dimension passes, according to the updated configuration information, single-machine operation detection is performed on the computing node, and after the single-machine operation test passes, the computing node is put online.

[0089] Exemplarily, refer to Figure 2Schematic diagram of the full-cycle state of the computing node shown. When it is detected that the computing node is in the "managed" state, it is necessary to verify the integrity of the necessary fields of the computing node (such as information like IP address, etc.), and when the necessary fields are complete, place the computing node in the "to be debugged" state. Among them, detecting that the computing node is in the "managed" state means that the computing node joins the computing cluster for the first time, which can be manually added by the operation and maintenance personnel or detected by the corresponding port of the computing cluster.

[0090] For a computing node in the "to be debugged" state, it is necessary to perform network detection in the single-machine dimension and IB network detection in the online dimension on this computing node, and after the network detection in the single-machine dimension passes, set this computing node to the "debugging" state.

[0091] For a computing node in the "debugging" state, it is necessary to perform single-machine operation detection on this computing node according to the configuration information of this computing node, and after the single-machine operation test passes, perform online processing on this computing node, and at this time set this computing node to the "in service" state.

[0092] For a computing node in the "in service" state, the running state of this computing node is detected in real time, and when a failure occurs to this computing node, fault handling is performed on this computing node. After that, the computing node with the fault handling completed is reset to the "to be debugged" state to perform pre-online detection again.

[0093] In the embodiment of the present application, by performing pre-online network detection and operation detection on both the computing node newly added to the computing cluster and the computing node that has completed fault handling, the reliability of the computing node during operation after going online can be effectively guaranteed.

[0094] In order to ensure the reliability of the online operation of the computing node, on the basis of the above embodiment, in the embodiment of the present application, an optional method for optimizing the online operation parameters is provided, as Figure 3 shown, specifically including the following steps:

[0095] S301, determine the online parameter adjustment strategy of the computing node according to the online exception parameters of the computing node in the IB network detection result, as well as the current global operation parameters and standard global operation parameters of the computing cluster in the online dimension.

[0096] Among them, the so-called IB network detection result is the detection result obtained after detecting the running state of the computing node under the IB network; the so-called online abnormal parameter is the relevant parameter with abnormality when the computing node is running online; the so-called current global running parameter is the relevant parameter of the overall running of the computing cluster at the current moment; the so-called standard global running parameter is the relevant parameter of the overall running of the computing cluster when it is in a normal running state; the so-called online parameter adjustment strategy is the relevant strategy for adjusting the online abnormal parameters.

[0097] In an alternative implementation, when the IB network detection in the online dimension fails, the online abnormal parameters with abnormalities can be located according to the difference between the IB network detection result of the computing node and the standard IB network detection result.

[0098] Furthermore, based on the difference between the current global running parameter and the standard global running parameter of the computing cluster where the computing node is located in the dimension of online abnormal parameters, the adjustment direction of the online abnormal parameters can be determined, so as to generate the online parameter adjustment strategy of the computing node.

[0099] Exemplarily, when it is determined that the online abnormal parameter is the bandwidth occupied by the computing node, if the current global bandwidth occupancy of the computing cluster where the computing node is located is greater than the standard global bandwidth occupancy (it should be noted that the standard global bandwidth occupancy here is not the maximum global bandwidth occupancy), the online parameter adjustment strategy is to reduce the bandwidth occupancy of the computing node.

[0100] In another alternative implementation, the IB network detection result, the current global running parameter and the standard global running parameter of the computing cluster in the online dimension can be input into the trained first strategy generation model at the same time. The first strategy generation model outputs the online parameter adjustment strategy of the computing node according to the online abnormal parameters, the current global running parameter, the standard global running parameter and the model parameters in the IB network detection result.

[0101] S302. Adopt the online parameter adjustment strategy to optimize the online running parameters of the computing node.

[0102] In an alternative implementation, the computer device can directly adjust the online running parameters of the computing node based on the online parameter adjustment strategy, so that the computing node adapts to the running of the computing cluster after going online.

[0103] In another alternative implementation, in order to more intuitively monitor the optimization process of the online running parameters, the operation and maintenance work order containing the online parameter adjustment strategy can be sent to the corresponding operation and maintenance personnel, and the operation and maintenance personnel optimize the online running parameters of the computing node based on the online parameter adjustment strategy.

[0104] It is understandable that, due to the large volume of operation and maintenance processing of the computing cluster, in order to prevent operation and maintenance personnel from missing exception handling, an exception handling count table can be configured for the computing cluster. That is, each time an exception occurs, the exception handling count table will be updated to the exception calculation under the corresponding exception type to prompt the corresponding operation and maintenance personnel to perform real-time monitoring.

[0105] Exemplarily, 9 common exception types can be preset, namely, management network detection exception; abnormal BMC network inspection result; abnormal IB network inspection result; in the baseline inspection, it is found that the system versions do not match; the configuration baseline is not bound; the test baseline is not bound; in the baseline inspection, other mismatches except for the system version mismatch; the single-machine operation test fails; Agent inspection exception. For the above abnormal situations, when any one of the exceptions is found, the corresponding exception count is incremented by 1. At this time, when the operation and maintenance personnel handling this type of exception find that the count has changed, they can immediately process this type of exception based on the corresponding operation and maintenance work order, and after the processing is completed, decrement the corresponding exception count by 1.

[0106] The exception content detected in the above process will be synchronously recorded in the list of computing nodes and overwrite the historical exception content at the previous inspection, thereby completing the log update of the computing nodes.

[0107] In the embodiment of the present application, by determining the online parameter adjustment strategy of the computing node according to the online exception parameter, the current global operation parameter, and the standard global operation parameter, it is possible to optimize the online operation parameters of the computing node from the perspective of the global operation of the computing cluster, thereby ensuring the rationality of the optimization of the online operation parameters.

[0108] After the computing node is online, if the online operation parameters of the computing node have not been optimized, in order to ensure the rationality of task allocation, based on the above embodiment, in the embodiment of the present application, an optional method for computing task allocation is provided. Specifically, after the computing node is online, if the online operation parameters of the computing node have not been optimized, then when there are unprocessed single-machine tasks in the computing cluster, single-machine tasks are allocated to the computing node; when there are no unprocessed single-machine tasks in the computing cluster and there are unprocessed online tasks, online tasks are allocated to the computing node according to the task calculation amount of the online tasks.

[0109] Among them, the so-called single-machine task is a computing task that can be completed independently by a computing node; the so-called online task is a computing task that requires multiple computing nodes to cooperate to complete; the so-called task calculation amount is the data resource situation required for the computing task.

[0110] In an alternative embodiment, various types of computing tasks to be allocated in the computing cluster can be obtained, and it can be determined whether there are single-machine tasks; if there are unprocessed single-machine tasks in the computing cluster, the single-machine tasks can be sent to the computing nodes for processing by the computing nodes, and after the online operation parameters of the computing nodes are optimized, online tasks can be allocated to the computing nodes.

[0111] If there are no unprocessed single-machine tasks in the computing cluster and only unprocessed online tasks exist, the online tasks can be sorted based on the task computation amounts of the online tasks in ascending order of the computation amounts, and a preset number of the online tasks ranked ahead can be sent to the computing nodes for processing by the computing nodes.

[0112] In the embodiments of the present application, after the computing nodes are online and the online operation parameters are not yet optimized, by sending computing tasks to the computing nodes according to the types of computing tasks to be allocated in the computing cluster, the rationality of computing task allocation can be ensured.

[0113] To ensure the efficiency of the computing node running environment configuration, based on the above embodiments, in the embodiments of the present application, an alternative method for node configuration is provided, as Figure 4 shown, which specifically includes the following steps:

[0114] S401, according to the node type of the computing node, select a target container image adapted to the computing node from each candidate container image.

[0115] Among them, the so-called node type is used to represent the type of computing resources included in the computing node. For example, it can include, but is not limited to, homogeneous nodes and heterogeneous nodes, etc. The so-called candidate container image is a container image stored locally that can be pulled up on the computing node; the so-called target container image is a container image adapted to the computing node.

[0116] Optionally, to ensure the efficiency of the running environment configuration, the container images corresponding to various node types can be pre-saved as candidate container images. Then, using the node type of the computing node as an index, query from each candidate container image to determine the target container image adapted to the computing node.

[0117] S402, pull up the target container image in the computing node and generate configuration information of the computing node according to the target container image.

[0118] Optionally, after selecting the computing node, the target container image can be directly pulled up on the computing node to create the running environment of the computing node; and, according to the configuration information of the target container image, generate the configuration information of the computing node.

[0119] In the embodiments of the present application, by pre - constructing each candidate container image and selecting a target container image adapted to the computing node from each candidate container image according to the node type of the computing node, it is not necessary to re - configure the operating environment of the computing node. Only by pulling up the target container image in the computing node can the operating environment configuration be completed, which can effectively improve the efficiency of operating environment configuration.

[0120] To ensure the reliability of node failure handling, on the basis of the above - mentioned embodiments, in the embodiments of the present application, an optional method for failure handling is provided, such as Figure 5 shown, and specifically includes the following steps:

[0121] S501. After the computing node goes online, if it is detected that the computing node is running abnormally, obtain the running abnormal information of the computing node.

[0122] Among them, the so - called running abnormal information is the running information generated when the computing node runs abnormally, which may include hardware abnormal information and / or software abnormal information. Further, the so - called hardware abnormal information is the running information generated when there is a hardware failure in the computing node; the so - called software abnormal information is the running information generated when there is a software abnormality in the computing node.

[0123] Optionally, after the computing node goes online, the running of the computing node can be detected in real - time, and when it is detected that the computing node is running abnormally, obtain the running abnormal information generated when the computing node runs abnormally.

[0124] S502. Determine the abnormal hardware in the computing node according to the hardware abnormal information in the running abnormal information, and process the abnormal hardware in the computing node according to the hardware abnormal information, the task information of the computing node, and the influence degree of the abnormal hardware on the computing tasks of the computing node; and / or, update the software of the computing node according to the software abnormal information in the running abnormal information, and update the configuration information of the computing node according to the software update situation.

[0125] Among them, the so - called abnormal hardware is the hardware with a fault; the so - called task information is the relevant information of each computing task in the computing node, which may include but is not limited to the relevant information of the computing resources required for the computing task; the so - called influence degree of the abnormal hardware on the computing tasks of the computing node is the importance degree of the abnormal hardware in the computing node. The so - called software update situation is the situation such as software version upgrade and re - configuration in the computing node.

[0126] In an alternative implementation, if only hardware exception information exists in the running exception information, the abnormal hardware in the computing node can be located according to the difference between the hardware exception information in the running exception information and the standard hardware information; thereafter, according to the hardware exception information, the task information of the computing node, and the impact degree of the abnormal hardware on the task processing of the computing node, a repair strategy for the abnormal hardware can be determined, and the repair strategy for the abnormal hardware can be used to process the abnormal hardware in the computer node.

[0127] Exemplarily, the hardware exception information, the task information of the computing node, and the impact degree of the abnormal hardware on the task processing of the computing node can be input into a trained second policy generation model at the same time. The second policy generation model outputs a repair strategy for the abnormal hardware according to the hardware exception information, the task information of the computing node, the impact degree of the abnormal hardware on the task processing of the computing node, and the model parameters; thereafter, an operation and maintenance work order containing the repair strategy can be sent to the corresponding operation and maintenance personnel to prompt the operation and maintenance personnel to monitor the progress of the hardware repair process.

[0128] In another alternative implementation, if only software exception information exists in the running exception information, the abnormal software in the computing node can be upgraded and reconfigured according to the software exception information; thereafter, the configuration information of the computing node can be updated according to the configuration information after the software update.

[0129] For example, as Figure 2 shown, when the computing node is in the "maintenance" state and it is determined that the cause of the exception of the computing node is software exception, the abnormal software can be determined according to the difference between the software exception information and the standard software information, and the abnormal software can be upgraded and reconfigured. Thereafter, the computing node is set to the "changing" state, and the configuration information of the computing node is updated according to the configuration information after the software update. Finally, after the configuration information is updated, the computing node is set to the "to be debugged" state to re-perform the detection process before the computing node goes online.

[0130] In yet another alternative implementation, if both hardware exception information and software exception information exist in the running exception information, it is necessary to determine the abnormal hardware in the computing node according to the hardware exception information in the running exception information, and process the abnormal hardware in the computer node according to the hardware exception information, the task information of the computing node, and the impact degree of the abnormal hardware on the task processing of the computing node. Also, the software of the computing node is updated according to the software exception information in the running exception information, and the configuration information of the computing node is updated according to the software update situation.

[0131] In the embodiments of the present application, by adopting corresponding methods to handle the faulty hardware and / or software in the computing node according to the type of running exception information, the flexibility and reliability of fault handling can be ensured.

[0132] To ensure the reliability of computing node repair, based on the above embodiments, in the embodiments of the present application, an optional method for repairing faulty hardware is provided, as Figure 6 shown, which specifically includes the following steps:

[0133] S601, determine the repairability of the computing node according to the hardware exception information.

[0134] Optionally, the repairability of the computing node can be determined according to the hardware fault type in the hardware exception information. Exemplarily, if the hardware exception information indicates that the faulty hardware is a replaceable component in the computing node, or the faulty component can be directly repaired, at this time, it is determined that the computing node is repairable, and step S602 is executed. If the hardware exception information indicates that the faulty hardware is a non-replaceable component in the computing node, at this time, it is determined that the computing node is not repairable, and step S603 is executed.

[0135] S602, when the repairability is repairable, select a replacement hardware from the alternative hardwares corresponding to the faulty hardware according to the task information of the computing node and the impact degree of the faulty hardware on the computing node's processing tasks, and use the replacement hardware to replace the faulty hardware in the computing node.

[0136] Among them, the so-called alternative hardware is other model hardwares that can replace the faulty hardware; the so-called replacement hardware is the hardware selected to replace the faulty hardware.

[0137] In an optional implementation manner, when it is determined that the computing node is repairable, it can be directly determined according to the task information of the computing node and the impact degree of the faulty hardware on the computing node's processing tasks whether it is necessary to upgrade the hardware model on the original hardware model of the faulty hardware; then, according to the determined hardware type, select the required replacement hardware from the alternative hardwares corresponding to the faulty hardware, and use the replacement hardware to replace the faulty hardware in the computing node.

[0138] In another optional implementation manner, referring to Figure 2 , the damage degree of the faulty hardware can be determined according to the hardware exception information. If it is determined according to the damage degree of the faulty hardware that there is no need to replace the hardware, the operation and maintenance work order can be directly sent to the corresponding operation and maintenance personnel according to the hardware exception information to prompt the operation and maintenance personnel to handle the faulty hardware.

[0139] If it is determined that the hardware needs to be replaced according to the damage degree of the defined abnormal hardware, after determining the replacement hardware, an operation and maintenance work order including the abnormal hardware and the replacement hardware is sent to the corresponding operation and maintenance personnel to prompt the operation and maintenance personnel to process the abnormal hardware. After replacing the abnormal hardware, the configuration information of the computing node can be updated according to the configuration information of the replacement hardware.

[0140] S603. In the case where the repairability is non-repairable, the computing node is taken offline.

[0141] Optionally, in the case where it is determined that the computing node is non-repairable, the computing node can be directly taken offline. It should be noted that the offline processing here is a non-rebootable scrapping process. Further, referring to Figure 2 , in practical applications, due to the need for energy-saving / liquid cooling maintenance, the computing nodes in the "to be debugged" state and the "under maintenance" state may be powered off, and these powered-off computing nodes can be restarted.

[0142] In the embodiments of the present application, by determining the repairability of the computing node according to the hardware exception information and adopting different repair methods to process the computing nodes in different repair situations, the reliability of the computing node repair can be ensured.

[0143] Figure 7 For the flowchart of the computing node processing method in another embodiment, on the basis of the above embodiments, this embodiment provides an optional example of the computing node processing method. Combining Figure 7 , the specific implementation process is as follows:

[0144] S701. When it is monitored that the computing node is a newly added node in the computing cluster, or when it is monitored that the computing node has completed the fault handling, a single-machine dimension network detection and an IB network detection in the online dimension are performed on the computing node.

[0145] Optionally, in the case where the computing node is a newly added node in the computing cluster, according to the node type of the computing node, a target container image adapted to the computing node is selected from each candidate container image; the target container image is pulled up in the computing node, and the configuration information of the computing node is generated according to the target container image.

[0146] S702. In the case where the single-machine dimension network detection passes and the IB network detection in the online dimension fails, a single-machine operation detection is performed on the computing node according to the configuration information of the computing node.

[0147] S703. According to the online exception parameters of the computing node in the IB network detection result, as well as the current global operation parameters and the standard global operation parameters of the computing cluster in the online dimension, an online parameter adjustment strategy for the computing node is determined.

[0148] S704 adopts an online parameter adjustment strategy to optimize the online operation parameters of the computing node.

[0149] S705, after the single-machine operation test is passed, the computing node is put online.

[0150] Optionally, after the computing node is put online, if the online operation parameters of the computing node are not optimized yet, when there are unprocessed single-machine tasks in the computing cluster, single-machine tasks are assigned to the computing node; when there are no unprocessed single-machine tasks in the computing cluster and there are unprocessed online tasks, online tasks are assigned to the computing node according to the task computing amount of the online tasks.

[0151] S706, after the computing node is put online, if it is detected that the computing node is running abnormally, the running abnormal information of the computing node is obtained.

[0152] S707, according to the hardware abnormal information in the running abnormal information, determine the abnormal hardware in the computing node and the repairability of the computing node.

[0153] S708, when the repairability is repairable, according to the task information of the computing node and the influence degree of the abnormal hardware on the task processing of the computing node, select a replacement hardware from the alternative hardwares corresponding to the abnormal hardware, and use the replacement hardware to replace the abnormal hardware in the computing node.

[0154] S709, when the repairability is not repairable, the computing node is taken offline.

[0155] S710, according to the software abnormal information in the running abnormal information, update the software of the computing node, and update the configuration information of the computing node according to the software update situation.

[0156] For the specific processes of the above S701 - S710, reference can be made to the description of the method embodiments above. Their implementation principles and technical effects are similar and will not be elaborated here.

[0157] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless specifically stated herein, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turns with at least a part of other steps or steps or stages in other steps.

[0158] Based on the same inventive concept, an embodiment of the present application further provides a computing node processing device for implementing the above-mentioned computing node processing method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following computing node processing devices can refer to the limitations on the computing node processing method in the foregoing, and will not be repeated here.

[0159] In an exemplary embodiment, as Figure 8 shown, a computing node processing device 1 is provided, including: a network detection module 10, an operation detection module 20, and a node online module 30, where:

[0160] The network detection module 10 is configured to perform a single-machine dimension network detection and an infinite bandwidth IB network detection in the online dimension on the computing node when a network detection event for the computing node is detected;

[0161] The operation detection module 20 is configured to perform a single-machine operation detection on the computing node and optimize the online operation parameters of the computing node according to the configuration information of the computing node when the network detection in the single-machine dimension passes and the IB network detection in the online dimension fails;

[0162] The node online module 30 is configured to perform an online processing on the computing node after the single-machine operation test passes.

[0163] In an exemplary embodiment, the network detection module 10 specifically includes:

[0164] It is detected that the computing node is a node newly added to the computing cluster; or, it is detected that the computing node has completed a fault handling.

[0165] In an exemplary embodiment, the operation detection module 20 specifically includes:

[0166] Determine the online parameter adjustment strategy for the computing node based on the online exception parameters of the computing node in the IB network detection result, as well as the current global operating parameters and standard global operating parameters of the computing cluster in the online dimension; adopt the online parameter adjustment strategy to optimize the online operating parameters of the computing node.

[0167] In an exemplary embodiment, the computing node processing device 1 further includes a task allocation module, wherein the task allocation module is specifically configured to:

[0168] After the computing node goes online, if the online operating parameters of the computing node have not been optimized, then when there are unprocessed single-machine tasks in the computing cluster, allocate single-machine tasks to the computing node; when there are no unprocessed single-machine tasks in the computing cluster and there are unprocessed online tasks, allocate online tasks to the computing node according to the task calculation amount of the online tasks.

[0169] In an exemplary embodiment, when the computing node is a node newly added to the computing cluster, the computing node processing device 1 further includes an environment configuration module, wherein the environment configuration module is specifically configured to:

[0170] Select a target container image adapted to the computing node from each candidate container image according to the node type of the computing node; pull up the target container image in the computing node, and generate configuration information of the computing node according to the target container image.

[0171] In an exemplary embodiment, the computing node processing device 1 further includes a fault handling module, wherein the fault handling module is specifically configured to:

[0172] An information acquisition unit, configured to, after the computing node goes online, if it detects that the computing node is operating abnormally, acquire the operating abnormal information of the computing node;

[0173] A hardware repair unit, configured to determine the abnormal hardware in the computing node according to the hardware abnormal information in the operating abnormal information, and process the abnormal hardware in the computing node according to the hardware abnormal information, the task information of the computing node, and the influence degree of the abnormal hardware on the task processing of the computing node; and / or,

[0174] A software upgrade unit, configured to update the software of the computing node according to the software abnormal information in the operating abnormal information, and update the configuration information of the computing node according to the software update situation.

[0175] In an exemplary embodiment, the hardware repair unit is specifically configured to:

[0176] Determine the repairability of the computing node according to the hardware exception information; in the case where the repairability is repairable, select a replacement hardware from the alternative hardwares corresponding to the abnormal hardware according to the task information of the computing node and the influence degree of the abnormal hardware on the task processing of the computing node, and use the replacement hardware to replace the abnormal hardware in the computing node; in the case where the repairability is not repairable, perform a decommissioning process on the computing node.

[0177] Each module in the above computing node processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0178] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 9 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store node operation data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a computing node processing method.

[0179] Those skilled in the art can understand that Figure 9 the structure shown in

[0180] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0181] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.

[0182] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.

[0183] It should be noted that the data involved in this application (including but not limited to node operation data, etc.) are all data authorized by users or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0184] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, Resistive Random Access Memory (ReRAM), Magnetoresistive Random Access Memory (MRAM), Ferroelectric Random Access Memory (FRAM), Phase Change Memory (PCM), graphene memory, etc. Volatile memory can include Random Access Memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, Artificial Intelligence (AI) processors, etc., without limitation.

[0185] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0186] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A computing node processing method, characterized in that, The method includes: In the case of detecting a network detection event for a computing node, performing network detection in the single-machine dimension and infinite bandwidth IB network detection in the online dimension on the computing node; In the case where the network detection in the single-machine dimension passes and the IB network detection in the online dimension fails, performing single-machine operation detection on the computing node according to the configuration information of the computing node, and optimizing the online operation parameters of the computing node; After the single-machine operation test passes, performing online processing on the computing node.

2. The method according to claim 1, wherein The detecting a network detection event for a computing node includes: Detecting that the computing node is a node newly added to the computing cluster; or, Detecting that the computing node has completed fault handling.

3. The method according to claim 1, wherein The optimizing the online operation parameters of the computing node includes: Determining an online parameter adjustment strategy for the computing node according to the online abnormal parameters of the computing node in the IB network detection result, and the current global operation parameters and standard global operation parameters of the computing cluster in the online dimension; Using the online parameter adjustment strategy to optimize the online operation parameters of the computing node.

4. The method according to claim 1, wherein The method further includes: After the computing node goes online, if the online operation parameters of the computing node are not optimized, then in the case where there are unprocessed single-machine tasks in the computing cluster, allocating the single-machine tasks to the computing node; In the case where there are no unprocessed single-machine tasks in the computing cluster and there are unprocessed online tasks, allocating online tasks to the computing node according to the task calculation amount of the online tasks.

5. The method according to claim 1, characterized in that In the case where the computing node is a node newly added to the computing cluster, the method further includes: Selecting a target container image adapted to the computing node from each candidate container image according to the node type of the computing node; Pulling up the target container image in the computing node and generating the configuration information of the computing node according to the target container image.

6. The method according to claim 1, wherein The method further includes: After the computing node goes online, if it is detected that the computing node is operating abnormally, obtaining the operation abnormal information of the computing node; Determining the abnormal hardware in the computing node according to the hardware abnormal information in the operation abnormal information, and processing the abnormal hardware in the computing node according to the hardware abnormal information, the task information of the computing node, and the influence degree of the abnormal hardware on the task processing of the computing node; and / or, Updating the software of the computing node according to the software abnormal information in the operation abnormal information, and updating the configuration information of the computing node according to the software update situation.

7. The method according to claim 6, wherein The processing the abnormal hardware in the computing node according to the hardware abnormal information, the task information of the computing node, and the influence degree of the abnormal hardware on the task processing of the computing node includes: Determining the repairability of the computing node according to the hardware abnormal information When the repairability is repairable, select a replacement hardware from the alternative hardwares corresponding to the abnormal hardware according to the task information of the computing node and the impact degree of the abnormal hardware on the task processing of the computing node, and use the replacement hardware to replace the abnormal hardware in the computing node; When the repairability is non-repairable, take the computing node offline.

8. A computing node processing device, characterized in that The device includes: A network detection module, configured to perform a network detection in a single-machine dimension and an Infiniband (IB) network detection in an online dimension on the computing node when a network detection event for the computing node is detected; An operation detection module, configured to perform a single-machine operation detection on the computing node and optimize the online operation parameters of the computing node according to the configuration information of the computing node when the network detection in the single-machine dimension passes and the IB network detection in the online dimension fails; A node online module, configured to bring the computing node online after the single-machine operation test passes.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.