Node management method and related device

By using fault prediction information and neural network models to isolate high-risk computing nodes in advance in the HPC environment, the job exit problem caused by computing node failure is solved, and efficient resource utilization and safe operation completion are achieved.

WO2025148491A1PCT designated stage expired Publication Date: 2025-07-17XFUSION DIGITAL TECH CO LTD

Patent Information

Application Number
PCT/CN2024/129162
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-11
Filing Date
2024-10-31
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

In the field of high performance computing (HPC), memory failures of computing nodes cause job exit, resulting in waste of resources and delayed calculation results, which is difficult to effectively prevent in the existing technology.

Method used

By obtaining the failure prediction information of the computing node, we determine whether it meets the isolation conditions, and isolate the computing nodes in advance when there are high risks, avoid job allocation, and use the fault prediction neural network model and management nodes to evaluate and isolate the failure risk.

Benefits of technology

It effectively avoids job exit caused by computing node failure, reduces resource waste, and improves the reliability of computing nodes and job security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024129162_17072025_PF_FP_ABST
    Figure CN2024129162_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in embodiments of the present application are a node management method and a related device. The method is applied to a management node in a server cluster. The method comprises: acquiring failure prediction information of a computing node in the server cluster; on the basis of the failure prediction information, determining whether the computing node satisfies an isolation condition; if the computing node satisfies the isolation condition, isolating the computing node; and when the computing node is in an isolated state, stopping allocating jobs to the computing node and / or allocating, to other computing nodes in an unisolated state in the server cluster, a job that the computing node is performing. By determining, on the basis of the failure prediction information of the computing node, whether the computing node satisfies the isolation condition, a computing node having a high failure risk is isolated in advance, so that the situation where the job of the server cluster exits due to a failure of the computing node can be effectively avoided, thereby reducing the loss caused by the failure of the computing node.
Need to check novelty before this filing date? Find Prior Art

Description

Node management method and related equipment

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on January 11, 2024, with application number 202410047061.7 and application name “Node Management Method and Related Equipment”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of servers, and in particular to a node management method and related equipment. Background Art

[0003] With the advancement of computer technology, the memory capacity required by computing nodes has increased. This increased memory also leads to an increasing rate of basic failure of computing nodes. In the field of high-performance computing (HPC), HPC jobs require a large amount of computing resources, often requiring multiple computing nodes to run for days or even dozens of days to complete. If a memory failure occurs in one of these multiple computing nodes, the entire HPC job can be aborted, resulting in the invalidation and waste of the computing resources already invested, and delayed calculation results.

[0004] Summary of the Invention

[0005] The embodiments of the present application provide a node management method and related equipment, which can effectively avoid the situation where the server cluster's jobs are exited due to computing node failures, thereby avoiding resource waste and reducing losses caused by computing node failures.

[0006] A first aspect of an embodiment of the present application provides a node management method, which is applied to a management node in a server cluster, where the server cluster also includes computing nodes. The method includes:

[0007] Obtain fault prediction information of the computing node, where the fault prediction information is used to indicate the risk of failure of the computing node; determine whether the computing node meets the isolation conditions based on the fault prediction information; if the computing node meets the isolation conditions, isolate the computing node; when the computing node is in an isolated state, stop assigning jobs to the computing node and / or assign the ongoing jobs of the computing node to other computing nodes in the server cluster that are not in an isolated state.

[0008] Optionally, the management node may receive fault prediction information sent by the computing node.

[0009] Optionally, the management node may obtain fault prediction information from a computing device running the management platform software.

[0010] In an embodiment of the present application, by determining whether a computing node meets the isolation conditions based on the fault prediction information of the computing node, the computing node with a high risk of failure is isolated in advance, which can effectively avoid the situation where the server cluster's jobs are exited due to computing node failure, thereby avoiding resource waste and reducing losses caused by computing node failure.

[0011] In one possible implementation, obtaining the fault prediction information of the computing node includes: obtaining the fault prediction information multiple times; and determining whether the computing node meets the isolation condition based on the fault prediction information includes: determining whether the computing node meets the isolation condition based on the fault prediction information obtained multiple times.

[0012] Optionally, obtaining the fault prediction information multiple times includes: periodically obtaining the fault prediction information multiple times.

[0013] In an embodiment of the present application, by comprehensively judging whether a computing node meets the isolation conditions based on fault prediction information obtained multiple times, it is possible to avoid isolation errors caused by single fault prediction information corresponding to short-term occasional working conditions or algorithm fluctuations of the computing node, thereby avoiding waste of resources caused by unnecessary isolation.

[0014] In one possible implementation, the determination of whether the computing node meets the isolation condition is based on the fault prediction information obtained multiple times, including: after each acquisition of the fault prediction information, if the fault prediction information indicates that the fault risk level of the computing node is high, adding one to the risk cycle count of the computing node; after the multiple acquisitions of the fault prediction information, if the risk cycle count is greater than or equal to a preset maximum value, determining that the computing node meets the isolation condition; if the risk cycle count is less than the preset maximum value, determining that the computing node does not meet the isolation condition.

[0015] In an embodiment of the present application, by accumulating risk cycle counts and determining that the computing node meets the isolation condition when the risk cycle count is greater than or equal to a preset maximum value, a computing node with a high fault risk level for a long time can be determined to meet the isolation condition and thus isolated; this can improve the accuracy of isolation and avoid waste of resources caused by unnecessary isolation.

[0016] In one possible implementation, the determination of whether the computing node meets the isolation condition is based on the fault prediction information obtained multiple times, including: after each acquisition of the fault prediction information, if the fault prediction information indicates that the fault risk level of the computing node is not high, clearing the risk cycle count of the computing node; after the multiple acquisitions of the fault prediction information, if the risk cycle count is greater than or equal to a preset maximum value, determining that the computing node meets the isolation condition; if the risk cycle count is less than the preset maximum value, determining that the computing node does not meet the isolation condition.

[0017] In the embodiment of the present application, by clearing the risk cycle count of the computing node when the failure risk level of the computing node is not high, the computing node can perform operations in more time, thereby speeding up the operation progress.

[0018] In one possible implementation, the method determines whether the computing node meets the isolation condition based on the fault prediction information obtained multiple times, including: after each acquisition of the fault prediction information, if the fault prediction information indicates that the fault risk level of the computing node decreases after the fault is repaired, and the fault risk level is not high, then clearing the risk cycle count of the computing node; after the multiple acquisitions of the fault prediction information, if the risk cycle count is greater than or equal to a preset maximum value, then determining that the computing node meets the isolation condition; if the risk cycle count is less than the preset maximum value, then determining that the computing node does not meet the isolation condition.

[0019] Optionally, if the fault prediction information indicates that the fault risk level of the computing node has decreased due to failure to repair the fault, and the fault risk level is not high, the risk cycle count of the computing node is reduced by one.

[0020] In an embodiment of the present application, the risk cycle count is cleared only when it is determined that the fault risk level of the computing node has decreased due to fault repair. A stricter clearing condition is set, which can make it easier for the computing node to meet the isolation conditions and improve the safety of the operation.

[0021] In one possible implementation, before determining whether the computing node meets the isolation condition based on the fault prediction information, the method also includes: determining whether the fault isolation function is turned on; if it is turned on, triggering the step of determining whether the computing node meets the isolation condition based on the fault prediction information.

[0022] In an embodiment of the present application, when it is determined that the fault isolation function is turned on, the corresponding fault isolation judgment and specific isolation actions are performed, and the switch of the fault isolation function can be controlled according to actual needs to more efficiently utilize computing resources.

[0023] In one possible implementation, the fault prediction information includes the probability that a memory failure occurs in the computing node, resulting in job exit; after obtaining the fault prediction information each time, the method further includes: if the probability is greater than or equal to a first threshold, determining that the fault risk level of the computing node is high; if the probability is less than the first threshold, determining that the fault risk level of the computing node is not high.

[0024] In the embodiment of the present application, compared with the scheme based on the failure probability of other hardware, there are more samples of memory failure, and in actual applications, the proportion of job exits such as downtime due to memory failure is higher. Therefore, the scheme of determining the failure risk level of the computing node based on the memory failure probability and the first threshold is more applicable.

[0025] A second aspect of an embodiment of the present application provides a node management method, which is applied to a computing node in a server cluster, wherein the server cluster also includes a management node; the method further includes:

[0026] Obtaining fault information of the computing node itself; determining fault prediction information based on the fault information, the fault prediction information being used to indicate the risk level of the computing node failing; and sending the fault prediction information to the management node.

[0027] In the embodiment of the present application, the fault information is obtained and the fault prediction information is determined by the computing node itself, and then the fault prediction information is sent to the management node. The process is simpler and more direct, and no additional computing equipment is required, thus saving costs.

[0028] In a possible implementation, determining fault prediction information based on the fault information includes: inputting the fault information into a fault prediction neural network model to obtain the fault prediction information.

[0029] In the embodiment of the present application, fault prediction information is obtained using a trained fault prediction neural network model, which has a higher accuracy rate.

[0030] In one possible implementation, obtaining the fault information of the computing node itself includes: periodically obtaining the fault information; before sending the fault prediction information to the management node, the method also includes: periodically obtaining node resource information; sending the fault prediction information to the management node includes: periodically sending node status information to the management node, and the node status information includes the node resource information and the fault prediction information of the same period.

[0031] In an embodiment of the present application, communication resources can be saved by periodically acquiring node resource information and fault prediction information of a computing node and including the two in node status information and sending the information to a management node.

[0032] In one possible implementation, the fault information includes fault information of the memory of the computing node. In the embodiment of the present application, the solution of determining fault prediction information based on memory fault information is more applicable.

[0033] A third aspect of the present application further provides a node management device, the device comprising:

[0034] A message sending and receiving module is used to obtain fault prediction information of a computing node, which is used to indicate the risk of failure of the computing node; a fault isolation module is used to determine whether the computing node meets the isolation conditions based on the fault prediction information; a node management module is used to isolate the computing node if the computing node meets the isolation conditions; and a scheduling engine is used to stop allocating jobs to the computing node and / or allocate the ongoing jobs of the computing node to other computing nodes that are not isolated if the computing node is in an isolated state.

[0035] The fourth aspect of the present application also provides a management node, which includes a processor and a memory; the processor and the memory are coupled; the memory is used to store program instructions; the processor is used to run the program instructions, so that the management node executes any possible implementation method in the first aspect.

[0036] The fifth aspect of the present application also provides a computing node, which includes a processor and a memory; the processor and the memory are coupled; the memory is used to store program instructions; the processor is used to run the program instructions, so that the computing node executes any possible implementation method in the second aspect.

[0037] The sixth aspect of the present application also provides a server cabinet, which includes a cabinet and a node, and the node is arranged in the cabinet. The node includes the management node as described in the fourth aspect and the computing node as described in the fifth aspect; the management node is communicatively connected with the computing node, and the management node is used to manage and schedule the computing resources of the computing node to perform operations.

[0038] The seventh aspect of the present application also provides a server cluster, which includes a management node as described in the fourth aspect and a computing node as described in the fifth aspect; the management node is communicatively connected to the computing node, and the management node is used to manage and schedule the computing resources of the computing node to perform operations.

[0039] It should be understood that the beneficial effects of the above aspects can be referenced to each other. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] FIG1 is a schematic diagram of a system architecture of a server cluster;

[0041] FIG2 is a schematic diagram of a system architecture of a server cluster provided in an embodiment of the present application;

[0042] FIG3a is an application scenario diagram of a server cluster provided in an embodiment of the present application;

[0043] FIG3 b is an application scenario diagram of a server cluster provided in an embodiment of the present application;

[0044] FIG4 is a schematic diagram of the structure of a management node provided in an embodiment of the present application;

[0045] FIG5 is a schematic diagram of the structure of a computing node provided in an embodiment of the present application;

[0046] FIG6 is a flow chart of a node management method provided in an embodiment of the present application;

[0047] FIG7 is a flow chart of another node management method provided in an embodiment of the present application;

[0048] FIG8 is a flow chart of another node management method provided in an embodiment of the present application;

[0049] FIG9 is a schematic diagram of a process for determining whether a computing node meets isolation conditions according to an embodiment of the present application;

[0050] FIG10 is a schematic diagram of a flow chart of another node management method provided in an embodiment of the present application;

[0051] FIG11 is a schematic structural diagram of a node management device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present application, rather than all the embodiments. Those skilled in the art will appreciate that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0053] The terms "first," "second," and the like in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.

[0054] Please refer to Figure 1, which is a schematic diagram of the system architecture of a server cluster. As shown in Figure 1, the server cluster includes a management node 10 and computing nodes 20a to 20c (hereinafter collectively referred to as computing nodes 20); the management node 10 is communicatively connected to each computing node 20. Optionally, the server cluster also includes a storage node 30, and each computing node 20 is communicatively connected to a storage node 30. Optionally, the server cluster is a high performance computing (HPC) cluster.

[0055] The management node 10 may be a computing device, and the computing node 20 may also be a computing device. The computing device may be a server, or may be a smart terminal such as a personal computer (PC), a laptop, or a tablet computer. For example, the computing device may be a server node.

[0056] The storage node 30 may be a storage device in which a database is running.

[0057] Management node 10 may include a central processing unit (CPU) for running a scheduler master process. The scheduler master process allocates resources for running jobs, i.e., allocates compute nodes for running jobs. Specifically, the scheduler master process may send job messages to compute nodes 20, instructing them to perform the job.

[0058] The computing node 20 includes a CPU, which can be used to run an agent process. The agent process can be used to communicate with the scheduler main process and receive job messages sent by the management node 10 through the scheduler main process. It can also be used to send its own node status information to the scheduler main process of the management node 10. Exemplarily, the node status information includes the resource usage of the computing node 20 itself, such as CPU resource usage and memory usage.

[0059] The storage system 30 is used to store data required by each computing node 20 during the operation. In other possible implementations, the data required by the computing node 20 during the operation can also be stored in the memory inside the computing node 20.

[0060] The management node 10 can also be used to receive job requests submitted by users, and its scheduler main process can schedule the computing nodes 20 to perform the job according to the job request. Specifically, the scheduler main process can determine the computing node 20 with sufficient idle resources to complete the job and send the corresponding job message to the computing node 20.

[0061] A job is a general term for the work that a server cluster is required to do in a job request submitted by a user.

[0062] Optionally, the computing node 20 includes a computing master node and computing sub-nodes. When the scheduler main process on the management node 10 schedules computing node resources, it sends a job message to the agent process of the corresponding computing master node. After receiving the job message sent by the management node 10, the agent process of the computing master node starts the job monitoring process to monitor the job status; and generates a host file (host file) based on the scheduling result contained in the job message. Exemplarily, the scheduling result contained in the job message can be computing node 1 and computing node 2, and the generated host file can include computing node 1 and computing node 2; then, the computing master node can use the mpirun command to pull up the local job process based on the host file, and pull up the job process on the computing sub-node indicated in the scheduling result through the secure shell protocol (SSH). The job processes can communicate with each other through the message passing interface (MPI).

[0063] It is understandable that the number of management nodes 10, computing nodes 20 and storage nodes 30 shown in FIG1 is for example only and not for limitation. In the specific implementation process, they can be set according to actual needs.

[0064] In the server cluster shown in Figure 1, the scheduler main process of the management node 10 periodically receives and updates the node status information sent by each computing node 20, so that when receiving a new job request, the management node 10 can accurately determine the computing device that can complete the corresponding job. When a computing node 20 is unable to report status information to the management node 10 due to a fault or other reasons, the management node 10 does not perceive the specific circumstances of the fault. Instead, it determines whether to isolate the computing node 20 based on settings such as the node status information reception period and the maximum waiting time. In other words, if the management node 10 does not receive the node status information reported by the computing node 20 within a certain period of time, the computing node 20 will be isolated.

[0065] In this node management method, the management node 10 performs isolation processing after a communication anomaly occurs in the computing node 20. If the communication anomaly is caused by a fault, the corresponding job may have already exited when the management node 10 performs the isolation processing, resulting in a waste of computing resources already invested in the job. In view of this, the embodiments of the present application provide a node management method and related equipment that can isolate computing nodes 20 with a high failure risk in advance, effectively avoiding the situation where the server cluster's jobs are exited due to the failure of the computing node 20, reducing the losses caused by the failure of the computing node 20, and avoiding the waste of computing resources.

[0066] The node management method provided in an embodiment of the present application is applied to a server cluster as shown in Figure 2. Please refer to Figure 2. Figure 2 is a structural schematic diagram of a server cluster provided in an embodiment of the present application. The server cluster includes a management node 10 and a computing node cluster. The computing node cluster includes multiple computing nodes 20, and the management node 10 is communicated with the multiple computing nodes 20 respectively.

[0067] Among them, the scheduler main process running on the CPU in the management node 10 includes a request processing module, a job management module, a scheduling engine, a fault isolation module, a node management module and a message sending and receiving module; the computing node 20 includes a CPU and a baseboard management controller (BMC), and the processes running on the CPU include an agent process, a monitoring process and a job process.

[0068] The request processing module is responsible for receiving job requests sent by users through the submission machine and sending the job information contained in the job request to the job management module. Specifically, users can input job information into the submission machine through a browser on a smart terminal such as a PC, server, or tablet, thereby controlling the submission machine to send the job request to the management node 10. Alternatively, users can directly operate the submission machine to send the job request to the management node 10 via a command line.

[0069] Specifically, the submission machine can be a server, or a smart terminal such as a PC, a laptop or a tablet computer.

[0070] The job management module is used to request the scheduling engine to allocate a computing node 20 based on the job information, and send job messages to the corresponding computing node 20 through the message transceiver module according to the scheduling results of the scheduling engine; it is also used to receive job recovery requests or job status update requests returned by the message transceiver module, and manage and monitor the status of each job in the corresponding computing node 20.

[0071] The scheduling engine is used to interact with the job management module and schedule jobs to available, non-isolated compute nodes 20 in the server cluster based on real-time job requirements in the job management module. Optionally, the scheduling engine can assign jobs to compute nodes 20 with a healthy health status (i.e., compute nodes 20 in a non-isolated state), and not assign jobs to compute nodes 20 with an unhealthy health status (i.e., compute nodes 20 in an isolated state).

[0072] The message transceiver module is used to implement information transmission and reception with each computing node 20 in the server cluster, and is specifically used to receive node status information reported by the agent process of each computing node 20. The node status information includes node resource information and fault prediction information.

[0073] The node management module is used to receive the node status information of the computing node 20 returned by the message sending and receiving module, update the node resource information in the node status information to the scheduling engine, so that the scheduling engine can perform job allocation based on the real-time resource usage of each computing node 20; and send the fault prediction information in the node status information to the fault isolation module.

[0074] The fault isolation module can be embedded in the framework of the scheduler's main process as a plug-in and intervene in the management process of the management node 10 via an electronic switch. Specifically, when the fault isolation function is enabled, the fault isolation module is used to receive fault prediction information for each computing node 20 from the node management module; based on this fault prediction information, it determines whether the corresponding computing node 20 meets the isolation condition. Then, based on whether the isolation condition is met, it instructs the node management module and the scheduling engine to update the isolation or non-isolation status of each computing node 20. When the fault isolation function is disabled, the node management module and the scheduling engine directly perform node management and scheduling based on the node resource information reported by the computing node 20.

[0075] Among them, the agent process in the computing node 20 is used to receive the job message sent by the message sending and receiving module, and to pull up the corresponding job process according to the job message, as well as the monitoring process for monitoring the status of the job process; the agent process is also used to collect various node resource information used for jobs in the computing node 20, as well as the fault prediction information calculated by the BMC, and then generate the node status information of the computing node 20 using the node resource information and fault prediction information and send it to the message sending and receiving module.

[0076] The BMC is used to collect fault information of one or more components in the computing node 20 that may cause the job to exit, and input the fault information into a trained fault prediction neural network model for calculation to obtain fault prediction information indicating the level of failure risk of the computing node 20.

[0077] It is understandable that more processes executed by the management node 10 and the computing node 20 in the server cluster shown in FIG. 2 will be described in the node management method section below and will not be repeated here.

[0078] It is understandable that, in order to implement the above functions, the management node 10 and the computing node 20 include at least one of the hardware structures and software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the present application can be implemented in the form of hardware, or in the form of a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0079] The embodiment of the present application can divide the scheduler main process into functional units according to the above functional description. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of software functional units. It should be noted that the division of units in the embodiment of the present application is schematic and is only a logical functional division. In actual implementation, there may be other division methods.

[0080] The node management method and related devices provided in the embodiments of the present application can be applied in different scenarios. The application scenarios of the embodiments of the present application are exemplarily described below.

[0081] In one possible implementation, the node management method and device provided in the embodiments of the present application can be applied to a distributed server cluster, where the management node 10 and the computing node 20 can be deployed in different locations. As shown in Figure 3a, the management node 10 and the computing node 20 can communicate with each other via a communication network composed of one or more communication devices.

[0082] The communication device may be a router or a switch. Exemplarily, the communication device may be a layer 2 switch or a layer 3 switch.

[0083] In another possible implementation, the node management method provided in the embodiments of the present application can be applied to a server cabinet or a multi-node server, and the management node 10 and the computing node 20 can be located in the same cabinet. As shown in Figure 3b, the embodiments of the present application can be applied to a server cabinet 40, which includes a management node 10, a computing node 20, and a cabinet 42. The management node 10 and the computing node 20 are located together in the cabinet 42, and the management node 10 and the computing node 20 can be communicatively connected via a cable backplane or bus of the cabinet 42.

[0084] Please refer to Figures 4 and 5, which are schematic diagrams of the structures of the management node and the computing node provided in the embodiments of the present application.

[0085] As shown in FIG. 4 , the management node 10 provided in an embodiment of the present application includes a processor 110 , a memory 120 , and a communication interface 130 .

[0086] Among them, the processor 110 may include one or more processing cores. The processor 110 uses various interfaces and lines to connect the various parts within the management node 10, and executes the part implemented by the management node in the node management method provided in the embodiment of the present application by running or executing instructions, programs, code sets or instruction sets stored in the memory 320, and calling data stored in the memory 320. Optionally, the processor 110 can be implemented in the form of at least one hardware of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 310 can integrate one or more combinations of a CPU, a graphics processing unit (GPU), and a modem. It is understandable that the above-mentioned modem may not be integrated into the processor 310, but may be implemented separately through a communication chip.

[0087] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 may include a program storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the method of the embodiment of the present application, etc.

[0088] The communication interface 130 is used to communicate with the computing node 20 or other management nodes 10 .

[0089] The processor 110, memory 120, and communication interface 130 are communicatively connected via a bus within the management node 10. This bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This bus can be categorized as an address bus, a data bus, a control bus, etc. For ease of illustration, FIG4 shows only one thick line, but this does not imply that there is only one bus or only one type of bus.

[0090] As shown in FIG5 , the computing node 20 provided in the embodiment of the present application includes a processor 210 , a memory 220 , and a communication interface 230 . Optionally, the computing node 20 further includes a management device 240 .

[0091] The types and functions of the processor 210 , memory 220 , and communication interface 230 are similar to those of the processor 110 , memory 120 , and communication interface 130 in the management node 10 shown in FIG. 4 , and are not described in detail here.

[0092] The management device 240 may be a management unit of a non-business module, and may also be referred to as an out-of-band management device 240. For example, the management device 240 may remotely maintain and manage the computing node 20 through a dedicated data channel; the management device 240 is completely independent of the operating system of the computing node 20.

[0093] The management device 240 can specifically be a management unit for the running status of the computing node 20, a management unit built into the processor 210, a management system in a management chip outside the processor, a BMC, a system management module (SMM), a management unit built into a business unit, or a combination of one or more management units such as a device management system in an operating system. The embodiments of the present application do not limit the specific form of the management device, which is only illustrative here. The management device 240 is used to execute the portion of the node management method provided by the embodiment of the present application together with the processor 210 implemented by the computing node.

[0094] In addition, those skilled in the art will understand that the structure of the management node or computing node shown in the above figures does not constitute a limitation. In actual applications, the management node or computing node may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here.

[0095] The various related devices provided in the embodiments of the present application are described above. The node management method provided in the embodiments of the present application will be described below with reference to Figures 6 to 10.

[0096] Please refer to Figures 6 and 7, which illustrate the node management method provided by the embodiment of the present application from the perspectives of the computing nodes and management nodes in the server cluster, respectively. Specifically, Figure 6 is a flow chart of a node management method provided by the embodiment of the present application. As shown in Figure 6, the method is applied to the computing nodes in the server cluster and specifically includes steps S101 to S103.

[0097] S101: A computing node obtains fault information of the computing node itself.

[0098] Among them, after the computing node is powered on, the computing node can execute preset program codes or instructions, communicate with the management node, and thus join the server cluster; after joining the server cluster, the computing node can periodically collect its own node status information and report it to the management node.

[0099] The node status information includes fault prediction information. The computing node can collect fault information about one or more of its components and then determine the fault prediction information based on the fault information. It is understood that the computing node collects information about components that may cause a job to exit if a fault occurs.

[0100] Optionally, the fault information includes fault information of the memory of the computing device.

[0101] Understandably, as memory technology upgrades—process size reduction, higher frequencies, and larger capacities—memory failure rates are also increasing. According to incomplete statistics, 61% of server downtimes are caused by memory issues, with memory module failures accounting for 83% of non-sudden downtimes. Among servers experiencing correctable memory errors (CE), the probability of uncorrectable errors (UCE) is 5.2 times higher. Therefore, memory failure status can serve as an important indicator for predicting the failure risk level of compute nodes.

[0102] Optionally, the fault information also includes fault information of one or more components of the computing node: a processor, a power supply, a hard disk, and a network card.

[0103] Optionally, the fault information also includes software fault information and network fault information. Exemplarily, the fault information includes operating system fault information and driver software fault information.

[0104] The fault information includes historical fault information. For example, the fault information includes corresponding information of faults that occurred in the past cycle, or corresponding information of faults that occurred after the computing node was powered on, or corresponding information of all faults that have occurred and are stored in the computing node.

[0105] This fault information includes specific information about the fault, such as fault type, duration, number of recurrences, and system status at the time of the fault. It also includes information about the fault's repair, such as the compute node's self-healing or maintenance personnel's response. It's understood that for common component faults, the compute node can self-heal by performing methods such as local power-on / off, isolation, and adjusting driver parameters.

[0106] Taking memory as an example, memory CE information can specifically include CE status, CE occurrence time, CE error count, CE physical address information, memory patrol error count, memory patrol error row address, memory patrol error column address, and the row address with the most memory patrol errors. The following description uses memory fault information as an example.

[0107] For hardware fault information, the management device of the computing node is used to monitor the operating status of each component in the computing node; when one or more components fail, the management device can collect and store the corresponding hardware fault information.

[0108] For software and network fault information, the operating system running on the processor of the computing node can collect corresponding fault information when these faults occur, and store it in the memory of the computing node or the storage node.

[0109] An agent process runs on the CPU of the computing node, and the agent process can obtain hardware fault information through the out-of-band management interface of the management device; and obtain software and network fault information from the memory or storage node.

[0110] S102: The computing node determines fault prediction information according to the fault information.

[0111] The fault prediction information is used to indicate the risk of a computing node failing. It is understood that the risk of a computing node failing here refers to the risk of a job exiting due to a computing node failure.

[0112] Specifically, the fault prediction information can be a classification of the compute node's current fault risk level, such as low, medium, or high; it can also be the compute node's current health or health score; or it can be the probability of a job exiting due to one or more hardware faults, such as the probability of a job exiting due to a memory fault or a network card fault. It is understood that when the fault prediction information is a probability, the magnitude of the fault prediction information is in the range [0, 1].

[0113] In a possible implementation, the computing node inputs the fault information into the fault prediction neural network model to obtain the fault prediction information output by the fault prediction neural network model.

[0114] The fault prediction neural network model is a pre-trained neural network model.

[0115] Optionally, the fault prediction neural network model is software running in the management device. The management device can directly input the collected fault information into the fault prediction neural network model for calculation to obtain the fault prediction information output by the fault prediction neural network model without occupying the computing resources of the processor of the computing node.

[0116] Exemplarily, the fault prediction neural network model can combine the currently input CE information and historical CE information to determine in turn whether the current CE information meets the conditions of a certain fault characteristic pattern, and generate a fault characteristic pattern code for the current computing node. The fault pattern code is used to indicate which fault characteristic pattern conditions the current computing node meets; and based on one or more fault characteristic pattern codes, a machine learning algorithm is used to predict the characteristic pattern of the computing node failure, as well as the probability of each fault characteristic pattern causing the job to exit, and the fault risk level or health of the computing node is determined based on the probability of each fault characteristic pattern causing the job to exit.

[0117] Exemplarily, the aforementioned machine learning algorithms may include threshold-based decision algorithms, decision tree algorithms, supervised machine learning algorithms, unsupervised machine learning algorithms, memory pin link detection algorithms, and the like. For example, the management module may determine the fault characteristic pattern of a computing node based on a decision tree algorithm, a random forest algorithm, or a neural network algorithm. The embodiments of this application do not limit the specific characteristic pattern of the machine learning algorithm used to determine the fault characteristic pattern; this is provided for illustrative purposes only.

[0118] In another possible implementation, a correspondence between fault information and fault prediction information is preset in the computing node; when the fault information collected by the computing node in step S101 can be classified as one of the preset fault information, the computing node can obtain the corresponding fault prediction information.

[0119] For example, when the fault information includes that the number of CEs of the memory is less than 10 times and the number of UCEs is less than 2 times, the computing node can obtain the corresponding fault prediction information: the fault risk level of the job exit caused by the memory failure is low; when the fault information includes that the number of CEs of the memory is greater than 20 times and the number of UCEs is greater than 10 times, the computing node can obtain the corresponding fault risk level as high.

[0120] S103: The computing node sends fault prediction information to the management node.

[0121] After determining the fault prediction information, the computing node may send the fault prediction information to the management node.

[0122] After the agent process of the computing node obtains the fault prediction information from the management device, it may send the fault prediction information to the message transceiver module of the management node.

[0123] Optionally, the computing node may periodically perform fault prediction and send corresponding fault prediction information to the management node.

[0124] Optionally, the computing node may periodically obtain its own node resource information and fault prediction information, and send the fault prediction information and node resource information together as the computing node status information to the management node.

[0125] The node resource information includes the usage information of the processor resources, memory resources, and other hardware resources of the computing node.

[0126] The fault prediction information sent by each computing node in the server cluster to the management node is in the same format. For example, the fault prediction information is a health value or a failure probability or a high, medium, or low risk level based on a unified evaluation scale (e.g., a percentage or a ten-point scale).

[0127] In the embodiment of the present application, the collection of fault information and the calculation of fault prediction information are performed through the management device of the computing node itself, which is more convenient and direct, and does not require additional computing equipment, thus saving costs.

[0128] Please refer to Figure 7, which is a flow chart of another node management method provided in an embodiment of the present application. As shown in Figure 7, the method is applied to a management node in a server cluster, and specifically includes steps S104 to S107.

[0129] S104: The management node obtains fault prediction information of the computing node.

[0130] Among them, the management node needs to obtain the fault prediction information of the computing node to determine the failure risk of the computing node, so as to isolate the computing node with high failure risk in advance to avoid the job exit caused by computing node failure, resulting in waste of computing resources.

[0131] In a possible implementation, the computing nodes in the server cluster execute steps S101 to S103 in the embodiment shown in FIG6 , and the management node can receive the fault prediction information sent by each computing node.

[0132] In another possible implementation, the computing node does not execute steps S101 to S103 in the embodiment shown in FIG6 , and the management node may obtain the fault prediction information of the computing node from a computing device running management platform software.

[0133] The computing device running the management platform software is connected to the management node and each computing node respectively. The management platform software is used to obtain fault information from each computing node and determine fault prediction information based on the fault information.

[0134] Optionally, the management platform software may input the fault information into a trained fault prediction neural network model to obtain the fault prediction information output by the fault prediction neural network model.

[0135] After determining the fault prediction information, the computing device running the management platform software may proactively send the fault prediction information to the management node, or may send the fault prediction information to the management node in response to an acquisition request from the management node.

[0136] Optionally, the management platform software may run on a management device, that is, the management device obtains fault information of the computing node, and then determines fault prediction information based on the fault information; and then performs fault isolation operations based on the fault prediction information.

[0137] In an embodiment of the present application, when the workload is large and the CPU resources of the computing node are insufficient, a computing device can be set up to run the management platform software, and the computing device can calculate the fault prediction information based on the fault information to save the CPU resources of the computing node.

[0138] S105: The management node determines whether the computing node meets the isolation condition based on the fault prediction information.

[0139] The management node may receive fault prediction information sent by each computing node in the computing node cluster, and process the fault prediction information of the computing nodes one by one to determine whether each computing node meets the isolation condition.

[0140] The isolation condition may be a failure risk condition of a computing node indicated by a single failure prediction information, or a failure risk change trend of a computing node indicated by multiple failure prediction information.

[0141] In a possible implementation, the management node may determine whether the computing node meets the isolation condition based on the fault prediction information currently received.

[0142] Exemplarily, the fault prediction information includes a fault risk level; when the fault risk level of a computing node is high, the management node may directly determine that the computing node meets the isolation condition.

[0143] Exemplarily, the fault prediction information includes a health score or a fault probability; when the failure probability of a computing node is greater than or equal to a first threshold, or the health score is lower than or equal to a second threshold, the management node can directly determine that the computing node meets the isolation condition.

[0144] In one possible implementation, the management node obtains fault prediction information from a computing node multiple times and then determines whether the computing node meets the isolation condition based on the fault prediction information obtained multiple times. In other words, for a computing node, the management node comprehensively determines whether the computing node meets the isolation condition based on the fault prediction information sent by the computing node multiple times.

[0145] Wherein, the multiple times is a preset number of times. Exemplarily, the management node may obtain fault prediction information of the same computing node three, five, or ten times, and determine whether the computing node meets the isolation condition based on the fault prediction information obtained three, five, or ten times.

[0146] In one possible implementation, the management node can periodically determine whether the computing node meets the isolation condition; specifically, for a computing node, the management node can obtain fault prediction information of the computing node multiple times within a cycle, and then determine whether the computing node meets the isolation condition at the end of the cycle based on all fault prediction information of the computing node obtained within the cycle.

[0147] For example, the management node can determine whether the computing node meets the isolation condition once every hour and obtain the fault prediction information of the computing node once every 5 minutes; for a computing node, the management node determines whether the computing node meets the isolation condition every hour based on all the fault prediction information of the computing node obtained in the past hour.

[0148] Optionally, the management node may periodically receive the fault prediction information and determine whether the computing node meets the isolation condition based on multiple periods of fault prediction information. For example, the management node may use a period between 5 and 20 minutes as the period length, and obtain fault prediction information once every period.

[0149] Optionally, the management node may determine whether the computing node meets the isolation condition based on fault prediction information received in the most recent M cycles, where M is a positive integer greater than 1.

[0150] Exemplarily, the fault prediction information includes a fault risk level; when the fault risk level of a computing node is high for M consecutive cycles, the management node may directly determine that the computing node meets the isolation condition.

[0151] Exemplarily, the fault prediction information includes a fault risk level; when a computing node has a high fault risk level for N cycles accumulated in the last M cycles, the management node can directly determine that the computing node meets the isolation condition, where N is a positive integer less than M and greater than 0.

[0152] Exemplarily, the fault prediction information includes a fault probability; when the fault probability of a computing node is greater than or equal to a first threshold for M consecutive cycles, the management node may determine that the computing node meets the isolation condition.

[0153] Exemplarily, the fault prediction information includes a health score; when an average value of the health score of a computing node within M cycles is less than or equal to a second threshold, the management node may determine that the computing node meets the isolation condition.

[0154] In one possible implementation, after each time fault prediction information is obtained, if the fault prediction information indicates that the fault risk level of the computing node is high, the management node increases the risk cycle count of the computing node by one; after multiple times of obtaining fault prediction information, if the risk cycle count reaches a preset maximum value, it is determined that the computing node meets the isolation condition; if the risk cycle count is less than the preset maximum value, it is determined that the computing node does not meet the isolation condition.

[0155] Among them, the management node obtains the fault prediction information multiple times in succession, and adjusts the risk cycle count of the computing node according to the fault prediction information each time after obtaining the fault prediction information; then, after obtaining the fault prediction information for a preset number of times in succession and adjusting the risk cycle count accordingly, it is determined whether the computing node meets the isolation condition according to the risk cycle count.

[0156] Optionally, if the fault prediction information indicates that the fault risk level of the computing node is not high, the management node clears the risk cycle count of the computing node.

[0157] In the embodiment of the present application, by clearing the risk cycle count of the computing node when the failure risk level of the computing node is not high, the computing node can perform operations in more time, thereby speeding up the operation progress.

[0158] Optionally, if the fault prediction information indicates that the fault risk level of the computing node decreases after the fault is repaired, and the fault risk level is not high, the risk cycle count of the computing node is reset to zero.

[0159] Compute nodes can sense the status of internal fault repairs. For example, maintenance personnel manually replace hardware such as memory and hard drives; the compute nodes themselves use self-healing algorithms to softly isolate faulty memory particles.

[0160] When a computing node senses that the hardware related to the fault prediction information has been repaired, the computing node will add a corresponding indication to the fault prediction information the next time it determines the fault prediction information, to inform the management node that the main reason for the change in the fault prediction information is that the hardware has been repaired, rather than a decrease in the fault risk level caused by algorithm fluctuations or occasional working conditions.

[0161] When the fault prediction information received by the management node includes this indication and the fault risk level is not high, the management node can determine that the computing node is in a relatively healthy state, and therefore the risk cycle count of the computing node can be cleared to zero; when the fault prediction information received by the management node does not include this indication and the fault risk level is not high, the management node can believe that the decline in the fault risk level may be caused by algorithm fluctuations or occasional working conditions, and at this time the management node can reduce the risk cycle count by one.

[0162] In an embodiment of the present application, the risk cycle count is cleared only when it is determined that the fault risk level of the computing node has decreased due to fault repair. A stricter clearing condition is set, which can make it easier for the computing node to meet the isolation conditions and improve the safety of the operation.

[0163] In one possible implementation, the fault prediction information includes the probability that a memory failure occurs in the computing node, causing the job to exit; after obtaining the probability each time, if the probability is greater than or equal to a first threshold, the fault risk level of the computing node is determined to be high; if the probability is less than the first threshold, the fault risk level of the computing node is determined to be not high.

[0164] S106. If the computing node meets the isolation condition, the management node isolates the computing node.

[0165] When a computing node meets the isolation condition, the management node can determine that the computing node is in an unhealthy state. Optionally, the fault isolation module in the management node can update the unhealthy state of the computing node to the node management module and the scheduling engine.

[0166] S107: When the computing node is in an isolated state, the management node stops allocating jobs to the computing node and / or allocates the ongoing jobs of the computing node to other computing nodes in the server cluster that are not in an isolated state.

[0167] The management node will not assign jobs to compute nodes that are in an unhealthy state.

[0168] In a possible implementation, after isolating a computing node, the management node receives a new job request; and then sends a corresponding job message to a computing node that has not been isolated according to the job request.

[0169] In one possible implementation, when a job is in progress when a computing node is isolated, the management node can wait for the completion result of the job. If the job is successful, the job result is obtained and returned to the user. If the job fails, the job is reallocated to a computing node in a non-isolated state in the server cluster.

[0170] In another possible implementation, the job performed by the isolated computing node can be transferred to a non-isolated computing node for continued operation through checkpoint / restart (C / R) technology. The embodiment of the present application does not limit the processing method of the job in the isolated computing node.

[0171] In an embodiment of the present application, by judging whether the computing node meets the isolation conditions based on the fault prediction information of the computing node, the computing node with a high failure risk level is isolated in advance, which can effectively avoid the situation where the server cluster job is exited due to the computing node failure, thereby avoiding the waste of computing resources and reducing the losses caused by the computing node failure.

[0172] The above is an overall description of the node management method provided in the embodiment of the present application. The following will further illustrate the node management method provided in the embodiment of the present application in combination with two embodiments.

[0173] Please refer to Figure 8, which is a flow chart of a node management method provided in an embodiment of the present application. The method is applied to a management node in a server cluster. As shown in Figure 8, the method includes S201 to S209.

[0174] S201: The management node obtains node status information of the computing node.

[0175] The node status information at least includes node resource information of the computing node.

[0176] The management node is in communication connection with a plurality of computing nodes in the server cluster; the management node receives node status information of the plurality of computing nodes and processes them one by one.

[0177] After receiving the node status information, the management node may update various resource information of the computing node according to the node status information, and allocate jobs according to the node resource information of each computing node.

[0178] In this embodiment, description is made by taking the example of a management node periodically acquiring node status information.

[0179] In a possible implementation, the management node receives corresponding node status information sent by the computing node.

[0180] In another possible implementation, the management node obtains the node status information from a computing device running management platform software.

[0181] The specific implementation methods of the above two possible implementations are similar to the implementation methods in step S103 in the embodiment shown in FIG6 , and are not described again here.

[0182] It is understandable that frequently running the fault prediction neural network model to obtain fault prediction information by the management device in the computing node may affect the performance of the computing node. Therefore, the computing node can control the on / off of this fault prediction function through a preset control policy to reduce the impact on computing node performance. Therefore, the node status information at one point in time may contain fault prediction information, while the node status information at another point in time may not contain fault prediction information.

[0183] S202: The management node updates the resource information of the computing node according to the node status information.

[0184] After the message transceiver module of the management node receives the status message, the node management module may update the resource information of the computing node in the node status information to the scheduling engine.

[0185] S203: The management node determines whether the fault isolation function is enabled.

[0186] The management node can conserve its computing resources by intermittently enabling and disabling the fault isolation function. Optionally, the management node can agree with each compute node to enable and disable both the management node's fault isolation function and the compute node's fault prediction function simultaneously. Specifically, when the fault isolation function is enabled, the sampling and reporting period for the compute node's node status information is the same as the period during which the management node determines whether the compute node meets the isolation conditions based on fault prediction information.

[0187] Optionally, the management node periodically receives node status information including fault prediction information, and after receiving P cycles of fault prediction information, enables the fault isolation function and determines whether the computing node needs to be isolated based on the fault risk of the computing node within the P cycles. P is a positive integer greater than or equal to 1.

[0188] Among them, the management node can control the shutdown and startup of the fault isolation function through a physical switch or a virtual switch; specifically, the management node can trigger a first trigger signal and a second trigger signal according to the status of the physical switch or the virtual switch, the first trigger signal is used to indicate that the fault isolation function is turned off, and the second trigger signal is used to indicate that the fault isolation function is turned on.

[0189] Optionally, a fault prediction scheduling switch fi is set in the management node; when fi=1, it indicates that the fault isolation function is turned on, and when fi=0, it indicates that the fault isolation function is turned off.

[0190] Among them, the management node can execute the above optional method to control the opening and closing of the fault prediction scheduling switch fi based on time; it can also control the opening and closing of the fault prediction scheduling switch fi based on the workload, resource usage or job importance in the server cluster.

[0191] For example, if the server cluster is about to execute a job with the highest importance level, the management node can turn on the fault prediction scheduling switch fi to ensure the safety of the job; and then turn off the fault prediction scheduling switch fi after the highest-level job is completed.

[0192] If the fault isolation function is disabled, the management node may execute step 209 ; if enabled, the management node may execute steps S204 to S209 .

[0193] It can be understood that when the fault isolation function is turned on, step S202 can be executed before step 203, between step 203 and step 208, or after step 208, and this embodiment of the present application does not specifically limit this.

[0194] S204: If the fault isolation function is enabled, the management node updates the fault prediction information of the computing node according to the node status information.

[0195] After the message transceiver module of the management node receives the status message, the node management module may update the fault prediction information of the computing node in the node status information to the fault isolation module, and the fault isolation module executes step S205.

[0196] S205: The management node determines whether the computing node meets the isolation condition based on the fault prediction information.

[0197] The management node may determine the state of the computing node in the current cycle through the fault prediction information of the computing node, and determine whether the computing node in this state can complete the job safely.

[0198] Among them, the isolation condition is the basis for the management node to determine whether the computing node can complete the job safely.

[0199] For details, please refer to Figure 9, which is a schematic diagram of a process for determining whether a computing node meets the isolation condition according to an embodiment of the present application. As shown in Figure 9, the process includes steps S301 to S306. The fault prediction information includes the fault probability of the computing node.

[0200] S301: The management node determines whether the failure probability is greater than a first threshold.

[0201] After receiving the fault prediction information, the management node may determine whether the fault probability is greater than or equal to a first threshold.

[0202] The management node may preset a corresponding threshold value corresponding to the form of the fault prediction information reported by the computing node. In this embodiment, the fault prediction information reported by the computing node includes the fault probability, and the management node may preset a first threshold value. When the value of the fault prediction information is greater than the first threshold value, the management node may determine that the fault risk level of the computing node in the current cycle is high.

[0203] By unifying the form in which computing nodes report fault prediction information, it is possible to embed the fault prediction mechanism into the scheduler main process while decoupling the scheduling process and the underlying fault prediction method of the computing nodes. In this way, when the server cluster includes computing nodes from different manufacturers and these computing nodes use their own fault prediction neural network models, the management node can perform unified fault isolation management on these computing nodes.

[0204] If the failure probability is greater than the first threshold, the management node executes step 302 ; if the failure probability is less than the first threshold, the management node executes step 303 .

[0205] S302: If the failure probability is greater than or equal to the first threshold, the management node increases the risk cycle count of the computing node by one.

[0206] The management node is provided with a risk cycle counter of each computing node. Whenever the received failure probability is greater than or equal to a first threshold, the risk cycle counter of the corresponding computing node is incremented by one.

[0207] In this embodiment, for a computing node, if the failure probability of the computing node obtained by the management node at one time is greater than or equal to the first threshold, the risk cycle count of the computing node is increased by one.

[0208] After adjusting the risk cycle count, the management node may execute step S304.

[0209] S303: If the failure probability is less than the first threshold, the management node reduces the risk cycle count of the computing node.

[0210] Optionally, whenever the received failure probability is less than a first threshold, the management node reduces the risk cycle count of the computing node by one.

[0211] Optionally, whenever the received failure probability is less than a first threshold, the management node clears the risk cycle count of the computing node.

[0212] In one possible implementation, the above two optional methods can also be combined: when the fault prediction information also indicates that the fault probability is the fault probability determined after the computing node is repaired, and the fault probability is less than the first threshold, the management node can clear the risk cycle count of the computing node; when the fault prediction information also indicates that the fault probability is not the failure probability determined after the computing node is repaired, and the failure probability is less than the first threshold, the management node clears the risk cycle count of the computing node.

[0213] After self-repair technology or manual repair, the failure probability of the computing node may be less than the first threshold. At this time, the management node can clear the risk cycle count to zero and determine that the computing node no longer meets the isolation conditions. If self-repair technology or manual repair has not been used, even if the failure probability is less than the first threshold, the management node cannot directly determine that the computing node is in a healthy state. At this time, the management node decrements the risk cycle count by one.

[0214] After adjusting the risk cycle count, the management node may execute step S304.

[0215] S304: The management node determines whether the risk cycle count of the computing node is greater than or equal to a preset maximum value.

[0216] When the risk cycle count of the computing node changes, the management node may further determine whether the risk cycle count is greater than or equal to a preset maximum value.

[0217] When a computing node fails, the computing node can use its self-healing function to repair the CE that occurs, or through timely maintenance and adjustment by maintenance personnel, the computing node can eliminate the failure risk before the failure occurs. Therefore, the management node can preset the maximum value of the risk cycle count based on historical data and actual experience, which not only gives the computing node time to adjust, but also prevents the self-healing time from being delayed too long, causing CE to evolve into UCE, and the computing node to crash and cause the job to exit.

[0218] If the risk cycle count is greater than or equal to the preset maximum value, the management node may execute step S305; if the risk cycle count is less than the preset maximum value, the management node may execute step S306.

[0219] S305: The management node determines that the computing node meets the isolation condition.

[0220] When the risk cycle count is greater than or equal to a preset maximum value, the management node may determine that the corresponding compute node has a high risk of job exit and cannot complete the job safely, so the job should not be assigned to it. At this point, the management node may determine that the compute node meets the isolation conditions and perform an isolation operation on the compute node.

[0221] Specifically, the management node changes or maintains the state of the computing node to an unhealthy state.

[0222] S306: The management node determines that the computing node does not meet the isolation condition.

[0223] When the risk cycle count is less than a preset maximum, the management node may determine that the compute node's job exit risk is low and the job can be completed safely, thereby determining that the compute node does not meet the isolation conditions. At this point, the management node may change or maintain the compute node's status to healthy.

[0224] In the embodiment of the present application, different risk cycle counting methods can be used to make it easier or more difficult for computing nodes to meet isolation conditions, thereby flexibly balancing the security and speed of computing node operations.

[0225] S206: If the computing node meets the isolation condition, the management node isolates the computing node.

[0226] The specific implementation of step S206 in this embodiment is similar to the specific implementation of step S106 in the embodiment shown in FIG6 , and will not be repeated here.

[0227] S207: If the computing node does not meet the isolation condition, the management node determines whether the computing node is in an isolated state.

[0228] When the computing node does not meet the isolation conditions, the management node can assign a job to the computing node. However, before assigning a job, the management node needs to first determine whether the computing node is in an isolated state to avoid incorrect assignment or failure to assign a job to the computing node.

[0229] Specifically, if the computing node is in an isolated state, the management node may execute step S208 to release the isolation state of the computing node so as to allocate jobs to the computing node; if the computing node is not in an isolated state, the management node may directly execute step S209.

[0230] S208: If the computing node is in an isolated state, the management node releases the isolated state of the computing node.

[0231] The node management module of the management node can obtain the isolation status of the computing node from itself. Optionally, the node management module can obtain the health status of the computing node based on the identifier of the computing node; if the health status is unhealthy, it can be determined that the computing node is currently in an isolated state; if the health status is healthy, it can be determined that the computing node is in a non-isolated state.

[0232] If the computing node is in an isolated state, the node management module may release the isolated state of the computing node. Optionally, the node management module may update the scheduling engine with information that the health state of the computing node is healthy, so that the scheduling engine can assign a job to the computing node.

[0233] S209: The management node determines whether there is any unprocessed node status information.

[0234] When the management node completes step S206 or S208, or determines in step S207 that the computing node is not in an isolated state, the management node can determine that the current node status information has been processed. At this point, the management node can further determine whether there is any unprocessed node status information. If so, step S202 is executed to enter a new round of node status information processing; if not, step S201 is executed to wait for receiving new node status information.

[0235] In an embodiment of the present application, the form of fault prediction information reported by the computing nodes can be unified, and the fault prediction mechanism can be embedded in the scheduler main process while decoupling the scheduling process and the underlying fault prediction method of the computing node. In this way, when the computing nodes of the server cluster adopt different fault prediction neural network models, the management node can directly know the failure risk of the computing node, which makes it more convenient for the management node to uniformly manage computing nodes of different manufacturers and models.

[0236] Please refer to Figure 10, which is a flow chart of a node management method provided in an embodiment of the present application. The method is applied to computing nodes in a server cluster; as shown in Figure 10, the method includes S401 to S406.

[0237] S401: The computing node collects its own node resource information.

[0238] The agent process running on the computing node can periodically collect its own node resource information and report it to the management node, so that the management node can allocate jobs according to the resource usage of the computing node.

[0239] S402: The computing node determines whether the fault prediction function is enabled.

[0240] The fault prediction function based on the fault prediction neural network model consumes a lot of computing resources. Frequently enabling and using this function will affect the performance of the computing node. Therefore, the computing node can enable this function intermittently.

[0241] S403: If the fault prediction function is enabled, the computing node collects fault information of the computing node itself.

[0242] Among them, when the fault prediction function is turned on, the computing node will calculate its own fault prediction information based on the fault prediction neural network model, and report the fault prediction information as its own node status information to the management node.

[0243] Before obtaining fault prediction information, the computing node needs to first collect its own fault information as input to the fault prediction neural network model. Optionally, the computing node can periodically collect its own fault information.

[0244] The fault information may include various fault information of components that may cause the computing node to crash or other components that may cause the running job to exit. For example, the fault information includes fault information of the computing node's memory, hard disk, processor, and network card.

[0245] The fault information may include fault information that occurred in the latest preset period of the corresponding component, may include fault information that occurred after the last power-on, and may also include all fault information that has occurred in these components.

[0246] S404: The computing node inputs the fault information into a fault prediction neural network model to obtain fault prediction information output by the model.

[0247] The fault prediction neural network model is a pre-trained neural network model. Optionally, the fault prediction neural network model is deployed in a management device of a computing node.

[0248] S405: The computing node generates its own node status information.

[0249] When the fault prediction function is turned off, the computing node can generate the node status information based on the collected node resource information; when the fault prediction function is turned on, the computing node can generate the node status information based on the collected node resource information and the fault prediction information.

[0250] S406: The computing node sends the node status information to the management node.

[0251] Optionally, the computing node can periodically send node status information to the management node. This node status information includes node resource information and fault prediction information for the same period. The periods for collecting node resource information, collecting fault information, and sending node status information can be the same or different. When the periods for these three actions are different, the period for sending node status information can be greater than the periods for collecting node resource information and fault information.

[0252] Specifically, the node status information includes node resource information and fault prediction information acquired by the computing node during a most recent cycle of sending status information.

[0253] After sending the node status information to the management node, the computing node may return to execute step S401.

[0254] In an embodiment of the present application, by performing fault prediction on the computing node itself and sending the corresponding fault prediction information to the management node, the management node can directly know the level of fault risk of the computing node. Preliminary isolation management can be performed based on the failure risk of the computing node without the need for the management node to perform fault prediction, which makes it more convenient for the management node to uniformly manage computing nodes of different manufacturers and models.

[0255] In conjunction with the node management method provided in the embodiments of the present application, the present application further provides a node management device 1000. For details, please refer to FIG11 , which is a schematic diagram of the structure of the node management device 1000. The node management device 1000 may specifically include:

[0256] The message sending and receiving module 1001 is used to obtain fault prediction information of the computing node, and the fault prediction information is used to indicate the risk of failure of the computing node; the fault isolation module 1002 is used to determine whether the computing node meets the isolation condition based on the fault prediction information; the node management module 1003 is used to isolate the computing node if the computing node meets the isolation condition; the scheduling engine 1004 is used to stop allocating jobs to the computing node and / or allocate the ongoing jobs of the computing node to other computing nodes in a non-isolated state when the computing node is in an isolated state.

[0257] In a possible implementation, the message transceiver module 1001 is specifically used to obtain the fault prediction information multiple times; the fault isolation module 1002 is specifically used to determine whether the computing node meets the isolation condition based on the fault prediction information obtained multiple times.

[0258] In one possible implementation, the fault isolation module 1002 is specifically used to add one to the risk cycle count of the computing node after each acquisition of the fault prediction information, if the fault prediction information indicates that the fault risk level of the computing node is high; after obtaining the fault prediction information multiple times, if the risk cycle count is greater than or equal to a preset maximum value, it is determined that the computing node meets the isolation condition; if the risk cycle count is less than the preset maximum value, it is determined that the computing node does not meet the isolation condition.

[0259] In one possible implementation, the fault isolation module 1002 is specifically used to clear the risk cycle count of the computing node after each acquisition of the fault prediction information, if the fault prediction information indicates that the fault risk level of the computing node is not high; after obtaining the fault prediction information multiple times, if the risk cycle count is greater than or equal to a preset maximum value, it is determined that the computing node meets the isolation condition; if the risk cycle count is less than the preset maximum value, it is determined that the computing node does not meet the isolation condition.

[0260] In one possible implementation, the fault isolation module 1002 is specifically used to clear the risk cycle count of the computing node after each acquisition of the fault prediction information, if the fault prediction information indicates that the fault risk level of the computing node decreases after the fault is repaired, and the fault risk level is not high; after obtaining the fault prediction information multiple times, if the risk cycle count is greater than or equal to a preset maximum value, it is determined that the computing node meets the isolation condition; if the risk cycle count is less than the preset maximum value, it is determined that the computing node does not meet the isolation condition.

[0261] In a possible implementation, the fault isolation module 1002 is specifically configured to determine whether the fault isolation function is enabled; if enabled, triggering the step of determining whether the computing node meets the isolation condition based on the fault prediction information.

[0262] In one possible implementation, the fault prediction information includes the probability that a memory failure occurs in the computing node, resulting in job exit; the fault isolation module 1002 is specifically used to determine that the fault risk level of the computing node is high when the probability is greater than or equal to a first threshold; and to determine that the fault risk level of the computing node is not high when the probability is less than the first threshold.

[0263] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of this application.

[0264] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0265] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0266] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0267] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0268] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. A node management method, characterized in that, Applied to a management node in a server cluster, the server cluster further including computing nodes; the method further includes: Obtaining failure prediction information of the computing nodes, the failure prediction information being used to indicate the level of risk of the computing nodes having failures; Determining, according to the failure prediction information, whether the computing nodes meet isolation conditions; If the computing nodes meet the isolation conditions, isolating the computing nodes; when the computing nodes are in an isolated state, stopping assigning jobs to the computing nodes and / or assigning the jobs that the computing nodes are currently performing to other non-isolated computing nodes in the server cluster.

2. The method according to claim 1, wherein The obtaining the failure prediction information of the computing nodes includes: Obtaining the failure prediction information multiple times; The determining, according to the failure prediction information, whether the computing nodes meet isolation conditions includes: Determining, according to the failure prediction information obtained multiple times, whether the computing nodes meet the isolation conditions.

3. The method according to claim 2, wherein The determining, according to the failure prediction information obtained multiple times, whether the computing nodes meet the isolation conditions includes: After each obtaining of the failure prediction information, if the failure prediction information indicates that the failure risk level of the computing node is high, incrementing the risk cycle count of the computing node by one; After the failure prediction information has been obtained multiple times, if the risk cycle count is greater than or equal to a preset maximum value, determining that the computing nodes meet the isolation conditions; If the risk cycle count is less than the preset maximum value, determining that the computing nodes do not meet the isolation conditions.

4. The method according to claim 2 or 3, characterized in that, The determining, according to the failure prediction information obtained multiple times, whether the computing nodes meet the isolation conditions includes: After each obtaining of the failure prediction information, if the failure prediction information does not indicate that the failure risk level of the computing node is high, clearing the risk cycle count of the computing node; After the failure prediction information has been obtained multiple times, if the risk cycle count is greater than or equal to a preset maximum value, determining that the computing nodes meet the isolation conditions; If the risk cycle count is less than the preset maximum value, determining that the computing nodes do not meet the isolation conditions.

5. The method according to claim 2 or 3, characterized in that The determining, according to the failure prediction information obtained multiple times, whether the computing nodes meet the isolation conditions includes: After each obtaining of the failure prediction information, if the failure prediction information indicates that the failure risk level of the computing node has decreased after a failure repair and the failure risk level is not high, clearing the risk cycle count of the computing node; After the failure prediction information has been obtained multiple times, if the risk cycle count is greater than or equal to a preset maximum value, determining that the computing nodes meet the isolation conditions; If the risk cycle count is less than the preset maximum value, determining that the computing nodes do not meet the isolation conditions.

6. The method according to any one of claims 3-5, characterized in that, The failure prediction information includes the probability that a memory failure occurs in the computing node, resulting in a job exit; After each obtaining of the failure prediction information, the method further includes: If the probability is greater than or equal to the first threshold, determine that the failure risk level of the computing node is high; If the probability is less than the first threshold, determine that the failure risk level of the computing node is not high.

7. A node management method, characterized in that, Applied to a computing node in a server cluster, the server cluster further includes a management node; the method further includes: Obtain the failure information of the computing node itself; Determine failure prediction information according to the failure information, where the failure prediction information is used to indicate the level of risk of the computing node having a failure; Send the failure prediction information to the management node.

8. The method according to claim 7, wherein The obtaining the failure information of the computing node itself includes: Periodically obtain the failure information; Before sending the failure prediction information to the management node, the method further includes: periodically obtaining the node resource information of the computing node itself; The sending the failure prediction information to the management node includes: Periodically send node status information to the management node, where the node status information includes the node resource information and the failure prediction information in the same period.

9. A management node, characterized in that, The management node includes a processor and a memory; The processor and the memory are coupled; The memory is used to store program instructions; The processor is used to run the program instructions so that the management node executes the method according to claims 1-6.

10. A computing node, characterized in that, The computing node includes a processor and a memory; The processor and the memory are coupled; The memory is used to store program instructions; The processor is used to run the program instructions so that the computing node executes the method according to claim 7 or 8.

Citation Information

Patent Citations

  • Method and system for monitoring virtual machine cluster

    CN105357038A

  • Heterogeneous cloud storage cluster fault automatic repair method, system, medium and terminal

    CN113535474A

  • BMC-based Kubernetes cluster physical node fault processing method and system

    CN114218004A

  • Container migration method and server cluster

    CN116126457A

  • Node management method and related equipment

    CN118093277A

Cited By

  • Cluster fault processing method and related equipment

    CN120880886A