Fault detection method and device for computing cluster
By combining data from cloud tenants and the cloud management platform, comprehensive in-band and out-of-band fault detection is performed, solving the problems of low efficiency and poor accuracy in computing cluster fault diagnosis. This achieves efficient and accurate fault troubleshooting and improves the execution efficiency of computing tasks and model training.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2024-10-30
- Publication Date
- 2026-05-01
AI Technical Summary
In existing technologies, the fault diagnosis efficiency of computing clusters is low and the accuracy is poor, which means that computing clusters can only run with faults in the event of a fault, and are prone to interruption again, affecting the efficiency of model training.
By combining data from cloud tenants and the cloud management platform, comprehensive in-band and out-of-band fault detection is performed. Data interaction is achieved through plugins and interfaces, and knowledge graphs and artificial intelligence models are combined to improve the accuracy and efficiency of fault detection.
It enables accurate detection of computing cluster failures, reduces the probability of further interruptions, improves the execution efficiency of computing tasks and model training, and reduces the need for manual intervention.
Smart Images

Figure CN121967269A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud technology, and in particular to a fault detection method and apparatus for computing clusters. Background Technology
[0002] Currently, generative artificial intelligence (AI-generated content, AIGC) technology is developing rapidly. AIGC models can be used for text creation, image creation, video creation, audio editing, game development, and code generation, and have a wide range of applications.
[0003] Training AIGC models typically requires the collaborative effort of large-scale computing clusters. During the training process, interruptions may occur due to various reasons, such as equipment failure or defects in the training software. Because computing cluster resources are costly, cloud tenants often have a requirement to quickly resolve faults and restore business operations. For example, they are typically required to resume the training task from the breakpoint within 10 minutes, regardless of whether the cause of the failure has been determined.
[0004] Currently, fault diagnosis of computing clusters is typically performed through manual analysis. This method is not only inefficient but also inaccurate, often failing to identify the actual cause of the fault within the timeframe allowed by the cloud tenant. However, the computing cluster must resume training from the breakpoint, meaning it can only operate with faults. But operating with faults is highly susceptible to further interruptions, leading to low model training efficiency. Summary of the Invention
[0005] This application provides a fault detection method and apparatus for computing clusters, which improves the efficiency and accuracy of fault detection in computing clusters.
[0006] Firstly, a fault detection method for a computing cluster is provided, which can be applied to a cloud management platform. The cloud management platform manages infrastructure providing cloud services, including servers within at least one cloud data center. The method includes: the cloud management platform receiving a first detection result, which indicates a fault in a first computing node, wherein the first computing node is one of M computing nodes that caused the interruption of a computing task, M is a positive integer, and the M computing nodes are deployed on at least one server. The cloud management platform determines a target detection result based on the first detection result and a second detection result, where the second detection result indicates a fault in the infrastructure, and the target detection result identifies K unavailable computing nodes out of the M computing nodes, where K is an integer.
[0007] In this method, the cloud management platform can combine a first detection result and a second detection result to determine the fault causing the computing task interruption. The first detection result indicates a fault in the first computing node, and the second detection result indicates a fault in the infrastructure. This is equivalent to combining data from the cloud tenant and the cloud management platform for comprehensive fault detection. The data relied upon for fault detection is more comprehensive, improving the accuracy of fault detection and enabling accurate fault elimination. For example, it can help accurately remove unavailable computing nodes, while continuing to execute computing tasks based on available computing nodes, reducing the probability of further interruption of computing tasks and improving the execution efficiency of computing tasks. Furthermore, this method enables automatic fault detection, eliminating the need for manual fault detection and improving fault detection efficiency.
[0008] In one possible implementation, the first detection result is determined based on the operational data of some or all of the M computing nodes, and the second detection result is determined based on the operational data of some or all of the equipment included in the infrastructure. This can be understood as the first detection result being determined based on operational data collected within the computing cluster, i.e., based on cloud tenant data, while the second detection result is determined based on the operational data of the infrastructure, i.e., based on data from the cloud management platform. This implementation effectively breaks down the data isolation barrier between cloud tenants and the cloud management platform, making the data relied upon by the cloud management platform for fault detection more comprehensive and improving the accuracy of fault detection.
[0009] In one possible implementation, the first detection result is received through a first interface between the cloud management platform and a second computing node among the M computing nodes. At least one of the M computing nodes has a first interface with the cloud management platform, and the second computing node belongs to at least one computing node. The second computing node is either the first computing node or any other computing node among the M computing nodes besides the first computing node. Based on this implementation, computing nodes within the computing cluster can interact with the cloud management platform through the first interface, breaking down the data isolation barrier between cloud tenants and the cloud management platform. This allows the cloud management platform to rely on more comprehensive data for fault detection, improving the accuracy of fault detection.
[0010] In one possible implementation, each of the at least one computing node includes an operating system, and the operating system of some or all of the at least one computing node includes a first plugin for identifying faulty computing nodes among the M computing nodes. Based on this implementation, computing nodes within the computing cluster can interact with the cloud management platform through the first plugin, breaking down the data isolation barrier between cloud tenants and the cloud management platform. This allows the cloud management platform to rely on more comprehensive data for fault detection, improving the accuracy of fault detection.
[0011] In one possible implementation, the target detection results obtained by the cloud management platform can have multiple applications. For example, the cloud management platform can provide a fault display interface to show the target detection results. Based on this implementation, the cloud management platform's operations and maintenance personnel can view the target detection results through this interface, facilitating the repair of faulty equipment. As another example, the cloud management platform can send the target detection results to some or all of the M computing nodes, allowing the computing tasks to continue on the remaining computing nodes (excluding K nodes). Based on this implementation, the computing cluster can accurately eliminate unavailable computing nodes while continuing computing tasks on available nodes, reducing the probability of further interruptions and improving the efficiency of computing task execution.
[0012] In one possible implementation, the second detection result is used to indicate multiple faults in the infrastructure. When the cloud management platform determines the target detection result based on the first and second detection results, it can do so by considering the confidence levels of the multiple faults and the first detection result. The confidence levels of the multiple faults are determined based on the correlation between the multiple faults and the first detection result. Based on this implementation, the cloud management platform can determine the confidence level of out-of-band detected faults based on the correlation between in-band and out-of-band faults, thus improving the accuracy of fault detection through mutual verification between in-band and out-of-band faults.
[0013] In one possible implementation, the cloud management platform can determine the confidence level of an out-of-band detected fault based on whether the fault detected in-band is related to the fault detected out-of-band. For example, if the first fault among multiple faults indicated by the second detection result is related to some or all of the faults indicated by the first detection result, the confidence level of the first fault is a first value. Or, if the first fault among multiple faults indicated by the second detection result is unrelated to the faults indicated by the first detection result, the confidence level of the first fault is a second value. The first value is greater than the second value. Based on this implementation, when the fault detected in-band is related to the first fault detected out-of-band, the confidence level of the determined first fault is higher, which better aligns with the actual fault determination logic. That is, when such a fault is detected both in-band and out-of-band, the probability of the fault occurring is higher. Thus, subsequent fault determination based on the confidence level can be more accurate, improving the accuracy of fault detection.
[0014] In one possible implementation, the multiple faults include faults in at least one device within the infrastructure, which includes one or more of servers, storage devices, network devices, or server racks. When the cloud management platform determines the target detection result based on the confidence levels of the multiple faults and the first detection result, the cloud management platform can determine the target detection result based on J faults from the multiple faults and the first detection result. The J faults are faults corresponding to at least one device, where at least one of the J faults corresponds to a first type of device, and the confidence level corresponding to the at least one fault is the highest confidence level among the faults corresponding to the first type of device, where K is a positive integer. That is, the cloud management platform can determine the fault with the highest confidence level for each type of device separately, and then combine the highest confidence level faults from each type of device to comprehensively determine the target detection result. Based on this implementation, the cloud management platform can filter out faults with lower confidence levels by device type, reducing the amount of data for fault detection and improving fault detection efficiency.
[0015] In one possible implementation, the cloud management platform determines the target detection result based on J faults out of multiple faults and a first detection result. This can be achieved in several ways. One implementation involves the cloud management platform determining the target detection result based on the J faults, the first detection result, and a knowledge graph. The knowledge graph represents the relationships between multiple knowledge nodes, which correspond to the faults indicated by the first detection result and the J faults. Based on this implementation, the cloud management platform can determine the target detection result based on a pre-built knowledge graph, improving the efficiency of fault detection. Another implementation involves the cloud management platform inputting the J faults and the first detection result into an artificial intelligence model to obtain the target detection result output by the AI model. Based on this implementation, the cloud management platform can determine the target detection result through the AI model, improving the accuracy of fault detection.
[0016] In one possible implementation, the method further includes: a cloud management platform receiving second information. The second information indicates M computing nodes and / or the runtime of a computing task. The cloud management platform can determine at least one server where the M computing nodes indicated by the second information reside, and determine a second detection result based on runtime data associated with the at least one server. The runtime data associated with the at least one server is obtained based on the runtime. Based on this implementation, the cloud management platform can filter out runtime data unrelated to the computing cluster's execution of computing tasks based on the second information, reducing the amount of data required for fault detection and improving the efficiency of fault detection.
[0017] In one possible implementation, the operational data associated with at least one server includes operational data of at least one device associated with at least one server. When the cloud management platform determines a second detection result based on the operational data associated with at least one server, it can also determine the second detection result based on the operational data of at least one device. The second detection result includes sub-detection results corresponding to each of the at least one device, with each sub-detection result indicating a fault in one device. The at least one device includes one or more of servers, storage devices, network devices, and server racks. Based on this implementation, the cloud management platform can determine sub-detection results by device type, thereby improving the accuracy of sub-detection results for each type of device and ultimately improving the accuracy of the final fault detection result.
[0018] In one possible implementation, the operational data of each of the at least one devices includes one or more of alarm information, at least one performance indicator data, or operational logs. When the cloud management platform determines the second detection result based on the operational data of the at least one device, it may perform the following steps for each of the at least one devices: determine a first candidate fault for each device based on the fault mapping relationship between alarm information and the fault corresponding to each device, wherein the fault mapping relationship includes the mapping relationship between faults and alarm information; determine a second candidate fault for each device based on the change information of each performance indicator in the at least one performance indicator data, and / or operational logs; and determine a sub-detection result for each device based on the first candidate fault and the second candidate fault. Based on this implementation, the cloud management platform can obtain the final fault detection result by integrating faults determined from multiple operational data, which can improve the accuracy of fault detection.
[0019] In one possible implementation, the computational task is a model training task. Based on this implementation, the method can be applied to model training scenarios. By accurately detecting faults that cause the model training task to be interrupted, it can help to accurately eliminate unusable computing nodes, and continue to execute the model training task based on available computing nodes, thereby reducing the probability of the model training task being interrupted again and improving the model training efficiency.
[0020] Secondly, a fault detection method for a computing cluster is provided. This method can be applied to a second computing node, which belongs to M computing nodes used to execute computing tasks, where M is a positive integer, and the M computing nodes are deployed on at least one server. The method includes: the second computing node sending a first detection result to a cloud management platform, the first detection result indicating a fault in the first computing node, which is the node among the M computing nodes that caused the computing task interruption. The cloud management platform manages the infrastructure providing cloud services, and the infrastructure includes at least one server. The second computing node receives a target detection result from the cloud management platform, the target detection result determining K unavailable computing nodes among the M computing nodes, where K is an integer. The second computing node continues to execute the computing task based on the target detection result, wherein the second computing node is the computing node other than the K unavailable computing nodes among the M computing nodes.
[0021] In one possible implementation, the computational task is a model training task.
[0022] The beneficial effects of any of the embodiments in the second aspect above can be referred to the beneficial effects of the corresponding embodiments in the second aspect above, and this application will not elaborate on them one by one.
[0023] Thirdly, a fault detection method for a computing cluster is provided. This method is applied to a second computing node, which belongs to M computing nodes used for executing computing tasks, where M is a positive integer, and the M computing nodes are deployed on at least one server. The method includes: the second computing node receiving a second detection result from a cloud management platform, the cloud management platform being used to manage infrastructure providing cloud services, the infrastructure including the at least one server, the second detection result indicating a fault in the infrastructure causing the computing task to be interrupted; the second computing node determining a target detection result based on the second detection result and a first detection result, the first detection result indicating a fault in a first computing node, which is the node among the M computing nodes that caused the computing task to be interrupted, the target detection result determining K unavailable computing nodes among the M computing nodes, where K is an integer; and the second computing node continuing to execute the computing task based on the target detection result, wherein the second computing node is the computing node other than the K computing nodes among the M computing nodes.
[0024] In this method, the second computing node can determine the fault causing the computing task interruption by combining the first and second detection results. The first detection result indicates a fault in the first computing node, and the second detection result indicates a fault in the infrastructure. This is equivalent to combining data from cloud tenants and the cloud management platform for comprehensive fault detection. The data relied upon for fault detection is more comprehensive, improving the accuracy of fault detection and helping to accurately troubleshoot faults. For example, the computing cluster can accurately remove unavailable computing nodes and continue executing computing tasks based on available computing nodes, reducing the probability of further interruption of computing tasks and improving the execution efficiency of computing tasks. Furthermore, this method enables automatic fault detection, eliminating the need for manual fault detection and improving fault detection efficiency.
[0025] In one possible implementation, the first detection result is determined based on the operational data of some or all of the M computing nodes, and the second detection result is determined based on the operational data of some or all of the equipment included in the infrastructure. This can be understood as the first detection result being determined based on operational data collected within the computing cluster, i.e., based on cloud tenant data, while the second detection result is determined based on the operational data of the infrastructure, i.e., based on data from the cloud management platform. This implementation effectively breaks down the data isolation barrier between cloud tenants and the cloud management platform, making the data relied upon by the cloud management platform for fault detection more comprehensive and improving the accuracy of fault detection.
[0026] In one possible implementation, the first detection result is received through a first interface between the cloud management platform and a second computing node among the M computing nodes. At least one of the M computing nodes has a first interface with the cloud management platform, and the second computing node belongs to at least one computing node. The second computing node is either the first computing node or any other computing node among the M computing nodes besides the first computing node. Based on this implementation, computing nodes within the computing cluster can interact with the cloud management platform through the first interface, breaking down the data isolation barrier between cloud tenants and the cloud management platform. This allows the cloud management platform to rely on more comprehensive data for fault detection, improving the accuracy of fault detection.
[0027] In one possible implementation, each of the at least one computing node includes an operating system, and the operating system of some or all of the at least one computing node includes a first plugin for identifying faulty computing nodes among the M computing nodes. Based on this implementation, computing nodes within the computing cluster can interact with the cloud management platform through the first plugin, breaking down the data isolation barrier between cloud tenants and the cloud management platform. This allows the cloud management platform to rely on more comprehensive data for fault detection, improving the accuracy of fault detection.
[0028] In one possible implementation, the second computing node can send the obtained target detection results to the cloud management platform, so that the cloud management platform's operation and maintenance personnel can promptly repair the faulty equipment based on the target detection results.
[0029] In one possible implementation, the second detection result is used to indicate multiple faults in the infrastructure. When the second computing node determines the target detection result based on the first and second detection results, it can also determine the target detection result based on the confidence levels of the multiple faults and the first detection result. The confidence levels of the multiple faults are determined based on the correlation between the multiple faults and the first detection result. Based on this implementation, the second computing node can determine the confidence level of out-of-band detected faults based on the correlation between in-band and out-of-band faults, thus improving the accuracy of fault detection through mutual verification between in-band and out-of-band faults.
[0030] In one possible implementation, the second computing node can determine the confidence level of an out-of-band detected fault based on whether the in-band detected fault is related to the out-of-band detected fault. For example, if the first fault among multiple faults indicated by the second detection result is related to some or all of the faults indicated by the first detection result, the confidence level of the first fault is a first value. Or, if the first fault among multiple faults indicated by the second detection result is unrelated to the faults indicated by the first detection result, the confidence level of the first fault is a second value. The first value is greater than the second value. Based on this implementation, when the in-band detected fault is related to the out-of-band detected first fault, the confidence level of the determined first fault is higher, which better conforms to the actual fault determination logic. That is, when such a fault is detected both in-band and out-of-band, the probability of the fault occurring is higher. Thus, subsequent fault determination based on the confidence level can be more accurate, improving the accuracy of fault detection.
[0031] In one possible implementation, the multiple faults include faults in at least one device in the infrastructure, which includes one or more of servers, storage devices, network devices, or server racks. When the second computing node determines the target detection result based on the confidence levels of the multiple faults and the first detection result, the second computing node can determine the target detection result based on J faults from the multiple faults and the first detection result. The J faults are faults corresponding to at least one device, where at least one of the J faults corresponds to a first type of device, and the confidence level corresponding to the at least one fault is the highest confidence level among the faults corresponding to the first type of device, where K is a positive integer. That is, the second computing node can determine the fault with the highest confidence level for each type of device, and then combine the highest confidence level faults of each type of device to comprehensively determine the target detection result. Based on this implementation, the second computing node can filter out faults with low confidence levels according to device type, which can reduce the amount of data for fault detection and improve fault detection efficiency.
[0032] In one possible implementation, the second computing node determines the target detection result based on J faults out of a plurality of faults and the first detection result. This can be achieved in several ways. One implementation involves the second computing node determining the target detection result based on the J faults, the first detection result, and a knowledge graph. The knowledge graph represents the relationships between multiple knowledge nodes, which correspond to the faults indicated by the first detection result and the J faults. Based on this implementation, the second computing node can determine the target detection result based on a pre-built knowledge graph, improving the efficiency of fault detection. Another implementation involves the second computing node inputting the J faults and the first detection result into an artificial intelligence model to obtain the target detection result output by the model. Based on this implementation, the second computing node can determine the target detection result through the artificial intelligence model, improving both the accuracy and efficiency of fault detection.
[0033] In one possible implementation, the computational task is a model training task. Based on this implementation, the method can be applied to model training scenarios. By accurately detecting faults that cause the model training task to be interrupted, it can help to accurately eliminate unusable computing nodes, and continue to execute the model training task based on available computing nodes, thereby reducing the probability of the model training task being interrupted again and improving the model training efficiency.
[0034] Fourthly, a fault detection method for a computing cluster is provided, which can be applied to a cloud management platform. The cloud management platform manages the infrastructure providing cloud services, including servers within at least one cloud data center. The method includes: the cloud management platform sending a second detection result to some or all of M computing nodes, where M is a positive integer, deployed on at least one server, to indicate a fault in a device within the infrastructure that caused the computing task to be interrupted. The cloud management platform receives target detection results from some or all of the M computing nodes, where K is an integer, to determine K unavailable computing nodes among the M nodes. Based on the target detection results, the cloud management platform determines the faults of devices related to the K computing nodes within the infrastructure.
[0035] In one possible implementation, the method further includes: a cloud management platform receiving second information. The second information indicates M computing nodes and / or the runtime of a computing task. The cloud management platform can determine at least one server where the M computing nodes indicated by the second information reside, and determine a second detection result based on runtime data associated with the at least one server. The runtime data associated with the at least one server is obtained based on the runtime. Based on this implementation, the cloud management platform can filter out runtime data unrelated to the computing cluster's execution of computing tasks based on the second information, reducing the amount of data required for fault detection and improving the efficiency of fault detection.
[0036] In one possible implementation, the operational data associated with at least one server includes operational data of at least one device associated with at least one server. When the cloud management platform determines a second detection result based on the operational data associated with at least one server, it can also determine the second detection result based on the operational data of at least one device. The second detection result includes sub-detection results corresponding to each of the at least one device, with each sub-detection result indicating a fault in one device. The at least one device includes one or more of servers, storage devices, network devices, and server racks. Based on this implementation, the cloud management platform can determine sub-detection results by device type, thereby improving the accuracy of sub-detection results for each type of device and ultimately improving the accuracy of the final fault detection result.
[0037] In one possible implementation, the operational data of each of the at least one devices includes one or more of alarm information, at least one performance indicator data, or operational logs. When the cloud management platform determines the second detection result based on the operational data of the at least one device, it may perform the following steps for each of the at least one devices: determine a first candidate fault for each device based on the fault mapping relationship between alarm information and the fault corresponding to each device, wherein the fault mapping relationship includes the mapping relationship between faults and alarm information; determine a second candidate fault for each device based on the change information of each performance indicator in the at least one performance indicator data, and / or operational logs; and determine a sub-detection result for each device based on the first candidate fault and the second candidate fault. Based on this implementation, the cloud management platform can obtain the final fault detection result by integrating faults determined from multiple operational data, which can improve the accuracy of fault detection.
[0038] In one possible implementation, the computational task is a model training task.
[0039] The beneficial effects of any of the embodiments in the fourth aspect above can be referred to the beneficial effects of the corresponding embodiments in the third aspect above, and this application will not elaborate on them one by one.
[0040] Fifthly, a fault detection device for a computing cluster is provided. The device may include a module for performing the first aspect and any possible implementation thereof, or the device may include a module for performing the second aspect and any possible implementation thereof, or the device may include a module for performing the third aspect and any possible implementation thereof, or the device may include a module for performing the fourth aspect and any possible implementation thereof.
[0041] The descriptions in the first to fourth aspects or any one of the first to fourth aspects described above are applicable to the fifth aspect or any one of the fifth aspects, and will not be repeated here.
[0042] A sixth aspect provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, causing the computing device cluster to perform the methods disclosed in the first aspect and any possible implementation thereof, or to perform the methods disclosed in the second aspect and any possible implementation thereof, or to perform the methods disclosed in the third aspect and any possible implementation thereof, or to perform the methods disclosed in the fourth aspect and any possible implementation thereof.
[0043] In a seventh aspect, this application provides a computer program product containing instructions that, when executed by a cluster of computer devices, cause the cluster of computer devices to implement the method disclosed in the first aspect and any possible implementation of the first aspect, or to execute the method disclosed in the second aspect and any possible implementation of the second aspect, or to execute the method disclosed in the third aspect and any possible implementation of the third aspect, or to execute the method disclosed in the fourth aspect and any possible implementation of the fourth aspect.
[0044] Eighthly, this application provides a computer-readable storage medium including computer program instructions that, when executed by a computing device cluster, cause the computing device cluster to perform the method disclosed in the first aspect and any possible implementation of the first aspect, or to perform the method disclosed in the second aspect and any possible implementation of the second aspect, or to perform the method disclosed in the third aspect and any possible implementation of the third aspect, or to perform the method disclosed in the fourth aspect and any possible implementation of the fourth aspect.
[0045] The beneficial effects of the implementation methods in any of the sixth to eighth aspects above can be referred to the beneficial effects of the corresponding implementation methods in the first to fourth aspects above, and this application will not elaborate on them one by one. Attached Figure Description
[0046] Figures 1A to 1C This is a schematic diagram of the system architecture provided for an embodiment of this application;
[0047] Figure 2 A flowchart illustrating a method provided in an embodiment of this application;
[0048] Figure 3 A schematic diagram illustrating the determination of the cause of a fault, provided in an embodiment of this application;
[0049] Figure 4 A schematic diagram of an out-of-band detection process provided in an embodiment of this application;
[0050] Figure 5A and Figure 5B A flowchart illustrating the process of screening out-of-band faults provided in an embodiment of this application;
[0051] Figure 6 Another flowchart illustrating a method provided in an embodiment of this application;
[0052] Figure 7 A flowchart illustrating another method provided in an embodiment of this application;
[0053] Figures 8-11 This is a schematic diagram of the device provided in the embodiments of this application;
[0054] Figure 12 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0055] Figure 13 This is a schematic diagram of a computing device cluster provided in an embodiment of this application. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings. In the description of the embodiments of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature.
[0057] It should be understood that in the embodiments of this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0058] The following explanations of some terms used in the embodiments of this application are provided to facilitate understanding by those skilled in the art.
[0059] (1) In-band and out-of-band
[0060] In-band management typically refers to a management model that uses conventional data channels for management and control, where management and control information and user business information are transmitted through the same logical channel. In this application embodiment, "in-band" can be understood as corresponding to a cloud tenant, and in-band resource data belongs to the cloud tenant.
[0061] Out-of-band management typically refers to a management model that uses dedicated data channels for management and control, where management and control information and user business information are transmitted on different logical channels. In this application embodiment, "out-of-band" can be understood as corresponding to a cloud service provider, and out-of-band resource data belongs to the cloud service provider.
[0062] (2) Computational tasks, large models and model training tasks
[0063] Computational tasks refer to various tasks performed using computing devices, including big data analysis, scientific computing, image processing, model training, and inference. Model training is a type of computational task. In the field of machine learning, model training refers to the process of adjusting the parameters of an algorithm model to enable it to learn patterns from input data and use them to predict or classify unknown data.
[0064] Large models refer to machine learning models with a massive number of parameters and complex computational structures. These models are typically built from deep neural networks and have billions or even hundreds of billions of parameters. Large models are designed to improve the expressive power and predictive performance of models, enabling them to handle more complex tasks and data. Large models have wide applications in various fields, including natural language processing, computer vision, speech recognition, and recommender systems. By training on massive amounts of data to learn complex patterns and features, large models have stronger generalization capabilities and can make accurate predictions on unseen data. For example, a large model could be an AIGC model.
[0065] Training a large model requires a large computing cluster and takes a considerable amount of time. Resuming training from a standby point is a technique for saving and restoring the model's training state. When training is interrupted, this technique can save the current model state and restore it in subsequent training sessions. Resuming training from a standby point allows training to continue after an interruption, rather than starting from scratch, saving time and avoiding increased training duration due to restarting.
[0066] When training tasks are interrupted, fault detection in the computing cluster is typically performed manually. For example, after a training task is interrupted, cloud tenants first look for the cause of the error in the model training logs and platform logs. However, as the scale of the computing clusters required for model training gradually increases, manually finding the cause of the fault becomes a bottleneck and may not meet the requirements for quickly repairing faults and restoring business operations. Furthermore, since model training is usually costly, when a training task is interrupted, cloud tenants often want to locate the fault within a short time (e.g., 10 minutes), and regardless of whether the fault can be located within this time, they need to resume training from the breakpoint. That is, if the cause of the fault can be found, the faulty node is removed and model training is restarted; if the cause cannot be found, the model is restarted from the breakpoint and runs with the fault. However, this method of running with the fault is likely to result in further interruptions, and repeated interruptions lead to low training efficiency.
[0067] Therefore, in order to reduce the probability of interruption after restarting model training and improve training efficiency, it is necessary to accurately locate the fault and remove unusable nodes after the training task is interrupted.
[0068] Based on this, embodiments of this application provide a fault detection method for computing clusters. This method can perform preliminary fault detection within the band (or out-of-band), and then combine the preliminary fault detection results with the out-of-band (or in-band) detection results to determine the final detection result. Because it combines the in-band and out-of-band detection results, the location of the fault can be more comprehensively determined, improving the accuracy of fault detection. Furthermore, this method enables automatic fault detection, eliminating the need for manual fault detection and improving fault detection efficiency.
[0069] The method described in this application can be applied to computing clusters or scenarios containing computing clusters, and can be used to implement fault detection in these scenarios. Scenarios containing computing clusters may include cloud service scenarios (such as public cloud scenarios). In some embodiments, "computing cluster" can also be replaced with computer cluster, distributed computing cluster, cloud computing cluster, data center, cloud data center, or other names, without specific limitations. As an example, when the computing task is a training task, the computing cluster can also be called a training cluster. In some embodiments, the computing cluster may also use different names depending on the type of computing task. For example, in intelligent computing scenarios, computing clusters can be called intelligent computing clusters, intelligent computing centers, or intelligent computing systems. Intelligent computing scenarios typically refer to computing scenarios that use intelligent methods for decision-making and prediction based on technologies such as artificial intelligence and big data analysis. Intelligent computing scenarios focus more on intelligence and adaptability. In general computing scenarios, computing clusters can be called general computing clusters. General computing scenarios typically refer to more general or conventional computing scenarios used to perform routine computing tasks, such as large-scale data processing and computation, emphasizing efficient and standardized computing processing, and usually not involving complex intelligent algorithms. In supercomputing scenarios, computing clusters can be called supercomputing clusters. Supercomputing scenarios typically belong to the field of high-performance computing (HPC), involving large-scale computation of complex scientific and engineering problems, and usually using supercomputers to provide extremely high computing power and the ability to solve complex problems.
[0070] Please refer to Figure 1AThis diagram illustrates an application scenario of the method provided in this application embodiment. The scenario may include a cloud management platform 100, which manages multiple cloud data centers set up by a cloud service provider in different regions. The cloud management platform provides interfaces related to public cloud services, such as web pages or application programming interfaces (APIs), for cloud tenants (or users) to remotely access public cloud services. Cloud tenants can log in to the cloud management platform via a pre-registered account and password on the public cloud access page. After successful login, they can select and purchase public cloud services provided by the cloud data center in the region selected by the cloud tenant on the public cloud access page. For example, the public cloud service may be a cloud computing service. When a cloud tenant purchases a cloud computing service, the cloud service provider can provide the cloud tenant with cluster resources for performing computing tasks.
[0071] The cloud management platform 100 manages the infrastructure providing cloud services. This infrastructure includes multiple regions, each containing at least one cloud data center. Cloud services run on servers located in at least one cloud data center within one of these regions. A region refers to the geographical location of a cloud service. Cloud services are categorized by region based on geographical location and network latency. Cloud services within the same region use infrastructure located in the same geographical area. For example, if a cloud service selects South China as its region, then it will use cloud data centers in the South China region to provide that service. A region may include one or more Availability Zones (AZs). An AZ is a collection of one or more cloud data centers with independent water and electricity supply. Within an Availability Zone, computing, network, and storage resources are logically further divided into multiple clusters. Multiple Availability Zones within a region are connected via high-speed fiber optic cables to meet the needs of cloud tenants building high-availability systems across AZs.
[0072] See Figure 1A As shown, the cloud management platform 100 can provide a client 400, which can be used by cloud tenants to purchase and manage cloud services. For example, the client 400 can display relevant interfaces for cloud computing services. These interfaces may include interfaces for displaying operational data of the cloud computing services, and interfaces for displaying fault detection results when a cloud computing service fails. Cloud tenants can manage and configure cloud computing services through these interfaces. Optionally, the client 400 can be implemented through a terminal device or a functional module (e.g., a webpage or application) within the terminal device. For example, the terminal device can be a personal computer (PC), a tablet computer (PAD), or a mobile phone, etc., with no specific limitations.
[0073] For example, Figure 1AIn the scenario shown, a cloud tenant can access the cloud management platform 100 via the Internet 300 through a client 400. The cloud tenant can purchase cloud computing services on the cloud management platform 100. The cloud tenant can input the relevant configuration information and service requirements for the cloud computing services they need on the cloud management platform 100 through the client 400. The cloud management platform 100 can configure and run a computing cluster 200 based on the configuration information and service requirements input by the cloud tenant. The computing cluster 200 may include one or more computing nodes, which can be deployed on the servers included in the cloud management platform. Figure 1A As shown, based on the cloud computing service configuration information and service requirements input by the client 400, the cloud management platform 100 assigns the cloud tenant's cloud computing service to cloud data center A in region 1 and cloud data center B in region 2. In other words, the computing nodes of the computing cluster 200 are deployed in cloud data center A in region 1 and cloud data center B in region 2. Therefore, the client 400 can access cloud data center A in region 1 and cloud data center B in region 2 via the Internet through the cloud management platform 100. It should be noted that the cloud tenant can choose which region and cloud data center provides the cloud computing service through the cloud management platform 100, or the cloud tenant can choose not to, and the cloud management platform 100 will randomly assign a cloud data center to provide the cloud computing service. There is no limitation on the allocation method of cloud data centers providing the service.
[0074] The cloud management platform 100 can also provide a client 500. The client 500 can be used by the management users of the cloud management platform 100 to manage the cloud management platform, or the operation and maintenance personnel of the cloud management platform 100 can use the client 500 to perform operation and maintenance management and repair of the cloud management platform.
[0075] For example, after the cloud management platform 100 configures a computing cluster 200 for a cloud tenant, the cloud tenant's computing tasks can be executed on the computing cluster 200. During the execution of the computing task, the computing task may be interrupted for various reasons, such as the computing task itself or the hardware resources of the computing cluster 200. When the computing task is interrupted, it is necessary to perform fault detection on the computing cluster 200 to eliminate the faulty node and continue the execution of the computing task. In this embodiment, preliminary fault detection can be performed within the computing cluster 200 (i.e., on the cloud tenant's side) or on the cloud management platform. The final detection result is then determined by combining the preliminary fault detection results with the detection results from the other party. Because the detection results from both parties are combined, the location of the fault can be more comprehensively determined, improving the accuracy of fault detection.
[0076] The client-side 400 can also be used to display fault detection results, such as in-band detection results, out-of-band detection results, or final detection results, to help cloud tenants understand the location of the fault and perform fault repair. For example, when the interruption of a computing task is caused by the program of the computing task itself, the cloud tenant can analyze the faulty program in a timely manner and repair it.
[0077] Client 500 can also be used to display fault detection results, such as in-band detection results, out-of-band detection results, or final detection results, to facilitate timely fault location and repair by the cloud management platform 100's operations and maintenance personnel. For example, when the interruption of computing tasks is caused by hardware failure, the cloud management platform 100's operations and maintenance personnel can promptly repair the faulty hardware.
[0078] In this embodiment, the cloud management platform 100 and / or computing cluster 200 can be implemented in software or in hardware. For example, in a software implementation, the cloud management platform 100 may include code running on computing instances. These computing instances can be at least one of physical hosts (computing devices), virtual machines, containers, or other computing devices. Optionally, there may be one or more computing devices. For example, the cloud management platform 100 may include code running on multiple hosts / virtual machines / containers. As another example, in a hardware implementation, the cloud management platform 100 may include at least one computing device, such as a server. Alternatively, the cloud management platform 100 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. The computing devices included in the cloud management platform 100 can be referred to as fault detection devices.
[0079] As a possible example, please refer to Figure 1B This is a schematic diagram of a cloud management platform 100 provided in an embodiment of this application. The cloud management platform 100 is used to provide infrastructure for cloud services, and the infrastructure may include the following types of devices:
[0080] (1) A server, also known as a computing device or host, provides computing resources for performing computing tasks. The server is the core component for performing computing tasks; a computing cluster typically consists of computing nodes on multiple servers, such as... Figure 1B The computing nodes shown are 1 to n. A server can deploy one or more computing nodes, and each computing node can contain one or more processors, memory, and other computing resources. The processor can be one or more of a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA). The cloud management platform 100 can provide multiple nodes to process computing tasks simultaneously, improving computing efficiency, and can add or remove computing nodes as needed to adapt to different working scenarios.
[0081] In some embodiments, the computational task may be the training task of an AI model. The server deploys 200 computing nodes of a computing cluster, which collaborate to perform the training task of the AI model. For example, these computing nodes can be used to perform processes such as forward inference and backpropagation of the AI model.
[0082] (2) Storage devices are used to store data and files related to computing tasks. For example, storage devices can be network attached storage (NAS), storage area network (SAN), or storage devices within a Hadoop Distributed File System (HDFS). When a compute node executes a computing task, it can retrieve data related to the subtask it needs to execute from the relevant storage device, and execute its corresponding subtask based on this data. After the subtask is completed, the execution result can also be stored in the storage device.
[0083] For example, an AI model may include multiple data processing layers. A computing node is used to execute the computation process of one of the multiple data processing layers. The computing node can load training data of the training batch allocated to itself or the output data of the previous data processing layer in the same training batch from the storage device, as well as load the model parameters of the data processing layer to be executed from the storage device, so as to execute the computation process of the data processing layer according to the loaded data and store the computation results in the storage device.
[0084] (3) Network devices are responsible for data transmission between servers, storage devices, and external networks, providing high-speed data transmission channels. The performance of network devices directly affects the overall performance of the computing cluster 200. Network devices may include one or more of switches, routers, or network interface cards (NICs). Network devices are used to connect the various nodes within the computing cluster 200.
[0085] For example, computing nodes need to load training data and model parameters from storage devices via network devices, and computing nodes also need to transfer computing results between each other via network devices.
[0086] (4) A server rack is a physical structure used to physically install and house servers, storage devices, and network equipment. Server racks provide power supplies, cooling devices, etc., to ensure the physical security of servers, storage devices, and network equipment. Server racks typically have thermal management capabilities, meaning that ventilation and heat dissipation performance must be considered during the design and selection of the rack to keep the equipment inside within its optimal operating temperature range. In addition, they should also have power management capabilities, meaning that the rack needs to provide a stable power supply and can integrate an uninterruptible power supply (UPS) to ensure system reliability.
[0087] In addition, the cloud management platform 100 may include other devices besides those mentioned above, without specific limitations. Through the cooperation of the aforementioned servers, storage devices, network devices, and racks, the cloud management platform 100 can configure a computing cluster 200 for cloud tenants to achieve high-efficiency computing. That is, the computing cluster 200 may include one or more of the following: servers, storage devices, network devices, and racks.
[0088] In this embodiment, the computing cluster 200 includes computing nodes that can be implemented in software or hardware. For example, one example of a computing node being implemented in software is that it may include code running on multiple hosts / virtual machines / containers. For example, a computing node may include an operating system (OS) running on a server. For example, a server may include one or more OSs, and an OS can be considered a computing node. As an example, when training an AI model, a cloud tenant can deploy training software within the various OSs included in the computing cluster. When the training software runs within an OS, it can be used to train a specified AI model. Another example of a computing node being implemented in hardware is that it may include at least one server; for example, a computing node is a server. Alternatively, a computing node may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0089] As mentioned above, computational tasks may be interrupted during execution. The interruption may be due to issues with the computational task itself or hardware resources. Therefore, it is necessary to accurately identify the fault in order to troubleshoot it promptly and ensure the smooth execution of the computational task. If the fault cannot be accurately located, it may be impossible to troubleshoot it correctly, and the computational task will run with the fault, which will increase the probability of the computational task being interrupted again and reduce the task execution efficiency.
[0090] Please see Figure 1C The diagram shown illustrates an architecture for fault detection provided in an embodiment of this application. The computing cluster 200 is built upon the physical resources provided by the cloud management platform 100. Cloud tenants can deploy their own computing programs within these computing clusters 200; for example, training programs (or training software) for training AI models can be deployed within the computing cluster 200. During the execution of computing tasks by the computing nodes within the computing cluster 200, data such as the resource status of the computing nodes can be collected. Since this data is collected by software deployed by the cloud tenant, it belongs to the cloud tenant itself and can be referred to as in-band resource data. See also... Figure 1CAs shown, an operating system (OS) can be installed on a compute node; in other words, an OS can be considered a compute node. Besides deploying computing software to execute computing tasks, cloud tenants can also deploy their own task monitoring programs. These task monitoring programs can collect data from the OS while the compute node is executing computing tasks. For example, they can collect data on CPU, hard drive, memory, and circuit boards, as well as runtime logs and error messages from the computing program. When a computing task is interrupted, the compute node can perform fault detection based on the information collected by the task monitoring program and the fault detection configuration information configured by the cloud tenant. For example, the fault detection configuration information can include expert experience or rule models used to implement fault detection. Since this fault detection method relies entirely on the cloud tenant's in-band resource data, it can also be called an in-band detection method.
[0091] In addition, the cloud management platform 100 also monitors the status of hardware resources to ensure timely repair of faulty hardware, thereby providing a better service experience for cloud tenants. See also Figure 1C As shown, the cloud management platform 100 can monitor data such as temperature, voltage, current, boards, fans, load, and status of the device through out-of-band firmware (e.g., a baseboard management controller (BMC)). Since this data is collected by the cloud management platform 100, it belongs to the cloud management platform and can be referred to as out-of-band resource data. The cloud management platform 100 can perform fault detection based on the collected out-of-band resource data and its fault detection configuration information to identify the faulty hardware device. For example, the fault detection configuration information may include alarm thresholds and fault detection rules to assist in fault detection. Because this fault detection method relies entirely on the out-of-band resource data of the cloud management platform 100, it can also be called an out-of-band detection method.
[0092] Because in-band resource data belongs to the cloud tenant, while out-of-band resource data belongs to the cloud management platform 100, and the data of the cloud tenant and the cloud management platform 100 are isolated from each other, neither the cloud tenant nor the cloud management platform 100 can fully locate the fault. For example, the in-band detection method mentioned above can usually only detect the cause of some in-band data errors, or the out-of-band detection method can usually only detect single points of failure in out-of-band hardware, resulting in low fault detection accuracy. Therefore, to improve the accuracy of fault detection, it may be necessary to perform global fault detection by combining in-band and out-of-band data.
[0093] One possible approach is to deploy a fault detection system either in-band or out-of-band, capable of simultaneously collecting both in-band and out-of-band resource data for comprehensive fault detection. However, this approach essentially exposes the resource data between the cloud tenant and the cloud management platform 100 to each other, posing significant data risks to either party and making implementation difficult.
[0094] Another possible approach is to configure a data interface between in-band and out-of-band systems. This interface can be used to transmit fault detection-related data, such as the results of fault detection performed in-band or out-of-band. This would allow for the integration of in-band and out-of-band fault detection results to obtain a more accurate fault detection result.
[0095] As an example, see Figure 1C As shown, in-band plugins can be configured on the OS, while out-of-band intelligent operation and maintenance (O&M) systems can be deployed on the cloud management platform 100. These intelligent O&M systems can also be called fault detection systems or fault diagnosis systems. Typically, the intelligent O&M system can be implemented using the device resources included in the O&M area of the cloud management platform 100. In some embodiments, the intelligent O&M system may include an engine for performing out-of-band fault detection, such as a cluster computing autonomous engine (CCAE). The in-band plugin and CCAE have an open API, through which they can exchange data. In some embodiments, the intelligent O&M system can be associated with CCAE. The in-band plugin and CCAE have an open API, through which they can exchange data. In this case, it can be understood that CCAE and the intelligent O&M system are independent. CCAE can obtain out-of-band resource data from the intelligent O&M system and in-band detection results from the in-band plugin via the open API, thus combining out-of-band and in-band data for fault detection.
[0096] The in-band plugin can acquire the in-band resource data required for fault detection during the operation of compute nodes, perform fault detection based on this data, and obtain in-band detection results. It can also send relevant data to CCAE via the open API, including information about compute cluster 200 (such as a list of OSes included in the cluster) and the in-band detection results. The in-band plugin can also be called a plugin probe, plugin system, or other names without restriction. CCAE can be used for the operation and maintenance of cloud computing-related systems and services, such as device data monitoring, fault detection, and troubleshooting. For example, CCAE can filter devices based on compute cluster 200 information, select data from out-of-band resource data related to compute cluster 200, perform out-of-band fault detection based on this data, obtain out-of-band detection results, and finally obtain the final fault detection result based on the correlation between the in-band and out-of-band detection results, thus determining a more accurate fault location. In this way, the entire compute cluster can be globally detected by combining out-of-band and in-band fault detection results, improving the accuracy of fault detection results.
[0097] In some embodiments, the in-band plugin may be published by a cloud management platform. Developers of the computing software on the computing cluster can write the address of the cloud management platform accessible to the in-band plugin when developing the software, so that the in-band plugin can send data to the cloud management platform based on that address. Alternatively, when the cloud management platform publishes an in-band plugin, it can broadcast the publication message of the in-band plugin in the cloud management platform. The computing cluster can load the in-band plugin according to the publication message, or load the in-band plugin with the confirmation of the cloud tenant.
[0098] After obtaining the final fault detection results, CCAE can send them to the cloud tenant or compute cluster 200 via OpenAPI or other methods. This allows compute cluster 200 to continue executing computing tasks after identifying the faulty node. Because the fault can be located more accurately, the faulty node can be precisely removed, reducing the probability of subsequent computing task interruptions. Additionally, CCAE can promptly notify relevant operations and maintenance personnel on the cloud management platform 100, enabling them to promptly repair the fault.
[0099] In some embodiments, a fault detection system can also be deployed in-band. In this case, after performing out-of-band fault detection, the out-of-band CCAE can send the out-of-band detection results to the in-band fault detection system via openAPI. The in-band fault detection system then combines the out-of-band and in-band detection results to obtain the final fault detection result.
[0100] Please see Figure 2The diagram shown is a flowchart illustrating a fault detection method for a computing cluster provided in an embodiment of this application. This method can be applied to... Figures 1A to 1C In the system architecture. The method may include the following steps:
[0101] Step 201: The second computing node within the computing cluster determines the first detection result, which indicates a fault in the first computing node, where the first computing node is the node among the M computing nodes that caused the computing task to be interrupted. The first detection result can also be replaced by an in-band detection result.
[0102] The computing cluster may include M computing nodes, where M is a positive integer. These M computing nodes are used to execute computing tasks, and they are deployed on at least one server. The at least one server is a device resource managed by a cloud management platform; in other words, the cloud management platform manages the infrastructure providing cloud services. This infrastructure includes servers within at least one cloud data center, and these servers may include the at least one server hosting the M computing nodes. For example, the cloud management platform could be... Figures 1A to 1C The cloud management platform 100 shown can have a computing cluster that can be Figures 1A to 1C The computing cluster shown is 200.
[0103] In some embodiments, the computational task can be a model training task, which refers to collaboratively training an AI model using the computing nodes included in the computing cluster. In this embodiment, the computing nodes can also be referred to as training nodes.
[0104] In some embodiments, cloud tenants can purchase cloud services through an interface provided by the cloud management platform. For example, a cloud service could be a cloud computing service, whereby a cloud tenant requests that its computing tasks be executed through infrastructure managed by the cloud management platform. In this case, the cloud management platform can select appropriate device resources for the computing task and build a computing cluster to execute it. Alternatively, a cloud service could be a service that provides infrastructure to cloud tenants, essentially allowing them to lease device resources from the cloud management platform. The leased device resources can then be used to build a computing cluster on which the cloud tenant can execute computing tasks.
[0105] During the execution of computing tasks by cloud tenants, task interruptions may occur, necessitating the detection of faults causing these interruptions. Within the computing cluster (i.e., within the band), a second computing node can perform in-band fault detection to determine the first detection result. This first detection result indicates a fault in the first computing node, which may include the first computing node that caused the task interruption and / or fault information of that first computing node. For example, the first detection result may include a list of faulty nodes and the specific fault occurring at each node in that list; the faulty nodes in the list are the first computing nodes.
[0106] Optionally, the number of second computing nodes can be one or more. If the number of second computing nodes is one, then the second computing node can be the first computing node, or the second computing node can be one of the other nodes among the M computing nodes besides the first computing node. If the number of second computing nodes is multiple, the multiple second computing nodes can include one or more first computing nodes, and / or one or more other nodes among the M computing nodes besides the first computing node.
[0107] As an example, each compute node within a compute cluster can serve as a secondary compute node.
[0108] In one possible implementation, the first detection result can be determined in the following way:
[0109] In method a1, the computing cluster can adopt a distributed architecture, meaning there may be no centralized management node. All computing nodes within the cluster can have equal permissions, and the second computing node can be any one or more of the M computing nodes. Each second computing node can perform fault detection and obtain corresponding fault detection results. These results can then be used as the first detection result. For example, each second computing node's fault detection result can be used as the first detection result; that is, each second computing node can perform fault detection and obtain its own first detection result. Different second computing nodes may obtain the same or different first detection results. Alternatively, the first detection result can be obtained by combining the fault detection results of multiple second computing nodes. For instance, each second computing node can broadcast its own fault detection results within the computing cluster. In this way, each second computing node can obtain at least one fault detection result, including its own, and thus obtain the first detection result based on at least one fault detection result. Typically, the first detection result obtained by each second computing node is the same.
[0110] In method a2, the computing cluster can also include a management node that can manage the computing cluster. In this case, the second computing node can be the management node, meaning that the first detection result can be obtained by the management node within the computing cluster performing fault detection.
[0111] In this embodiment, a fault detection program can be deployed on each second computing node, and the fault detection programs on all second computing nodes within the computing cluster can be constructed as an in-band fault detection system. For example, the fault detection program can be deployed by the cloud tenant itself, or it can be provided to the cloud tenant by the cloud management platform, and the cloud tenant can choose whether to use the fault detection program.
[0112] When a computing task is interrupted, the second computing node can acquire runtime data between the start time of the computing task and the time of the interruption, and determine a first detection result based on this runtime data. The runtime data can include runtime data from some or all of the M computing nodes. Each second computing node can collect its own runtime data or obtain runtime data from other computing nodes. The runtime data can include runtime status data of the computing node's CPU, hard disk, memory, and boards, and can also include runtime logs from the computing node's operation, such as error information during the execution of the computing task. For example, each second computing node can start searching from the computing node that directly reported an error until all error information is found, and then collect all the error information.
[0113] As an example, the second computing node determines the first detection result based on the runtime data and fault rules. For instance, the fault rules could be configured with corresponding parameter thresholds for one or more parameters, and a fault can be determined based on whether each parameter exceeds the parameter threshold. Alternatively, a mapping relationship can be pre-built between error information, runtime status data, and fault causes; the fault rules could include this mapping relationship. Then, the second computing node can query the fault cause corresponding to the current error information and runtime status data based on the mapping relationship, thereby obtaining the first detection result.
[0114] For example, see Figure 3 The diagram shown is a schematic representation of determining the cause of a fault according to an embodiment of this application. Taking the mapping relationship between error information and fault causes as an example, a mapping relationship table can be constructed based on expert experience, connecting different error messages with fault causes. This mapping relationship table represents the mapping relationship between error information and fault causes, such as... Figure 3The error message 1 corresponds to fault cause 1, error message 2 corresponds to fault cause 2, and so on. Therefore, we can search this mapping table for all error messages actually generated during the execution of the computation task until the fault cause corresponding to each error message is determined. Then, from the determined fault causes, we can identify one or more of the most likely fault causes. Similarly, a mapping table can be constructed based on other data such as runtime status data. Alternatively, a mapping table can be constructed between multiple types of data (such as runtime status data and error messages) and fault causes.
[0115] Optional, see Figure 3 As shown, the mapping table can also include the confidence level for each fault cause. The confidence level can also be called a weight value or credibility, etc. The confidence level in the mapping table can be set based on expert experience, such as... Figure 3 Each mapping relationship in the mapping table shown can be an error message - fault cause - confidence level. In this way, each fault cause corresponding to each error message found according to the mapping table has a corresponding confidence level. Then, one or more fault causes with the highest confidence level can be determined from all fault causes as the fault causes for in-band fault detection, and the first detection result can be obtained. The first detection result can include the first computing node that may fail and the cause of the failure during the time period from the start of the computing task (or when continuing execution, it can be a restart) until the interruption.
[0116] In the above description, the interruption of the computing task can also be understood as a failure of the computing task or an abnormality in the computing task.
[0117] Step 202: The second computing node sends the first detection result to the cloud management platform. Correspondingly, the cloud management platform receives the first detection result.
[0118] In this embodiment, at least one of the M computing nodes in the computing cluster has a first interface with the cloud management platform. The at least one computing node includes a second computing node, or in other words, the second computing node belongs to the at least one computing node. The second computing node can then send a first detection result to the cloud management platform through the first interface, and the cloud management platform can also receive the first detection result through the first interface. The first interface serves as a data bridge between the cloud management platform and the cloud tenant's computing cluster, enabling data transfer between in-band and out-of-band. For example, when a computing node in the computing cluster needs to send data to the cloud management platform, it can call the first interface to send the data; conversely, the cloud management platform can also call the first interface to transfer data when it needs a computing node to send data. This breaks down the data isolation barrier between the cloud tenant and the cloud management platform, making it possible to combine in-band and out-of-band fault detection results, thus improving the accuracy of fault detection.
[0119] As one implementation of implementation method a1 above, each second computing node within the computing cluster can independently determine a first detection result, and then each second computing node can send the first detection result to the cloud management platform through a first interface. That is, each second computing node can perform fault detection based on its own acquired operational data, and then send the first detection result to the cloud management platform, which can then aggregate the first detection results sent by all the second computing nodes. As another implementation of implementation method a1 above, each second computing node within the computing cluster can independently determine a fault detection result, and then the second computing nodes within the computing cluster can broadcast the fault detection results to each other. Each second computing node can aggregate its own obtained fault detection results and the received fault detection results to obtain a first detection result, and then send the first detection result to the cloud management platform. Optionally, the number of second computing nodes performing the aggregation can be one or more, which can be set according to the actual scenario requirements and is not limited in this regard.
[0120] In this embodiment of the application, each of the M computing nodes may include an OS. The OS can provide a basic operating environment for the execution of computing tasks or the operation of computing software. Therefore, the description of computing nodes can also be replaced with OS, or an OS can be understood as a computing node.
[0121] In one possible implementation, the operating system of at least some or all of the M computing nodes includes a first plugin, which can be used to implement data transfer with the cloud management platform. For example, a second computing node includes the first plugin, and the second computing node can use the first plugin to call a first interface to send a first detection result to the cloud management platform.
[0122] Optionally, the first plugin can also be used to identify faulty computing nodes among the M computing nodes. For example, if the second computing node contains the first plugin, the second computing node can perform fault detection through the first plugin to obtain a first detection result, and can send the first detection result to the cloud management platform by calling the first interface through the first plugin.
[0123] For example, the first plugin could be Figure 1C The in-band plugin shown can have its first interface as follows: Figure 1C The openAPI shown.
[0124] In one possible implementation, to improve the data security of cloud tenants, when the second computing node sends data to the cloud management platform, the second computing node can perform data anonymization processing on the data to be sent, such as replacing the cloud tenant's private data with publicly available strings, so as to avoid directly exposing the cloud tenant's private data to the cloud management platform.
[0125] In one possible implementation, the second compute node sends the first detection result to the cloud management platform with the permission of the cloud tenant. For example, the cloud tenant can pre-configure to allow comprehensive fault detection combining in-band and out-of-band methods, essentially allowing certain in-band data to be sent to the cloud management platform. In this case, the second compute node can send the first detection result to the cloud management platform. Alternatively, the second compute node can send the first detection result to the cloud management platform after confirmation from the cloud tenant. For instance, the cloud tenant can be queried before sending, and the second compute node sends the first detection result after confirmation. Yet another example is that the cloud tenant can actively operate on the interface provided by the cloud management platform, calling the open API on the interface to send the first detection result to the cloud management platform. It is understood that this same approach can be used when sending other data besides the first detection result.
[0126] Step 203: The cloud management platform determines the second detection result, which is used to indicate the failure of the infrastructure.
[0127] In this embodiment, the cloud management platform manages the infrastructure providing cloud services. The platform can perform out-of-band fault detection based on the operational data of some or all of the devices included in the infrastructure, obtaining a second detection result, which can also be called an out-of-band detection result. The fault indicated by the second detection result is a fault in the infrastructure providing cloud services to the cloud tenant. The infrastructure may include multiple devices, and the second detection result may indicate a fault in one of the devices included in the infrastructure, or in other words, a fault in one of the devices included in the infrastructure that causes the interruption of computing tasks.
[0128] The infrastructure included in the cloud management platform may include a variety of devices, such as servers, storage devices, network devices, and server racks.
[0129] In general, to ensure the service experience of cloud tenants, the cloud management platform can set up out-of-band firmware (such as BMC) to monitor the status of the device. Therefore, the cloud management platform can determine the second detection result based on the operating data collected by the out-of-band firmware.
[0130] In some embodiments, the second computing node can also send second information to the cloud management platform, and the cloud management platform can receive the second information. This second information can be sent together with the first detection result, or it can be sent separately from the first detection result. The second information can be used to indicate one or more of the following:
[0131] (1) Indicates the M computing nodes included in the computing cluster, that is, indicates which computing nodes the computing cluster specifically includes. For example, the second information may contain the identifiers of the M computing nodes. As an example, the identifier of a computing node may be the identifier of the OS on the computing node. For example, the second information may contain a list of OSes, which contains the identifiers of the OSes participating in performing the computing tasks.
[0132] Based on the M computing nodes indicated by the second information, the cloud management platform can determine at least one server where the M computing nodes reside. Therefore, the cloud management platform only needs to perform fault detection on the server-related devices involved in the computing cluster, narrowing the scope of fault detection. This not only improves the accuracy of fault detection but also reduces the data processing load on the cloud management platform, thus increasing the efficiency of fault detection.
[0133] As an example, a cloud management platform can store the mapping between compute nodes and servers. Based on this mapping and the M compute nodes indicated by the second information, the cloud management platform can then determine at least one server containing those M compute nodes. Since the operation of a server typically depends on other devices, such as network equipment, storage devices, and server racks, the platform can further search for information on the other devices that the determined at least one server depends on. For ease of description, servers, network equipment, storage devices, and server racks will be referred to as out-of-band devices in the following text.
[0134] (2) Indicates the execution time of the computation task. For example, the second information includes the first time the computation task started and the second time the computation task was interrupted. Or, for example, the second information may include the total execution time of the computation task, or the time since the computation task last started.
[0135] The cloud management platform can determine the operational data related to at least one server containing M computing nodes within the operational time indicated by the second information, and determine the second detection result based on the operational data related to at least one server. Therefore, the cloud management platform only needs to perform fault detection based on relevant data within the operational time of the computing cluster, reducing the data processing load and improving fault detection efficiency. As an example, the second information includes a first time and a second time. The cloud management platform can filter out the operational data of out-of-band devices within this time period based on the first and second times, and then determine the second detection result based on the operational data of the out-of-band devices.
[0136] As described above, an out-of-band device may include at least one device associated with at least one server, and the operational data associated with at least one server may include the operational data of the at least one device, which includes one or more of servers, storage devices, network devices, and cabinets.
[0137] Therefore, the cloud management platform can determine the second detection result based on the operating data of at least one device. For example, the cloud management platform can determine the fault of each of the at least one devices based on the operating data of each device, and obtain the sub-detection result corresponding to each device. In other words, the second detection result can include the sub-detection results corresponding to each of the at least one device, and each sub-detection result can be used to indicate the fault of a device.
[0138] In one possible implementation, see Figure 4 The diagram illustrates an out-of-band detection process provided in an embodiment of this application. The second information includes a list of computing nodes, which contains M computing nodes in the computing cluster that execute computing tasks. The cloud management platform can use this list and the correspondence between computing nodes and servers to locate the N servers where the M computing nodes reside. M and N can be the same or different. Furthermore, the cloud management node can locate other out-of-band devices related to the N servers based on the topology, such as... Figure 4 The diagram shows K storage devices, I network devices, and P server racks. Additionally, the cloud management platform can filter the operational data of the out-of-band devices identified above based on the runtime indicated by the second information, such as... Figure 4 The operating data of N servers, K storage devices, I network devices, and P server racks shown can be used to obtain sub-detection results for each type of out-of-band device. Figure 4 The results shown are the sub-detection results for the server, storage device, network device, and server rack.
[0139] In some embodiments, the operational data of each device may include multiple types of information, such as alarm information, at least one performance metric, or operational logs. Alarm information refers to abnormal operation triggered during the operation of each device; alarm information may also be called error information, anomaly information, or alert information. At least one performance metric includes the value of at least one performance metric for the corresponding device during the execution time of the computing task. The at least one performance metric may include parameters that measure device performance, such as CPU, memory, and temperature. Operational logs are used to record events during device operation. It is understood that the various types of information included in the operational data are all related to device operation, or measure the device's operating status from different perspectives or data sources. Essentially, each type of information can determine a device fault, and there are correlations between these various types of information. These correlations can be used to achieve more accurate fault detection; that is, faults determined by multiple types of information can be cross-checked to improve the accuracy of fault detection.
[0140] As an example, the cloud management platform can pre-build the correspondence between the operating data and faults of each type of device. Then, after obtaining the operating data of a device, it can query the correspondence based on the actual obtained operating data, obtain the corresponding fault, and generate the sub-detection results of that type of device.
[0141] As another example, when the cloud management platform determines the corresponding sub-detection results based on the operating data of each device, it can, on the one hand, determine the first candidate fault for that device based on the fault mapping relationship between alarm information and the corresponding fault for that device. The fault mapping relationship includes the mapping relationship between faults and alarm information. For example, similar to in-band fault detection, a fault mapping relationship can be generated in advance based on experience for alarm information and faults of out-of-band devices, i.e., the correspondence between alarm information and faults. Then, when alarm information is actually obtained, the corresponding first candidate fault can be obtained by querying the fault mapping relationship based on the actual alarm information. When there are multiple alarm information, the fault mapping relationship can be queried for each alarm information one by one to obtain the fault corresponding to each alarm information. The first candidate fault can contain the faults corresponding to multiple alarm information. Optionally, the fault mapping relationship can also be implemented using a mapping relationship table. For example, a mapping relationship table can be configured for each device. A mapping relationship table can contain the correspondence between alarm information and faults for a device, where the faults in the mapping relationship table can also be referred to as fault causes or fault types.
[0142] On the other hand, changes in performance metrics can also characterize changes in the operating status of equipment, and operating logs can reflect whether there are any abnormalities in the equipment's operation. Therefore, the cloud management platform can also determine the second candidate fault for each type of equipment based on the change information of each performance metric in at least one performance metric data and / or operating logs. Then, the cloud management platform can determine the sub-detection result corresponding to that type of equipment based on the first and second candidate faults. For example, the cloud management platform can select the same fault in the first and second candidate faults as the final sub-detection result.
[0143] As another example, when determining the corresponding sub-detection results based on the operational data of each device, the cloud management platform can first determine the first candidate fault for that device based on the mapping relationship between alarm information and the corresponding fault. Then, the cloud management platform can determine the confidence level of each fault in the first candidate faults using change information of each performance indicator in at least one performance metric and / or operational logs. If the confidence level of a certain fault is low, it indicates that the fault identified by the alarm information may be inaccurate, and the fault can be removed from the first candidate faults. In other words, the cloud management platform can enhance the fault identified by the alarm information through abrupt changes in performance indicators and the assistance of operational logs, thereby improving the accuracy of the final sub-detection results.
[0144] In addition to the methods mentioned above, an AI model for fault detection can be pre-trained. The AI model can determine the fault location of a certain device based on its operating data. The operating data of a certain device can be input into the AI model to obtain the corresponding sub-detection results. This application does not limit the method of determining the fault.
[0145] Step 204: The cloud management platform determines the target detection result based on the first and second detection results. The target detection result is used to identify K unusable computing nodes out of the M computing nodes. Here, K is an integer.
[0146] The first detection result is determined based on the operational data collected within the computing cluster, while the second detection result is determined based on the operational data collected by the cloud management platform. Since the computing cluster is built using the device resources of the cloud management platform, both results are from different data sources targeting the same device resources for fault detection. Therefore, the first and second detection results should be correlated. Based on this correlation, the first and second detection results can be fused to obtain the final target detection result. Because the target detection result incorporates in-band and out-of-band fault detection results, it can more comprehensively detect faults in the computing cluster, improving the accuracy of fault detection.
[0147] In this embodiment of the application, the first detection result and the second detection result are related, which can be understood as the fault indicated by the first detection result being related to the fault indicated by the second detection result.
[0148] In some embodiments, the faults output by in-band fault detection may include several fault types such as storage faults, network faults, server faults, and software faults. Each type may also include sub-fault types. The fault indicated by the first detection result may be one or more of the fault types, or it may be a sub-fault type under certain fault types. It is understood that storage faults, network faults, and server faults are all associated with out-of-band devices. For example, a storage fault detected in-band indicates that an out-of-band storage device may also be faulty. Therefore, the out-of-band fault detection result can be combined to verify whether the storage device is actually faulty. If the out-of-band detection result also indicates that a certain storage device is faulty, then the probability that it is a storage device fault is higher.
[0149] For example, if the second detection result can be used to indicate multiple faults in the infrastructure, the cloud management platform can determine the confidence level of each fault based on its correlation with the first detection result. That is, the confidence level of a fault determined based on out-of-band operational data can be verified by using the first detection result obtained from in-band operational data.
[0150] In one possible implementation, see Figure 5A The diagram illustrates a flowchart for screening out-of-band faults according to an embodiment of this application. When the cloud management platform determines the second detection result based on out-of-band operational data, it can query a mapping table based on alarm information in the operational data (or in combination with performance index data, operational logs, etc.) to identify various faults. The mapping table can also include the confidence level for each fault, referred to as the first confidence level. In other words, each mapping relationship in the mapping table can be alarm information-fault-confidence level. Thus, each fault corresponding to each alarm information retrieved from the mapping table has a corresponding first confidence level, such as... Figure 5A The faults x1 to xk shown correspond to their respective first confidence levels. It should be noted that although each fault is referred to as having a first confidence level, the first confidence level values can be different for different faults.
[0151] Furthermore, for each of the identified faults, a correlation analysis can be performed based on the relationship between the fault and the first detection result to determine the confidence level of the fault; this confidence level is called the second confidence level. Then, the first and second confidence levels can be combined to obtain the third confidence level; or, the first confidence level can be updated based on the second confidence level to obtain the third confidence level, and this third confidence level is used as the latest confidence level for the fault. Figure 5A As shown, the third confidence levels corresponding to faults x1 to xk can be obtained respectively. Optionally, the third confidence level can be proportional to both the first and second confidence levels. As an example, the product of the first and second confidence levels can be used as the third confidence level.
[0152] As an example, for each of multiple faults, the confidence level (i.e., the second confidence level) of the fault can be determined based on whether it is related to the fault indicated by the first detection result. Taking the first fault among multiple faults as an example, the first fault can be any of the multiple faults. If the first fault is related to some or all of the faults indicated by the first detection result, then the confidence level of the first fault is the first value. For example, if there are faults related to the first fault in the first detection result, then the confidence level of the first fault is the first value; that is, if similar faults are detected both in-band and out-of-band, then the confidence level of the fault is high. As another example, if the first fault is related to at least two faults in the first detection result, then the confidence level of the first fault is the first value. If the first fault is unrelated to the fault indicated by the first detection result, then the confidence level of the first fault is the second value. The first value is greater than the second value. The first fault being unrelated to the fault indicated by the first detection result can be understood as the first fault being unrelated to any fault in the first detection result.
[0153] In this embodiment of the application, determining whether two faults are related can be achieved through the following methods:
[0154] Method b1 can determine whether the two faults are of the same type, and / or whether the computing node indicated in the in-band is associated with the device indicated out of band.
[0155] In this context, whether two faults are of the same type refers to the fact that the two faults belong to the same fault type or subtype. For example, if the first fault is a storage fault, and the in-band indication fault is also a storage fault, then the two faults are determined to be of the same type.
[0156] Determining whether a compute node indicated by an in-band indicator is associated with a device indicated by an out-of-band indicator can be understood as whether the operation of the compute node depends on the corresponding device. For example, if a compute node indicated by an in-band indicator is deployed on a server indicated by an out-of-band indicator, then the compute node is associated with that device. As another example, if a compute node indicated by an in-band indicator is deployed on a server indicated by an out-of-band indicator, and that server is connected to network device 1, then the compute node is associated with network device 1.
[0157] Method b2 can measure the correlation between the first fault and the fault indicated by the first detection result by the similarity between the first fault and the fault indicated by the first detection result. When the similarity is greater than or equal to a specified threshold, it is determined that the first fault and the fault indicated by the first detection result are related. Otherwise, when the similarity is less than the specified threshold, it is determined that the first fault and the fault indicated by the first detection result are not related.
[0158] Besides the methods mentioned above, other methods can also be used to determine relevance, and there are no specific restrictions on these methods.
[0159] Using the methods described above, the cloud management platform can obtain the confidence levels (such as first confidence level, second confidence level, or third confidence level) for multiple faults. Then, based on these confidence levels and the first detection result, the cloud management platform can determine the target detection result. For example, the cloud management platform can identify one or more faults with the highest confidence levels from among multiple faults, and then obtain the final target detection result based on these one or more faults and the fault indicated by the first detection result.
[0160] In this embodiment, the infrastructure managed by the cloud management platform may include at least one device, and the implementation of a computing cluster depends on this at least one device. Therefore, the multiple faults indicated by the second detection result determined by the cloud management platform may include faults of at least one device in the infrastructure. For example, the at least one device may include one or more of servers, storage devices, network devices, or server racks.
[0161] Therefore, the cloud management platform can filter out the most likely faults for each of the at least one devices, and then combine the most likely faults of the at least one device with the first detection result to determine the final target detection result.
[0162] Specifically, J faults can be identified from the multiple faults indicated by the second detection result. These J faults correspond to at least one device, where at least one of the J faults corresponds to a first type of device within the at least one device. The confidence level of this at least one fault is the highest among the faults corresponding to the first type of device. Here, J is a positive integer. In other words, at least one fault with the highest confidence level (such as a first confidence level, a second confidence level, or a third confidence level) can be identified from the faults corresponding to the first type of device. The first type of device can be any type of device within the at least one device. Thus, J faults are obtained, and the cloud management platform can then determine the target detection result based on these J faults and the first detection result.
[0163] The number of at least one fault can be set according to actual needs. For example, the number of at least one fault can be set to 1, 2, or 3, etc., without any specific limitation. Taking the number of at least one fault as an example, if the first type of device is a server, then at least one fault can include the fault with the highest confidence among the faults corresponding to the server. Similarly, if the first type of device is a storage device, then at least one fault can include the fault with the highest confidence among the faults corresponding to the storage device. In this way, the fault with the highest confidence can be determined from each type of device, and then the final target detection result can be determined based on the fault with the highest confidence for each type of device and the first detection result.
[0164] For example, see Figure 5B The diagram illustrates another process for screening out-of-band faults according to an embodiment of this application. At least one device may include a server, storage device, network device, and server rack. The cloud management platform can determine candidate faults corresponding to each device, and each device may have one or more candidate faults. Furthermore, for each device, at least one fault with the highest confidence level can be selected, such as... Figure 5B The J faults are defined as follows: the highest confidence level fault of the server, the highest confidence level fault of the storage device, the highest confidence level fault of the network device, and the highest confidence level fault of the rack.
[0165] The cloud management platform can combine pre-defined rules to determine the target detection result based on one or more out-of-band faults and the first in-band detection result. The pre-defined rules can be used to adjust the confidence levels of both the out-of-band and in-band faults, and then sort all the adjusted faults according to their confidence levels. The one or more faults with the highest confidence levels are then output as the final target detection result. It can be understood that the one or more faults here can be the one or more faults determined based on the correlation relationship, the faults with the highest confidence levels for each of the aforementioned devices (i.e., J faults), or a combination thereof.
[0166] Taking one or more faults as J faults as an example, the cloud management platform can determine the target detection results in the following ways:
[0167] In method c1, the cloud management platform can use a knowledge graph to determine the target detection result. Generally, a knowledge graph is a structured semantic knowledge base that can be used to describe concepts and their relationships in the physical world in symbolic form. The basic building blocks of a knowledge graph are "entity-relationship-entity" triples, as well as entities and their related attribute-value pairs. Entities are interconnected through relations, forming a network-like knowledge structure. In this embodiment, the knowledge graph may include multiple knowledge nodes, which can correspond to the faults involved in this embodiment. One knowledge node can correspond to one type of fault. For example, multiple knowledge nodes can correspond to the fault indicated by the first detection result and J faults. The knowledge graph can also represent the association relationships between multiple knowledge nodes. Therefore, based on the association relationships between the J faults in the knowledge graph and the fault indicated by the first detection result, the most likely fault can be determined, and the target detection result can be obtained.
[0168] As an example, the cloud management platform can adjust the confidence levels of J faults and the faults indicated by the first detection result based on the correlation between J faults in the knowledge graph and the faults indicated by the first detection result. Finally, it sorts all faults according to their confidence levels and selects one or more faults with the highest confidence levels as the final target detection result. For example, it can start searching the knowledge graph from a specified fault in the first detection result. The confidence levels of other faults with closer correlations to the specified fault can be increased, while the confidence levels of other faults with less correlations can be decreased. Similarly, this process can be performed on other faults indicated by the first detection result until all faults in the first detection result have been traversed. Optionally, when adjusting the confidence levels of other faults based on different faults in the first detection result, the adjustment magnitude can be different. For example, when the confidence level of a fault is higher, the adjustment magnitude for other faults can be larger. Alternatively, the adjustment magnitude for each fault can be pre-configured based on experience.
[0169] As another example, knowledge graphs can also be used for reasoning based on existing knowledge. For instance, in this embodiment, a knowledge graph can be used to reason about the most likely fault given a fault condition. Then, the cloud management platform can determine the target detection result based on J faults and the fault indicated by the first detection result, combined with the reasoning capabilities of the knowledge graph.
[0170] In method c2, the cloud management platform can use an AI model to determine the target detection result. That is, a large number of training samples can be collected in advance to train the AI model. Each training sample can, for example, contain multiple faults, including faults that can be identified in-band and / or out-of-band. The AI model can be trained in a supervised manner, in which case each training sample can be labeled with the most likely fault given the presence of multiple faults. Alternatively, the AI model can be trained unsupervised, without restriction. After training, the AI model can be used to determine the target detection result. This is achieved by inputting J faults and the fault indicated by the first detection result into the AI model to obtain the target detection result output by the AI model.
[0171] In this embodiment, the target detection result can be used to determine K unusable computing nodes out of M computing nodes in a computing cluster. The target detection result may include one or more of the following information:
[0172] (1) Information indicating faulty devices in the cloud management platform. For example, this information could be the identifier of the faulty device.
[0173] The faulty device could be one or more of the following: a server, a storage device, a network device, or a server rack. For a specific type of device, it could also indicate which device experienced what type of fault. For example, it could indicate fault A in server 1, fault B in storage device, etc.
[0174] Understandably, unavailable devices can include not only the faulty device itself, but also other devices that depend on it. For example, if network device 1 fails, then servers that rely on network device 1 for data transmission will also be unavailable. Therefore, after identifying the faulty device, other devices that depend on it for operation can be identified as unavailable based on the network topology. Ultimately, unavailable servers can be identified based on the unavailable devices, and the computing nodes deployed on these servers will also be unavailable, thus identifying the unavailable computing nodes.
[0175] (2) Information indicating unavailable devices in the cloud management platform. For example, this information could be an identifier of an unavailable device. Unavailable devices can be identified, for example, in the manner described above.
[0176] (3) Information indicating unavailable compute nodes. For example, this information could be the identifiers of K unavailable compute nodes. Alternatively, if a compute node may contain an operating system, then this information could also be the identifier of the operating system on the unavailable compute nodes.
[0177] Based on the above implementation methods, the cloud management platform can achieve automatic fault detection without manual intervention, thus improving fault detection efficiency. Furthermore, data can be transferred between the cloud management platform and cloud tenants, breaking down the data isolation barriers between them. This allows the cloud management platform to combine the first in-band detection result with the second out-of-band detection result to determine the final target detection result, resulting in more comprehensive fault detection data and improved accuracy.
[0178] After obtaining the target detection results, the cloud management platform can be used to assist both cloud tenants and the cloud management platform in quickly resolving their respective problems. For example, for cloud tenants, steps 205 and 206 can be executed, allowing the computing cluster to continue executing computing tasks after removing unavailable computing nodes.
[0179] Step 205: The cloud management platform sends the target detection results to the computing cluster. Correspondingly, the computing cluster receives the target detection results.
[0180] Sending target detection results from the cloud management platform to the computing cluster can be understood as sending target detection results to some or all of the M computing nodes in the computing cluster.
[0181] As an example, the cloud management platform can send target detection results to the computing cluster through a first interface. For example... Figure 1C In the system architecture, after the out-of-band CCAE performs fault detection, the CCAE calls the openAPI to send the target detection results to the in-band plugin, so as to send the target detection results to the computing nodes of the computing cluster.
[0182] Step 206: The computing cluster continues to execute computing tasks based on the target detection results.
[0183] The computing cluster can identify K unavailable computing nodes based on object detection results, and then remove these K unavailable computing nodes from the M computing nodes before continuing the computing task. For example, the computing cluster can continue the computing task using the remaining (MK) computing nodes. In this case, the computing cluster can reallocate the computing subtasks to be executed by each computing node. Alternatively, the computing cluster can add K new computing nodes to take over the computing task from the previously unavailable K computing nodes. As an example, if the computing task is a model training task, then the model can be resumed from the breakpoint after removing the K unavailable computing nodes.
[0184] In some embodiments, if only some computing nodes receive the target detection result, one or more of these computing nodes can determine whether to continue executing the computing task based on the target detection result, and notify other computing nodes that need to continue executing the computing task to continue executing the computing task, or notify the computing nodes that have stopped executing the computing task (i.e., the K unavailable computing nodes) to stop executing the computing task.
[0185] In some embodiments, if only some computing nodes receive the target detection results, these computing nodes can broadcast the target detection results in the computing cluster. Then each computing node can determine whether to continue executing the computing task based on whether it is an unavailable node.
[0186] In some embodiments, some or all of the computing nodes in the computing cluster may include a second computing node, meaning the cloud management platform can send target detection results to the second computing node. When the second computing node is one of the M computing nodes excluding the K computing nodes, i.e., the second computing node is an available computing node, it can continue to execute computing tasks based on the target detection results. It is understood that some or all of the computing nodes in the computing cluster may also include other nodes besides the second computing node. If the second computing node is an unavailable node, then the second computing node will no longer continue to execute computing tasks, or an alarm may be triggered, such as notifying the cloud tenant to perform relevant checks (e.g., in the case of a software-related fault).
[0187] Based on the above embodiments, since the fault detection method of this application is more efficient and reduces the time required for fault detection, the computing task can be resumed as soon as possible, minimizing the losses to cloud tenants caused by task interruption. Furthermore, the fault detection method of this application has higher accuracy, allowing for more precise identification of unavailable computing nodes, thereby reducing the probability of further interruption of the computing task.
[0188] For example, for a cloud management platform, the following step 207 can be performed so that the cloud management platform's operations and maintenance personnel can repair the faulty equipment in a timely manner.
[0189] Step 207: The cloud management platform provides a fault display interface, which is used to display the target detection results.
[0190] For example, the cloud management platform can send target detection results to the accounts of management users or operation and maintenance personnel on the cloud management platform. Management users or operation and maintenance personnel can open the fault display interface in the corresponding client to view the target detection results. Then, management users can assign operation and maintenance personnel to repair the fault, or operation and maintenance personnel can view the target detection results and repair the faulty equipment in a timely manner.
[0191] It is understandable that some of the above steps are optional, therefore... Figure 2 It is shown in dashed lines.
[0192] The following describes the solution provided in the embodiments of this application, taking intelligent computing as an example. (See also...) Figure 6 The diagram shown illustrates another flowchart of the method provided in this application embodiment. Depending on the functions of the cloud management platform, the cloud management platform may include an in-band fault detection subsystem, an out-of-band fault detection subsystem, and an intelligent computing cluster. The intelligent computing cluster is used to execute training tasks; for example, it can be the aforementioned computing cluster. Here, we take an intelligent computing cluster comprising M OSs as an example, where one OS can be considered a training node. The in-band fault detection subsystem is used to implement in-band fault detection. For example, this in-band fault detection subsystem can be used to implement… Figure 2 The embodiments shown illustrate the functionality implemented by the computing nodes (such as the second computing node). As an example, the in-band fault detection subsystem may consist of fault detection software within the OS included in the intelligent computing cluster, or it may consist of in-band plugins within the OS included in the intelligent computing cluster. The out-of-band fault detection subsystem is used to implement in-band fault detection; for example, this out-of-band fault detection subsystem can be used to implement… Figure 2 The cloud management platform implements the functions shown in the embodiment.
[0193] See Figure 6 As shown, the method may include the following steps:
[0194] Step 601: The in-band fault detection subsystem determines the first detection result.
[0195] Cloud tenants or cloud management platform operations personnel can start an AI model training task on the intelligent computing cluster. If the training task is interrupted due to an anomaly, the in-band fault detection subsystem can detect the interruption. The in-band fault detection subsystem can then determine the first detection result based on the in-band collected runtime data.
[0196] For example, the in-band fault detection subsystem can search for faulty OSes starting from those that directly report errors, based on the time of training task interruption, the start time of training task, and information about the OSes participating in training (such as an OS list), until all error information is found and summarized. This allows the system to identify potentially faulty OSes within the intelligent computing cluster and their error information. Furthermore, a mapping table can be pre-built based on expert experience, establishing a relationship between different error messages and root causes. This mapping table includes the relationships between error messages and root causes, allowing the identification of the root cause corresponding to each error message. A confidence level can be configured for each mapping relationship based on expert experience. The confidence level of each root cause can then be obtained from the mapping table, and one or more root causes with the highest confidence levels, along with their corresponding OSes, can be selected to obtain the first detection result.
[0197] Step 602: The in-band fault detection subsystem sends the first detection result to the out-of-band fault detection subsystem. Additionally, the in-band fault detection subsystem can also send the time of training task interruption, the time of training task start, and information about the OSs participating in the training to the out-of-band fault detection subsystem. For example, the in-band fault detection subsystem can... Figure 1C The openAPI shown sends data to the out-of-band fault detection subsystem.
[0198] Step 603: The out-of-band fault detection subsystem obtains information about the servers where the participating OSs reside based on their information. This allows it to locate operational data related to these services (referring to out-of-band collected operational data). Furthermore, by analyzing the time of training task interruption and the start time of the training task, it can filter out the server's operational data during the training task's execution time. For example, this operational data may include alarm information, performance metrics, and log information.
[0199] Step 604: The out-of-band fault detection subsystem compares in-band and out-of-band data, that is, the server's operating data obtained in step 603 is compared with the OS's fault root cause received in step 602, to determine the first sub-detection result related to the server.
[0200] In one example, the out-of-band fault detection subsystem can first perform fault detection based on the server's operating data, determine the root causes of the faults related to the server (such as the faulty server and the root causes of these server faults), and then compare these root causes of the faults with the root causes of the faults of the OS received in step 602 to obtain the first sub-detection result related to the server.
[0201] The out-of-band fault detection subsystem can first perform fault detection based on the server's operating data. This can be done by querying the mapping table corresponding to the server based on the operating data to determine the server that has failed, the root cause of the failure, and the first confidence level of each root cause. For each root cause of failure on each server, the second confidence level can be determined by comparing it with the root causes of failure in the operating system and determining whether there is a correlation between the two. Finally, the first and second confidence levels are combined to obtain the final confidence level of each root cause of failure on each server.
[0202] In one example, the server's operational data can also be directly compared with the OS's root cause of failure received in step 602. The comparison result can be one of two things: either the server's operational data contains data related to the OS's root cause of failure (or OS error messages), in which case the confidence (or weight) of this related data is high; or the server's operational data does not contain data related to the OS's root cause of failure (or OS error messages), in which case the confidence (or weight) of this unrelated data is low.
[0203] In some embodiments, the process of step 604 can also be executed by the device resources included in the intelligent computing cluster. If it is implemented by the device resources included in the intelligent computing cluster, then step 605 can be executed. Figure 6 This will be illustrated using this example.
[0204] Step 605: Send the first sub-detection result to the out-of-band fault detection subsystem.
[0205] Step 606: The out-of-band fault detection subsystem obtains the server information of the participating OSs based on their information. By using the server information as input, it can find the information of the network devices associated with these servers, and then find the operational data of these network devices. Furthermore, by using the time when the training task was interrupted and the time when the training task started, it can filter out the operational data of these network devices during the training task's execution time. For example, this operational data may include alarm information, performance indicator data, and log information.
[0206] Step 607: The out-of-band fault detection subsystem compares in-band and out-of-band data, that is, compares the network device's operating data obtained in step 606 with the OS's fault root cause received in step 602, to determine the second sub-detection result related to the network device.
[0207] In one example, the out-of-band fault detection subsystem can first perform fault detection based on the network device's operating data, determine the root cause of the fault related to the network device (such as the network device that failed and the root cause of the fault in these network devices), and then compare these root causes of the fault with the root cause of the fault in the OS received in step 602 to obtain the second sub-detection result related to the network device.
[0208] The out-of-band fault detection subsystem can first perform fault detection based on the network device's operational data. This can be done by querying the mapping table corresponding to the network devices based on the operational data to determine the faulty network devices, the root causes of these network device faults, and the first confidence level of each root cause. For each root cause of fault in each network device, a comparison can be made with the root causes of fault in the operating system. Based on the correlation between the two, a second confidence level for each root cause of fault in each network device can be determined. Finally, by combining the first and second confidence levels, the final confidence level for each root cause of fault in each network device can be obtained.
[0209] In one example, the network device's operational data can also be directly compared with the OS's root cause of failure received in step 602. The comparison result can be one of two possibilities: either the network device's operational data contains data related to the OS's root cause of failure (or OS error messages), in which case the confidence (or weight) of this related data is high; or the network device's operational data does not contain data related to the OS's root cause of failure (or OS error messages), in which case the confidence (or weight) of this unrelated data is low.
[0210] In some embodiments, the process of step 607 can also be executed by the device resources included in the intelligent computing cluster. If it is implemented by the device resources included in the intelligent computing cluster, then step 608 can continue to be executed.
[0211] Step 608: Send the second sub-detection result to the out-of-band fault detection subsystem.
[0212] Step 609: The out-of-band fault detection subsystem obtains information about the servers hosting the participating OSes based on their information. By using the server information as input, it can locate the information of the storage devices associated with these servers, and then find the operational data of these storage devices. Furthermore, by using the time the training task was interrupted and the time the training task started, it can filter out the operational data of these storage devices during the training task's execution time. For example, this operational data may include alarm information, performance indicator data, and log information.
[0213] Step 610: The out-of-band fault detection subsystem compares in-band and out-of-band data, that is, compares the operating data of the storage device obtained in step 609 with the root cause of the OS fault received in step 602, to determine the third sub-detection result related to the storage device.
[0214] In one example, the out-of-band fault detection subsystem can first perform fault detection based on the operating data of the storage device, determine the root cause of the fault related to the storage device (such as the storage device that failed and the root cause of the fault of these storage devices), and then compare these root causes of the fault with the root cause of the fault of the OS received in step 602 to obtain the third sub-detection result related to the storage device.
[0215] The out-of-band fault detection subsystem can first perform fault detection based on the operating data of the storage devices. This can be done by querying the mapping table corresponding to the storage devices based on the operating data to determine the faulty storage devices, the root causes of these storage devices' faults, and the first confidence level of each root cause. For each root cause of fault in each storage device, a comparison can be made with the root causes of faults in the operating system. Based on the correlation between the two, a second confidence level for each root cause of fault in each storage device can be determined. Finally, by combining the first and second confidence levels, the final confidence level for each root cause of fault in each storage device can be obtained.
[0216] In one example, the operating data of the storage device can also be directly compared with the root cause of the OS failure received in step 602. The comparison result can be one of two possibilities: either the operating data of the storage device contains data related to the root cause of the OS failure (or error information of the OS), in which case the confidence (or weight) of this related data is high; or the operating data of the storage device does not contain data related to the root cause of the OS failure (or error information of the OS), in which case the confidence (or weight) of this unrelated data is low.
[0217] In some embodiments, the process of step 610 can also be executed by the device resources included in the intelligent computing cluster. If it is implemented by the device resources included in the intelligent computing cluster, then step 611 can be executed.
[0218] Step 611: Send the third sub-detection result to the out-of-band fault detection subsystem.
[0219] Step 612: The out-of-band fault detection subsystem obtains the server information of the participating OSs based on their information. By using the server information as input, it can find the information of the racks associated with these servers, and then find the operational data of these racks. Furthermore, by using the time when the training task was interrupted and the time when the training task started, it can filter out the operational data of these racks during the training task's execution time. For example, this operational data may include alarm information, performance indicator data, and log information.
[0220] Step 613: The out-of-band fault detection subsystem compares the in-band and out-of-band data, that is, the cabinet operation data obtained in step 612 is compared with the OS fault root cause received in step 602, to determine the fourth sub-detection result related to the cabinet.
[0221] In one example, the out-of-band fault detection subsystem can first perform fault detection based on the rack's operating data, determine the root causes of faults related to the rack (such as the rack that failed and the root causes of these racks' faults), and then compare these root causes of faults with the root causes of faults of the OS received in step 602 to obtain the fourth sub-detection result related to the rack.
[0222] The out-of-band fault detection subsystem can first perform fault detection based on the rack's operational data. This can be done by querying the mapping table corresponding to the rack based on the operational data to determine the rack that has failed, the root cause of the failure in these racks, and the first confidence level of each root cause. For each root cause of failure in each rack, it can be compared with the root cause of failure in the OS. Based on the correlation between the two, the second confidence level of each root cause of failure in each rack can be determined. Finally, by combining the first and second confidence levels, the final confidence level of each root cause of failure in each rack can be obtained.
[0223] In one example, the rack's operational data can also be directly compared with the OS's root cause of failure received in step 602. The comparison result can be one of two possibilities: either the rack's operational data contains data related to the OS's root cause of failure (or OS error messages), in which case the confidence (or weight) of this related data is high; or the rack's operational data does not contain data related to the OS's root cause of failure (or OS error messages), in which case the confidence (or weight) of this unrelated data is low.
[0224] In some embodiments, the process of step 613 can also be executed by the device resources included in the intelligent computing cluster. If it is implemented by the device resources included in the intelligent computing cluster, then step 614 can be executed.
[0225] Step 614: Send the fourth sub-detection result to the out-of-band fault detection subsystem.
[0226] It is understandable that steps 603-605, 606-608, 609-611, and 612-614 above correspond to the fault detection processes for servers (i.e., out-of-band computing devices), network devices, storage devices, and server racks, respectively. In specific implementations, fault detection can be selectively performed on one or more of the servers, network devices, storage devices, and server racks. Therefore, the fault detection processes of steps 603-605, 606-608, 609-611, and 612-614 above can be combined accordingly. For example, when fault detection is required for servers, network devices, storage devices, and server racks, the fault detection processes of steps 603-605, 606-608, 609-611, and 612-614 above can all be executed. For example, when fault detection is required for servers and cabinets, the fault detection process described in steps 603-605 and 612-614 above can be executed. As another example, when fault detection is required for servers, network devices, and cabinets, the fault detection process described in steps 603-605, 606-608, and 612-614 above can be executed.
[0227] Step 615: The out-of-band fault detection subsystem determines the target detection result based on the first detection sub-result, the second detection sub-result, the third detection sub-result, and the fourth detection sub-result.
[0228] In one implementation, as described above, if the first detection sub-result includes the failed servers, their respective root causes, and confidence levels, then they can be sorted by confidence level to obtain the root cause with the highest confidence level and its corresponding server. Similarly, if the second detection sub-result includes the failed network devices, their respective root causes, and confidence levels, then they can be sorted by confidence level to obtain the root cause with the highest confidence level and its corresponding network device. This process can also be repeated to obtain the root causes with the highest confidence levels for storage devices and server racks.
[0229] In one implementation, if the first detection sub-result includes one or more operational data points and their confidence levels (or weights) for each server, they can be sorted by confidence level, and the operational data of the server with the highest confidence level can be selected. Then, based on the operational data of the server with the highest confidence level and the corresponding mapping table, the root cause of the fault can be determined. For example, the operational data may include alarm information, performance indicator data, and operational logs, and the mapping table may include the mapping relationship between alarm information and the root cause of the fault. In one example, the corresponding root cause of the fault can be obtained by querying the mapping table based on the alarm information included in the operational data, and the root cause of the fault determined by the alarm information can be filtered using changes in performance indicators and operational logs to select one or more root causes of the fault with the highest probability. In another example, the alarm information can be enhanced using changes in performance indicators and operational logs to select one or more alarm information points with higher accuracy, and then the corresponding root cause of the fault can be obtained by querying the mapping table based on these one or more alarm information points.
[0230] Similarly, the root causes of network device failures, storage device failures, and server rack failures can be identified using the methods described above.
[0231] Furthermore, the root causes of server, network device, storage device, and rack failures obtained above, along with the first detection results obtained within the band, can be combined to make a comprehensive judgment and obtain the target detection result. For example, the target detection result can be obtained through the aforementioned knowledge graph or AI model.
[0232] The out-of-band fault detection subsystem can determine the faulty device and / or the faulty training node (or OS) in the intelligent computing cluster based on the root causes of the faults of the server, network device, storage device and cabinet obtained above, and the first detection results obtained in the in-band. It can then determine the unusable training node (or OS) in the intelligent computing cluster.
[0233] Step 616: The out-of-band fault detection subsystem sends the target detection result to the in-band fault detection subsystem. This target detection result can be used to identify unavailable training nodes (or operating systems) within the intelligent computing cluster. For example, if the target detection result indicates an unavailable training node (or operating system) within the intelligent computing cluster, the intelligent computing cluster can resume training after removing the unavailable training node. Optionally, the target detection result sent by the out-of-band fault detection subsystem to the in-band fault detection subsystem may include part or all of the information in the target detection result determined by the out-of-band fault detection subsystem.
[0234] In addition, based on the information on the root cause of the fault and the faulty equipment obtained from the out-of-band fault detection subsystem, maintenance personnel can promptly repair the faults of the equipment related to the intelligent computing cluster.
[0235] In the above descriptions, the focus is on transmitting in-band fault detection results to out-of-band for further fault detection. In some embodiments, out-of-band fault detection results can also be transmitted to in-band for further fault detection. See [link to documentation]. Figure 7 The diagram shown is a flowchart of a fault detection method for a computing cluster provided in an embodiment of this application. The method may include the following steps:
[0236] Step 701: The cloud management platform determines the second detection result, which is used to indicate the failure of the infrastructure that caused the computing task to be interrupted.
[0237] Step 702: The cloud management platform sends the second detection result to the second computing node. Correspondingly, the second computing node receives the second detection result. The second computing node can be some or all of the M computing nodes in the computing cluster.
[0238] Step 703: The second computing node determines the first detection result, which is used to indicate the fault of the first computing node. The first computing node is the node among the M computing nodes that caused the computing task to be interrupted.
[0239] Step 704: The second computing node determines the target detection result based on the second detection result and the first detection result. The target detection result is used to determine the K unusable computing nodes among the M computing nodes.
[0240] Step 705: The second computing node continues to execute the computing task based on the target detection result. For example, if the second computing node is one of the M computing nodes excluding the K computing nodes, the second computing node continues to execute the computing task.
[0241] In one possible implementation, after removing K unavailable compute nodes, the computation task can continue to be executed based on the remaining compute nodes, or new compute nodes can be added, or the computation task can continue to be executed based on the added compute nodes and the remaining compute nodes. The number of added compute nodes can be one or more, and there is no limitation on this. For example, a hot standby node can be configured for the compute cluster, and new compute nodes can be added from the hot standby nodes.
[0242] Step 706: The second computing node sends the target detection results to the cloud management platform.
[0243] Step 707: Based on the target detection results, the cloud management platform determines the faults of the devices related to the K computing nodes among the devices included in the infrastructure.
[0244] In this way, the cloud management platform's operations and maintenance personnel can promptly repair faults in the computing cluster's related equipment.
[0245] It is understood that the steps performed by the aforementioned second computing node can be executed by a single second computing node or by multiple second computing nodes working collaboratively. The aforementioned steps are similar to... Figure 2 The processing in the corresponding embodiments is similar; therefore, for the specific implementation and beneficial effects of the above steps, please refer to [link / reference needed]. Figure 2 The descriptions in the corresponding embodiments will not be repeated here.
[0246] In summary, this application addresses the high failure rate and persistent faults when resuming training after interruption by providing a fault detection method. This method combines in-band and out-of-band methods for comprehensive fault detection, improving accuracy. For example, in a public cloud scenario, an open API can be published on the public cloud user interface. Public cloud users can call this open API to transmit general information to the backend operations and maintenance platform. The operations and maintenance platform automatically performs global indexing and fault detection, identifies the root cause of the fault, and automatically returns the information to the public cloud user, achieving automatic root cause identification.
[0247] Based on the same inventive concept, embodiments of this application also disclose a fault detection device for a computing cluster, which can be applied to a cloud management platform. This cloud management platform is used to manage the infrastructure providing cloud services, including servers within at least one cloud data center. Figure 8 This is a schematic diagram of the structure of a fault detection device for a computing cluster provided in an embodiment of this application, as shown below. Figure 8 As shown, the device includes:
[0248] The receiving module 801 is used to receive a first detection result, which is used to indicate a fault in the first computing node. The first computing node is the node that caused the interruption of the computing task among M computing nodes. The M computing nodes are used to execute computing tasks, M is a positive integer, and the M computing nodes are deployed on at least one server.
[0249] The determination module 802 is used to determine the target detection result based on the first detection result and the second detection result. The second detection result is used to indicate the failure of the infrastructure, and the target detection result is used to determine the K unavailable computing nodes out of the M computing nodes, where K is an integer.
[0250] In one possible implementation, the first detection result is determined based on the operating data of some or all of the M computing nodes, and the second detection result is determined based on the operating data of some or all of the equipment included in the infrastructure.
[0251] In one possible implementation, the first detection result receiving module 801 receives the result through a first interface between the cloud management platform and the second computing node among the M computing nodes. The first interface is between at least one of the M computing nodes and the cloud management platform. The second computing node belongs to at least one computing node and is the first computing node, or the second computing node is any of the other computing nodes among the M computing nodes besides the first computing node.
[0252] In one possible implementation, each of the at least one computing node includes an operating system, and the operating system of some or all of the at least one computing node includes a first plugin for identifying a faulty computing node among the M computing nodes.
[0253] In one possible implementation, the target detection results obtained by the determination module 802 can have various applications. For example, the device further includes a display module 803 for providing a fault display interface for displaying the target detection results. As another example, the device further includes a sending module 804 for sending the target detection results to some or all of the M computing nodes, so that the computing task can continue to be executed by the other computing nodes besides K of the M computing nodes.
[0254] In one possible implementation, the second detection result is used to indicate multiple faults in the infrastructure. When the determining module 802 determines the target detection result based on the first and second detection results, the determining module 802 can determine the target detection result based on the confidence levels of the multiple faults and the first detection result. The confidence levels of the multiple faults are determined based on the correlation between the multiple faults and the first detection result.
[0255] In one possible implementation, the determining module 802 can determine the confidence level of an out-of-band detected fault based on whether the in-band detected fault is related to the out-of-band detected fault. For example, if the first fault among the multiple faults indicated by the second detection result is related to some or all of the faults indicated by the first detection result, the confidence level of the first fault is a first value. Or, if the first fault among the multiple faults indicated by the second detection result is unrelated to the faults indicated by the first detection result, the confidence level of the first fault is a second value. Wherein, the first value is greater than the second value.
[0256] In one possible implementation, the multiple faults include faults in at least one device in the infrastructure, which includes one or more of a server, storage device, network device, or cabinet. When determining the target detection result based on the confidence levels of the multiple faults and the first detection result, the determining module 802 can determine the target detection result based on J faults from the multiple faults and the first detection result. The J faults are faults corresponding to at least one device, wherein at least one of the J faults corresponds to a first type of device among the at least one device, and the confidence level corresponding to the at least one fault is the highest confidence level among the faults corresponding to the first type of device, where K is a positive integer.
[0257] In one possible implementation, the determining module 802 determines the target detection result based on J faults out of a plurality of faults and the first detection result, which can be implemented in various ways. One implementation is that the determining module 802 can determine the target detection result based on the J faults, the first detection result, and a knowledge graph, where the knowledge graph represents the relationships between multiple knowledge nodes, and the multiple knowledge nodes correspond to the faults indicated by the first detection result and the J faults. Another implementation is that the determining module 802 can input the J faults and the first detection result into an artificial intelligence model to obtain the target detection result output by the artificial intelligence model.
[0258] In one possible implementation, the receiving module 801 is further configured to receive second information. The second information indicates M computing nodes and / or indicates the runtime of a computing task. The determining module 802 can determine at least one server where the M computing nodes indicated by the second information reside, and determine a second detection result based on runtime data associated with the at least one server. The runtime data associated with the at least one server is obtained based on the runtime.
[0259] In one possible implementation, the operational data associated with at least one server includes operational data of at least one device associated with at least one server. When the determining module 802 determines the second detection result based on the operational data associated with at least one server, it can determine the second detection result based on the operational data of at least one device. The second detection result includes sub-detection results corresponding to each of the at least one device, and each sub-detection result is used to indicate a fault of one device. The at least one device includes one or more of servers, storage devices, network devices, and server racks.
[0260] In one possible implementation, the operating data of each of the at least one devices includes one or more of alarm information, at least one performance indicator data, or operating logs. When the determining module 802 determines the second detection result based on the operating data of the at least one device, it may perform the following steps for each of the at least one devices: determine a first candidate fault corresponding to each device based on the fault mapping relationship between alarm information and the fault corresponding to each device, wherein the fault mapping relationship includes the mapping relationship between faults and alarm information; determine a second candidate fault corresponding to each device based on the change information of each performance indicator in the at least one performance indicator data, and / or operating logs; and determine a sub-detection result corresponding to each device based on the first candidate fault and the second candidate fault.
[0261] In one possible implementation, the computational task is a model training task.
[0262] This device can be used to perform, for example Figure 2 The method steps performed by the cloud management platform in the illustrated embodiment are described below. Therefore, for a description of the device, please refer to [link / reference needed]. Figure 2 The description of the illustrated embodiments will not be repeated here.
[0263] Based on the same inventive concept, this application also discloses another fault detection device for a computing cluster. This device can be applied to a second computing node, which belongs to M computing nodes used to perform computing tasks, where M is a positive integer and the M computing nodes are deployed on at least one server. Figure 9 This is a schematic diagram of the structure of a fault detection device for a computing cluster provided in an embodiment of this application, as shown below. Figure 9 As shown, the device includes:
[0264] The sending module 901 is used to send a first detection result to the cloud management platform. The first detection result is used to indicate the failure of the first computing node. The first computing node is the node among M computing nodes that caused the interruption of the computing task. The cloud management platform is used to manage the infrastructure that provides cloud services. The infrastructure includes at least one server.
[0265] The receiving module 902 is used to receive the target detection results sent by the cloud management platform. The target detection results are used to determine the K unusable computing nodes out of the M computing nodes, where K is an integer.
[0266] The execution module 903 is used to continue the calculation task based on the target detection result, wherein the second calculation node is the calculation node other than the K calculation nodes among the M calculation nodes.
[0267] In one possible implementation, the computational task is a model training task.
[0268] This device can be used to perform, for example Figure 2 The method steps performed by the computing cluster or the second computing node in the illustrated embodiment are described below. Therefore, a description of the device can be found by referring to [reference needed]. Figure 2 The description of the illustrated embodiments will not be repeated here.
[0269] Based on the same inventive concept, this application also discloses a fault detection device for a computing cluster. This device can be applied to a second computing node, which belongs to M computing nodes used to perform computing tasks, where M is a positive integer and the M computing nodes are deployed on at least one server. Figure 10 This is a schematic diagram of the structure of a fault detection device for a computing cluster provided in an embodiment of this application, as shown below. Figure 10 As shown, the device includes:
[0270] The receiving module 1001 is used to receive a second detection result sent by a cloud management platform, the cloud management platform being used to manage the infrastructure providing cloud services, the infrastructure including the at least one server, and the second detection result being used to indicate a failure of the infrastructure that caused the computing task to be interrupted.
[0271] The determination module 1002 is used to determine a target detection result based on the second detection result and the first detection result. The first detection result is used to indicate a fault in the first computing node, which is the node among the M computing nodes that caused the interruption of the computing task. The target detection result is used to determine K unusable computing nodes among the M computing nodes, where K is an integer.
[0272] The execution module 1003 is used to continue executing the calculation task based on the target detection result, wherein the second calculation node is a calculation node other than the K calculation nodes among the M calculation nodes.
[0273] In one possible implementation, the first detection result is determined based on the operating data of some or all of the M computing nodes, and the second detection result is determined based on the operating data of some or all of the equipment included in the infrastructure.
[0274] In one possible implementation, the first detection result is received through a first interface between the cloud management platform and a second computing node among the M computing nodes, wherein at least one of the M computing nodes has a first interface with the cloud management platform, the second computing node belongs to at least one computing node, the second computing node is the first computing node, or the second computing node is another computing node among the M computing nodes besides the first computing node.
[0275] In one possible implementation, each of the at least one computing node includes an operating system, and the operating system of some or all of the at least one computing node includes a first plugin for identifying a faulty computing node among the M computing nodes.
[0276] In one possible implementation, the device may further include a sending module 1004 for sending the target detection results obtained by the determining module 1002 to the cloud management platform.
[0277] In one possible implementation, the second detection result is used to indicate multiple faults in the infrastructure. When the determining module 1002 determines the target detection result based on the first and second detection results, it can determine the target detection result based on the confidence levels of the multiple faults and the first detection result. The confidence levels of the multiple faults are determined based on the correlation between the multiple faults and the first detection result.
[0278] In one possible implementation, the determining module 1002 can determine the confidence level of an out-of-band detected fault based on whether the in-band detected fault is related to the out-of-band detected fault. For example, if the first fault among the multiple faults indicated by the second detection result is related to some or all of the faults indicated by the first detection result, the confidence level of the first fault is a first value. Or, if the first fault among the multiple faults indicated by the second detection result is unrelated to the faults indicated by the first detection result, the confidence level of the first fault is a second value. Wherein, the first value is greater than the second value.
[0279] In one possible implementation, the multiple faults include faults in at least one device in the infrastructure, which includes one or more of a server, storage device, network device, or cabinet. When determining the target detection result based on the confidence levels of the multiple faults and the first detection result, the determining module 1002 can determine the target detection result based on J faults among the multiple faults and the first detection result. The J faults are faults corresponding to at least one device, wherein at least one of the J faults corresponds to a first type of device among the at least one device, and the confidence level corresponding to the at least one fault is the highest confidence level among the faults corresponding to the first type of device, where K is a positive integer.
[0280] In one possible implementation, the determining module 1002 determines the target detection result based on J faults out of a plurality of faults and the first detection result, which can be implemented in various ways. One implementation is that the determining module 1002 can determine the target detection result based on the J faults, the first detection result, and a knowledge graph, where the knowledge graph represents the relationships between multiple knowledge nodes, and the multiple knowledge nodes correspond to the faults indicated by the first detection result and the J faults. Another implementation is that the determining module 1002 can input the J faults and the first detection result into an artificial intelligence model to obtain the target detection result output by the artificial intelligence model.
[0281] In one possible implementation, the computational task is a model training task.
[0282] This device can be used to perform, for example Figure 7 The method steps performed by the computing cluster or the second computing node in the illustrated embodiment are described below. Therefore, a description of the device can be found by referring to [reference needed]. Figure 7 The description of the illustrated embodiments will not be repeated here.
[0283] Based on the same inventive concept, embodiments of this application also disclose a fault detection device for a computing cluster, which can be applied to a cloud management platform. This cloud management platform is used to manage the infrastructure providing cloud services, including servers within at least one cloud data center. Figure 11 This is a schematic diagram of the structure of a fault detection device for a computing cluster provided in an embodiment of this application, as shown below. Figure 11 As shown, the device includes:
[0284] The sending module 1101 is used to send a second detection result to some or all of the M computing nodes, where M is a positive integer and the M computing nodes are deployed on at least one server. The second detection result is used to indicate the failure of the infrastructure that caused the computing task to be interrupted.
[0285] The receiving module 1102 is used to receive target detection results sent by some or all of the M computing nodes. The target detection results are used to determine K unusable computing nodes among the M computing nodes, where K is an integer.
[0286] The determination module 1103 is used to determine the faults of the infrastructure associated with K computing nodes based on the target detection results.
[0287] In one possible implementation, the receiving module 1102 is further configured to receive second information. The second information indicates M computing nodes and / or indicates the runtime of a computing task. The determining module 1103 is further configured to determine at least one server where the M computing nodes indicated by the second information reside, and to determine a second detection result based on runtime data associated with the at least one server. The runtime data associated with the at least one server is obtained based on the runtime.
[0288] In one possible implementation, the operational data associated with at least one server includes operational data of at least one device associated with at least one server. When determining the second detection result based on the operational data associated with at least one server, the determining module 1103 can determine the second detection result based on the operational data of at least one device. The second detection result includes sub-detection results corresponding to each of the at least one device, and each sub-detection result is used to indicate a fault of one device. The at least one device includes one or more of servers, storage devices, network devices, and server racks.
[0289] In one possible implementation, the operating data of each of the at least one devices includes one or more of alarm information, at least one performance indicator data, or operating logs. When the determining module 1103 determines the second detection result based on the operating data of the at least one device, it may perform the following steps for each of the at least one devices: determine a first candidate fault corresponding to each device based on the fault mapping relationship between alarm information and the fault corresponding to each device, wherein the fault mapping relationship includes the mapping relationship between faults and alarm information; determine a second candidate fault corresponding to each device based on the change information of each performance indicator in the at least one performance indicator data, and / or operating logs; and determine a sub-detection result corresponding to each device based on the first candidate fault and the second candidate fault.
[0290] In one possible implementation, the computational task is a model training task.
[0291] This device can be used to perform, for example Figure 7 The method steps performed by the cloud management platform in the illustrated embodiment are described below. Therefore, for a description of the device, please refer to [link / reference needed]. Figure 7 The description of the illustrated embodiments will not be repeated here.
[0292] It should be noted that each of the above modules can be implemented in software or hardware. When implemented in software, each module can be an application or code block running on a computer device. The computer device can be at least one of a physical host, virtual machine, container, or other computing device. Furthermore, there can be one or more computer devices. For example, module 802 can be an application running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers running the application can be distributed within the same availability zone (AZ) or in different AZs. Similarly, the multiple hosts / virtual machines / containers running the application can be distributed within the same region or in different regions. Typically, a region can include multiple AZs.
[0293] Similarly, multiple hosts / virtual machines / containers used to run the application can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a region can include multiple VPCs, and a VPC can include multiple Availability Zones (AZs).
[0294] When implemented in hardware, the aforementioned modules may include at least one computing device, such as a server. Alternatively, the aforementioned modules may be devices implemented using application-specific integrated circuits (ASICs) or programmable logic devices (PLDs). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof. The multiple computing devices included in the aforementioned modules may be distributed within the same Availability Zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the aforementioned modules may be distributed within the same region or in different regions. Likewise, the multiple computing devices included in the aforementioned modules may be distributed within the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0295] It should be noted that each of the above modules can be used to perform some or all of the steps in the fault detection method of the computing cluster.
[0296] The fault detection device for computing clusters disclosed in this application has clearly defined roles and close cooperation among its various modules. The modules work together to efficiently complete the fault detection process for the computing cluster.
[0297] This application also provides a computing device, which will be described below. Figure 12 , Figure 12 This is a schematic diagram of a computing device 1200 providing a fault detection method for a computing cluster according to an embodiment of this application. The computing device 1200 includes a bus 1201, a processor 1203, a memory 1202, and a communication interface 1204. The processor 1203, memory 1202, and communication interface 1204 communicate via the bus 1201. The computing device 1200 can be the aforementioned cloud management platform or a second computing node; in other words, the computing device 1200 can be used to execute the method steps performed by the cloud management platform or the second computing node. It should be understood that this application does not limit the number of processors and memories in the computing device 1200.
[0298] Bus 1201 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 12 The bus 1201 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 1201 may include a path for transmitting information between various components of the computing device 1200 (e.g., memory 1202, processor 1203, communication interface 1204).
[0299] The processor 1203 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0300] The memory 1202 may include volatile memory, such as random access memory (RAM). The processor 1203 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0301] The memory 1202 stores executable program code, which, when executed, enables a fault detection method for the computing cluster. In other words, the memory 1202 contains instructions for the cloud management platform or computing nodes to execute the fault detection method for the computing cluster.
[0302] The communication interface 1204 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1200 and other devices or communication networks.
[0303] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0304] Please see below Figure 13 This section uses the method and steps of using a computing device cluster to execute a cloud management platform as an example. Figure 13 This is a schematic diagram of the structure of a computing device cluster that executes a fault detection method for a computing cluster according to an embodiment of this application. The computing device cluster includes at least one computing device 1200. The memory 1202 of one or more computing devices 1200 in the computing device cluster may store the same cloud management platform instructions for the fault detection method of the computing cluster.
[0305] In some possible implementations, one or more computing devices 1200 in the computing device cluster can also be used to execute some instructions of the fault detection method of the computing cluster. In other words, a combination of one or more computing devices 1200 can jointly execute the instructions of the fault detection method of the computing cluster.
[0306] It should be noted that the memory 1202 in different computing devices 1200 within the computing device cluster can store different instructions for executing some functions of the cloud management platform. That is, the instructions stored in the memory 1202 of different computing devices 1200 can achieve the aforementioned... Figure 9or Figure 10 The functionality of one or more modules within it.
[0307] This application also provides a computer program product containing instructions. This computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any available medium. When the computer program product is run on at least one computer device, it causes the at least one computer device to perform the aforementioned fault detection method applied to a cloud management platform for computing clusters.
[0308] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the aforementioned fault detection method applied to a cloud management platform for executing a computing cluster.
[0309] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.
Claims
1. A fault detection method for a computing cluster, characterized in that, The method is applied to a cloud management platform, which manages infrastructure providing cloud services, the infrastructure including at least one server within a cloud data center; the method includes: Receive a first detection result, which is used to indicate a failure of a first computing node, wherein the first computing node is the node that caused the interruption of the computing task among M computing nodes, the M computing nodes are used to execute the computing task, M is a positive integer, and the M computing nodes are deployed on at least one server; The target detection result is determined based on the first detection result and the second detection result, where the second detection result is used to indicate a fault in the infrastructure, and the target detection result is used to determine K unavailable computing nodes out of the M computing nodes, where K is an integer.
2. The method according to claim 1, characterized in that, The first detection result is determined based on the operating data of some or all of the M computing nodes, and the second detection result is determined based on the operating data of some or all of the equipment included in the infrastructure.
3. The method according to claim 1 or 2, characterized in that, The first detection result is received through a first interface between the cloud management platform and the second computing node among the M computing nodes, wherein at least one of the M computing nodes has the first interface with the cloud management platform, the second computing node belongs to the at least one computing node, the second computing node is the first computing node, or the second computing node is another computing node among the M computing nodes besides the first computing node.
4. The method according to any one of claims 1 to 3, characterized in that, Each of the at least one computing node includes an operating system, and the operating system of some or all of the at least one computing node includes a first plugin for identifying faulty computing nodes among the M computing nodes.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: A fault display interface is provided, which is used to display the target detection results; and / or, The target detection result is sent to some or all of the M computing nodes so that the computing task can continue to be executed by the other computing nodes besides the K computing nodes.
6. The method according to any one of claims 1 to 5, characterized in that, The second detection result is used to indicate multiple faults in the infrastructure. The target detection result is determined based on the first and second detection results, including: The target detection result is determined based on the confidence levels of the plurality of faults and the first detection result, wherein the confidence levels of the plurality of faults are determined based on the correlation between the plurality of faults and the first detection result.
7. The method according to claim 6, characterized in that, If the first fault among the plurality of faults is related to some or all of the faults indicated by the first detection result, the confidence level of the first fault is a first value; or, If the first fault among the plurality of faults is unrelated to the fault indicated by the first detection result, the confidence level of the first fault is the second value; Wherein, the first value is greater than the second value.
8. The method according to claim 6 or 7, characterized in that, The multiple failures include failures of at least one device in the infrastructure, which includes one or more of servers, storage devices, network devices, or server racks; Based on the confidence levels of the multiple faults and the first detection result, the target detection result is determined, including: Based on J faults among the plurality of faults and the first detection result, the target detection result is determined, wherein the J faults are faults corresponding to the at least one device, wherein at least one of the J faults corresponds to a first type of device among the at least one device, and the confidence level corresponding to the at least one fault is the highest confidence level among the faults corresponding to the first type of device, wherein K is a positive integer.
9. The method according to claim 8, characterized in that, Based on J faults out of the plurality of faults and the first detection result, the target detection result is determined, including: Based on the J faults, the first detection result, and the knowledge graph, the target detection result is determined. The knowledge graph is used to represent the relationships between multiple knowledge nodes, and the multiple knowledge nodes correspond to the faults indicated by the first detection result and the J faults; or, The J faults and the first detection result are input into the artificial intelligence model to obtain the target detection result output by the artificial intelligence model.
10. The method according to any one of claims 1 to 9, characterized in that, The method further includes: Receive second information, which is used to indicate the M computing nodes and / or indicate the running time of the computing task; Identify at least one server where the M computing nodes reside; The second detection result is determined based on operational data associated with the at least one server, wherein the operational data associated with the at least one server is obtained based on the running time.
11. The method according to claim 10, characterized in that, The operational data associated with at least one server includes operational data of at least one device associated with the at least one server; Determining the second detection result based on operational data associated with the at least one server includes: The second detection result is determined based on the operating data of the at least one device. The second detection result includes sub-detection results corresponding to each of the at least one device, and each sub-detection result is used to indicate a fault of one device. The at least one device includes one or more of servers, storage devices, network devices, and server racks.
12. The method according to claim 11, characterized in that, The operational data for each of the at least one devices includes one or more of the following: alarm information, at least one performance indicator data, or operational logs: Determining the second detection result based on the operating data of the at least one device includes: For each of the at least one device, perform the following steps: Based on the fault mapping relationship between alarm information and each device, a first candidate fault corresponding to each device is determined, wherein the fault mapping relationship includes the mapping relationship between fault and alarm information; Based on the change information of each performance indicator in the at least one performance indicator data, and / or the operation log, determine the second candidate fault corresponding to each device; Based on the first candidate fault and the second candidate fault, determine the sub-detection result corresponding to each device.
13. The method according to any one of claims 1 to 12, characterized in that, The computational task is a model training task.
14. A fault detection method for a computing cluster, characterized in that, The method is applied to a second computing node, which belongs to M computing nodes used to perform computing tasks, where M is a positive integer, and the M computing nodes are deployed on at least one server; the method includes: Send a first detection result to the cloud management platform. The first detection result is used to indicate a failure of the first computing node. The first computing node is the node among the M computing nodes that caused the interruption of the computing task. The cloud management platform is used to manage the infrastructure that provides cloud services. The infrastructure includes the at least one server. Receive target detection results from the cloud management platform, the target detection results being used to determine K unusable computing nodes out of the M computing nodes, where K is an integer; The computation task is executed based on the target detection result, wherein the second computation node is a computation node other than the K computation nodes among the M computation nodes.
15. A fault detection method for a computing cluster, characterized in that, The method is applied to a cloud management platform, which manages infrastructure providing cloud services, the infrastructure including at least one server within a cloud data center; the method includes: Send a second detection result to some or all of the M computing nodes used to perform computing tasks, where M is a positive integer, and the M computing nodes are deployed on at least one server. The second detection result is used to indicate the failure of the device that caused the computing task to be interrupted among the devices included in the infrastructure. Receive target detection results sent by some or all of the M computing nodes, wherein the target detection results are used to determine K unusable computing nodes among the M computing nodes, where K is an integer; Based on the target detection results, the faults of the devices associated with the K computing nodes among the devices included in the infrastructure are determined.
16. A fault detection method for a computing cluster, characterized in that, The method is applied to a second computing node, which belongs to M computing nodes used to perform computing tasks, where M is a positive integer, and the M computing nodes are deployed on at least one server; the method includes: Receive a second detection result from a cloud management platform, the cloud management platform being used to manage the infrastructure providing cloud services, the infrastructure including the at least one server, the second detection result being used to indicate a failure of the infrastructure that caused the interruption of the computing task; The target detection result is determined based on the second detection result and the first detection result. The first detection result is used to indicate the failure of the first computing node, which is the node among the M computing nodes that caused the interruption of the computing task. The target detection result is used to determine the K unusable computing nodes among the M computing nodes, where K is an integer. The computation task continues to be executed based on the target detection result, wherein the second computation node is a computation node other than the K computation nodes among the M computation nodes.
17. A fault detection device for a computing cluster, characterized in that, It includes a module for performing the method as described in any one of claims 1 to 13 or 15, or includes a module for performing the method as described in claim 14 or 16.
18. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the computing device cluster to perform the method as described in any one of claims 1 to 13, or to perform the method as described in claim 15, or to perform the method as described in claim 14, or to perform the method as described in claim 16.
19. A computer program product containing instructions, characterized in that, When the instructions are executed by a cluster of computer devices, the cluster of computer devices performs the method as described in any one of claims 1 to 13, or performs the method as described in claim 15, or performs the method as described in claim 14, or performs the method as described in claim 16.
20. A computer-readable storage medium, characterized in that, The method includes computer program instructions that, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 13, or perform the method as described in claim 15, or perform the method as described in claim 14, or perform the method as described in claim 16.