Fault detection method and device based on cloud service
By leveraging the fault detection service of the cloud management platform and integrating fault detection results from different levels through a layered architecture and topology relationships, the problem of multiple fault delimitation methods being unable to make comprehensive decisions in complex distributed architectures is solved, thereby improving the accuracy and flexibility of fault detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-11
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies cannot select the final fault delimitation result when different fault information is determined by various fault delimitation methods, resulting in insufficient fault detection accuracy. This is especially true in application systems with complex distributed architecture topologies, where integrated decision-making methods cannot effectively integrate different fault information.
The cloud management platform provides fault detection services, receives user configuration information and service requests, and leverages a layered architecture and topology to integrate fault detection results from different levels. It also combines user-input priorities and expert rules to achieve comprehensive decision-making for multiple fault detection methods.
It improves the accuracy and flexibility of fault detection results, meets users' personalized needs, and enhances the accuracy and reliability of the final detection results through layer-by-layer fusion and information complementarity.
Smart Images

Figure CN121664631A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cloud technology, and in particular to a fault detection method and apparatus based on cloud services. Background Technology
[0002] An application system is a computing cluster that provides backend services for enterprise applications; for example, an application system might be the backend service system for a financial system. Fault isolation refers to the process of determining the scope of a fault when an application system fails. For example, fault isolation can identify which specific subsystem within the application system is malfunctioning, causing the application system to malfunction and allowing for rapid repair.
[0003] Currently, in order to improve the accuracy of fault delimitation, various methods can be used for fault delimitation. However, when the fault information determined by the various methods is different, it is impossible to select the result obtained by any one method as the final fault delimitation result. Summary of the Invention
[0004] This application provides a cloud-based fault detection method and apparatus, which is used to make comprehensive decisions based on fault detection results obtained from multiple fault delimitation methods, thereby improving the accuracy of fault detection results.
[0005] Firstly, a fault detection method based on cloud services is provided, which can be applied to a cloud management platform. The cloud management platform manages the infrastructure providing cloud services. This infrastructure includes multiple regions, each region comprising at least one cloud data center. The cloud services run on at least one server in at least one cloud data center located in one of the multiple regions. The method includes: providing a configuration interface for the fault detection service. The configuration interface receives system configuration information and a first service requirement from a user-input application system. The application system adopts a layered architecture, comprising multiple layers and multiple topology nodes. Each topology node is located in one of the multiple layers. The system configuration information indicates the topological relationships between the multiple topology nodes, and the first service requirement indicates fault detection for at least two layers. The method also includes: acquiring runtime data from the multiple topology nodes during application system operation, which is used to determine the fault detection results corresponding to at least two layers; and displaying the fused detection results of the application system on the result interface of the fault detection service. The fused detection results are obtained by fusing the fault detection results corresponding to at least two layers based on the topological relationships.
[0006] In this method, the user can configure the system configuration information of the application system to be detected, and configure at least two levels for fault detection. The cloud management platform can perform fault detection on these at least two levels based on the runtime data of the application system, and fuse the fault detection results corresponding to these at least two levels according to the topology indicated by the system configuration information to obtain the final fused detection result. It is evident that this embodiment of the application achieves the fusion of fault detection results from different levels. Furthermore, since this method performs fault detection and fusion from different levels, it can fully utilize the complementarity of information between different levels, which is beneficial to improving the accuracy of the final fused detection result.
[0007] In one possible implementation, the configuration interface is further configured to receive a second service request input by the user. This second service request indicates N candidate faults of the application system, where the level of the N candidate faults is higher than any of the at least two levels, or the level of the N candidate faults is the highest of the at least two levels. Based on the second service request input by the user, the fusion detection results can be displayed on the results interface. The fusion detection results include at least one candidate fault from the N candidate faults, which is obtained by fusing the fault detection results corresponding to at least two levels layer by layer according to the topological relationship, using a hierarchical approach from low to high.
[0008] In this implementation, the user can also configure N candidate faults for the final output. The level corresponding to these N candidate faults is higher than or equal to the highest of at least two levels. Thus, during layered fusion, the cloud management platform can fuse layer by layer from low to high until the level corresponding to the N candidate faults is reached. On one hand, this provides users with a way to configure the final output, meeting their personalized needs and improving the flexibility of the cloud management platform. On the other hand, the layer-by-layer fusion approach gradually integrates information from each level, which helps improve the degree of information fusion between levels and enhances the accuracy of the final fused detection results.
[0009] In one possible implementation, at least two levels include a first level and a second level. The first level includes multiple first topology nodes, and the second level includes multiple second topology nodes. The fault detection results corresponding to the at least two levels include the first fault detection result corresponding to the first level and the second fault detection result corresponding to the second level. The fused detection result includes a third fault detection result, which is determined based on the topology relationship and the second fault detection result, and the third fault detection result corresponds to the first level; and / or, the fused detection result includes a fourth fault detection result, which is obtained by fusing the third fault detection result and the first fault detection result.
[0010] In this implementation, the layered fusion process involves some intermediate data. When displaying the fusion detection results on the results interface, this intermediate data can also be presented. This allows users to perceive more information related to the layered fusion, improving their service experience. Furthermore, for two different layers, during the layered fusion process, the fault detection results from one layer can be transferred to the other layer before fusion. The new fault detection results incorporate information from both layers, improving the accuracy of fault detection.
[0011] In one possible implementation, the third fault detection result includes multiple confidence levels, and the first fault detection result is used to indicate fault information of multiple first topology nodes, with each confidence level corresponding one-to-one with a single first topology node. In this implementation, converting a fault detection result at one level into a confidence level for a fault detection result at another level is equivalent to using the detection result at one level as a verification of the fault detection result at another level, thereby enhancing the fault detection result at the other level and improving the accuracy and reliability of fault detection.
[0012] In one possible implementation, the second fault detection result indicates the fault probability of each of the multiple second topology nodes. The third fault detection result includes one or more of the following: a confidence level that each first topology node is fault-free, determined based on the fault probabilities of the multiple second topology nodes; a confidence level that each first topology node belongs to a first fault among N candidate faults, determined based on the fault probabilities of the multiple second topology nodes; a confidence level that each first topology node belongs to a second fault among N candidate faults, determined based on the fault probabilities of second topology nodes of a first type among the multiple second topology nodes, where the second fault affects the operation of the first type of second topology nodes; or, a confidence level that each first topology node belongs to a third fault among N candidate faults, determined based on the fault probabilities of second topology nodes of a second type among the multiple second topology nodes, where the third fault affects the operation of the second type of second topology nodes. In this implementation, the first fault detection result is used to indicate the fault category to which each first topology node belongs. The first fault detection result can be enhanced by converting the fault probability of the second topology nodes that have a topological relationship with each first topology node into a confidence level that each first topology node belongs to a fault category, thereby improving the accuracy and reliability of fault detection.
[0013] In one possible implementation, the configuration interface is also used to receive user-inputted data indicator configuration information. This information indicates the data indicators to be collected for each topology node across multiple levels. Subsequently, runtime data from multiple topology nodes can be obtained based on the data indicator configuration information. In this implementation, a method for configuring the required data indicators is provided to the user. Since users have a better understanding of their application system, user-configured data indicators can improve accuracy, and combined with more accurate runtime data, this can further enhance the accuracy of fault detection.
[0014] In one possible implementation, the configuration interface is further configured to receive a data indicator configuration file input by the user, the data indicator configuration file including data indicator configuration information; alternatively, the configuration interface is further configured to display multiple data indicators and receive a data indicator selected by the user from the multiple data indicators, the data indicator configuration information including the selected data indicator. This implementation provides users with multiple ways to configure the data indicators to be collected, allowing users to choose the specific method according to their needs, thus improving the user experience.
[0015] In one possible implementation, the configuration interface is used to display multiple levels and receive at least two levels selected by the user from the multiple levels; alternatively, the configuration interface is used to receive at least two fault detection methods input by the user, wherein each fault detection method is used to obtain a fault detection result corresponding to one of the at least two levels. This implementation provides the user with multiple ways to configure the required detection levels, allowing the user to choose the specific method according to their needs, thus improving the user experience.
[0016] In one possible implementation, the configuration interface is used to receive user input of at least two fault detection methods, including one or more of the following: the configuration interface is used to receive user input of at least one expert rule, wherein each expert rule is used to obtain a fault detection result corresponding to a level; or, the configuration interface is used to receive user input of at least one fault detection model, wherein after training each fault detection model using a corresponding training method, each fault detection model is used to obtain a fault detection result corresponding to a level according to the corresponding fault detection method. In this implementation, configuration can be performed separately according to the different types of fault detection methods, improving the accuracy of the configuration process.
[0017] In one possible implementation, the configuration interface is used to receive system configuration information of the application system input by the user, including: the configuration interface is used to receive a system configuration file input by the user, the system configuration file including system configuration information; or, the configuration interface is used to receive editing operations input by the user, the editing operations being used to edit topology nodes and edit the topological relationships between topology nodes. This implementation provides users with multiple ways to configure system configuration information, allowing users to choose the specific method according to their needs, thus improving the user experience.
[0018] In one possible implementation, the fusion detection results of the application system are displayed on the result interface of the fault detection service, including: displaying a topology diagram of the application system on the result interface; and marking the faulty topology nodes in the application system on the topology diagram based on the fusion detection results. In this implementation, the fault results can be marked on the topology diagram, allowing users to more intuitively perceive the location of the application system's faults and improving the user experience.
[0019] Secondly, a fault detection method based on cloud services is provided, which can be applied to a cloud management platform. The cloud management platform manages the infrastructure providing cloud services. This infrastructure includes facilities located in multiple regions, each region including at least one cloud data center. The cloud services run on at least one server in at least one cloud data center located in one of the multiple regions. The method includes: providing a configuration interface for fault detection services, the configuration interface receiving user input of the application system's topology type and a third service requirement, the third service requirement indicating N fault types of the application system, where N is a positive integer greater than 1; acquiring runtime data of the application system; determining at least two fault detection results based on the runtime data using at least two fault detection methods, each fault detection result corresponding to one fault detection method; wherein, under each fault type of the topology type, the at least two fault detection methods have corresponding priorities; and displaying the fused detection results of the application system on the result interface of the fault detection service, the fused detection results being obtained by fusing the at least two fault detection results according to the respective priorities of the at least two fault detection methods.
[0020] In this method, for each fault type within the topology of the application system, at least two fault detection methods have corresponding priorities. The cloud management platform can combine these priorities to fuse the at least two fault detection results to obtain the final fused detection result. On one hand, this method enables the cloud management platform to make comprehensive decisions based on at least two fault detection results, solving the problem of inability to make comprehensive decisions when multiple fault detection results differ. On the other hand, the priority is comprehensively determined by the topology type, fault type, and fault detection method, which is equivalent to a targeted evaluation of the applicability of each fault detection method to the current application system's fault detection. Therefore, fusing the fault detection results based on this evaluation result can improve the accuracy of the final fused detection result.
[0021] In one possible implementation, the configuration interface is also used to display multiple faults from the fault mode library, and to receive M faults selected by the user from the multiple faults, where M is a positive integer greater than 1; wherein each of the N fault types corresponds to at least one fault in the fault mode library. Fault drill data is obtained when multiple fault examples are run in the application system, with the multiple fault examples corresponding to the M faults. The fault drill data is used to determine the reference detection results of the application system. The fault drill results are displayed on the configuration interface, including the priorities of at least two fault detection methods under each fault type of the topology type. The fault drill results are determined based on the reference detection results corresponding to each of the at least two fault detection methods.
[0022] In this implementation, the application system can be drilled using M types of faults selected by the user from a fault mode library. The results of these drills determine the priority of at least two fault detection methods for each fault type within the topology type. This is equivalent to using prior experience to determine the priority of each fault detection method, which helps improve the accuracy of the final fusion detection results. Furthermore, the M types of faults are configured by the user. Since users have a better understanding of their own application system and needs, user configuration enhances the relevance of the fault drills, thereby more accurately determining the priority of each fault detection method.
[0023] In one possible implementation, each of the at least two fault detection results indicates the probability of the fault type to which the application system belongs. The fused detection result includes one or more of the following: the weight of each fault type among the multiple fault types indicated by the at least two fault detection results, wherein the weight of each fault type is determined based on the probability of the fault type indicated by the at least two fault detection results and the corresponding priority; or, the fault type with the highest weight among the multiple fault types.
[0024] In this implementation, some intermediate data is involved in the fusion process. When displaying the fusion detection results on the results interface, this intermediate data can also be presented. This allows users to perceive more information related to the layered fusion, which helps improve the user's service experience. Furthermore, at least two fault detection results can be fused based on the probability of the fault detection result itself (which can be understood as the reliability of the fault detection method itself) combined with a priority based on prior experience, which helps improve the accuracy of fault detection.
[0025] In one possible implementation, the configuration interface is further configured to receive user-inputted data indicator configuration information, which indicates the data indicators required to be collected for the application system. Obtaining runtime data of the application system includes: acquiring runtime data of the application system based on the data indicator configuration information. In this implementation, a method for configuring the required data indicators is provided to the user. Since users have a better understanding of their application system, user-configured methods can improve the accuracy of the required data indicators, and combined with more accurate runtime data, can further enhance the accuracy of fault detection.
[0026] In one possible implementation, the configuration interface is also used to receive at least two fault detection methods configured by the user. In this implementation, providing the user with a way to configure the required fault detection methods helps improve the accuracy of fault detection.
[0027] Thirdly, a cloud service-based fault detection device is provided. This device can be applied to a cloud management platform for managing the infrastructure that provides cloud services. The infrastructure includes multiple regions, each region including at least one cloud data center, and the cloud services run on at least one server in at least one cloud data center located in one of the multiple regions.
[0028] The device includes: a service configuration module for providing a configuration interface for fault detection services. The configuration interface receives system configuration information and a first service requirement from the user-input application system. The application system adopts a layered architecture, comprising multiple layers and multiple topology nodes. Each topology node is located in one of the multiple layers. The system configuration information indicates the topological relationships between the multiple topology nodes, and the first service requirement indicates fault detection for at least two layers. A data acquisition module acquires runtime data from the multiple topology nodes during application system operation. This runtime data is used to determine the fault detection results corresponding to at least two layers. A result display module displays the fused detection results of the application system on the fault detection service result interface. The fused detection results are obtained by fusing the fault detection results corresponding to at least two layers based on the topological relationships.
[0029] In one possible implementation, the configuration interface is further configured to receive a second service request input by the user. This second service request indicates N candidate faults of the application system, where the level corresponding to the N candidate faults is higher than any of the at least two levels, or the level corresponding to the N candidate faults is the highest of the at least two levels. The result display module is specifically configured to display the fusion detection result on the result interface based on the second service request input by the user. The fusion detection result includes at least one candidate fault from the N candidate faults, and this at least one candidate fault is obtained by fusing the fault detection results corresponding to at least two levels layer by layer according to the topological relationship, using a hierarchical approach from low to high.
[0030] In one possible implementation, at least two levels include a first level and a second level. The first level includes multiple first topology nodes, and the second level includes multiple second topology nodes. The fault detection results corresponding to the at least two levels include the first fault detection result corresponding to the first level and the second fault detection result corresponding to the second level. The fused detection result includes a third fault detection result, which is determined based on the topology relationship and the second fault detection result, and the third fault detection result corresponds to the first level; and / or, the fused detection result includes a fourth fault detection result, which is obtained by fusing the third fault detection result and the first fault detection result.
[0031] In one possible implementation, the third fault detection result includes multiple confidence levels, and the first fault detection result is used to indicate fault information of multiple first topology nodes, with each confidence level corresponding to one of the multiple first topology nodes.
[0032] In one possible implementation, the second fault detection result indicates the fault probability of each of the plurality of second topology nodes. The third fault detection result includes one or more of the following: a confidence level that each first topology node is fault-free, determined based on the fault probabilities of the plurality of second topology nodes; a confidence level that each first topology node belongs to a first fault among N candidate faults, determined based on the fault probabilities of the plurality of second topology nodes; a confidence level that each first topology node belongs to a second fault among N candidate faults, determined based on the fault probabilities of second topology nodes of a first type among the plurality of second topology nodes, the second fault affecting the operation of the first type of second topology nodes; or, a confidence level that each first topology node belongs to a third fault among N candidate faults, determined based on the fault probabilities of second topology nodes of a second type among the plurality of second topology nodes, the third fault affecting the operation of the second type of second topology nodes.
[0033] In one possible implementation, the configuration interface is further configured to receive user-inputted data indicator configuration information, which indicates the data indicators to be collected for each topology node in multiple layers. The data acquisition module is specifically configured to acquire runtime data from multiple topology nodes during application system operation based on the data indicator configuration information.
[0034] In one possible implementation, the configuration interface is further configured to receive a data indicator configuration file input by the user, the data indicator configuration file including data indicator configuration information; or, the configuration interface is further configured to display multiple data indicators and receive a data indicator selected by the user from the multiple data indicators, the data indicator configuration information including the data indicator selected by the user.
[0035] In one possible implementation, the configuration interface is used to display multiple levels and receive at least two levels selected by the user from the multiple levels; or, the configuration interface is used to receive at least two fault detection methods input by the user, wherein each fault detection method is used to obtain a fault detection result corresponding to one of the at least two levels.
[0036] In one possible implementation, the configuration interface is used to receive at least one expert rule input by the user, wherein each expert rule is used to obtain a fault detection result corresponding to a level; or, the configuration interface is used to receive at least one fault detection model input by the user, wherein after each fault detection model is trained using the corresponding training method for each fault detection model, each fault detection model is used to obtain a fault detection result corresponding to a level according to the corresponding fault detection method.
[0037] In one possible implementation, the configuration interface is used to receive system configuration files input by the user, the system configuration files including system configuration information; or, the configuration interface is used to receive editing operations input by the user, the editing operations being used to edit topology nodes and edit the topological relationships between topology nodes.
[0038] In one possible implementation, the result display module is specifically used to display the topology diagram of the application system on the result interface; and to mark the faulty topology nodes in the application system in the topology diagram according to the fusion detection results.
[0039] The third aspect or any embodiment of the third aspect is a step implementation of the apparatus corresponding to the first aspect or any embodiment of the first aspect. The description in the first aspect or any embodiment of the first aspect is applicable to the third aspect or any embodiment of the third aspect, and will not be repeated here.
[0040] Fourthly, a cloud service-based fault detection device is provided. This device can be applied to a cloud management platform for managing the infrastructure that provides cloud services. The infrastructure includes multiple regions, each region including at least one cloud data center. The cloud services run on at least one server in at least one cloud data center located in one of the multiple regions.
[0041] The device includes: a service configuration module for providing a configuration interface for fault detection services, the configuration interface for receiving user input of the application system's topology type and third service requirements, the third service requirements indicating N fault types of the application system, where N is a positive integer greater than 1; a data acquisition module for acquiring runtime data of the application system; a detection execution module for determining at least two fault detection results based on the runtime data using at least two fault detection methods, each fault detection result corresponding to one fault detection method; wherein, under each fault type of the topology type, the at least two fault detection methods have corresponding priorities; and a result display module for displaying the fused detection results of the application system on the result interface of the fault detection service, the fused detection results being obtained by fusing the at least two fault detection results according to the respective priorities of the at least two fault detection methods.
[0042] In one possible implementation, the configuration interface is further configured to display multiple faults from a fault mode library, and to receive M faults selected by the user from the multiple faults, where M is a positive integer greater than 1; wherein each of the N fault types corresponds to at least one fault in the fault mode library. The device also includes a fault drill module for acquiring fault drill data when multiple fault examples are run in the application system, where the multiple fault examples correspond to the M faults, and the fault drill data is used to determine the reference detection results of the application system. The service configuration module is further configured to display fault drill results on the configuration interface, the fault drill results including the priorities of at least two fault detection methods under each fault type of the topology type, and the fault drill results are determined based on the reference detection results corresponding to the at least two fault detection methods.
[0043] In one possible implementation, each of the at least two fault detection results indicates the probability of the fault type to which the application system belongs. The fused detection result includes one or more of the following: the weight of each fault type among the multiple fault types indicated by the at least two fault detection results, wherein the weight of each fault type is determined based on the probability of the fault type indicated by the at least two fault detection results and the corresponding priority; or, the fault type with the highest weight among the multiple fault types.
[0044] In one possible implementation, the configuration interface is also used to receive user-inputted data indicator configuration information, which indicates the data indicators required to be collected for the application system. The data acquisition module is specifically used to acquire runtime data of the application system based on the data indicator configuration information.
[0045] In one possible implementation, the configuration interface is also used to receive at least two fault detection methods configured by the user.
[0046] The fourth aspect or any embodiment of the fourth aspect is a step implementation of the apparatus corresponding to the second aspect or any embodiment of the second aspect. The description in the second aspect or any embodiment of the second aspect is applicable to the fourth aspect or any embodiment of the fourth aspect, and will not be repeated here.
[0047] Fifthly, a computing device cluster is provided, including at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, such that the computing device cluster performs the methods disclosed in the first aspect and any possible implementation of the first aspect, or performs the methods disclosed in the second aspect and any possible implementation of the second aspect.
[0048] In a sixth aspect, this application provides a computer program product containing instructions that, when executed by a cluster of computer devices, cause the cluster of computer devices to implement the method disclosed in the first aspect and any possible implementation thereof, or to perform the method disclosed in the second aspect and any possible implementation thereof.
[0049] In a seventh aspect, this application provides a computer-readable storage medium including computer program instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method disclosed in the first aspect and any possible implementation thereof, or to perform the method disclosed in the second aspect and any possible implementation thereof. Attached Figure Description
[0050] Figures 1A to 1D This is a schematic diagram of the system architecture provided for an embodiment of this application;
[0051] Figure 2 A flowchart illustrating a method provided in an embodiment of this application;
[0052] Figures 3A-3E This is a schematic diagram of the interface provided for an embodiment of this application;
[0053] Figure 4A and Figure 4B A schematic diagram illustrating the fusion of fault detection results provided in an embodiment of this application;
[0054] Figure 5 A schematic diagram of layered fusion provided for an embodiment of this application;
[0055] Figure 6 A flowchart illustrating another method provided in an embodiment of this application;
[0056] Figures 7A to 7D Example diagram of the topology type provided in the embodiments of this application;
[0057] Figure 8 A flowchart illustrating the priority determination process provided in the embodiments of this application;
[0058] Figure 9 A schematic diagram of the structure of an apparatus provided in an embodiment of this application;
[0059] Figure 10 A schematic diagram of another device provided in an embodiment of this application;
[0060] Figure 11 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0061] Figure 12 This is a schematic diagram of a computing device cluster provided in an embodiment of this application. Detailed Implementation
[0062] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings. In the description of the embodiments of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include one or more of that feature.
[0063] It should be understood that in the embodiments of this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, a and b, a and c, b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0064] An application system is a computer system specifically designed for a particular application area or business need, capable of performing specific tasks or solving specific problems. An application system typically consists of application programs (or software), system resources (or hardware), and a database. The application program is responsible for implementing the system's main functions, such as processing data, performing calculations, and generating reports. System resources refer to the physical resources required to support the operation of the application program, and the database is used to store, manage, and retrieve the data needed by the application program. Application systems can be widely used in various fields, such as enterprise management, financial services, healthcare, or education, helping to improve work efficiency, optimize business processes, and enable data analysis.
[0065] The widespread implementation of distributed technologies has led to increasingly large-scale application systems, increasing the complexity and uncertainty of the fault delimitation process and making fault delimitation more difficult. To improve the accuracy of application systems, multiple fault delimitation methods are often used separately, and the results of these methods are then combined to obtain the final fault delimitation result. For example, current methods often employ voting or training new models to fuse the results of different fault delimitation methods. However, these comprehensive decision-making methods are often based on the assumption that the output forms of multiple fault delimitation methods are consistent. For instance, current comprehensive decision-making methods are only applicable when the fault granularity indicated by the results of multiple fault delimitation methods is the same. Therefore, these comprehensive decision-making methods can only be used for application systems with relatively simple architectural topologies.
[0066] However, for application systems with more complex distributed architecture topologies, modeling from different perspectives is required. This can lead to different fault delimitation methods determining different fault information, such as differences in the output format of fault information (e.g., fault granularity) or conflicting fault information. Current integrated decision-making methods cannot synthesize the results of multiple fault delimitation methods.
[0067] In view of this, embodiments of this application provide a cloud service-based fault detection method to solve the comprehensive decision-making problem when using multiple fault delimitation methods to detect faults in application systems.
[0068] In some scenarios, the application system can also be replaced with computer cluster, distributed computing cluster, cloud computing cluster, data center, cloud data center or other names, without specific restrictions.
[0069] Please refer to Figure 1AThis is a schematic diagram illustrating an application scenario of the method provided in this application embodiment. The scenario may include a cloud management platform 100, which manages multiple cloud data centers set up by cloud vendors in different regions. The cloud management platform can provide interfaces related to public cloud services, such as web pages or application programming interfaces (APIs), for users to remotely access public cloud services. Users can log in to the cloud management platform on the public cloud access page using a pre-registered account and password. After successful login, users can select and purchase public cloud services provided by the cloud data center in the selected region on the public cloud access page, such as object storage services, virtual machine services, container services, or the fault detection service involved in this application embodiment.
[0070] The cloud management platform 100 manages the infrastructure providing cloud services. This infrastructure includes facilities located in multiple regions, each containing at least one cloud data center. Cloud services run on at least one server in at least one cloud data center located in one of the regions. A region refers to the geographical location of a cloud service. Cloud services can be categorized by region based on geographical location and network latency. Cloud services within the same region use infrastructure located in the same geographical location. For example, if a cloud service selects South China as its region, then it will use cloud data centers in the South China region to provide that cloud service. A region may include one or more Availability Zones (AZs). An AZ is a collection of one or more cloud data centers with independent water and electricity supply. Within an Availability Zone, computing, network, and storage resources are logically further divided into multiple clusters. Multiple Availability Zones within a region are connected via high-speed fiber optic cables to meet users' needs for building high-availability systems across AZs.
[0071] See Figure 1A As shown, the cloud management platform 100 can provide a client 400, which can be used to display the relevant interface of the fault detection service. Users 500 can configure and view the fault detection service of the application system 200 through this interface. For example, they can configure service requirements for the fault detection service and the system configuration information of the application system 200 through the fault detection service interface. Additionally, users can view the fault detection results through the fault detection service interface. Optionally, the client 400 can be implemented through a terminal device or a functional module (such as a webpage or application) within the terminal device. For example, the terminal device can be a personal computer (PC), tablet computer (PAD), or mobile phone, etc., with no specific limitations.
[0072] For example, Figure 1AIn the scenario shown, user 500 can access cloud management platform 100 via client 400 through internet 300. User 500 can purchase device fault detection services for application system 200 on cloud management platform 100. User 500 can input / select relevant configuration information and service requirements of application system 200 on cloud management platform 100 through client 400. Cloud management platform can run fault detection services for application system 200 according to user 500's relevant configuration information and service requirements. Figure 1A As shown: Taking cloud data center A-region 1 as an example, user 500 inputs the relevant configuration information and service requirements of application system 200 through cloud management platform 100. Application system 200 will then be assigned to cloud data center A in region 1 for fault detection services. Specifically, application system 200 accesses cloud data center A in region 1 via the Internet through cloud management platform 100. It should be noted that the choice of which cloud data center in which region service user 500's application system 200 is assigned is not limited here. User 500 can select a service data center for application system 200 through cloud management platform 100, or user 500 can leave it unselected, with cloud management platform 100 randomly assigning a service data center.
[0073] The cloud management platform 100, also known as a fault detection platform or fault detection system, provides fault detection services for the application system 200. Optionally, the fault detection service can be a fault localization service for the application system 200, and the cloud management platform 100 can also be called a fault localization system. In some scenarios, the cloud management platform 100 can also provide other fault detection services, such as fault location and root cause localization, and therefore the cloud management platform 100 can also be called a fault location system or root cause localization system, without specific limitations.
[0074] In some embodiments, the cloud management platform 100 may be independent of the application system 200, meaning that the resources relied upon by the cloud management platform 100 and the application system 200 are unrelated. The cloud management platform 100 can obtain the data required for fault detection of the application system 200, such as the operational data of the application system 200, to perform fault detection on the application system 200 based on this data. Optionally, the cloud management platform 100 may be a dedicated cloud management platform for the application system 200, or it may be a shared cloud management platform for multiple application systems 200, meaning that the cloud management platform 100 can provide fault detection services to multiple application systems 200.
[0075] In some embodiments, the cloud management platform 100 and the application system 200 may belong to the same system, meaning they may rely on the same system resources. For example, the cloud management platform 100 may be part of the application system 200, or vice versa. In this embodiment, the cloud management platform 100 can directly obtain the data required for fault detection in the application system 200, such as the application system 200's operational data, to perform fault localization on the application system 200 based on this data.
[0076] For example, when the cloud management platform 100 performs fault detection on the application system 200 based on the obtained operating data of the application system 200, it can use a variety of fault detection methods to obtain corresponding fault detection results, so as to achieve a comprehensive decision based on the multiple fault detection results using the method provided in the embodiments of this application.
[0077] In this embodiment, the cloud management platform 100 can be implemented through software or hardware. For example, in a software implementation, the cloud management platform 100 may include code running on computing instances. These computing instances can be at least one of physical hosts (computing devices), virtual machines, containers, etc. Optionally, there may be one or more computing devices. For example, the cloud management platform 100 may include code running on multiple hosts / virtual machines / containers. As another example, in a hardware implementation, the cloud management platform 100 may include at least one computing device, such as a server. Alternatively, the cloud management platform 100 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. The computing devices included in the cloud management platform 100 can be referred to as fault detection devices. Fault detection equipment can also be called fault location equipment, fault location device, fault determination device, or root cause location device, etc., without any specific restrictions.
[0078] Application system 200 provides the infrastructure required for applications (or services), and this infrastructure is located in at least one region, where each region includes at least one cloud data center (or database). In other words, application system 200 provides backend services for applications (or services). The application system 200 involved in the embodiments of this application can adopt any possible architectural topology.
[0079] As a possible example, please refer to Figure 1B This is a schematic diagram of the topology of the application system 200 provided in an embodiment of this application. (See reference...) Figure 1B The architecture topology of the application system 200 may include a combination of resource vertical topology and application horizontal topology. Resource vertical topology can be understood as the dependencies between system resources; that is, it is an architecture topology divided according to the hierarchy of system resources. Application horizontal topology can be understood as an architecture topology divided according to the hierarchy of application programs (or components).
[0080] According to the resource vertical topology, application system 200 may include the following:
[0081] (1) Region: A region is a collection of geographically located data centers. Regions help users distribute system resources across different geographical locations to improve disaster recovery capabilities and reduce latency. For example, for an application system 200 with multiple regions, the region of the application system 200 can refer to the geographical location of the system resources that the application system 200 provides services to. Backend services within the same region use infrastructure located in the same geographical location. For example, if a service within an application system 200 selects the South China region as its region, then the service will be provided using system resources from the South China region.
[0082] (2) Availability Zone (AZ): Each region can contain one or more AZs. Each AZ is a cluster of one or more data centers with independent power, network, and cooling systems to ensure high availability and fault tolerance. By distributing applications across different AZs, users can achieve greater redundancy and reliability.
[0083] (3) Host Groups: Each AZ may include one or more host groups. Each host group is a collection of physical machines, typically used for centralized management and configuration to simplify resource deployment and maintenance.
[0084] (3) Physical machine (PM): Each host group includes one or more physical machines. For example, a physical machine can be an actual hardware server. Unlike virtual machines, a physical machine does not have a virtualization layer but runs the operating system and applications directly. Physical machines can provide complete hardware resources and provide direct hardware access and control.
[0085] (4) Virtual Machine (VM): A physical machine may contain one or more VMs. A VM is a virtualized instance created and run using virtualization technology, allowing multiple VMs to share the same hardware resources. Each VM has its own independent operating system and application environment. One or more containers can run on each VM, such as... Figure 1B The application container shown can be used to run the application corresponding to application system 200.
[0086] (5) Bare Metal Server (BMS): Also known as a single-tenant physical server, a BMS consists of one or more independent servers. It is a computing service that combines the performance of an elastic cloud server and a physical machine. It can provide enterprises with dedicated physical servers in the cloud, offering excellent computing performance and data security for core databases, critical application systems, high-performance computing, big data, and other businesses. Each BMS can run one or more containers, such as... Figure 1B The database container shown can be used to implement the database-related functions corresponding to application system 200.
[0087] According to the application-level topology, application system 200 may include the following:
[0088] (1) Application services: An application system 200 may provide one or more application services. For example, if the application system 200 is a financial system, an application service may be a deposit service or a lending service, etc.; or, if the application system 200 is an education system, an application service may be a course selection service or an online course service, etc.; or, if the application system 200 is a food delivery system, an application service may be a food delivery service or an errand service, etc., without any restrictions.
[0089] (2) Microservices, also known as service groups, middleware, etc., are a combination of one or more microservices to implement an application's business logic. Microservices are essentially a software architecture pattern where application logic is broken down into small, independently deployable services, each responsible for specific business functions. Each microservice typically has its own database and business logic, and can be developed, deployed, and scaled independently.
[0090] (3) Components, also known as units, business units, service units, or microservice units, can be one or more components. A component usually refers to a modular part within a single application, and a component can be understood as a function under a microservice.
[0091] (4) Basic modules: A component may include one or more basic functional modules. A basic module can be understood as the smallest functional unit in an application system 200.
[0092] (5) Containers. Containers are a lightweight, portable virtualization technology used to package and run applications and their dependencies. Containers provide a consistent runtime environment, allowing applications to run consistently across different systems and environments. For example... Figure 1B As shown, in this embodiment, containers can be divided into database containers and application containers based on their different functions. Database containers can be used to implement database-related functions corresponding to application system 200, and different database containers can be divided according to the different types of data they store. For example, database container A is used to store and detect data for users 1-100, and database container B is used to store and detect data for users 101-200. Application containers can be used to run applications corresponding to application system 200; different application containers can run the same or different applications.
[0093] The above-mentioned A including B can be understood as B being a component of A, or the implementation of A depending on B, or the implementation of B depending on A, or B being deployed on A, etc.
[0094] See Figure 1C This is a schematic diagram of a system architecture for a cloud management platform 100 provided in an embodiment of this application. The cloud management platform 100 may include service purchase-related functions, service configuration-related functions, and service usage-related functions. These will be described in detail below.
[0095] I. Service Purchase Related Functions: These functions allow users to purchase fault detection services. For example... Figure 1C As shown, service purchase-related functions may include a service purchase module. For example, the service purchase module can provide users with a service purchase interface, where users can purchase fault detection services according to the prompts on the service purchase interface. Correspondingly, after the user purchases the fault detection service, the service purchase module will also perform corresponding processing, such as storing the user's order information.
[0096] II. Service configuration-related functions can be used to provide users with functions related to configuring the fault detection service, or to configure functions related to the fault detection service according to user instructions. After configuration, the fault detection service can execute the fault detection process according to the configured content. As a possible example, such as... Figure 1C As shown, service configuration-related functions may include the following modules:
[0097] (1) Resource Configuration Module: This module is used to configure the physical resources used by the fault detection service, which may include computing resources and storage resources. After configuration through the resource configuration module, the fault detection service can be deployed on the corresponding resources and utilize these resources to implement the fault detection service.
[0098] Optionally, the resource configuration module can provide a resource configuration page through which users can input resources for the fault detection service. For example, users can input resource requirements for the fault detection service, such as the quantity or type of resources. Alternatively, the resource configuration page can also provide users with selectable resources from which they can choose the resources needed to deploy the fault detection service.
[0099] (2) Basic Configuration Module: This module is used to configure the basic information of the fault detection service. This basic information can also be understood as the basic configuration information of the application system. Optionally, the basic information may include the faults of the application system, that is, the faults of the application system that the fault detection ultimately needs to output. The faults of the application system can be multiple faults or multiple fault types. Optionally, the faults of the application system can be configured according to the actual situation of the application system and the user's own needs. For example, fault detection refers to fault delimitation for the application system, so the faults of the application system can be determined based on the fault range for rapid fault recovery or based on the means of rapid fault recovery.
[0100] Optionally, the basic configuration module can provide a basic configuration page through which users can input various faults of the application system. For example, users can input various fault-related information, such as the fault name and description. Alternatively, the basic configuration page can provide users with a selection of faults to choose from. Or, the basic configuration page can display a topology diagram of the application system, allowing users to specify topology nodes where the application system may fail; a failure of a topology node can be considered a fault.
[0101] (3) The indicator management module: The fault detection process relies on the collected operational data of the application system. The indicator management module can be used to configure and manage the data indicators required by the application system. During application system operation, operational data can be collected based on the configured data indicators, and subsequent fault detection can be performed based on this data. Data indicators can be understood as the operational characteristics of the application system; these characteristics can reflect the faults of corresponding topology nodes in the application system. For example... Figure 1B As shown, the application system can adopt a layered architecture, that is, the architecture of the application system can contain multiple layers, and the topology nodes contained in the application system can be distributed on different layers. Then, the indicator management module can be used to configure or manage the data indicators to be collected at different layers. The topology nodes at different layers can be different, therefore, the data indicators to be collected at different layers can be different.
[0102] Optionally, the metrics management module can provide users with a configuration file upload interface, through which users can upload data metric configuration files, which may include the data metrics required by the application system. Optionally, the metrics management module can provide a metrics configuration page, through which users can select or configure the data metrics required by the application system.
[0103] (4) Rule Management Module: Also known as the rule engine, this module configures and manages expert rules used in fault detection. Expert rules are rules used to make an overall judgment on the application system's faults. Specifically, expert rules are rules formed based on past experience; by combining operational data with expert rules, faults in the application system can be identified. For example, expert rules can be formed based on park-level business indicators and the experience of operation and maintenance experts, thereby judging the application system's faults from an overall perspective.
[0104] Optionally, the rule management module can provide a rule selection page, which displays expert rules that users can choose from. Users can select from these expert rules and configure the expert rules that the application system needs to use when detecting faults based on the user's selection.
[0105] (5) Model management module, which can be used to configure and manage the fault detection models required for fault detection. Optionally, the fault detection model can be a neural network model, and each fault detection model can output corresponding fault detection results based on the input application system's operating data. Optionally, different fault detection models can be used to output fault detection results corresponding to different levels. Optionally, different fault detection models use different fault detection methods.
[0106] One implementation involves a model management module providing a model selection page that displays available fault detection models for the user to choose from. Another implementation allows users to input model requirements, indicating their specific needs for the required model or the fault detection requirements of the application system. The model management module can then match the most suitable fault detection model to the user based on these requirements. A further implementation establishes a correspondence between models and application system levels. For example, different fault detection models can be used for fault detection at different levels. The user can specify the required fault detection level, and the corresponding fault detection model can be determined based on this correspondence. When multiple applicable fault detection models exist for each level, the user can specify one or more models as the fault detection models for that level, or determine one or more models based on model metrics (such as accuracy).
[0107] (6) A topology management module, which can be used to configure and manage the topology relationships between topology nodes included in the application system. Optionally, the topology management module can provide a configuration file upload interface, through which users can upload system configuration files. These configuration files include system configuration information that indicates the topology relationships between topology nodes in the application system. Alternatively, the topology management module can also provide a system configuration editing page, where users can edit topology nodes and the topology relationships between them, such as adding topology nodes and editing the topology relationships between pairs of topology nodes.
[0108] This application does not limit the specific configuration of the above modules in the embodiments.
[0109] III. Relevant functions during service use, which can be used to provide users with relevant functions when using the fault detection service. For example... Figure 1C As shown, service configuration-related functions may include the following modules:
[0110] (1) A detection execution module is used to execute the fault detection process. For example, the detection execution module can perform fault detection on the application system according to the method provided in the embodiments of this application, obtain multiple fault detection results, and fuse the multiple fault detection results to obtain the final fused detection result of the application system. Specifically, when the obtained multiple fault detection results are fault detection results of different levels, the embodiments of this application can fuse the multiple fault detection results according to the topological relationship between the topological nodes in the application system, and finally obtain the fused detection result corresponding to one level. In this way, by fusing fault detection results of different levels, the accuracy of fault detection of the application system can be improved.
[0111] (2) The detection results module can be used to display the detection process when providing fault detection services to users, as well as the fault detection results, such as displaying the final fused detection results or intermediate detection results obtained during the fault detection process. Figure 1C As shown, the relevant functions when using the service may include a detection results module. For example, the detection results module can provide users with a display interface of the fault detection process, showing the detection progress and other data involved in the fault detection process. Alternatively, the detection results module can also provide users with a display interface of the fault detection results, allowing users to view the fault detection results.
[0112] The functions of the above-mentioned modules can be further divided into different sub-modules, or the functions of multiple modules in the above-mentioned modules can be merged into one block. This application embodiment does not impose specific limitations on this.
[0113] See Figure 1D This is a schematic diagram of another architecture of the cloud management platform 100 provided in this application embodiment. The cloud management platform 100 may include service purchase-related functions, service configuration-related functions, and service usage-related functions. These will be described in detail below.
[0114] I. Service Purchase Related Functions: These functions allow users to purchase fault detection services. For an introduction to these functions, please refer to [link to relevant documentation]. Figure 1C The introduction is only a partial one, so I won't go into too much detail.
[0115] II. Service configuration-related functions can be used to provide users with functions related to configuring the fault detection service, or to configure functions related to the fault detection service according to user instructions. After configuration, the fault detection service can execute the fault detection process according to the configured content. As a possible example, such as... Figure 1D As shown, service configuration-related functions may include the following modules:
[0116] (1) Resource Configuration Module: This module is used to configure the resources used by the fault detection service, which may include computing resources and storage resources. After configuration through the resource configuration module, the fault detection service can be deployed on the corresponding resources and utilize these resources to implement the fault detection service.
[0117] (2) Basic Configuration Module: This module is used to configure the basic information of the fault detection service. This basic information can also be understood as the basic configuration information of the application system. For example, this basic information may include the application system's topology mode (or topology type). The application system's topology mode refers to the topology mode adopted by the application system's system architecture. Different topology modes result in different distributions of system resources, and correspondingly, the topology relationships relied upon during fault detection and the output detection results may differ. Optionally, the application system's topology mode may include single-region-single AZ-single database, single-region-multiple AZ-multiple database, single-region-dual AZ-primary / standby database, or dual-region-multiple AZ-multiple database (dual-active) topology modes, etc. Each topology mode will be introduced in the relevant sections below, and will not be elaborated on here.
[0118] Optionally, multiple topology modes can be displayed in the configuration interface, allowing users to select the topology mode for the application system to be fault-detected. Alternatively, if the displayed topology modes do not include the application system's topology mode, users can create a new topology mode and edit its related information.
[0119] Optionally, the basic information may also include other information about the application system. For example, the basic information may include candidate fault types of the application system, that is, one or more of these candidate fault types that the fault detection ultimately needs to output. There are no restrictions on the specific information included in the basic information.
[0120] (3) Indicator Management Module: The fault detection process depends on the collected application system operation data. The indicator management module can be used to configure and manage the data indicators required by the application system. Then, when the application system is running, the operation data can be collected according to the configured data indicators, and fault detection can be performed according to the operation data.
[0121] (4) The method management module can be used to configure and manage the fault detection methods used in fault detection. Fault detection methods can also be called fault detection models or fault detection schemes, etc. The method management module provides a variety of fault detection methods for users to choose from, and users can select the fault detection method they need.
[0122] (5) Fault Drill Module: This module can be used to configure relevant information during fault drills, such as configuring the fault mode library used during fault drills and which faults in the fault mode library will be drilled. After the user configures the relevant information during fault drills, the fault drill module can execute the fault drill process according to the user's configuration. Fault drills can also be called fault simulation or chaos engineering, which refers to simulating the occurrence of various faults in the application system. By collecting fault drill data during fault drills, the fault detection data can be detected by the fault detection method configured by the user to obtain the corresponding reference detection results. Then, based on the reference detection results, the priority of each fault detection method under a specific fault type in a specific topology mode (or topology type) can be determined. Subsequently, when the fault detection results output by multiple fault detection methods are different, the priority of the fault detection methods can be combined to fuse the fault detection results output by multiple fault detection methods to obtain the final fused detection result.
[0123] This application does not limit the specific configuration of the above modules in the embodiments.
[0124] III. Relevant functions during service use, which can be used to provide users with relevant functions when using the fault detection service. For example... Figure 1D As shown, service configuration-related functions may include the following modules:
[0125] (1) A detection execution module is used to execute the fault detection process. For example, the detection execution module can perform fault detection on the application system according to the method provided in the embodiments of this application, obtain multiple fault detection results, and fuse the multiple fault detection results to obtain the final fused detection result of the application system. Specifically, when the obtained multiple fault detection results are different, the embodiments of this application can combine the priority of the fault detection methods to fuse the fault detection results output by the multiple fault detection methods to obtain the final fused detection result. In this way, the priority determined by prior experience obtained through actual fault drills is more accurate, and the fused detection result obtained according to the priority is more accurate, which helps to improve the accuracy of fault detection in the application system.
[0126] (2) The detection result module can be used to display the detection process when providing fault detection services to users, as well as the fault detection result display function, such as displaying the final fused detection result or the intermediate detection result obtained during the fault detection process.
[0127] The functions of the aforementioned modules can be further broken down into different sub-modules, or the functions of multiple modules can be merged into one block. This application does not impose specific limitations on this. Figure 1D and Figure 1C For similar or identical modules, please refer to Figure 1C The details of that part will not be repeated here.
[0128] In some embodiments, considering the increasingly complex topology of application systems, the complexity of fault detection also increases, making it necessary for application systems to configure corresponding fault detection services. For example, such fault detection services need to have fault delimitation capabilities to quickly pinpoint the location of the fault when an application system malfunctions, thereby enabling rapid recovery of the application system's services.
[0129] To improve the accuracy of fault detection in application systems, fault detection often needs to be performed from different dimensions. This leads to different fault detection methods determining different dimensions of the fault detection results. Dimension can also be understood as the output format of the fault detection results (e.g., fault granularity). To address this issue, this application provides a cloud service-based fault detection method. This method can perform fault detection from different layers for layered application systems, and it can also fuse the fault detection results from different layers to obtain a final fused detection result. In this application, fault detection can refer to fault delimitation, fault location, or root cause localization, etc. The following mainly uses fault detection for fault delimitation as an example for description.
[0130] refer to Figure 2 This is a flowchart illustrating a method provided in an embodiment of this application, which can be applied to... Figure 1A Application scenarios or Figure 1C In the system architecture, the method may include the following steps:
[0131] Step 201: Provide a configuration interface for the fault detection service. The configuration interface receives system configuration information and a first service requirement from the user. The application system adopts a layered architecture, which includes multiple layers and multiple topology nodes. Each topology node is located in one of the multiple layers. The system configuration information indicates the topological relationships between the multiple topology nodes, and the first service requirement indicates that fault detection should be performed on at least two of the multiple layers.
[0132] The fault detection service can be understood as a cloud service provided by the cloud management platform. This platform offers a variety of cloud services, and users can choose to purchase the services that meet their needs. The cloud management platform provides a purchase interface for the fault detection service. After logging into the cloud management platform with their account and password, users can purchase the fault detection service through this interface. After purchasing the service, it will be available in the user's list of services.
[0133] Before performing fault detection on the application system based on this fault detection service, the service needs to be configured. The cloud management platform can provide users with a configuration interface for the fault detection service. Users can configure the service through this interface, and the information they input during configuration can help implement the fault detection service. If the user is an existing user (i.e., the user has used the fault detection service before), their historical configuration information can be retrieved and reused. Alternatively, the user can adjust the historical configuration information to use as the current configuration information for the fault detection service. For example, historical configuration information can be displayed to the user, allowing them to make adjustments. If the user is a new user (i.e., the user has never used the fault detection service before), they need to add their own configuration information. Optionally, to help users configure quickly, multiple configuration templates can be provided. Users can choose one configuration module or adjust a configuration template before use. It should be noted that the configuration of the fault detection service can be a one-time process; once configured, the information does not need to be configured again during subsequent service use. Therefore, step 201 is not a mandatory step.
[0134] Combination Figure 2 As shown, the cloud management platform provides a configuration interface for fault detection services. Users can input the system configuration information and primary service requirements of the application system through the client on this configuration interface. The client can then send the system configuration information and primary service requirements of the application system input by the user to the cloud management platform.
[0135] In this embodiment, the configuration interface can be used to receive system configuration information of the application system input by the user. The application system may adopt a layered architecture, which may include multiple levels, each level comprising multiple topology nodes, with each topology node located in one of the multiple levels. Alternatively, it can be described as the application system's topology including multiple topology nodes distributed across multiple levels, with each level comprising a subset of topology nodes. This system configuration information can be used to indicate the topological relationships between multiple topology nodes; that is, the system configuration information can be used to indicate the topological relationships between any two topology nodes. For example, based on the system configuration information, other topology nodes contained in (or associated with, or dependent on) a topology node can be determined, such as see [reference needed]. Figure 1BIn an application system, application services can be a layer, and microservices can be a layer. The system configuration information identifies the microservices included in an application service. Alternatively, components can be a layer, and containers can be a layer; the system configuration information identifies the containers included in a component. Optionally, the system configuration information can be a topology graph, or it can be a topology hierarchy table, which may include the topology nodes subordinate to each top-level node. Alternatively, the system configuration information can be represented in other forms without limitation.
[0136] The configuration of system configuration information can be implemented in a variety of ways.
[0137] One implementation involves a configuration interface that receives system configuration files input by the user. This configuration file includes system configuration information. For example, the configuration interface may include an upload interface for the system configuration file, allowing the user to upload it. The cloud management platform can then receive the uploaded configuration file. See also [example description]. Figure 3A The configuration interface 301 shown includes a file selection control 3011 and an upload control 3012. Users can select the system configuration file to be uploaded by operating the file selection control 3011, and users can upload the selected system configuration file to the cloud management platform by operating the upload control 3012.
[0138] Another implementation involves a configuration interface that receives user input for editing operations. These operations are used to edit topology nodes and the topological relationships between them. In other words, the configuration interface allows users to edit the topology structure of the application system. For example, the configuration interface may include a topology drawing page (or a drawing sub-interface) where users can draw the application system's topology. See also [example...] for an example. Figure 3A The configuration interface 301 shown includes an add node control 3013, a connect node control 3014, and a drawing area 3015. Users can add topology nodes in the drawing area 3015 by using the add node control 3013, connect topology nodes by using the connect node control 3014, and preview and edit the topology structure in the drawing area 3015. Circles and squares represent different levels of hierarchy.
[0139] The two implementation methods described above can be as follows: Figure 3A The above two implementation methods can be implemented on the same page or through different configuration interfaces. In other words, the two implementation methods can correspond to different pages, and there are no specific restrictions on this.
[0140] The cloud management platform provides users with the ability to view the topology of configured application systems. (See [link]). Figure 3A The "Topology View" option allows users to view the topological relationships within the application system. This can be done through a topology diagram, a topology hierarchy table, or other possible methods without restriction.
[0141] The configuration interface in this embodiment can also be used to receive a first service request from the application system input by the user. The first service request indicates that fault detection should be performed on at least two of the multiple levels, that is, the levels that the fault detection service needs to detect. Different levels correspond to topological nodes of different granularities (dimensions) in the application system, so the first service request can also be understood as indicating the granularity (or dimension) of the fault to be detected.
[0142] The configuration of the first service requirement can be implemented in a variety of ways.
[0143] One implementation involves a configuration interface that can display multiple levels. At least two levels can be selected from these levels, and the configuration interface can then receive the user's selection of at least two levels, thus obtaining the first service request. For example, see [link to example]. Figure 3B The configuration interface 302 shown may include example 3021, which corresponds to Figure 1B The configuration interface 302 shows the hierarchical structure of the application system's resource vertical topology, from which the user can select at least two hierarchical levels. The configuration interface 302 may also include example 3022, which corresponds to example 3021. Figure 1B The application system's application-level topology, as shown, includes layers from which the user can select at least two layers. Alternatively, the configuration interface 302 may also include all the layers in Examples 3021 and 3022, allowing the user to select at least two layers. Furthermore, the configuration interface 302 may include layers beyond those indicated by the system configuration information. For example, these layers could be derived from the topology relationships indicated by the system configuration information, such as a layer including topology nodes like hardware, software, and external dependencies. Alternatively, they could be user-defined layers, such as... Figure 3B The system can provide users with a custom interface 3023, allowing users to customize the required levels.
[0144] Another implementation involves implicitly selecting the level by choosing a fault detection method, since different levels can employ different fault detection methods, or in other words, different fault detection methods detect faults at different levels. Different fault detection methods differ in their fault detection principles, required input data, and output fault detection results. The configuration interface can receive at least two fault detection methods from the user. Each fault detection method can be used to obtain a fault detection result corresponding to one of at least two levels; that is, each fault detection method can be used to perform fault detection for a specific level and output the corresponding fault detection result for that level. In this embodiment, performing fault detection for a specific level means that the fault indicated by the final output fault detection result is located at that level, but it is not necessarily limited to fault detection based solely on data related to that level.
[0145] Optionally, the configuration interface can display multiple optional fault detection methods, allowing users to select at least two. Alternatively, the configuration interface can provide a loading interface for fault detection methods, enabling users to load the required at least two methods. Furthermore, the configuration interface can provide a custom interface for fault detection methods, allowing users to customize the required at least two methods. These various implementation methods can also be combined to implement input for at least two fault detection methods.
[0146] Optionally, fault detection methods can be categorized into different types, and different types of fault detection methods can be input separately. For example, fault detection methods can be divided into expert rules and fault detection models. Expert rules refer to rules formed based on past experience; by combining the application system's operational data with expert rules, faults in the application system can be identified. In other words, the configuration interface can receive at least one expert rule input by the user, where each expert rule is used to obtain a fault detection result corresponding to a specific level. Optionally, the configuration interface can display multiple expert rules for the user to choose from, allowing the user to select at least one expert rule. See the example below. Figure 3C As shown in the configuration interface 303, this configuration interface 303 includes an expert rule selection area 3031, where users can select the desired expert rules. Alternatively, the configuration interface can also provide a custom interface for expert rules, through which users can customize expert rules. For example, see [link to example]. Figure 3CAs shown in configuration interface 303, this interface includes an expert rule customization area 3032. Users can edit relevant information about custom expert rules in this area, such as the rule name, description, and content. Finally, they can submit the confirmation to generate a new expert rule. Alternatively, other configuration methods can be used. For example, configuration interface 303 can also provide an expert rule upload interface, allowing users to upload configuration files to input expert rules. There are no specific restrictions on the configuration method. Configuration interface 303 is only one possible example; there are no limitations on the specific configuration interface for expert rules.
[0147] Optionally, the cloud management platform can provide users with a viewing function for expert rules that have already been entered by the user; see [link to relevant documentation]. Figure 3C In the "Rule View" section, users can view the expert rules they have entered, including the rule name, rule description, and specific rule content.
[0148] A fault detection model is a model that needs to be trained using a specific model training method before it can be used. Each fault detection model can output corresponding fault detection results based on the operational data of the input application system. For example, a fault detection model can be a deep learning model or a neural network model. In other words, the configuration interface can be used to receive at least one fault detection model input by the user. Specifically, after training each fault detection model using its corresponding training method, each fault detection model is used to obtain a fault detection result corresponding to a certain level based on the corresponding fault detection method. Optionally, the configuration interface can display multiple fault detection models for the user to choose from, allowing the user to select at least one. For example, see [link to example]. Figure 3D As shown in the configuration interface 304, this interface includes a model selection area 3041. The model selection area 3041 can display some or all models in the current model library. Users can select the desired fault detection models in the model selection area 3041. Alternatively, the configuration interface can also provide a file upload interface for fault detection models, allowing users to upload model files. For example, see [link to example]. Figure 3D As shown in configuration interface 304, this interface includes a file addition interface 3042 and a file upload interface 3043. Users can add model files through the file addition interface 3042 and upload model files through the file upload interface 3043. Alternatively, other configuration methods can be used, without specific restrictions. Configuration interface 304 is only one possible example, and there are no restrictions on the configuration interface for specific fault detection models.
[0149] Fault detection models can be categorized into supervised training and unsupervised training methods based on their training approaches. Supervised training methods utilize collected training data (including labeled training data) for supervised training. Once the model reaches convergence, it can be used for fault detection in real-world processes. Unsupervised training methods do not require labeled training data, making them less complex than supervised methods. However, they are limited in accuracy, and the application scenarios of supervised and unsupervised training models may differ; for example, the types of fault detection results they output may differ.
[0150] In addition to the above configuration, other possible configurations may be included, and this application embodiment does not limit them.
[0151] Step 202: Obtain runtime data from multiple topology nodes during application system operation. For example... Figure 2 As shown, the cloud management platform can obtain runtime data of the application system during its operation.
[0152] Obtaining runtime data from multiple topology nodes during application system operation can be based on user-inputted system configuration information and primary service requirements. In other words, after the user inputs system configuration information and primary service requirements, the fault detection service for the application system has been configured and can then be used to detect faults in the application system by collecting runtime data from the application system.
[0153] Operational data from multiple topology nodes can be acquired at various times during application system runtime. For example, application system runtime data can be collected synchronously during application system operation. When an application system failure occurs, fault detection can be performed directly based on this runtime data, reducing latency in waiting for runtime data acquisition. Additionally, it reduces the probability of failing to collect runtime data when the application system fails, ensuring sufficient data support for fault detection. Another example is acquiring runtime data from multiple topology nodes when the application system fails. In other words, the application system failure is the trigger condition for runtime data acquisition. This eliminates the need for additional runtime data acquisition during application system operation, reducing the processing burden on the cloud management platform and saving storage space by avoiding excessive storage of runtime data.
[0154] There are several ways to obtain runtime data from multiple topology nodes during the application system's operation.
[0155] One implementation involves the cloud management platform sending a data acquisition request to the application system. This request requests runtime data from multiple topology nodes within the application system. The cloud management platform then receives the runtime data from the application system. The data acquisition request can specify the desired runtime data. For example, the cloud management platform can send a data acquisition request to the application system's management device. After acquiring the corresponding runtime data, the management device then sends the runtime data back to the cloud management platform.
[0156] Another implementation method is that the application system can provide a data acquisition interface to the cloud management platform, and the cloud management platform can obtain the running data of multiple topology nodes of the application system during runtime through the data acquisition interface.
[0157] Another implementation involves the cloud management platform including one or more agent modules. These agent modules run within the application system and collect runtime data from multiple topology nodes. The cloud management platform can then obtain this runtime data through these agent modules. The number of agent modules and their location within the application system can be configured based on the specific needs of the application system. For example, if the application system comprises two regions, there could be two agent modules, each located in the management device or gateway device of one region, capable of collecting runtime data from each region. Of course, the number of agent modules can also be increased; there is no limitation on this.
[0158] Another implementation involves retrieving runtime data from the database if the application system reports runtime data to the database. For example, the cloud management platform can subscribe to specific data in the database (i.e., runtime data required for fault detection), and retrieve the updated data when the database is updated.
[0159] Optionally, the required runtime data can be user-configurable, meaning the user can specify which data metrics need to be collected. The configuration interface in this embodiment can also receive user-inputted data metric configuration information. This information indicates the data metrics to be collected for each topology node across multiple layers. Based on this configuration information, runtime data from multiple topology nodes can be obtained during application system operation. Data metrics for any two layers can be the same or different. For example, data metrics can be configured based on the characteristics of each topology node. Data metrics can be understood as parameters or feature dimensions, and runtime data includes the values of these parameters or feature dimensions. Data metrics can include common metrics across all layers and metrics unique to specific layers. For example, common metrics might be central processing unit (CPU) utilization or memory usage. When a topology node is a container, unique metrics might be the number of times the container is called (or run).
[0160] The configuration of data indicator information can be implemented in various ways.
[0161] One implementation involves a configuration interface that receives user-inputted data metric configuration files, which include data metric configuration information. For example, the configuration interface may include an upload interface for the data metric configuration files, allowing users to upload such files. See also [example...] Figure 3E The configuration interface 305 shown includes a file addition interface 3051 and a file upload interface 3052. Users can add data indicator configuration files through the file addition interface 3042 and upload them through the file upload interface 3052. The upload interface can also use other formats; no specific restrictions are placed on this.
[0162] Another implementation involves a configuration interface that displays multiple data metrics, allowing users to select one. This interface can also receive the selected metrics from the user's selection and generate corresponding data metric configuration information, including the user-selected metrics. For example, see [link to example]. Figure 3EThe configuration interface 305 shown includes an indicator selection area 3053. This area displays selectable data indicators for each level, allowing users to select the desired data indicators for each level. Alternatively, the indicator selection area 3053 can display all data indicators, allowing users to bind specific indicators to different levels as the data indicators required for each level. Or, the indicator selection area 3053 can also display general data indicators available for user selection; the selected indicator will be collected for all levels, and corresponding unique data indicators will be displayed for each level, allowing users to choose from these options.
[0163] The two implementation methods described above can be used individually or in combination. Alternatively, other configuration methods can be adopted, such as providing users with custom data metrics. This application embodiment does not impose specific limitations on this. Configuration interface 305 is only one possible example, and there are no limitations on the configuration interface for specific data metrics.
[0164] Step 203: Determine the fault detection results corresponding to at least two levels based on the operating data of multiple topology nodes.
[0165] In this embodiment, the collected operational data from multiple topology nodes can be used to determine fault detection results corresponding to at least two levels. The determination of the fault detection result for each level can be based on the operational data of some or all of the multiple topology nodes. Alternatively, for each fault detection method, the required input operational data can be some or all of the collected operational data.
[0166] For example, the fault detection results corresponding to at least two levels can include one or more implementation methods:
[0167] The first implementation method is based on fault detection results using expert rules. This method takes a holistic approach to the application system, combining the experience of operations and maintenance experts to use rules for judgment to determine faults in the application system.
[0168] For example, see Figure 4A As shown in (a), an expert rule can be formed based on park-level business indicators and the experience of operation and maintenance experts to determine the fault category from the perspective of the overall application system. The formed expert rule can be used to detect faults based on the operational data of the application system in actual application scenarios. By inputting operational data and combining it with the expert rule, fault detection results based on the expert rule can be obtained. Figure 4A Figure (a) illustrates a fault detection result based on expert rules, which outputs one or more fault types belonging to the application system, such as... Figure 4AThe diagram illustrates hardware failures, software failures, and external dependency failures. Hardware failures refer to failures in the hardware system of an application system, that is, failures in the system resources upon which the application system depends for operation; therefore, they can also be called resource failures or system failures. Software failures refer to failures in the software system of an application system, that is, failures in the application programs running within the application system; therefore, they can also be called application failures. External dependency failures refer to failures in external resources that the application system depends on for operation; for example, external resources may include database resources, and therefore, they can also be called database failures. Of course, failure types can also be classified in other ways, and there are no specific restrictions on this.
[0169] The second implementation method utilizes fault detection results from a fault detection model trained using supervised training. This method learns from the application's historical faults. Specifically, it collects data metrics from the application system's normal and faulty phases during fault drills, labels these metrics, and then performs supervised training. The labels represent fault detection results for a specific level of topology nodes. These labels are used to calculate the model's loss value and adjust parameters during supervised training. Once the fault detection model reaches convergence, it can be used to detect faults in that level of topology nodes within the application system. For example, to detect faults at the basic module level, fault labels for the basic module can be created from the collected runtime data. Supervised training using these labels is then performed, and once the fault detection model reaches convergence, it can be used to detect faults at the basic module level of the application system.
[0170] For example, see Figure 4A As shown in (b), after supervised training is completed and the running data is input, the fault detection results of each basic module in the application system can be obtained, such as... Figure 4A The fault detection results of basic modules 1 to P shown can include the probability that the basic module belongs to each fault type, such as... Figure 4A The diagram shows the probability of basic module 1 being fault-free, the probability of it being a hardware fault, the probability of it being a software fault, and the probability of it being an external dependency fault. It should be noted that... Figure 4A The fault detection result shown in (b) is only an example, and the embodiments of this application do not limit it.
[0171] The third implementation method is based on the fault detection results of a fault detection model trained using unsupervised training. This method combines the historical operational data of topology nodes in the application system during normal (i.e., fault-free) phases with unsupervised training (i.e., unsupervised learning) to learn the characteristics of the application system during normal phases. This can be understood as reconstructing a model of the application system during normal phases, and then detecting whether topology nodes are faulty based on the reconstruction error between the characteristics of the actual scenario (i.e., the characteristics obtained from the actual collected operational data) and the learned characteristics of the normal phases.
[0172] Taking a topology node as a container node as an example, it can be based on the container node (such as...) Figure 1B The historical normal operation data of the application container and database container (shown) are used to build an unsupervised normal state reconstruction model. Finally, the reconstruction error between the actual operating data of the container node and the normal operation data is used to determine whether the container node is faulty. For example, see [link to example]. Figure 4A As shown in (c), after inputting the runtime data, the fault detection results for each container can be obtained, which may include the fault probability or the probability of no fault for that container, such as... Figure 4A The failure probabilities of containers 1 to L are shown in (c).
[0173] It should be noted that whether a topology node is faulty can also be detected by other methods besides the unsupervised training model. For example, a supervised training model can also be used for detection, and there are no specific restrictions on this.
[0174] Step 204: Based on the topological relationship, fuse the fault detection results corresponding to at least two levels to obtain the fused detection result. For a detailed implementation of step 204, please refer to the following sections.
[0175] Step 205: Display the fusion detection results of the application system on the result interface of the fault detection service.
[0176] In this embodiment, the cloud management platform can provide a result interface for fault detection services. The final fused detection result can be displayed in the result interface, allowing users to view the fused detection result and understand the faults existing in the current application system. The fused detection result refers to the detection result obtained by fusing the fault detection results corresponding to at least two levels. That is, the method in this embodiment can fuse fault detection results from different levels and output them.
[0177] Optionally, when displaying the fusion detection results on the results interface, a topology diagram of the application system can be shown, and faulty topology nodes in the application system can be marked on the topology diagram based on the fusion detection results. This allows users to more intuitively view the faulty topology nodes. Alternatively, other methods can be used to display the fusion detection results; there are no specific limitations on this.
[0178] Optionally, the user can specify the level at which the final output fusion detection result is located. That is, the configuration interface is also used to receive a second service request input by the user, which indicates N candidate faults of the application system. The level corresponding to the N candidate faults is higher than any of the at least two levels; that is, the level at which the final output fusion detection result is located is not in the level specified in the first service request, but the level at which the fault indicated by the second service request is located is higher than the at least two levels specified in the first service request. Alternatively, the level corresponding to the N candidate faults is the highest level among the at least two levels.
[0179] Furthermore, based on the second service request input by the user, the fusion detection results can be displayed on the results interface. These fusion detection results include at least one candidate fault from N candidate faults. This candidate fault is obtained by fusing the fault detection results of at least two levels according to the topological relationship, using a hierarchical approach from low to high. In other words, the user can specify N candidate faults, and the final output fault can be one or more of these N candidate faults. The user-specified faults belong to the highest level. During the fault detection process, fault detection is also performed from other levels. The fault detection results from other levels are fused layer by layer from low to high until the fault detection result of the highest level is obtained, which is the final fusion detection result.
[0180] Optionally, the user's input of a second service request can be implemented in several ways. One implementation is that the configuration interface can display multiple candidate faults, from which the user can select N candidate faults. Another implementation is that the user does not need to input anything additionally and can determine the highest-level candidate fault from at least two levels as N candidate faults. Yet another implementation is that the user can define N candidate faults. Alternatively, other implementation methods can be used, without specific restrictions.
[0181] Optionally, the results interface can also display detailed information about the fault detection process. For example, the results interface can provide a details viewing interface, which users can trigger to display detailed information.
[0182] Optionally, in addition to displaying the fusion detection results, the results interface can also display other fault detection results. For example, other fault detection results may include fault detection results corresponding to at least two of the above levels, as well as intermediate results generated during the fusion process.
[0183] In one example, at least two levels may include a first level and a second level. The first level includes multiple first topology nodes, and the second level includes multiple second topology nodes. A first topology node may include one or more second topology nodes; that is, the first level is higher than the second level. The fault detection results corresponding to the aforementioned at least two levels may include the first fault detection results corresponding to the first level and the second fault detection results corresponding to the second level. The first level and the second level can be any two of the multiple levels included in the application system. The following examples illustrate several possible scenarios, but are not limited to these:
[0184] Case 1: The first level is the subsystem level. A subsystem can be, for example, a hardware subsystem, a software subsystem, or an externally dependent subsystem. That is, a first topology node can be one of these three subsystems. The second level is... Figure 1B The basic module hierarchy shown, i.e., a second topology node, can be a basic module. The first fault detection result can indicate a subsystem fault, and the second fault detection result can indicate a basic module fault. For example, the first fault detection result can indicate whether the faulty subsystem is a hardware subsystem, a software subsystem, or an externally dependent subsystem. Figure 4A The first fault detection result shown in (a) may include the fault probabilities corresponding to the hardware subsystem, software subsystem, and externally dependent subsystem, respectively. The second fault detection result may indicate the fault of each basic module, i.e., whether each basic module is faulty, and if so, the specific fault type, such as... Figure 4A The second fault detection result shown in (b) includes the probability that each basic module is fault-free, the probability that it is a hardware fault, the probability that it is a software fault, and the probability that it is an external dependency fault.
[0185] Scenario 2: The first level is the subsystem level, meaning a first topology node can be one of a hardware subsystem, a software subsystem, or an externally dependent subsystem. The second level is... Figure 1B The container hierarchy shown means that a second topology node can be a container. A first fault detection result can indicate a subsystem fault, and a second fault detection result can indicate a container fault. For example, the first fault detection result can indicate which subsystem—hardware subsystem, software subsystem, or externally dependent subsystem—is experiencing the fault. Figure 4AThe first fault detection result shown in (a) may include the fault probabilities corresponding to the hardware subsystem, software subsystem, and externally dependent subsystem, respectively. The second fault detection result may indicate the fault of each container, that is, whether each basic module is faulty, for example... Figure 4A The second fault detection result shown in (c) includes the probability of failure or the probability of no failure for each container.
[0186] Case 3, the first level is Figure 1B The basic module hierarchy shown indicates that a first topology node can be a basic module. The second level is... Figure 1B The container hierarchy shown means that a second topology node can be a container. The first fault detection result can indicate a fault in a basic module, and the second fault detection result can indicate a fault in a container. For example, the first fault detection result can indicate a fault in each basic module, that is, whether each basic module is faulty, and if so, the specific fault type. Figure 4A The second fault detection result shown in (b) includes the probability that each basic module is fault-free, the probability that it is a hardware fault, the probability that it is a software fault, and the probability that it is an external dependency fault. The second fault detection result can indicate the fault of each container, that is, whether each basic module is faulty, for example... Figure 4A The second fault detection result shown in (c) includes the probability of failure or the probability of no failure for each container.
[0187] It should be noted that the first and second layers can differ depending on the system architecture, and can be configured according to actual needs. Furthermore, this example only uses the first and second layers; however, in real-world scenarios, more layers can be included. When other layers are included, the same principle can be applied to their integration, and this embodiment does not impose any limitations on this.
[0188] One possible approach to fusion is to transfer the fault detection results corresponding to the second level to the first level. (See also...) Figure 4B As shown, since there is a topological relationship between the second topological node and the first topological node, if the second topological node fails, the first topological node with which it has a topological relationship will inevitably be affected, and its failure probability will be greatly increased. In other words, the first topological node will also fail. Therefore, the second fault detection result can be converted to the first level through this topological relationship to obtain the third fault detection result corresponding to the first level. The third fault detection result can include the converted fault information of each first topological node in the first level.
[0189] In one possible implementation, the first fault detection result is used to indicate fault information of multiple first topology nodes, and the third fault detection result includes multiple confidence levels, each corresponding one-to-one with a first topology node. That is, the transformed fault information of each first topology node included in the third fault detection result is the confidence level of the fault information of the first topology node. Confidence level, also known as reliability or confidence probability, is used to characterize the reliability of the fault information of the first topology node. Converting the second-level fault detection result into the confidence level of the first-level fault detection result can be understood as using the second-level fault detection result as a verification of the first-level fault detection result to verify its reliability, thereby improving the reliability of the final fault detection result.
[0190] In one example, the first fault detection result indicates the fault to which each first topology node belongs. The first fault detection result can indicate that the first topology node is fault-free, or it can indicate which of N candidate faults the fault of the first topology node belongs to, including a first fault, a second fault, and a third fault. For example, the first fault could be a hardware fault, the second fault could be a software fault, and the third fault could be an external dependency fault. The second fault detection result indicates the fault probability of multiple second topology nodes respectively. The second fault detection result can then be converted into a confidence level of the fault information of the first topology nodes at the first level. The specific conversion method can be as follows:
[0191] First, the confidence level of each first topology node being fault-free is determined based on the fault probabilities of multiple second topology nodes. That is, for a first topology node, the confidence level of the first topology node being fault-free is determined based on the fault probabilities of the second topology nodes that have a topological relationship with the first topology node. In other words, the confidence level of the first topology node being fault-free is determined based on the fault probabilities of the second topology nodes included in the first topology node. Accordingly, the third fault detection result may include the confidence level of each first topology node being fault-free.
[0192] Optionally, the confidence level of a first topology node being fault-free can be determined based on the maximum fault probability of all second topology nodes included in the first topology node, or it can be determined based on the minimum fault-free probability of all second topology nodes included in the first topology node. For example, the confidence level of a first topology node being fault-free can be the difference between the upper probability limit and the maximum fault probability of all second topology nodes included in the first topology node. For example, if the maximum fault probability of all second topology nodes is X, then the confidence level of the first topology node being fault-free is 1-X. Alternatively, the confidence level of a first topology node being fault-free can be the minimum fault-free probability of all second topology nodes included in the first topology node.
[0193] Second, the confidence level of each first topological node belonging to the first fault is determined based on the fault probabilities of multiple second topological nodes. That is, for a first topological node, the confidence level of the first topological node belonging to the first fault can be determined based on the fault probabilities of the second topological nodes that have a topological relationship with the first topological node. In other words, the confidence level of the first topological node belonging to the first fault can be determined based on the fault probabilities of the second topological nodes included in the first topological node. Accordingly, the third fault detection result may include the confidence level of each first topological node belonging to the first fault.
[0194] Optionally, the confidence level of a first topology node belonging to the first fault can be determined based on the maximum fault probability of all second topology nodes included in that first topology node. Taking the first fault as a hardware fault as an example, when a hardware fault occurs, the operation of all topology nodes in the application system may be affected. Therefore, the confidence level of belonging to the first fault can be determined based on the fault probabilities of all second topology nodes. For example, the confidence level of a first topology node belonging to the first fault can be the maximum fault probability of all second topology nodes included in that first topology node.
[0195] Third, the confidence level of each first topology node belonging to the second fault is determined based on the respective fault probabilities of the first type of second topology nodes among multiple second topology nodes. The second fault affects the operation of the first type of second topology nodes. This can be understood as some or all of the first type of second topology nodes failing to operate normally when the third fault occurs. Therefore, for a given first topology node, the confidence level of the first topology node belonging to the second fault can be determined based on the respective fault probabilities of the first type of second topology nodes that have a topological relationship with it. Alternatively, the confidence level of the first topology node belonging to the second fault can be determined based on the respective fault probabilities of the first type of second topology nodes included in the first topology node. Accordingly, the third fault detection result can include the confidence level of each first topology node belonging to the second fault.
[0196] Optionally, the confidence level of a first topology node belonging to the second fault can be determined based on the maximum fault probability of second topology nodes of the first type included in the first topology node. Taking the aforementioned second fault as a software fault as an example, when a software fault occurs, the operation of software-related topology nodes (such as container nodes) within the application system may be affected. Therefore, the confidence level of belonging to the second fault can be determined based on the fault probability of second topology nodes of the first type. For example, the confidence level of a first topology node belonging to the second fault can be the maximum fault probability of second topology nodes of the first type included in the first topology node.
[0197] Fourth, the confidence level of each first topology node belonging to the third fault is determined based on the failure probabilities of the second type of second topology nodes among multiple second topology nodes. The third fault affects the operation of the second type of second topology nodes. This can be understood as some or all of the second type of second topology nodes failing to operate normally when the third fault occurs. Therefore, for a first topology node, the confidence level of the first topology node belonging to the third fault can be determined based on the failure probabilities of the second type of second topology nodes that have a topological relationship with it. Alternatively, the confidence level of the first topology node belonging to the third fault can be determined based on the failure probabilities of the second type of second topology nodes included in the first topology node. Accordingly, the third fault detection result can include the confidence level of each first topology node belonging to the third fault.
[0198] Optionally, the confidence level of a first topology node belonging to the third fault can be determined based on the maximum fault probability of the second type of second topology nodes included in that first topology node. Taking the aforementioned third fault as an external dependency fault (such as a database fault) as an example, when an external dependency occurs, the operation of the topology nodes (such as database nodes) related to the external dependency within the application system may be affected. Therefore, the confidence level of belonging to the third fault can be determined based on the fault probability of the second type of second topology nodes. For example, the confidence level of a first topology node belonging to the third fault can be the maximum fault probability of the second type of second topology nodes included in that first topology node.
[0199] See also Figure 4B As shown, after obtaining the third fault detection result, the first and third fault detection results can be fused to obtain the fourth fault detection result. The fourth fault detection result can indicate the fused fault information of each first topology node.
[0200] In this case, both the first fault detection result and the third fault detection result correspond to the first level, therefore their fusion is a fusion of fault detection results at the same level. The fusion of the first fault detection result and the third fault detection result can be implemented in various ways.
[0201] One implementation is that the fault detection method corresponding to the first fault detection result and the fault detection method corresponding to the third fault detection result can have different priorities. Then, the first fault detection result and the third fault detection result can be fused by referring to the different priorities.
[0202] Another implementation method is to fuse the first fault detection result and the third fault detection result using a voting method. For example, the first fault detection result can indicate the probability of each first topology node belonging to each of the multiple faults, and the third fault detection result can indicate the confidence level of the probability of each first topology node belonging to each of the multiple faults. Then, for a first topology node, the confidence probability of the first topology node belonging to a fault can be determined based on the probability of the first topology node belonging to a fault and the corresponding confidence level. For example, see Table 1, which is a fusion example of a first topology node. A1 to A4 are the probabilities of the first topology node belonging to each category (i.e., the first fault detection result), and B1 to B4 are the confidence levels of the first topology node belonging to each category (i.e., the third fault detection result). The corresponding confidence probability can be obtained based on the probability and confidence level of each category. An*Bn in Table 1 is only a calculation example and is not specifically limited. For each first topology node, the confidence probability shown in Table 1 can be obtained, that is, the fused fault information of each first topology node can be several confidence probabilities as shown in Table 1.
[0203]
[0204]
[0205] Table 1
[0206] Another implementation method is to input the first fault detection result and the third fault detection result into a pre-trained fusion model to obtain the fourth fault detection result output by the model.
[0207] Alternatively, other possible fusion methods may be adopted, and this application embodiment does not limit this.
[0208] Optionally, the fourth fault detection result can also indicate the confidence probability of each category, that is, the confidence probability of each first topology node belonging to each category is further fused. For example, the confidence probability shown in Table 1 can be obtained for each first topology node, and finally the confidence probability of each category can be determined. The confidence probability of a category is the sum of the confidence probabilities of all first topology nodes in that category, thus obtaining the fourth fault detection result.
[0209] Therefore, when displaying the fusion detection results on the results interface, the displayed fusion detection results may also include the third fault detection results and / or the fourth fault detection results.
[0210] The following section will use specific examples to illustrate the layered fusion process described above.
[0211] See Figure 5As shown, this example employs three fault detection methods for fault detection. These three methods, based on different perspectives (i.e., levels or dimensions) and fault detection processes, yield different fault detection results. The three fault detection methods are as follows:
[0212] Method 1, such as Figure 4A As shown in (a), this method forms expert rules based on park-level business indicators and the experience of operation and maintenance experts. It judges the fault category from the perspective of the entire application and finally outputs the fault detection results based on the expert rules, that is, it outputs the fault type (or the probability of each fault type) to which the application system belongs. Figure 4A The hardware failures, software failures, and external dependency failures shown are as follows.
[0213] Method 2, such as Figure 4A As shown in Figure (b), this method uses a supervised multi-classification model for fault detection. That is, it combines historical data and data indicators of the normal and fault stages of the application system collected from fault drills, labels them, and then performs supervised training to detect faults in the basic modules. It outputs the fault classification result of each basic module, that is, the category to which each basic module belongs, or the probability of each basic module belonging to each category, namely, the probability of belonging to no fault, the probability of belonging to hardware fault, the probability of belonging to software fault, and the probability of belonging to external dependency fault.
[0214] Method 3, such as Figure 4A As shown in (c), this method uses an unsupervised model for fault detection, that is, it performs state detection on the deployed basic nodes, such as containers or VMs. Taking containers as an example, it performs unsupervised learning based on the historical normal data of containers to obtain a normal state reconstruction model of containers. Finally, it can determine whether the container is faulty based on the reconstruction error between the running data collected in the actual scenario and the normal state reconstruction model, and finally output the fault probability of each container.
[0215] Since the three fault detection methods described above yield fault ranges with varying accuracy and may also result in different fault categories, a comprehensive approach is required. In this example, a layered fusion strategy is employed, integrating fault detection results from different levels layer by layer to improve the accuracy and reliability of fault detection. First, the fault classification results of the basic module in Method 2 are fused with the container states in Method 3. Since the fault classification results of the basic module correspond to the container states, the node states in Method 3 are converted into the confidence levels of the fault types in the basic module of Method 2. Second, the results of all basic modules are fused using a majority vote. Finally, the resulting model is fused with the expert rule results from Method 1. If the confidence probability of the model result is low, the result of the expert rule from Method 1 is selected as the fallback result; otherwise, the model result is output.
[0216] See also Figure 5 As shown, the specific steps are as follows:
[0217] Step S10: Perform underlying fusion, which involves fusing the fault classification results of the basic modules with the state results of the containers.
[0218] In step S10, the node abnormal state detection results from method three are converted into the confidence levels of module fault categories from method two, achieving underlying data fusion. This conversion of confidence probabilities enhances the accuracy of the fault classification module and the reliability of the results.
[0219] In this embodiment, the application system may include multiple basic modules, each of which contains multiple application containers and database containers. Method 2 outputs the fault classification result for each basic module, which can be represented as: [probability of no fault, probability of hardware fault, probability of software fault, probability of external dependency fault]. Method 3 performs fault detection on each container to obtain container-level detection results, which can be represented as: [fault probability, no fault probability].
[0220] Specifically, the detection results at the container level of Method 3 are converted into the confidence levels of the fault types of the basic modules of Method 2. For any basic module A, the conversion method can be as follows:
[0221] (1) Confidence that the basic module A is fault-free: Take the minimum value of the fault-free probability of all containers in the basic module A.
[0222] (2) Confidence that basic module A is a hardware failure: take the maximum value of the failure probability of all containers in basic module A.
[0223] (3) Confidence of failure of basic module A as an external dependency: take the maximum value of the failure probability of the database container in basic module A.
[0224] (4) Confidence that basic module A is a software fault: take the maximum value of the failure probability of the application container in basic module A.
[0225] After the transformation, for basic module A, the product of the probability that basic module A belongs to each category and the corresponding confidence level can be calculated, and this product can be used as the confidence probability that basic module A belongs to each category.
[0226] Step S11: Perform intermediate layer fusion, that is, fuse the fault classification results of the basic modules.
[0227] The intermediate layer refers to the integration and processing layer of fault classification results from the basic modules fused from the bottom layer. In this layer, the results of different basic modules are merged according to certain rules to obtain more comprehensive fault detection results.
[0228] Optionally, the fusion of the intermediate layer can employ a voting mechanism. This involves determining the confidence probability of each category based on the confidence probabilities of all basic modules belonging to the same category, thus fusing the fault classification results of each basic module. For example, the sum of the confidence probabilities of all basic modules belonging to the same category can be used as the confidence probability of a single category. For instance, adding the confidence probabilities of all basic modules belonging to no faults yields the confidence probability of the application system being fault-free; adding the confidence probabilities of all basic modules belonging to hardware faults yields the confidence probability of the application system belonging to hardware faults, and so on. By fusing the fault detection results of multiple basic modules, the accuracy of decision-making is increased, and the risk of misjudgment by a single module is reduced.
[0229] Step S12: Perform top-level fusion, which involves fusing the model results with the expert rule results. The model results refer to the results obtained after the intermediate-level fusion described above, while the expert rule results refer to the results obtained using Method 1.
[0230] The top layer is the stage where the model results are fused with the expert rules from Method 1. The top layer integrates information from the bottom and intermediate layers and incorporates expert rule judgments to improve the accuracy of the final fault identification. During top-level fusion, if the confidence probability of the model results obtained from the intermediate layer fusion is low, the expert rule results from Method 1 will be prioritized as a fallback strategy to ensure that expert rules provide a final guarantee when the model results are unreliable.
[0231] See Figure 5 As shown, if the confidence probability of the model result obtained from the intermediate layer fusion is high (e.g., greater than or equal to a set threshold), the model result is directly output. The model result may include information such as fault type, the basic module of the fault, and the container of the fault. Therefore, the result interface can display information such as fault type, the basic module of the fault, and the container of the fault. For example, to facilitate users' understanding of the connections between faults at different levels, faults at different levels can be displayed together, for example, using a chain structure. This chain structure allows users to see the location of faults at each level from the top to the bottom. If the confidence probability of the model result obtained from the intermediate layer fusion is low (e.g., less than a set threshold), the expert rule result is directly output. The expert rule result may include the fault type.
[0232] Through the aforementioned decision-making scheme based on hierarchical fusion, from the direct processing of bottom-level data to the fusion of results in the middle layer, and finally to the top-level final decision, a comprehensive, multi-faceted, and highly reliable integrated delimitation and fault detection system is formed. This not only improves the accuracy and robustness of fault detection, but also makes full use of the complementarity of information between layers, achieving efficient location and handling of faults in complex systems.
[0233] In some embodiments, to improve the accuracy of fault detection in application systems, multiple fault detection methods can be used. However, the fault detection results determined by different fault detection methods may conflict. To address this issue, this application provides another fault detection method based on cloud services. This method can pre-perform fault drills for various fault types under various topology types, collect fault drill data, and then perform fault detection using different fault detection methods based on the fault drill data to determine the priority of each fault detection method for a specific fault type under a specific topology type. In practical application scenarios, when the fault detection results determined by different fault detection methods conflict, the different fault detection results can be fused according to the priority to obtain the final fused detection result. Here, fault detection in this application embodiment can refer to fault delimitation, fault location, or root cause location, etc. The following mainly uses fault detection for fault delimitation as an example for description.
[0234] refer to Figure 6 This is another flowchart illustrating the method provided in this application embodiment, which can be applied to... Figure 1A Application scenarios or Figure 1D In the system architecture, the method may include the following steps:
[0235] Step 601: Provide a configuration interface for fault detection services. The configuration interface is used to receive user input regarding the application system's topology type and third-party service requirements. The third-party service requirements indicate N types of faults in the application system, where N is a positive integer greater than 1.
[0236] The cloud management platform provides a purchase interface for fault detection services. After logging into the platform with their account and password, users can purchase fault detection services through this interface. After purchase, the service will be available in the user's list of services. For example, through... Figure 1D The service purchase module enables users to purchase fault detection services.
[0237] Combination Figure 6 As shown, the cloud management platform provides a configuration interface for fault detection services. Users can input the topology type of the application system and the third-party service requirements through the client on this configuration interface. The client can then send the user-input topology type of the application system and the third-party service requirements to the cloud management platform.
[0238] Before performing fault detection on the application system based on this fault detection service, the service needs to be configured. The cloud management platform can provide users with a configuration interface for the fault detection service. Users can configure the service through this interface, and the information entered during configuration can help implement the fault detection service. It should be noted that the configuration of the fault detection service can be a one-time process; that is, after the configuration is completed once, the information does not need to be configured again in subsequent service use. Therefore, step 601 is not a mandatory step.
[0239] In this embodiment of the application, the configuration interface can be used to receive user input of the application system's topology type, which can also be called architecture type, topology mode, or architecture topology mode, etc.
[0240] Optionally, the topology type can be divided from the perspective of the architecture adopted by the application system. For example, the application system may adopt a microservice architecture, a monolithic architecture, a service-oriented architecture, etc.
[0241] Optionally, the topology type can be categorized from the perspective of application system resource distribution. Different topology types result in different system resource distributions, and consequently, the topology relationships relied upon for fault detection and the output detection results can differ. Several topology types are illustrated below:
[0242] Topology type 1, single region - single AZ - single database topology type. See also Figure 7A The image shows an example of topology type 1, in which the application system's infrastructure is deployed in a single region, which includes an Availability Zone (AZ), and the AZ also contains only one database.
[0243] Topology type 2, a single-region, multi-AZ, multi-database topology. See also... Figure 7B The image shows an example of topology type 2, in which the application system's infrastructure is deployed in a single region that includes multiple Availability Zones (AZs), such as... Figure 7B As shown in AZ1 to AZ3, multiple AZs adopt a distributed database architecture.
[0244] Topology type 3, a single-region, dual-AZ, master-slave database topology. See also... Figure 7C The image shows an example of topology type 3, in which the application system's infrastructure is deployed in a single region that includes two AZs, such as... Figure 7B As shown in AZ1 to AZ2, multiple AZs deploy their own databases. The two databases adopt a master-slave database architecture, with one database as the master database and the other as the standby database.
[0245] Topology type 4, a dual-region, multi-AZ, multi-database (active-active) topology type. See also... Figure 7D The image shows an example of topology type 4, in which the application system's infrastructure is deployed in two regions, such as... Figure 7D The regions shown are 1 and 2, each comprising multiple AZs, such as... Figure 7D The area shown includes AZ1 to AZ3, and area 2 also includes AZ1 to AZ3. The AZs in the two areas are different. A distributed database architecture is used between multiple AZs in a region, and data consistency is maintained between databases in different regions using bidirectional replication technology of data replication service (DRS).
[0246] Topology type configuration can be implemented in various ways. For example, the configuration interface can display multiple topology types, allowing users to select the topology type corresponding to their application system. Alternatively, the configuration interface can provide an input interface through which users can input the topology type corresponding to their application system. Yet another option is to provide a topology type recognition interface, where users can input necessary information, such as system configuration information or a description of the application system, and the topology type corresponding to the application system can be identified based on the user's input.
[0247] In this embodiment, the configuration interface can also be used to receive third service requirements from the application system input by the user. These third service requirements indicate N types of faults in the application system, and the final output fusion detection result is one or more of these N fault types. That is, the cloud management platform can provide a configuration interface for users to input detection targets, thereby determining the user's needs based on the user's input, such as determining the fault type the user wants to detect and the level of the fault type.
[0248] In some scenarios, users do not need to input third-party service requests. The cloud management platform can identify possible fault types based on the application system's system configuration information. Alternatively, the cloud management platform can use the default fault type without making specific restrictions.
[0249] In addition to the above configuration, other possible configurations may be included, and this application embodiment does not limit them.
[0250] Step 602: Obtain runtime data from the application system. For example... Figure 6 As shown, the cloud management platform can obtain runtime data of the application system during its operation.
[0251] Obtaining runtime data of an application system can be based on the topology type and third-party service requirements input by the user. In other words, after the user inputs the topology type and third-party service requirements, the fault detection service for the application system has been configured. Fault detection of the application system can then be achieved through this fault detection service, thereby collecting runtime data of the application system for fault detection.
[0252] For the timing and implementation method of acquiring runtime data, please refer to the introduction in step 202 above, which will not be repeated here.
[0253] Optionally, the configuration interface can also be used to receive user-inputted data indicator configuration information. This information indicates the data indicators required to be collected for the application system, allowing the system to obtain runtime data based on the configuration information. For details on configuring data indicator information, please refer to step 202 above; it will not be repeated here.
[0254] Step 603: Determine at least two fault detection results based on the operating data using at least two fault detection methods, with each fault detection result corresponding to one fault detection method; wherein, under each fault type of the topology type, at least two fault detection methods have corresponding priorities.
[0255] Optionally, at least two fault detection methods can be the default methods used by the cloud management platform. Alternatively, the at least two fault detection methods can also be user-configured. That is, the configuration interface is also used to receive at least two fault detection methods configured by the user. For example, the configuration interface can display multiple optional fault detection methods, from which the user can select at least two. Alternatively, the configuration interface can also provide a fault detection method loading interface, through which the user can load the required at least two fault detection methods. Or, the configuration interface can also provide a fault detection method customization interface, through which the user can customize the required at least two fault detection methods. All of the above implementation methods can also be combined to implement the input of at least two fault detection methods.
[0256] This application does not limit the specific fault detection method. The fault detection method can be a detection method based on expert rules, a detection method based on a supervised learning model, a detection method based on an unsupervised learning model, or any other possible detection method.
[0257] Optionally, considering that different fault detection methods may have different accuracies, this application embodiment assigns different priorities to different fault detection methods, and the fusion of fault detection results can be performed with reference to the priority of the fault detection methods.
[0258] In this embodiment, the priority of each fault detection method can be obtained through a fault drill process. The process for determining the priority of each fault detection method is described below. (See also...) Figure 8 The diagram shown is a flowchart illustrating the priority determination process.
[0259] Step S20: The preparation phase of the fault drill. In this phase, an exhaustive approach can be used to identify all possible topology types, fault mode libraries, and fault types for the application system. This may include identifying and defining the application system's topology type, establishing a fault mode library (which may include all potential faults and their impact on the application system), and determining fault types. Fault type determination can be based on the fault's source, scope of impact, or nature, etc. One fault type can correspond to at least one fault in the fault mode library.
[0260] Table 2 below shows examples of variables that may be involved in fault drills. The number of schemes represents the number of at least two fault detection methods used in fault detection, denoted by N, where N takes the value 3, indicating three fault detection methods. Credibility refers to the probability (or reliability) of the fault detection result obtained using a particular fault detection method during the fault drill, denoted by T, where T ranges from 0 to 1. The credibility threshold is used in real-world scenarios to measure whether the current fault detection method is used. The values for topology type, fault, and fault type can be values from a code table containing candidate values for topology type, fault, and fault type. The code tables corresponding to topology type, fault, and fault type can be the same or different.
[0261] variable name Variable symbol Range of values Value examples Number of options N 3 3 Credibility T 0~1 0.5 Credibility threshold Θ 0.7 0.7 Priority P 1,2..N 3 Topology AM Code table Active-active architecture in the same city, primary and backup database Fault FM Code table Database connection interrupted Fault type FT Code table Database failure
[0262] Table 2
[0263] Step S21: Conduct fault drills for application systems requiring fault detection. During fault drills, faults can be selectively introduced into specific application systems to obtain fault drill data when fault samples are running in the application system.
[0264] Optionally, the fault simulation process can also be tailored to the user's specific needs. For example, the cloud management platform's configuration interface can also display various faults from the fault mode library, and receive M faults selected by the user from these faults, where M is a positive integer greater than 1. Each of the N fault types corresponds to at least one fault in the fault mode library.
[0265] After the user selects the M types of faults to be practiced in the application system, multiple fault examples corresponding to the M faults can be injected into the application system, and fault practice data can be obtained when running multiple fault examples in the application system. This fault practice data can also be called fault data samples, which can be used to determine the reference detection results of the application system. Specifically, the application system can run each of the M fault examples separately to obtain the corresponding fault practice data, that is, inject one fault at a time to collect the operational data when that fault occurs; or it can run multiple fault examples within the M fault examples to obtain the corresponding fault practice data, that is, inject multiple faults at a time to collect the operational data when that fault occurs; or it can combine these two methods to obtain richer fault practice data.
[0266] Optionally, the injected fault in this application embodiment can be a hardware fault, such as network latency or storage failure, or a software fault, such as service crash, or a database fault, such as database deadlock.
[0267] Step S22: Analyze the fault drill results. At least two fault detection methods can be used to determine the reference detection results of the application system based on the fault drill data. The reference detection results can indicate the fault type corresponding to a fault drill data point. Based on the reference detection results corresponding to the at least two fault detection methods, the priority of each fault type under the topology type of the application system is determined.
[0268] One implementation involves calculating scores for at least two fault detection methods for each fault type within the application system's topology, based on reference detection results, and then determining priorities based on these scores. For a fault type within a topology, the higher the accuracy of the reference detection results obtained by a fault detection method, the higher its corresponding score. Besides accuracy, the scoring can also consider other factors, such as the complexity of the fault detection method; generally, higher complexity should result in a lower score.
[0269] Another implementation is to display the reference detection results for each fault type of the application system's topology to the user, and the user can input the priority of each fault detection method.
[0270] The two implementation methods described above can be used individually or in combination. For example, after prioritizing the data using the first method, the user can confirm the priority a second time to improve the accuracy of the final priority.
[0271] Optionally, the priority can be determined based on the number of fault detection methods. For example, if there are 5 fault detection methods, the priority can include 5 levels, so each fault detection method can have a different priority. Optionally, the priority can be a fixed set of levels. For example, the priority can be divided into three levels: high, medium, and low. High is represented by the number 3, medium by the number 2, and low by the number 1. Priorities can be assigned to different fault detection methods based on the reference detection results of each method. Different fault detection methods can have the same or different priorities.
[0272] Step S23: Display the fault drill results on the configuration interface. The fault drill results include the priority of at least two fault detection methods under each fault type of the topology type.
[0273] The cloud management platform's configuration interface can also be used to display the final fault drill results. Specifically, the configuration interface can display the priorities of at least two fault detection methods for each fault type within the application system's topology. Optionally, the cloud management platform can provide a priority viewing interface. Upon receiving a user's trigger command for this interface, the configuration interface will display the priorities of at least two fault detection methods for each fault type within the application system's topology.
[0274] The priorities obtained above can be used to fuse multiple fault detection results when an application system actually fails.
[0275] Step 604: Based on the priority of at least two fault detection methods, fuse the results of at least two fault detection methods to obtain a fused detection result.
[0276] In this embodiment of the application, when an application system actually fails, at least two of the above-mentioned fault detection methods can be used for fault detection, resulting in at least two corresponding fault detection results. For example, when an application system actually fails, the fault type can be quickly identified based on the existing fault mode library and previous drill experience.
[0277] Optionally, for any one of the at least two fault detection results, it may include the fault type to which the application system belongs and the probability (or confidence level) of belonging to that fault type. When the probability is less than or equal to a certain threshold, the fault detection result can be considered invalid and can be screened out from the at least two fault detection results, meaning that the fault detection result does not need to participate in the subsequent fusion process. When the probability is greater than a certain threshold, the fault detection result can be considered valid and can participate in the subsequent fusion process.
[0278] Optionally, if at least two fault detection results are the same, that is, if the fault types indicated by at least two fault detection results are consistent, then the final fault detection result is output directly.
[0279] Optionally, if at least two fault detection results are different, that is, if at least two fault detection results are the same and indicate at least two fault types, or if at least two fault types are in complete or partial conflict, then at least two fault detection results need to be fused to output the final fused detection result.
[0280] For example, if each of the at least two fault detection results indicates the probability of a fault type belonging to the application system, the conflict can be resolved by calculating the weight of each fault detection result. That is, the weight of each fault type can be determined based on the probability of the fault type indicated by the at least two fault detection results and their corresponding priorities. The weight of each fault type is the product of the reliability of the fault detection result and the priority of the fault detection method.
[0281] For example, the weights of several fault detection results are calculated as follows:
[0282] The probability that fault detection result 1 indicates fault type A is 0.85. The priority of the fault detection method corresponding to fault detection result 1 under fault type A in the topology of this application system is 2 (i.e., medium). Therefore, the weight of fault detection result 1 is 0.85*2=1.7.
[0283] The probability of fault detection result 2 indicating fault type B is 0.7. The priority of the fault detection method corresponding to fault detection result 2 under fault type B in the topology of this application system is 1 (i.e., medium). Therefore, the weight of fault detection result 2 is 0.7*1=0.7.
[0284] The probability that fault detection result 3 indicates fault type B is 0.75. The priority of the fault detection method corresponding to fault detection result 3 under fault type A in the topology of this application system is 1 (i.e. low). Therefore, the weight of fault detection result 3 is 0.75*1=0.75.
[0285] For different solutions with the same fault detection result, their corresponding weights can be summed. As shown above, the three fault detection results indicate fault type A and fault type B. Fault type A has a final weight of 1.7, while fault type B has a final weight of 1.45. Since the final weight of fault type A is greater than that of fault type B, it can be determined that the application system belongs to fault type A. In other words, after obtaining the weight of each fault type, the fault type with the highest weight can be selected as the final detection result, allowing the user to choose the most suitable solution for fault repair based on the final detection result.
[0286] Step 605: Display the fusion detection results of the application system on the result interface of the fault detection service.
[0287] In this embodiment, the cloud management platform can provide a result interface for fault detection services. The final fused detection result can be displayed on the result interface, allowing users to view the fused detection result and understand the faults existing in the current application system. The fused detection result refers to the result obtained by fusing at least two fault detection results according to the priorities of at least two fault detection methods. That is, the method in this embodiment can fuse different fault detection results and output them based on the priorities obtained from fault drills.
[0288] Specifically, such as Figure 6 As shown, users can trigger the viewing of fault detection results through client operations. Correspondingly, the cloud management platform can send the required fusion detection results to the client based on this operation, and the client can display the fusion detection results in the results interface.
[0289] Optionally, when displaying the fusion detection results on the results interface, a topology diagram of the application system can be shown, and faulty topology nodes in the application system can be marked on the topology diagram based on the fusion detection results. This allows users to more intuitively view the faulty topology nodes. Alternatively, other methods can be used to display the fusion detection results; there are no specific limitations on this.
[0290] Optionally, the final output fusion detection result may include the weight of each of the multiple fault types indicated by at least two fault detection results, and / or the fault type with the highest weight among the multiple fault types.
[0291] Through the above implementation methods, combining a fault mode library, fault drills, and architectural topology patterns, a comprehensive decision is made based on the fault detection results obtained from multiple fault detection methods. This process embodies a systematic and scientific fault response mechanism. By pre-emptively "inoculating" faults through fault drills, the adaptability and resilience of the application system when encountering real faults are improved. Furthermore, the accuracy and efficiency of fault detection results are ensured through dual screening based on priority and credibility.
[0292] The two fault detection methods provided in this application can be implemented individually or in combination. When implemented in combination, the second fault detection method can be used to fuse multiple fault detection results at the same level during the layered fusion process described above, in order to further improve the accuracy of fault detection.
[0293] Based on the same inventive concept, this application also discloses a cloud service-based fault detection device, which can be applied to a cloud management platform for managing the infrastructure that provides cloud services. The infrastructure includes multiple regions, each region including at least one cloud data center, and the cloud services run on at least one server in at least one cloud data center located in one of the multiple regions. Figure 9 This is a schematic diagram of the structure of a cloud service-based fault detection device provided in an embodiment of this application, as shown below. Figure 9 As shown, the device includes:
[0294] Service configuration module 901 is used to provide a configuration interface for fault detection services. The configuration interface is used to receive system configuration information and first service requirements of the application system input by the user. The application system adopts a layered architecture, which includes multiple layers and multiple topology nodes. Each topology node is located in one of the multiple layers. The system configuration information is used to indicate the topological relationship between the multiple topology nodes. The first service requirement indicates that fault detection is performed for at least two of the multiple layers.
[0295] The data acquisition module 902 is used to acquire the running data of multiple topology nodes during the application system's operation. The running data of the multiple topology nodes is used to determine the fault detection results corresponding to at least two levels.
[0296] The result display module 903 is used to display the fusion detection results of the application system on the result interface of the fault detection service. The fusion detection results are obtained by fusing the fault detection results of at least two levels according to the topology relationship.
[0297] In one possible implementation, the configuration interface is further configured to receive a second service request input by the user, the second service request indicating N candidate faults of the application system, wherein the level corresponding to the N candidate faults is higher than any of the at least two levels, or the level corresponding to the N candidate faults is the highest level among the at least two levels.
[0298] The result display module 903 is specifically used to display the fusion detection results on the result interface based on the second service request input by the user. The fusion detection results include at least one candidate fault from N candidate faults. This at least one candidate fault is obtained by fusing the fault detection results corresponding to at least two levels layer by layer according to the topological relationship, using a hierarchical approach from low to high.
[0299] In one possible implementation, at least two levels include a first level and a second level. The first level includes multiple first topology nodes, and the second level includes multiple second topology nodes. The fault detection results corresponding to the at least two levels include the first fault detection result corresponding to the first level and the second fault detection result corresponding to the second level. The fused detection result includes a third fault detection result, which is determined based on the topology relationship and the second fault detection result, and the third fault detection result corresponds to the first level; and / or, the fused detection result includes a fourth fault detection result, which is obtained by fusing the third fault detection result and the first fault detection result.
[0300] In one possible implementation, the third fault detection result includes multiple confidence levels, and the first fault detection result is used to indicate fault information of multiple first topology nodes, with each confidence level corresponding to one of the multiple first topology nodes.
[0301] In one possible implementation, the second fault detection result indicates the fault probability of each of the plurality of second topology nodes. The third fault detection result includes one or more of the following: a confidence level that each first topology node is fault-free, determined based on the fault probabilities of the plurality of second topology nodes; a confidence level that each first topology node belongs to a first fault among N candidate faults, determined based on the fault probabilities of the plurality of second topology nodes; a confidence level that each first topology node belongs to a second fault among N candidate faults, determined based on the fault probabilities of second topology nodes of a first type among the plurality of second topology nodes, the second fault affecting the operation of the first type of second topology nodes; or, a confidence level that each first topology node belongs to a third fault among N candidate faults, determined based on the fault probabilities of second topology nodes of a second type among the plurality of second topology nodes, the third fault affecting the operation of the second type of second topology nodes.
[0302] In one possible implementation, the configuration interface is also used to receive user-inputted data indicator configuration information, which indicates the data indicators to be collected for each topology node in multiple layers. The data acquisition module 902 is specifically used to acquire runtime data from multiple topology nodes during application system operation based on the data indicator configuration information.
[0303] In one possible implementation, the configuration interface is further configured to receive a data indicator configuration file input by the user, the data indicator configuration file including data indicator configuration information; or, the configuration interface is further configured to display multiple data indicators and receive a data indicator selected by the user from the multiple data indicators, the data indicator configuration information including the data indicator selected by the user.
[0304] In one possible implementation, the configuration interface is used to display multiple levels and receive at least two levels selected by the user from the multiple levels; or, the configuration interface is used to receive at least two fault detection methods input by the user, wherein each fault detection method is used to obtain a fault detection result corresponding to one of the at least two levels.
[0305] In one possible implementation, the configuration interface is used to receive at least one expert rule input by the user, wherein each expert rule is used to obtain a fault detection result corresponding to a level; or, the configuration interface is used to receive at least one fault detection model input by the user, wherein after each fault detection model is trained using the corresponding training method for each fault detection model, each fault detection model is used to obtain a fault detection result corresponding to a level according to the corresponding fault detection method.
[0306] In one possible implementation, the configuration interface is used to receive system configuration files input by the user, the system configuration files including system configuration information; or, the configuration interface is used to receive editing operations input by the user, the editing operations being used to edit topology nodes and edit the topological relationships between topology nodes.
[0307] In one possible implementation, the result display module 903 is specifically used to display the topology diagram of the application system on the result interface, and to mark the faulty topology nodes in the application system in the topology diagram according to the fusion detection results.
[0308] This device can be used to perform, for example Figure 2 The method steps shown in the embodiments are therefore relevant to the description of the apparatus. Figure 2 The description of the illustrated embodiments will not be repeated here.
[0309] It should be noted that each of the above modules can be implemented in software or hardware. For example, the implementation of service configuration module 901 will be described below. Similarly, the implementation of other modules can refer to the implementation of service configuration module 901.
[0310] When implemented in software, the service configuration module 901 can be an application or code block running on a computer device. The computer device can be at least one of a physical host, virtual machine, container, or other computing device. Furthermore, there can be one or more computer devices. For example, the service configuration module 901 can be an application running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the application can be distributed within the same availability zone (AZ) or in different AZs. Similarly, the multiple hosts / virtual machines / containers used to run the application can be distributed within the same region or in different regions. Typically, a region can include multiple AZs.
[0311] Similarly, multiple hosts / virtual machines / containers used to run the application can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a region can include multiple VPCs, and a VPC can include multiple Availability Zones (AZs).
[0312] When implemented in hardware, the service configuration module 901 may include at least one computing device, such as a server. Alternatively, the cloud resource configuration service configuration module 901 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0313] The service configuration module 901 includes multiple computing devices that can be distributed within the same Availability Zone (AZ) or in different AZs. Similarly, the cloud resource configuration service configuration module 901 includes multiple computing devices that can be distributed within the same region or in different regions. Likewise, the service configuration module 901 includes multiple computing devices that can be distributed within the same VPC or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0314] It should be noted that each of the above modules can be used to perform some or all of the steps in the cloud service-based fault detection method.
[0315] The cloud-based fault detection device disclosed in this application has clearly defined roles and close cooperation among its various modules. The modules work together to efficiently complete fault detection services for application systems.
[0316] Based on the same inventive concept, this application also discloses another cloud service-based fault detection device, which can be applied to a cloud management platform for managing the infrastructure that provides cloud services. The infrastructure includes multiple regions, each region including at least one cloud data center, and the cloud services run on at least one server in at least one cloud data center located in one of the multiple regions. Figure 10 This is a schematic diagram of the structure of a cloud service-based fault detection device provided in an embodiment of this application, as shown below. Figure 10 As shown, the device includes:
[0317] Service configuration module 1001 is used to provide a configuration interface for fault detection services. The configuration interface is used to receive the topology type of the application system and the third service requirements input by the user. The third service requirements indicate N types of faults of the application system, where N is a positive integer greater than 1.
[0318] The data acquisition module 1002 is used to acquire runtime data of the application system during runtime.
[0319] The detection execution module 1003 is used to determine at least two fault detection results based on the running data using at least two fault detection methods, with each fault detection result corresponding to one fault detection method; wherein, under each fault type of the topology type, the at least two fault detection methods have corresponding priorities.
[0320] The result display module 1004 is used to display the fusion detection results of the application system on the result interface of the fault detection service. The fusion detection results are obtained by fusing at least two fault detection results according to the respective priorities of at least two fault detection methods.
[0321] In one possible implementation, the configuration interface is also used to display a variety of faults in the fault mode library, and to receive M faults selected by the user from the variety of faults, where M is a positive integer greater than 1; wherein each of the N fault types corresponds to at least one fault in the fault mode library.
[0322] The device also includes a fault simulation module 1005, which is used to acquire fault simulation data when running multiple fault samples in the application system. The multiple fault samples correspond to M types of faults, and the fault simulation data is used to determine the reference test results of the application system.
[0323] The service configuration module 1001 is also used to display the fault drill results on the configuration interface. The fault drill results include the priority of at least two fault detection methods under each fault type of the topology type. The fault drill results are determined based on the reference detection results corresponding to the at least two fault detection methods.
[0324] In one possible implementation, each of the at least two fault detection results indicates the probability of the fault type to which the application system belongs. The fused detection result includes one or more of the following: the weight of each fault type among the multiple fault types indicated by the at least two fault detection results, wherein the weight of each fault type is determined based on the probability of the fault type indicated by the at least two fault detection results and the corresponding priority; or, the fault type with the highest weight among the multiple fault types.
[0325] In one possible implementation, the configuration interface is also used to receive data indicator configuration information input by the user, which indicates the data indicators that need to be collected for the application system.
[0326] The data acquisition module 1002 is specifically used to acquire runtime data of the application system based on the data indicator configuration information.
[0327] In one possible implementation, the configuration interface is also used to receive at least two fault detection methods configured by the user.
[0328] This device can be used to perform, for example Figure 6 The method steps shown in the embodiments are therefore relevant to the description of the apparatus. Figure 6 The description of the illustrated embodiments will not be repeated here.
[0329] It should be noted that each of the above modules can be implemented in software or hardware. For example, the implementation method of each module can also refer to the implementation method of the service configuration module 901 above, which will not be elaborated further.
[0330] This application also provides a computing device, which will be described below. Figure 11 , Figure 11 This is a schematic diagram of a computing device 1100 that runs a cloud-based fault detection method according to an embodiment of this application. The computing device 1100 includes a bus 1101, a processor 1103, a memory 1102, and a communication interface 1104. The processor 1103, memory 1102, and communication interface 1104 communicate with each other via the bus 1101. The computing device 1100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1100.
[0331] Bus 1101 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 11 The bus 1104 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 1104 may include a path for transmitting information between various components of the computing device 1100 (e.g., memory 1102, processor 1103, communication interface 1104).
[0332] The processor 1103 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0333] The memory 1102 may include volatile memory, such as random access memory (RAM). The processor 1103 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0334] The memory 1102 stores executable program code, which, when executed, enables a cloud-based fault detection method. In other words, the memory 1102 contains instructions from the cloud management platform for executing the cloud-based fault detection method.
[0335] The communication interface 1104 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1100 and other devices or communication networks.
[0336] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0337] Please see below Figure 12 , Figure 12 This is a schematic diagram of the structure of a computing device cluster that executes a cloud-based fault detection method according to an embodiment of this application. The computing device cluster includes at least one computing device 1100. The memory 1102 of one or more computing devices 1100 in the computing device cluster may store the same cloud management platform instructions for the cloud-based fault detection method.
[0338] In some possible implementations, one or more computing devices 1100 in the computing device cluster can also be used to execute some instructions of the cloud service-based fault detection method. In other words, a combination of one or more computing devices 1100 can jointly execute the instructions of the cloud service-based fault detection method.
[0339] It should be noted that the memory 1102 in different computing devices 1100 within the computing device cluster can store different instructions for executing some functions of the cloud management platform. That is, the instructions stored in the memory 1102 of different computing devices 1100 can achieve the aforementioned... Figure 9 or Figure 10 The functionality of one or more modules within it.
[0340] This application also provides a computer program product containing instructions. This computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any available medium. When the computer program product is run on at least one computer device, it causes the at least one computer device to perform the aforementioned fault detection method applied to a cloud management platform for cloud-based services.
[0341] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the aforementioned fault detection method applied to a cloud management platform for performing cloud-based services.
[0342] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of this application.
Claims
1. A fault detection method based on cloud services, characterized in that, The method is applied to a cloud management platform for managing infrastructure that provides cloud services. The infrastructure includes facilities located in multiple regions, each region including at least one cloud data center. The cloud services run on at least one server located in at least one cloud data center in one of the multiple regions. The method includes: A configuration interface for providing fault detection services is provided. The configuration interface is used to receive system configuration information and a first service requirement of the application system input by the user. The application system adopts a layered architecture, which includes multiple layers and multiple topology nodes. Each topology node is located in one of the multiple layers. The system configuration information is used to indicate the topological relationship between the multiple topology nodes. The first service requirement indicates that fault detection is performed for at least two of the multiple layers. The running data of the multiple topology nodes during the operation of the application system is obtained, and the running data of the multiple topology nodes is used to determine the fault detection results corresponding to the at least two levels; The result interface of the fault detection service displays the fusion detection result of the application system, which is obtained by fusing the fault detection results corresponding to the at least two levels according to the topology relationship.
2. The method according to claim 1, characterized in that, The configuration interface is also used to receive a second service request input by the user. The second service request is used to indicate N candidate faults of the application system. The level corresponding to the N candidate faults is higher than any of the at least two levels, or the level corresponding to the N candidate faults is the highest level among the at least two levels. The display of the fusion detection results of the application system on the result interface of the fault detection service includes: Based on the second service requirement input by the user, the fusion detection result is displayed on the result interface; wherein, the fusion detection result includes at least one candidate fault among the N candidate faults, and the at least one candidate fault is obtained by fusing the fault detection results corresponding to at least two levels layer by layer according to the topological relationship in a hierarchical direction from low to high.
3. The method according to claim 1 or 2, characterized in that, The at least two levels include a first level and a second level, the first level includes a plurality of first topology nodes, and the second level includes a plurality of second topology nodes; The fault detection results corresponding to the at least two levels include the first fault detection result corresponding to the first level and the second fault detection result corresponding to the second level. The fusion detection result includes a third fault detection result, which is determined based on the topology and the second fault detection result, and the third fault detection result corresponds to the first level; and / or, the fusion detection result includes a fourth fault detection result, which is obtained by fusing the third fault detection result and the first fault detection result.
4. The method according to claim 3, characterized in that, The third fault detection result includes multiple confidence levels, and the first fault detection result is used to indicate the fault information of the multiple first topology nodes. The multiple confidence levels correspond one-to-one with the multiple first topology nodes.
5. The method according to claim 3 or 4, characterized in that, The second fault detection result indicates the fault probability of each of the plurality of second topology nodes; The third fault detection result includes one or more of the following: The confidence level of each first topology node being fault-free is determined based on the respective fault probabilities of the plurality of second topology nodes; The confidence level of each first topology node belonging to the first fault of N candidate faults is determined based on the respective fault probabilities of the plurality of second topology nodes. The confidence level of each first topology node belonging to the N candidate faults is determined based on the fault probability of the first type of second topology nodes among the plurality of second topology nodes, and the second fault affects the operation of the first type of second topology nodes; or, The confidence level of each first topology node belonging to the third fault among the N candidate faults is determined based on the respective fault probabilities of the second type of second topology nodes among the plurality of second topology nodes, and the third fault affects the operation of the second type of second topology nodes.
6. The method according to any one of claims 1 to 5, characterized in that, The configuration interface is also used to receive data indicator configuration information input by the user, which is used to indicate the data indicators to be collected for the topology nodes of each of the multiple levels. Obtaining the runtime data of the multiple topology nodes during the operation of the application system includes: obtaining the runtime data of the multiple topology nodes during the operation of the application system based on the data indicator configuration information.
7. The method according to claim 6, characterized in that, The configuration interface is also used to receive data indicator configuration information input by the user, including: The configuration interface is also used to receive the data indicator configuration file input by the user, the data indicator configuration file including the data indicator configuration information; or... The configuration interface is also used to display multiple data metrics and receive data metrics selected by the user from the multiple data metrics, wherein the data metric configuration information includes the data metrics selected by the user.
8. The method according to any one of claims 1 to 7, characterized in that, The configuration interface is used to receive the first service request of the application system input by the user, including: The configuration interface is used to display multiple levels and to receive at least two levels selected by the user from the multiple levels; or, The configuration interface is used to receive at least two fault detection methods input by the user, wherein each fault detection method is used to obtain a fault detection result corresponding to one of the at least two levels.
9. The method according to claim 8, characterized in that, The configuration interface is used to receive at least two fault detection methods input by the user, including one or more of the following: The configuration interface is used to receive at least one expert rule input by the user, wherein each expert rule is used to obtain a fault detection result corresponding to a level; or, The configuration interface is used to receive at least one fault detection model input by the user. After each fault detection model is trained using the corresponding training method, each fault detection model is used to obtain a fault detection result corresponding to a level according to the corresponding fault detection method.
10. The method according to any one of claims 1 to 9, characterized in that, The configuration interface is used to receive system configuration information of the application system input by the user, including: The configuration interface is used to receive system configuration files input by the user, the system configuration files including the system configuration information; or... The configuration interface is used to receive editing operations input by the user, and the editing operations are used to edit topology nodes and edit the topology relationships between topology nodes.
11. The method according to any one of claims 1 to 10, characterized in that, The display of the fusion detection results of the application system on the result interface of the fault detection service includes: The results interface displays a topology diagram of the application system. Based on the fusion detection results, the faulty topology nodes in the application system are marked in the topology graph.
12. A fault detection method for an application system, characterized in that, The method is applied to a cloud management platform for managing infrastructure that provides cloud services. The infrastructure includes facilities located in multiple regions, each region including at least one cloud data center. The cloud services run on at least one server located in at least one cloud data center in one of the multiple regions. The method includes: A configuration interface for providing fault detection services is provided. The configuration interface is used to receive the topology type of the application system and the third service requirement input by the user. The third service requirement indicates N types of faults of the application system, where N is a positive integer greater than 1. Obtain runtime data of the application system during runtime; At least two fault detection results are determined based on the operational data using at least two fault detection methods, with each fault detection result corresponding to one fault detection method; wherein, under each fault type of the topology type, the at least two fault detection methods have corresponding priorities; The result interface of the fault detection service displays the fusion detection result of the application system, which is obtained by fusing the at least two fault detection results according to the respective priorities of the at least two fault detection methods.
13. The method according to claim 12, characterized in that, The configuration interface is also used to display a variety of faults included in the fault mode library, and to receive M faults selected by the user from the variety of faults, where M is a positive integer greater than 1; wherein, each of the N fault types corresponds to at least one fault among the M faults; The method further includes: Acquire fault drill data when running multiple fault examples in the application system, wherein the multiple fault examples correspond to the M types of faults, and the fault drill data is used to determine the reference detection result of the application system; The configuration interface displays the fault drill results, which include the priority of the at least two fault detection methods under each fault type of the topology type. The fault drill results are determined based on the reference detection results corresponding to the at least two fault detection methods.
14. The method according to claim 12 or 13, characterized in that, Each of the at least two fault detection results indicates the probability of the fault type to which the application system belongs; The fusion detection results include one or more of the following: The weight of each fault type among the multiple fault types indicated by the at least two fault detection results, wherein the weight of each fault type is determined based on the probability and corresponding priority of the fault type indicated by the at least two fault detection results; or, The fault type with the highest weight among the various fault types.
15. The method according to any one of claims 12 to 14, characterized in that, The configuration interface is also used to receive data indicator configuration information input by the user, and the data indicator configuration information is used to indicate the data indicators that need to be collected for the application system. The acquisition of runtime data of the application system includes: The runtime data of the application system is obtained based on the data indicator configuration information.
16. The method according to any one of claims 12 to 15, characterized in that, The configuration interface is also used to receive the at least two fault detection methods configured by the user.
17. A fault detection device based on cloud services, characterized in that, The apparatus is applied to a cloud management platform for managing infrastructure that provides cloud services. The infrastructure includes facilities located in multiple regions, each region including at least one cloud data center. The cloud services run on at least one server located in at least one cloud data center in one of the multiple regions. The apparatus includes: The service configuration module is used to provide a configuration interface for fault detection services. The configuration interface is used to receive system configuration information and a first service requirement of the application system input by the user. The application system adopts a layered architecture, which includes multiple layers and multiple topology nodes. Each topology node is located in one of the multiple layers. The system configuration information is used to indicate the topological relationship between the multiple topology nodes. The first service requirement indicates that fault detection is performed for at least two of the multiple layers. The data acquisition module is used to acquire the running data of the multiple topology nodes during the operation of the application system. The running data of the multiple topology nodes is used to determine the fault detection results corresponding to the at least two levels. The result display module is used to display the fusion detection result of the application system on the result interface of the fault detection service. The fusion detection result is obtained by fusing the fault detection results corresponding to the at least two levels according to the topology relationship.
18. A fault detection device based on cloud services, characterized in that, The apparatus is applied to a cloud management platform for managing infrastructure that provides cloud services. The infrastructure includes facilities located in multiple regions, each region including at least one cloud data center. The cloud services run on at least one server located in at least one cloud data center in one of the multiple regions. The apparatus includes: The service configuration module is used to provide a configuration interface for fault detection services. The configuration interface is used to receive the topology type of the application system and the third service requirement input by the user. The third service requirement indicates N types of faults of the application system, where N is a positive integer greater than 1. The data acquisition module is used to acquire runtime data of the application system during runtime; The detection execution module is used to determine at least two fault detection results based on the running data using at least two fault detection methods, with each fault detection result corresponding to one fault detection method; wherein, under each fault type of the topology type, the at least two fault detection methods have corresponding priorities; The result display module is used to display the fusion detection results of the application system on the result interface of the fault detection service. The fusion detection results are obtained by fusing the at least two fault detection results according to the respective priorities of the at least two fault detection methods.
19. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 11, or to perform the method as described in any one of claims 12 to 16.
20. A computer program product containing instructions, characterized in that, When the instruction is executed by a cluster of computer devices, the cluster of computer devices performs the method as described in any one of claims 1 to 11, or performs the method as described in any one of claims 12 to 16.
21. A computer-readable storage medium, characterized in that, The method includes computer program instructions that, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 11, or perform the method as described in any one of claims 12 to 16.